Showing posts with label scale of fluctuation. Show all posts
Showing posts with label scale of fluctuation. Show all posts

Sunday, October 25, 2015

Unifying Machine Learning to create breakthrough perspectives



Machine Learning – a unifying perspective & new paths

PG Madhavan, Ph.D.
Chairman, Syzen Analytics, Inc., Seattle, WA, USA
pgmad@syzenanalytics.com

Dr. PG Madhavan is the Founder of Syzen Analytics, Inc. He developed his expertise in Analytics as an EECS Professor, Computational Neuroscience researcher, Bell Labs MTS, Microsoft Architect and startup CEO. PG has been involved in four startups with two as Founder.
Major Original Contributions:
·       Computational Neuroscience of Hippocampal Place Cell phenomenon related to the subject matter of 2014 Nobel Prize in Medicine.
·       Random Field Theory estimation methods, relationship to systems theory and industry applications.
·       Early Bluetooth, Wi-Fi, 2.5G/EDGE and Ultra-wideband wireless technology standards and products.
·       Currently developing Systems Analytics bringing model-based methods into current Analytics practice.
PG has 12 issued US patents and over 100 publications & platform presentations to Sales, Marketing, Product, Industry Standards and Research groups. More at www.linkedin.com/in/pgmad

Pedro Domingos in his new book, “The Master Algorithm”, has done us a huge favor. As is true of any emerging technology field, Machine Learning (ML) is a “bag of tricks” today; it takes a while for a unifying framework to emerge. Then, one can see various aspects of ML as special cases of a general theory rather than a grab-bag of tools and techniques.

Pedro has taken a great early step to such unification. He has collected all major ML initiatives into a taxonomy that makes sense; five schools of thought: the evolutionaries, connectionists, symbolists, Bayesians, and analogizers. I believe this does not go far enough in the unification of ML thought however . . .

From the early days of “ML”, I see Pattern Recognition and Classification as a better unifying perspective. In particular, the classic textbook of Duda & Hart, “Pattern Classification & Scene Analysis”, published in 1973 is my starting point!

Duda & Hart’s approach in simple terms is as follows. Given labelled samples, obtain a class description consisting of either a distance metric (Euclidean, intra-class, etc.) or a probability density function and then derive a decision rule (Maximum A-posteriori Probability, Bayes, etc.) from the description. The decision rule specifies a decision boundary in feature space among classes.

Alternatively, decision surface can be derived directly from labelled samples which is then called a “Discriminant Function”, perceptron being an example. Then, most if not all current ML techniques can be seen as dueling methods to derive Discriminant Functions!

Discriminant Functions can be linear or nonlinear (neural network with back-propagation, deep learning, support vector machines, kernel PCA, etc.) and outputs can be binary, integer or real valued. Various learning algorithms can be seen as belonging to the family of iterative/ recursive/ adaptive learning algorithms (Least Mean Square being a great old standby!) that update the parameters of the Discriminant Function as new data arrive.

In the discussion above, features were considered as “static” and not context-sensitive (for identifying a word within a sentence as an example). Context-sensitivity or Dynamics can be added to improve classification by incorporating Markov models (or Hidden Markov Models for tractable computations). Markov model is a special case of State Space Models which are well-studied in Systems Theory.

Setting aside Supervised Classification when labelled samples (or “desired signals”) are available, what can we do when there is no supervision? This is the realm of much harder Unsupervised Learning, which is very useful in transforming basic features into more and more meaningful ones. One usually brings in some overall desirable property to guide unsupervised learning. From the domain of “blind processing” (Radar signal processing, for example), Mutual Information among classes can be minimized as a learning process in the belief that the “best” classification happens when the classes have least overlapping information (better “efficiency” in representation).

Instead of entropy-related quantities that are hard to estimate, it is likely that Scale of Fluctuation which is related to “order” and “state space volume” may be a quantity to optimize for a new unsupervised learning process. (For more information on Scale of Fluctuation, refer to my papers, “Instantaneous Scale of Fluctuation Using Kalman-TFD and Applications in Machine Tool Monitoring”, 1997 & “Kalman Filtering and time-frequency distribution of random signals”, 1996).

In all of the existing ML bags of tricks, we are still staying at the surface level! We are modeling the attributes or data DIRECTLY. What if we went one level deeper? Model the SYSTEM that generates the data! Syzen Analytics, Inc., takes such an explicit approach in what we call “SYSTEMS” Analytics” which has already demonstrated significant value in business applications.

In Syzen’s retail commerce application, our Systems Analytics approach hypothesizes that there is a system, either explicit or implicit, behind the scenes generating customer purchase behaviors and purchase propensities. This ‘one-level-deeper system model parameters’ can be more effective for pattern recognition and classification purposes instead of the data that the model generates! There is a long history of model parameters providing better estimates (in power spectrum analysis, for example). Scale of Fluctuation mentioned earlier seems to have another desirable property of quantifying “coupling” among deeper-level model parameters.

Context-sensitivity dynamics is a very good avenue to exploit. The dynamics could be over any independent variable (time always comes to mind first but it is only one of the possibilities). As I noted in my recent blog (“SYSTEMS Analytics – the next big thing in Big Data & Analytics”), “Extensions to Systems Analytics in the future will be inspired by the insight that in reality, data exist in *embedded* forms in preference and influence networks which are distributed in time and space” AND other independent dimensions (shopper preference, for example).

Let me pull all of the notions discussed so far into a diagram.




Once the patterns have been recognized and classes identified, the resulting classes can be used for all sorts of applications such as Recommendation Engine, Language Translation, Fraud Detection and many others. The approach I outline above allows you to take a unified approach till the application development stage. In doing so, the unified approach also points out new paths ahead for ML!

Some readers would have noticed an undertow of dichotomies while reading this “opinion piece”: Theoretic vs Heuristic; Formal vs Ad hoc; Mathematics vs AI; Electrical Engineering vs Computer Science academic departmental affiliations! I am firmly in the former camps. However, as an engineer, I am personally happy to start with heuristic solutions but quickly put them on firm mathematical foundations before “gotchas” and unintended consequences of ad hoc methods catch up with me. 

The unification of ML proposed here opens up a multilane highway – join the journey and create more breakthroughs with us or on your own!

In this blog, I have not provided many references – web search will get you most; Pedro Domingos’ “The Master Algorithm” book is an excellent source of ML-related literature. For the newer and less familiar work, please contact me directly.



Sunday, November 24, 2013

X-Event Marketing


Dr. PG Madhavan developed his expertise in Data Analytics as an EECS Professor, Computational Neuroscience researcher, Bell Labs MTS, Microsoft Architect and startup CEO. Overall, he has extensive experience of 20+ years in leadership roles at major corporations such as Microsoft, Lucent, AT&T and Rockwell and startups, Zaplah Corp (Founder and CEO), Global Logic, Solavei and SymphonyEYC He is continually engaged hands-on in the development of advanced Analytics algorithms and all aspects of innovation (12 issued US patents with deep interest in adaptive systems and social networks). More at www.linkedin.com/in/pgmad


This blog is a coming together of a bunch of my past blogs and my reading of a recent book by John Casti called “X-Events”.

PG Blog: “What does ‘Emergent Properties in Network Dynamics’ have to do with Shopping?”; http://pgmadblog.blogspot.com/2012/10/what-does-emergent-properties-in.html?m=1
PG Blog: “Network Dynamics & Coupling: Shannon’s Reverie Reprised . . . “ http://pgmadblog.blogspot.com/2012/11/network-dynamics-coupling-shannons.html?m=1
Book by John Casti, "X-Events: The Collapse of Everything";

Reading X-Events book triggered ideas from some of my past academic work (http://www.jininnovation.com/SoF.PDF) combined with my present startup, “Syzen Analytics” (http://www.SyzenAnalytics.com/), leads me to suggest an interesting way to create “pocket” X-events in Retail Commerce via X-Event Marketing or “XM”!

Let us start at the beginning . . .
1.     Shannon understood the importance of rare events in conveying information. He sought a mathematical formulation to capture this important property of our brains and came up with the definition of “entropy”. As I had made clear in the original blog, this is a fictitious reverie I made up to give one the guts to turn other valuable physical/ practical insights into mathematical forms!

2.     Similarly, from my past research in Neuroscience, I had some strong impressions that have refused to go away despite the passage of decades. One is the activity that precedes the most significant X-Event in each of our lives, Death! Unreplicated and unpublished observations in our neurophysiology labs had shown that one of the surprising events that happen before a mouse with brain in-dwelling electrodes dies is the all-out firing of “complex spike cells” (which when the mouse is alive and performing “place cell” tasks are painfully difficult find and record). In other words, the “natural state” of the brain seems to be excitatory and the work of a healthy brain seems to be to assert control and coordination by keeping the *inhibitory tone* high so that all hell does not break loose (by cells firing away with no coordination or control). I speculate a regime-transition of the sort shown below in death: the right-most picture is indicative of brain cells firing away in an uncontrolled fashion (just before a major “X-Event”, death or epileptic seizure) with the left-most indicative of a properly functioning brain (“nicely” coupled). The next story will illuminate the idea of the middle picture where the “system is uncoupled” with the system poised for all its energy to pile into a single “mode” and create a major suppression of inhibition in the brain.

3.     Following “Shannon’s approach”, how do I make these intuitions mathematical? Let us look at the right-most picture. As I outlined in my “Shannon’s Reverie Reprised” blog, everything is firing away and the potential field will show up equally across the scalp, much like a single “wavefront” that reaches all the regions of the scalp simultaneously. Imagine you are at a beach looking out to the sea and gentle waves are rolling in – let us say in parallel to the beach. If you are standing knee-deep in the water and look in a direction parallel to the beachfront (i.e., up or down the beach), the spatial frequency in your “look direction” (or the frequency of “corrugation”) is 0 cycles/meter! Much like the ocean waves, the single wavefront of activity is nearly “constant” across the distributed neocortex and its relevant strength can be thought of as the power at zero frequency. I happen to know that power at zero frequency is called “Scale of Fluctuation” or “θ” in random field theory. From the previous work of Eric Vanmarcke (“Random Fields”, 1983, with an updated edition in 2010), θ is defined below. I refer you to my blog and academic paper mentioned at the outset for the gory details!
 
Is Theta as powerful and useful as Shannon’s entropy . . . time will tell. As you see in my academic publication, Theta does have some intriguing properties in the case of 2nd order linear time-invariant systems which may indeed prove useful in the future! Adding results from my simulation and data analysis of machine tool chatter, my summary observation is that:

Large value of Theta can be good or bad. It indicates “Coupling”: *good* coupling as a nicely functioning brain (or lathe) or a *bad* coupling as in death (or chatter)! The key here is that low value of theta can be a “predictor” of impending X-Event!

As you can imagine, if our speculation above holds up, we may be able to predict X-Events (or Taleb’s “black swans”) such as 2008 Great Recession or Tohoku earthquake and eventual tsunami. In the case of Tohoku, a few extra hours of warning may have helped Fukushima engineers reach a consensus and move the power generators to a higher location (thus avoiding the nuclear disaster).

XM:
In this blog on X-Event Marketing (“XM”), my purpose is different. I want to *create* desirable, “pocket” X-Events to beneficial ends!

One of the preconditions of “interesting” dynamics in systems is sufficient “complexity”. This can be taken to mean multiple actors with high degrees-of-freedom, dense interconnections with linear and non-linear coupling and in general, hard to analyze! One such natural system is the “shopper network”.
A modern-day shopper is enmeshed in an ever-varying network of preferences and influences (from friends and family) and embedded in a network distributed in time and space – just the type of complex systems where “interesting” dynamics can arise. We want to create desirable “pocket” X-Events in the shopper network with the purpose of increasing business value for the shopper or the seller or both.


This is the so-called “Purchase Funnel” that an individual shopper goes through when executing a product transaction. The marketing, merchandizing and offer activities are devised to impact the shopper at predetermined points of importance in the shopper’s decision making process. For example, a well-conceived offer delivered via shopper’s mobile phone at the point of “Desire” may precipitate a Purchase “Action” – a win for the seller (in this case both the retailer and the manufacturer of that product; for example, Walmart and Procter & Gamble, respectively).

This was one shopper. Now, you have to consider the shopper being embedded in a network. At this point, it may be useful for you to review my earlier blog, “Social Network Theory” http://pgmadblog.blogspot.com/2013/08/social-network-theory.html, where Barabasi’s work in Network Graphs and the following terms are discussed:
1. Clusters (of friends who “hang out”).
2. Weak links (to your long-forgotten high school classmates).
3. Hubs (politicians and others with massive number of contacts).

Clusters and hubs are some examples of the influence on an individual shopper. There are many such shoppers, all interconnected, when studying the dynamics with a view to creating a “pocket” X-Event. Even this does not complete the complexity picture – the shoppers are distributed in space (you are in NYC and your buddy from Seattle tweets about a cool product that he bought) and time (good and bad product reviews reaches you at various times and not in a synchronized fashion).



That the Shopper Network is sufficiently complex to engender “interesting” dynamics must be beyond a doubt by now! The picture above may be a reasonable representation of the Shopper Network (WITHOUT the time element captured – we will need a video showing the undulations in the ”heat map”). The “heat” shown as ‘yellow’ spots indicate the total dollars spent each day on a particular product on a particular day, aggregated to a county for all of USA. Clearly, this will vary from day-to-day.

How do we capture the dynamics of the Shopper Network picture shown above in a tractable manner so that we can understand the dynamics and then track the effect of any manipulations we do in the network so that we can close the loop to adaptively adjust our manipulations to achieve a desired effect?

This is the essential question of XM. Of course, there are other important questions such as the type and timing of manipulations (Offers to selected customers? A country-wide branding effort? Is it a new product or an existing one? And so on). As a systems engineer, I will find a partner who is more qualified than I am in answering the marketing questions; I will focus on Systems Analytics tools to create a quantitative execution infrastructure to deploy marketeers’ ingenuity!

Engineers love scalars! I suspect it is because having a single control variable is easier to manage than many. Probability theory is replete with reduction to scalars – I am talking about mean, variance, correlation coefficient, Entropy, Theta, . . . The full underlying information is available in the joint probability density functions but they are a bear to handle! So, we reduce it to a scalar – clearly, in the process, we have thrown away MUCH information but what is left is what we plan to use; so our justification for selecting and using a scalar for a particular engineering task is very *operational*.

We will use the scalar, Theta, to answer Shopper Network analysis and control question I asked a few paragraphs earlier.

Scenario: New Product Introduction
Procter & Gamble is going to introduce a new detergent and wants to create a ground-swell of purchase interest (our “pocket” X-Event).

·         P&G introduces the product quietly and collects the heat-map data from T-log data of their partner Retailers.
·         Sysan (our startup) processes the heat-map data: total dollar spent each day on the new detergent, aggregated to a county for all of USA.
·         Sysan calculates Theta (a single number) as the baseline measure.
·         P&G Marketing has developed 2 major tools for the launch of the new detergent: (1) nation-wide TV ad campaign and (2) a 10% cash discount for the first 1 million customers in US.
·         P&G runs a 1-week TV ad campaign and while collecting T-log data.
·         Sysan monitors Theta for the week.
·         P&G suspends the TV ad campaign and starts the cash discount offer.
·         Sysan monitors Theta and finds that Theta is entering a “Decoupled” regime with low values of Theta.
·         Sysan advises P&G to resume the TV ad campaign in an effort to trigger “bad” coupling among various counties of US.
·        This pushes Theta into a high-value, the coupling between counties get tight and a “pocket” X-Event occurs where there is a ground-swell of purchases of the new detergent country-wide!


NOTE that in this fictitious example, P&G would have pretty much done the various marketing campaigns that I speculate here on their own. The KEY point is that P&G does NOT have a scalar (or any) measure to get immediate feedback regarding their efforts, combined for the whole country. The ability to quantify “closed-loop” intervention is the value of our humble scalar, Theta, in X-Event Marketing!

Clearly, what manipulations are best applied to the network and when is an open question. Utilizing Theta, Sysan is developing methods using “massive simulation methods” (for background, please see: John Casti, “Would-Be Worlds: How Simulation is Changing the Frontiers of Science”) that can predict the effect of various interactions at various instances of time and points in space, thus providing guidance for the most effective use of vast sums of money required for ad campaigns and discount offers.

Stay tuned for our initial results . . .

PG

Thursday, November 8, 2012

Network Dynamics & Coupling: Shannon’s Reverie Reprised . . . updated


Shannon’s (fictional) Reverie . . .
Claude Shannon woke up one morning and said to himself, “Think of how our brains operate – it habituates to repeated stimuli but pays attention to a rare stimulus; things that are rare must carry a lot of information!” So, what do I know about quantifying rare or unlikely things? I know that things that are highly likely are highly probable; so unlikely-things or novelty can be thought of as the inverse of probability. But snap! Probability goes from 0 to 1; I need a “squashing” function around it so that the novelty measure does not blow up fast but at the same time, very low probability things (highly novel things) are highly weighted. How about . . .?




A bit of cleanup via taking expectations, probability density functions and some proper logarithms and we have Shannon’s famous equation for “information”,


When you reduce a (joint) probability density function to a scalar, there is a lot that you throw away; the “trick” is that the scalar that you come up with captures some aspect of reality that is *useful*. As the explosion of communication technologies in the past few decades shows, Shannon’s scalar sure did!

I am not claiming that this is how Shannon did his research but this is one way to approach new insights that you may have and their quantification.

Social and Other Networks:
In my recent blog on Social Networks, “What does ‘Emergent Properties in Network Dynamics’ have to do with Shopping?”, I noted the following: “Facebook connects us in a vast network – this is only a first step. The deep reason for the fascination with social networking can be understood from the shopper example. Shoppers are enmeshed in an ever-changing network of social interactions and preferences. Today in Retail Analytics, data are treated as isolated bits of information. In reality, data exist in *embedded* forms in preference and influence networks of the shopper as well as distributed in time and space.”

As you know, interest in understanding such Social Networks and controlling their “dynamics” via “influence functions” of nodes, etc., are at a fever pitch – advertising, retail and many other day-to-day eCommerce activities can benefit from a better conceptualization and quantification of social network dynamics.


We know that “coupling” in networks generate very interesting dynamics (see, Steven Strogatz, “Sync: The emerging science of spontaneous order”, 2003, for a very readable overview).  Consider the ultimate of all networks – the brain. When we do “brain mapping”, interesting patterns arise. The brain mapping pictures show scans of a “depressed” and a “non-depressed” person. In the Depressed case, the brain regions are NOT coupled whereas in the non-depressed or Normal case, there is significantly more coupling and more uniform activity across the entire brain. Note however that if we looked at such a scan for an epileptic patient during a seizure, the scan will be all “lit up” showing nearly-complete coupling – that is a degenerate case!

Reprising the Reverie . . .
Coming back to healthy coupling and following Shannon’s (fictional) thinking process outlined in the first paragraph, what do I know about quantifying “coupling”? I know that when the underlying sources are coupled, their “stimulation” of the neocortex is uniform and they show up equally across the network at the same time, much like a single “wavefront” that reaches all the regions simultaneously. Imagine you are at a beach looking out to the sea and gentle waves are rolling in – let us say in parallel to the beach. If you are standing knee-deep in the water and look in a direction parallel to the beachfront (i.e., up or down the beach), the spatial frequency in your “look direction” (or the frequency of “corrugation”) is 0 cycles/meter! Much like the ocean waves, the single waverfront of stimulation is nearly “constant” across the distributed neocortex and its relevant strength can be thought of as the power at zero frequency.

I happen to know that power at zero frequency is called “Scale of Fluctuation” or “θ” in random field theory. From the previous work of Eric Vanmarcke (“Random Fields”, 1983, with an updated edition in 2010),
For the initiated, the equation and the accompanying figures below are hugely meaningful! In the figure, the “height”, (g(0) times π ) and the “area” under the normalized autocorrelation function, ρ(τ), are marked in blue – this is “θ”! This is the graphical meaning of Vanmarcke’s equation for θ.


The result shown above is for a time series. For the brain mapping case, “θ” is 4π2g(0,0) of its 2-D normalized spectral density (same pattern follows for higher dimensional random fields). Calculations of θ for 1-D and 2-D cases are straight forward; ways to calculate “instantaneous” values of θ are also available using Kalman Filtering (introduced in my past publication, “Instantaneous Scale of Fluctuation Using Kalman-TFD & Applications in Machine Tool Monitoring”). Some curious properties of θ for 2nd order linear time invariant systems were also developed there. To recap the highlights –

From discrete-time linear time-invariant system principles, we know that constant damping ratio and undamped natural frequency contours in the z-plane are as shown on the left.
It is notable that for a second-order system, the constant θ contours shown on the right have remarkably simple geometric shapes. In fact, for θ = 1, the equation is quartic but very similar to a circle with origin at (0.5 + j0) and radius = 0.5!

Equation for θ = 1 contour is (x2 + y2) 2 + x2 + y2 – 2x = 0

There are more details in PG Madhavan, Theory and estimation techniques for Random Field Theory and "Theta" with practical applications: Instantaneous Scale ofFluctuation Using Kalman-TFD & Applications in Machine Tool Monitoring, SPIEProceedings, SPIE Vol. 3162, pp. 78-89, 1997.

Similar to Shannon’s scalar, H, θ reduces the joint probability density function to a scalar. Does θ capture some aspect of reality that is useful? The constant θ contours above seem to imply great significance as fundamental as natural frequency and damping – but at this time, such insights are not forthcoming!

Similar to Shannon’s scalar, H, θ reduces the joint probability density function to a scalar. Does θ capture some aspect of reality that is useful? Some real-world applications of θ from the past (see its use for machine tool chatter prediction) point to the following physical insights.



While highly speculative, previous studies and our “Shannon approach” suggest that θ is proportional to “coupling” and to “order” in a distributed node system whereas it is inversely related to “degrees of freedom (df)”. In the case of “df”, the concept is that more degree of freedom is a “dangerous precipice” for a distributed system where different parts are de-synchronized and they can spin off in different directions - pandemonium can ensue!

Large θ seems to indicate an ordered, widely-cooperative and well-functioning network; however, it is conceivable that a large θ may also indicate degenerate cases such as epileptic seizure or full-fledged chatter conditions (extreme cases of coupled sources and distributed action). An example is shown below.

Large θ on the left is a hallmark of “distributed order” whereas the large θ on the right, that of “locked-in order”. Low θ condition in the middle is visually indicative of disorder and the potential for degeneracy!


θ in Social Networks:
To help develop our intuition, let us consider some snapshots of geographcally distributed network maps. We have (1) Internet activity over continental United States, (2) LinkedIn infographic map and (3) Facebook social network map.

We notice strong coupling in the Eastern half of US among Internet nodes and similar features in the LinkedIn and Facebook networks. Our intuition is that the *bright spots* obvious in the Northeast US of the Internet map or the EU area in the Facebook map indicate more “correlated” activity. For simple network maps, we have developed primitive methods to estimate θ based on correlation functions.


Before we leave this blog, consider the brain map and the US map of the Internet. What we see in these pictures can be called “surface structure”, i.e., observed or measureable quantities. In the brain, the Surface Structure is created by activity deep within the brain (I am NOT referring to Chomskian linguistics model here). In the past, naïve physical modeling has conceptualized dipole oscillators in the “deep structure” of the brain giving rise to the Surface Structure as a starting point for theorizing. In the case of the brain, there are indeed Deep Structures (nuclei and ganglia and their dendritic potentials) giving rise to voltage variations on the scalp surface (earthquake tremors recorded on the surface of the earth and activity deep in the earth’s crust form a similar model).

Clearly, the Surface Structure of the Internet cannot be related to any actual Deep Structure in a physical model – there is no mechanical turk behind the Internet pulling the strings! However, even in such cases, conceptualizing observed activity as the resultant of implicit Deep Structure may be useful in developing analysis methods. The hope is that the physical model of network activity utilizing the concept of explicit or implicit deep structures with internal coupling will help advance our “analytics” tools for the extraction of patterns and information from spatially and temporally distributed networked systems.


www.JINinnovation.com
Dr. PG Madhavan was CTO Software Solutions at Symphony Teleca Corp. Previously, he was the CTO & VP Engineering for Solavei LLC and the Associate Vice President--Technical Advisory for Global Logic Inc.  PG has 20+ years of software products, platforms and framework experience in leadership roles at major corporations such as Microsoft, Lucent, AT&T and Rockwell and startups, Zaplah Corp (Founder and CEO) and Solavei. Application areas include mobile, Cloud, eCommerce, banking, retail, enterprise, consumer devices, M2M, digital ad media, medical devices and social networking in both B2B and B2C market segments.  He is an innovation leader driving invention disclosures and patents (12 issued US patents) with a Ph.D. in Electrical & Computer Engineering.  More about PG at www.linkedin.com/in/pgmad