Showing posts with label machine learning. Show all posts
Showing posts with label machine learning. Show all posts

Monday, September 9, 2013

Analytical software for analysts - are they way too complex?


Are there analytical software out there that actually make learning from data intuitive? I have experience with quite a few of these packages but none of them are intuitive for the average business analyst without making them useless after looking at data in one or two dimensions. While this is good for business, I must admit that it makes life difficult as the problems one has to tackle get quite mundane when responding to queries from the not so statistically literate. 

What would be the ideal requirements for one to actually be able to get ideas from data? Let us assume that the average user has a sense of the business he / she is dealing in. At the end of the analysis he should be able to get a sense of how to drive the business forward or at least has a good sense of what are some of the drivers that would explore further. Let us further assume that the average business user also has the ability to understand counter-intuitive results and can basically understand two dimension analysis and can possibly understand three dimension analysis but will be unable to move forward beyond that. 

Ideally when my business problems are well-defined (in the sense that I at least know what I want to solve initially even though I might realize that I need to solve something much larger later), then these tools should be able to at least drive some initial value for the analysts by incorporating these business requirements. But when I am sifting through data without a clue as to what I am looking for, how do I identify patterns that are meaningful and at the same time not require me to be in that business domain forever?

Regression analysis required significant understanding of the statistics to be able to confidently drive the analysis. CART / CHAID type algorithms are relatively easier to understand but I am not sure if there are decent implementations of a software that makes the learning from CHAID / CART intuitive. Bayesian networks or topological data analysis might be an answer but I have not worked enough with these to have a viewpoint on the implementation perspective. These are good with identifying patterns but do not necessarily make it easier for the business to get their reads better.

Ultimately I believe business problems need to be solved with the business context in mind and there are no general software that will enable that. Is it time for one to be created?

Tuesday, August 27, 2013

Machine learning ... will it tell me the future?

I recently got pulled into a conversation with one of our really smart analysts. He has a dreamy vision of getting into bureaucratic India and at the same time is really smart when it comes to coding. We got talking and he started describing this competition on Kaggle. While I heard him out and was also getting a bit excited about doing something hands on, I realized that the industry has really grown around me and I have not had the chance to appreciate the growth.

Kaggle is one of many websites that offer competitions. KD Nuggets, Analytic Bridge etc. are other sites that have these competitions. What is interesting is where the the solution approaches seeming to be heading to. A few years ago (or maybe many years ago - and it can be a separate blog topic), we (at grad school) discussed the coming of age of machine learning. Given large amounts of data, how can you get accurate predictions for different problems. With the focus being only on predictions, these algorithms were able to meet many statistical techniques purely due to the lack of any constraints that a data generating model would impose on a statistician. Why are we constraining ourselves from a hypothesis perspective? Has statistics lost out on the chance to be the next cool thing in the world and will machine learning take over? It makes sense to understand why is one even relevant in this day and age.

Many machine learning algorithms are inherently black box purely because of the way the characteristics relate to the object that needs to be predicted. While there are ways of understanding which characteristics are important and associated sensitivities, there is potential for it to be misleading if not diagnosed properly. Most machine learning techniques have a significant validation component to ensure that the algorithms are robust and can handle exception cases.

Where does this lead us to? One of the most interesting expectations from machine learning is we can live in a IRobot kind of environment where machines can predict survival rate based on their learning. Google has designed an algorithm that can identify cats (even though I am not sure what it would call it) and there is potential for machines to get smarter with time.

BTW here is a plug for one more analytics competition. Should be fun if you are in college!!! (It is quite rewarding from a financial perspective!)