Wednesday, March 5, 2008

Free data!

I have a longer post about selecting data for modeling, but for now, just know that the UN has put it's data statistical data online. Perhaps the nicest feature is that the site will search across all of their published datasets. I like adding macro level data into modeling and customer insight projects and the UN is a good source.

Tuesday, March 4, 2008

We're number 1!

Interestingly, a search for Business Analytics Blog, on Google, brings up Da Facto in the number 1 spot. It is a little niche-y, but still.

What is the secret sauce in direct marketing?

I am a big fan what I call tribal wisdom. Kevin Hillstrom put together a post of database marketing tribal wisdom. The item that most resonates for me? Segmentation and treating segments differently for marketing purposes. This was one of the first things I took from my McKinsey experience. Much of what he talks about are just specific instances of differential treatment of segments. Nice post.

More on data quality

I was speaking with someone about ways to assess data quality for predictive modeling. I have written on tactically how to ensure quality data, but here is a little framework you can use when thinking about data quality. Your data needs to be accurate, granular, and complete. Accuracy of data: Does the data accurately capture the attribute (e.g., income) that it was intended to? Granularity of data: In every case, the more granular the data, the better the predictive modeling can be. Also, you can usually roll up granular data (from individual to Households, Households to zip+4, etc) to higher level if you need to (for analytic or appending purposes.) Completeness of data: Any given dataset is going to have missing data. Missing data is a funny thing. Of the three, accuracy, granularity, and completeness, the later is the one that you can most influence. Obviously, the less missing data, obviously, the better. But before choosing an appropriate remediation method you need to understand why the data is missing. If the problem is that the datafeeds are broken, you are going to need to get the feeds fixed. If the data does not exist, but you have enough coverage to do some predictive modeling, you can predict the values of missing data. Or you may just need to fill missing values with the mean or median values.

Monday, March 3, 2008

Linking analytics and psychology

Wired has an article on the 1 Million dollar prize Netflix is offering to the person (or team) that can improve their recommendation algorithm by 10%. Most of the competitors rely on fairly advanced math, but one guy is implementing behavioral economic principles and is competing against the big boys.

Wednesday, February 27, 2008

Dan's book is making the news

Here is a link to a NYT article about Predictably Irrational. Fuqua is doing something very smart. While Dan is on his book tour, they are scheduling Alumni and recruiting events around his schedule. I went to one of the talks. If you have a chance to hear him talk about his research, you should go. The research on deception is fascinating...

Monday, February 25, 2008

Visualize!

Here is a chart created by the New York Times staff showing box office receipts over time for every movie release in 2007. It is a neat chart showing how "bursty" blockbusters are and how Oscar contenders have a longer tail. Neat, but you have to work for that insight. Some kind of filitering or making the horizontal access not from the calendar date but on weeks of distribution, would have made the point clearer. People go gaga for this stuff, but the nice presentation obscures the insight that you want to confer.

Friday, February 22, 2008

How should you manage your relationships with recruiters?

As the old joke says, carefully.

I get calls from recruiters at a fairly constant rate and have 3 principles that I use when talking with them. Note that most of the folks I speak with are retained search recruiters, they are paid, in advance, to fill a position. Companies tend to use retained search firms for more senior jobs and the search firm has an exclusive arrangement with the firm. I occasionally get a call from a contingency search firm. These firms are paid when they fill a job and are not exclusive. Most of what I am going to say is applicable to retained search folks who have a more relationship based business. Contingency folks are more transactional, so relationship building may not be as critical. Even still, making friends is always worthwhile. On to the principles.

First, always take the call. I speak with every recruiter that calls, even if they are recruiting for a position that I am not appropriate for. Someone gave me the advice that you should cultivate a recruiter network. It was good advice. In order to build that network, you actually need to speak to them. So, even if I am not right for the position, I chat with the recruiter. Often, especially for analytic jobs, they don't know the space, help them understand the job rec (truly!) or refer them to someone else. I always try to make the calls a positive experience for both of us. Even if we just chat about raising kids.

Second, if you are not interested in a job really try to pass on a referral. I almost always pass on a name. This means that I need to spend a couple of minutes looking through my contacts and see who might be appropriate for their job. One of my former direct reports is my go-to guy for referrals. If I am not interested in a job, he gets the referral. This helps him build his network and helps me deepen my relationship with the recruiter.

Third, be honest in your assessment of your interest level. If you are not the right person for the job or the job is too small, tell the recruiter. Don't try to get the interview for the practice. You'll mess up your relationship with the recruiter. Having said that, I have let a recruiter talk me into interviewing at a company, even though the scope was too small. The company agreed and then built a job around my skills. I wound up not taking the job, but I was up front about my concerns and they decided to proceed with the process, anyway.

I consider my recruiter network a real asset. Every job offer I received (I had 3) was through a recruiter. A good recruiter network will make your job search much easier.

Thursday, February 21, 2008

At least this time she did not hit me

Different kind of post. I noticed that all of the big bloggers are name droppers. Here is my most recent brush with greatness.

I ran into Barabra Minto a couple of weeks ago. She wrote the Pyramid Principle that I recommend on the right side of the page. I met her once before, at a McKinsey Alumni event. She introduced herself by hitting me fairly hard in the arm and calling me a jerk. She thought I looked like Ricky Gervaise from the British version of The Office; Since he plays the role of a jerk, she thought it would be funny to inflict some pain on him/me. As I said to her, before you hit someone, you might want to make sure they are who you think they are. We got it all straightened out and had a laugh. Fast forward 2 years. I saw her again at a McKinsey event and re-introduced myself, related the story, etc. When I got to the looking like Ricky part of the story she said "You do look like him." At least she is consistent.

Wednesday, February 20, 2008

Moving bubble charts

While I am doing analysis, I don't worry too much about visualization. I am very hypothesis driven and data visualization is a great compliment to a more exploratory approach. Having said that, I do give a lot of thought to how I present the data to others. I don't think I have ever used animation, but after watching this video, I may need to expand my horizons. You can play with the Trendalyzer shown in the video here.
Report Portal has a moving bubble chart type that you can use with your own data. Very neat. Thanks to Jonathan Salkoff for turning me on to Trendalyzer.

Predicatably Irrational

My friend, Dan, just wrote a book called Predictably Irrational. Dan is a Behavioral Economist and is the person I know who is most likely to win a Nobel Prize. Fascinating guy. How interesting? Check out this interview of him talking about his new book. Here is the link to the book. Predictably Irrational

Friday, February 15, 2008

Skip level meetings

For part 2 of the Dafacto meeting series I present: Skip level meetings.

I love having regular skip level meetings. For those who have never heard of a skip level, the term refers to the direct reports of your direct reports. I try to hold a 1-1 meeting with each of my skips every 4-6 weeks. In my last organization, I had regular skip level meetings with about 9 people (all of my US folks. I did not hold regular skips with my overseas staff, though I met with almost all of them when I went to visit). Including the weekly 1-1’s with my direct reports, I typically had 7 hours of meetings with various staff members. Obviously, this was an enormous investment in time. Was it worthwhile? Absolutely. Otherwise you are dependent on your directs for information regarding things like staff morale, organizational issues, project progress, etc. I know staff found them valuable.

What did we talk about? Though the skip owned the meeting agenda (notice a theme), the skips were really focused on two things. First, staff development. I wanted to learn the staffs’ career and personal goals. We would use the time to talk about if they were making progress against those goals and what could I do to help.The time was very focused on their careers. In fact, after my first meeting, each person had to put together a development plan so we would have a structured conversation about their development.

Second, we talked about their projects. We talked about what was going well and what was not. People knew that I was fair and that if a project wasn’t going well, they would often tell me in those meetings. There were a couple of times where one of my directs was not living up to his commitment to his staff and the folks on his team needed a channel to be able to voice their concerns. Also, I would get to learn what was going well and use that information to give folks special recognition (I gave bottles of wine and gift cards) and visibility both in the team and to more senior executives.

One last thing. As a manager, it is very easy to move or cancel these meetings. After all, these folks are on your staff and are going to understand that stuff comes up. Right? As a leader, it is a terrible idea. If I had no choice, I would sometime postpone, knowing that this was sending a bad message to the staff member. I tried would make sure that they knew I did not want to cancel the meeting, and reschedule as soon as practical. The worst thing is blowing off the meetings. You wind up alienating the staff instead of helping them.

Wednesday, February 13, 2008

Staff Meetings

I am amazed at how many managers don't have staff meetings. Do they think they are not necessary? I think regular staff meetings are a critical management practice for, well, good management and leadership. And don't get me started on weekly meetings with direct reports or skip meetings. First things first. Staff meetings.

When I first started managing multiple groups, I found that I was repeating the same news over and over again. Also, some of the groups were working on complimentary projects or were interdependent and I was increasingly acting as the communication bridge between groups. So, I started to have staff meetings. I have always taken the same general approach to my staff meetings. First some general practices.

First, who owns the agenda and runs the meeting? Not me. Never me. Typically, one of my direct reports who I am starting to think about promoting. I want to give them the experience of running the meetings. They are going to need to do it themselves soon enough. Also, I don't see any reason to control the agenda. If I want to talk about something, I'll ask to have it on the agenda or just bring it up in the meeting.

Second, who attends? This depends on the organization and the needs. If you are often talking sharing confidential information, then a small, senior staff meeting is the way to go. If you want to use the meeting for sharing information across groups, then invite the senior folks and maybe their direct reports. I have seen people invite their directs on odd weeks and include the skip level folks on even weeks. I tend to go with inviting the larger group and make it clear that what is discussed does not go outside the family.

Third, how long? Between an hour and an hour and a half.

Typical agenda?

1. Company Updates. I use this time to talk any big company or departmental news that is relevant. Typically, this time was spent explaining why the company, my boss, or myself was doing something that did not seem to make sense to the staff. Some senior folks are pretty command and control. like to have a pretty tight reign on the discussions. I would rather use the time to share information.

2. Ken Rona updates. I give the team a sense of what I am working on. I do this so the staff can act as effective agents on my behalf and bring up any items that would materially affect my work. In this way, the people on the team can proactively participate in helping me solve my problems. Also, I really liked discussing my work in front of the team. Not only were folks helpful and pushed my thinking, it is good for morale. people like having the transparency.

3. Direct Reports update. My directs share their project lists. The agenda keeper is responsible for putting together an update project list for the group for every meeting. Mostly I focused these discussions on time lines. Are we going to meet this commitment we made. I think having to affirm the commitments in public, every week, keeps people focused.

4. Information sharing. We share interesting team outputs. I am a big fan of sharing information across silos. I often find that someone would do an analysis and share with the team, only to find that either there was a better way to do the analysis or that we could reuse the analysis for another internal client.

Oh, another tip. Have someone bring food. We did not use catering. We rotated this responsibility and reimbursed the cost of the food. It would have been easier (and more expensive) to have it catered, but I liked that people could bring their own style to the catering.

Tuesday, February 12, 2008

The Analytic Value Chain - Do the analysis

Finally! Lets do the analysis! Actually, I want to spend more time on what not to do.

I really did not expect to write much about how to select the analysis required to solve your problem (what!). I have assumed that you know the appropriate analysis to conduct to answer your original research question. In retrospect, maybe that isn't a great assumption, but here was my thinking: My posts are designed for managers of analytic teams and the folks that work on those teams who are still developing their managerial skills. Those people (I thought) should know the right analyses for a given situation.

Increasingly, I am questioning this assumption. I have found that analysts who are well trained in advanced analytic techniques and remember their training are not the rule; those who understand their business, and can creatively apply their training to a new business problem are rare. Maybe 30 percent of the statisticians I have worked with wholly qualify under my criteria (mostly at A fOrmer empLoyer. Hiring a director is what inspired me to write this thread. I will do a post on how to identify a high potential statistician). Most analysts are technicians and they have a hard time suggesting analyses for problems they have not seen before. This is not an indictment of statisticians, just applying statistical tools to business problems is hard. How hard? Let me illustrate.

Conducting an Ordinary Least Squares regression when you should be using a logistic regression is a common mistake. By using the simpler OLS analysis, you can get totally wrong conclusions, leading to incorrect decisions. I am talking about answers that may not even be directionally correct. So, it is important that you use Logit, even though it is more complicated, when the situation warrants (when the thing you want to predict is a yes or no). I won't hire someone who does not know when to use logit.

This mistake is so basic and no one using regression should make it. And it happens all the time. I know a company (none that I have ever worked for) that were using OLS when they needed to use Logit in a production environment, affecting hundreds of jobs for their clients a month. Did the statisticians who built the system know better? I don't know. I do know that the company has advanced statisticians working for them and who know better and are advocating for change. This is why you need the talented experts. They stop you from doing really dumb things.

As much as I would like to, I can't map out the correct analysis for any given business situation. Having said that, there are some common situations that I have come across over the last couple of years that you should watch out for.

Regression:
If you are using regression, you need to pay attention to the frequency distribution of the dependent variable. Unless the dependent variable is continuous, has a relatively large range, and is normally distributed, Ordinary Least Squares regression is not going to give you the right answer. You may need a more sophisticated analytic technique. Some rules of thumb: If you are using OLS on a binary variable (think yes or no) you are going to need a more sophisticated technique. Also, watch out if the dependant variable has a natural floor or ceiling. So, income is a good example. Very few people make less than zero dollars. So, zero is the floor. If you have an floor, then you may need to go with a Tobit. Depends on the distro of your dependant variable. If the dependant is normally distributed, then you are probably ok. If not, Tobit...

Impact of Seasonality/Time:
Most folks come up with some arithmatic technique to model seasonality. I hate this. You can never unpack the drivers of your dependant variable. Instead, you should use some kind of time series technique; e.g., ARIMA. ARIMA will let you figure out what the real drivers of behavior over time, as well as taking time into account.

Segmentation:
Check out the Kenny's one rule post on segmentation. The short version: For segmentation, don't try to do too much with one segmentation scheme. I think of segmentation like regression. You build a segmentation scheme for specific purposes.

Designing Experiments
In a direct marketing context, at least done correctly, testing is continous. I have had a number of occasions where the tests are not readable due to a business ownerer needing more volume and killing control groups or specific test cells when doing complicated designs.


Winding down
The big statistics software vendors have not yet developed bullet proof tools to help inexperienced analysts to do the appropriate analyses. My advice is to hire an analytics expert and teach them your business. Don't try (too hard) to find someone who is an expert in both analytics and your industry. Industry knowledge is much easier to teach than anaytic expertise. Doing the analysis is hard. And can take a while (it once took someone on my team three months to build a data set and do the analysis for a difficult business problem. But he nailed it and it changed how our business partners thought about the drivers of their business.

A cool site for the analyticly minded

Have you checked out Propser.com? It is like ebay for loans. You can be a Borrower or a seller of Loans. David Ye turned me onto it a while ago and I just setup an account. I like the low risk intro into the world of credit...

Sunday, January 13, 2008

New baby

Posting has been slow with the job search (took a ton of time) and my wife adding a new (and healthy) baby boy, Doyle, to the family on December 31st!

I was fortunate enough to have multiple offers and have accepted an offer that keeps us in DC. The new job starts in early February. More details once I start.

The baby is healthy, though we had a little scare involving the NICU for a couple of days. Baby is home and way over his birth weight. He was the last baby born at Georgetown in 2007. Insert tax deduction joke here.

All in all, January has been a great month.

Tuesday, December 11, 2007

Kenny's 1 rule of segmentation

I was reading the segmentation post on Analytic Engine and it reminded me that I wanted to post my standard segmentation rant. First a story. If you don't know anything about Marketing Segmentation, check out the Wikipedia post.

When I first got to A fOrmer empoLoyer (think a large web portal), I found that we were big users of Prizm clustering, by Claritas. For those of you who have not seen the Prizm product, it is cool. The folks at Claritas have taken the whole US population and ran a cluster analysis on, well, us. They have classified each household into 66 unique segments. The segments group households that have similar lifestyle and socio-economic traits. They are a really neat way to do market to a target a pre-defined demographic and get some basic insight into who your customers are.

One problem, though, A fOrmer empoLoyer's targeting efforts were not based on simple demographics. We built fairly sophisticaed models to predict who would accept a given offer. In our case, Prizim clusters were rarely predictive, over an above the other variables we had available. People at the company had tried to use Prizm (and some of the other Claritas products) for all kinds of marketing-y things and they just did not provide value.

So, some bright person at A fOrmer empoLoyer said "we need to develop our own segmentation scheme. We'll provide a Universal Segmentation" (that is really what it was called) that can be used for any marketing or targeting activity." Strike two. The Universal Segmentation was a worse solution than Prizm and never made it out of the lab. Actually, A fOrmer empoLoyer took several runs at building a custom segmentation scheme using cluster analysis. None of them were found to be useful. As an aside, I mentioned the Universal Segmentation project to an expert in segmentation and he laughed and laughed. It was a bonding moment.

So why did A fOrmer empoLoyer have some much difficulty in using a tool that is in use by marketers everywhere? In both cases, the segments were not built with A fOrmer empoLoyer's needs in mind. They tried to do too much.

To finish the story. A fOrmer empoLoyer had been sending out a mass email to the whole customer base as a way of stimulating engagement and generating page views. The program was, uh, less than effective. I convinced the Customer Engagement folks to let us build a custom segmentation scheme for their email newsletter program that categorized each customer on the basis of the content they visited. In the spirit of transparency, Omniture did most of the work, they had the data. The thought was that the content someone viewed was a good proxy for their interests.

From the segmentation, we found that A fOrmer empoLoyer had less than 20 unique customer segments, but that very few segments described most of our customers. The Member Engagement staff started to create newsletters for each of the large segments. We had just started to use the segmentation scheme before A fOrmer empoLoyer stopped marketing, but the first campaign had much higher open rates than the mass mailing approach.

The lesson here is that when you initiate a segmentation project, you need to be really thoughtful about what you are trying to accomplish. Don't use every variable that you have available and just start cranking the k-means. Your segments will not be interpretable, and thus (you never see thus any more) won't be actionable. You need focus.

Instead, think about what you are trying to accomplish (e.g., be able to classify your customers into demographic groups) and build datasets that only include variables that are actionable or would give you insight about your customers. Build a segmentation scheme for one purpose. And when the guys at Claritas say "we have been doing this for 20 years. We have this nut cracked. Use our segmentation scheme", ask yourself, why do they have several segmentation products?

I could say some other stuff about creating segments, but if you focus on what you are trying to accomplish and build your dataset accordingly, you should have good results.

Monday, December 10, 2007

Break Down The Wall!

The other day, someone asked me what I do to break down silos between functional teams. First, let me say that I don't think breaking down silos is an academic exercise; you almost always generate value by talking to folks outside your group, department, business unit, whatever. I can think of at least 10 examples of really impactful projects that have come out of "Silo Busting." Second, creating an environment for silo busting is not hard. It just requires some time. However, the way you attack your silos varies depending on the construction material.

Just keep in mind, some silos are harder to bust than others. If structural conflicts (e.g., working together means that a rational person would do worse off by cooperating) or long standing animosity between silos exist, the problem becomes a lot tougher.

Set expectations.
When I first take responsibility for a team, I set the expectation that we are going to make things better. Every day. And that means innovating. And that in my experience, innovation often comes when two people who are very different talk to each other about their areas of expertise and their problems. So, there is just an incentive to getting to know your colleagues and what they do; Opportunity for improvement often results.

As an aside, I actually have an expectations document that I share with all of my staff; innovation in the service of continuous improvement, etc is one the core values listed in the document. Also, I set an example. I am often going out to lunch to meet with folks whom I have no obvious business connection and I encourage my staff to do the same.

Information sharing
In most cases, people are working in silos not becuase they enjoy it but becuase they are too busy to share information. There are some easy tactics that you can employ to facilitate information sharing. A couple of things I do is have whole team staff meetings where and monthly round tables. During these meetings, we discuss the projects we are working on if there are ways that the group can help. Barriers fall very naturally. Also, as I mentioned above, I try to go out to lunch with people outside my business unit and encourage my staff to take their clients (and anyone else they think they should get to know better) out to lunch. To provide an incentive, I offer to pay for it.

Personality or Competence
Other times, it is a personality or competence issue. Here, I run at the problem head on and really set the expectation that the staff has to work together effectively. I then put the staff members in situations where they have to work together to be successful (e.g., some special business initiative.) The expectation of being an effective collaborator also becomes part of their development plans and makes sure the problem gets fixed. If we get to the development plan stage, then I am going to be paying close attention and actively trying to help the staff member be an effective collaborator. Typically, by working together, people reach some accommodation, build some level of relationship, and break down the silos.

If it turns out that the problem is due to competence, well, no amount of relationship building is going to matter. If the staff member on my team is causing the problem, then a performance plan is going to be put in place. If it is a staff member outside my influence, I will likely have a conversation with their boss about their performance. Good luck.

Saturday, December 1, 2007

The Analytic Value Chain - Understand the relationships in your data

One the dataset is together and the data is QA'ed, New analysts want to dive right in and start mining the data or building statistical models . More experienced hands just putter. The run descriptives, they check to see relationships that they thought they would see actually exist (age and income are positively related, for example), creating some histograms to look at the frequency distribution of the variables.

Why all the putterage? An effective analyst needs to develop an intuition about the data. Without that intuition, they won't have the common sense required to make some of the judgment calls they are going to need to make when they get down to modeling. Data analysis requires judgement (which is why you hire experts to do it) and without a good intuitive understanding of how the data behaves, the judgement calls to come are going to be on an unstable foundation. And you get to understand the business dynamics (by understanding what variables drive outcomes) in a way that the people who are running the business can not.

Another point, you are going to find relationships that were unexpected. These unexpected relationships are things you should take note of, you are going to want to follow-up on them and understand if they are real. These unexpected relationships are the beginning of finding what I call "Game Changing Analyses", the home run of data analytics. What do I mean by game changing? Analyses that lead your business partners into changing their business strategy, that indicate that activities that are fundamental to the business need to be reassessed. A good analyst should be able to have this kind of impact multiple times a year.

For those data analysts who read the blog, developing the intuition is critical for your ability to have impact and your reputation. Imagine that you are presenting some work to a senior executive and they ask a question that you can't answer based on your analysis. If you have a good understanding of the data you can say "I can't answer your question definitively, but given what I know of the data, I believe this to be the answer to your question. Of course, I will check the answer as soon as I get back to my desk." To the extent that you consistently get these kinds of questions right, you be come a trusted resource for the executive. I know more than one analyst whose job was saved because they proved over and over again that they were an expert in the dynamics of a business unit.

Upshot, spend the time getting to know your data . If you are a manager, build time to explore data into your project plans. Understanding the relationships in the data is a critical part of the data analysts and managers job.

Monday, November 12, 2007

Change Managment - Creating a data driven boss

Avinash has a nice post on helping your boss become more data driven. I generally like lists of best practices and he is a gold mine of the things. While he is focused on Web Analytics, most of what he espouses is relevant in the business analytics world. Check it out.

The Analytic Value Chain - QA the data

Related posts
The Analytic Value Chain introduced
Defining the problem
Determine Data Requirements
Locate Relevant Data
Extract, Transform, and Load the data

Sorry for the long delay between posts. I have been focused on my job search and just got my head above water.

I won't do a big song and dance on this part of the value chain, I already did a post on QA'ing your data, here.

A fast story on how the prevalence of data quality problems . I was speaking with someone whom is an expert in the direct marketing industry about a brief sojourn in consulting. She went into consulting because she likes doing novel things and she thought that consulting was going to provide that variety. She was lamenting that all of the engagements she worked on were similar, with "correcting data hygiene" as how she spent most of her time. The problem is endemic.

No matter where you get your data, a third party or an internal data mart, you have to check the quality of the data. And regularly. Assume data is guilty until proven innocent.

Tuesday, October 23, 2007

AOL Layoffs

Posts are going to be a bit slow for a couple of weeks. AOL laid off about 80% of my business unit (the access business) and I was, as they say, impacted. When you stop doing marketing, you don't need folks in Marketing Analysis. If anyone has a need for business analysts, please contact me. A number of my staff are looking for work.

My resume is here.

Saturday, October 13, 2007

"Growing" a SAS analyst

The other day, someone asked me how to “grow” a SAS business analyst. My first thought was “Let Capital One do it for you”. I then got to thinking about what it means to “grow” someone. I think the question was really “what skills does a SAS analyst need?” I was talking to David Ye, who is a senior manager on my stats team about this problem and he noted that there is often a difference between SAS programmers and SAS business analyst. Just to be clear, I am talking about a business analyst role.

Putting aside things that make a good analyst (being a voracious and tenacious problem solver, have a good understanding of your business dynamics, good written communication skills, etc), an analyst who relies on SAS as their primary analytic tool needs to:

  • To be able to pull their own data (SQL skills and Proc SQL)
  • Know how to use SAS efficiently (can’t overstate the importance of this, think temp tables…)
  • Have a good understand the analytic procedures available (and options) they need for their job (I like Proc Means, Anova, Reg, Corr, Cluster, Factor, Chart, Plot, and Tabulate. Also, if you are doing serious experimental work, you need GLM and Mixed)
  • How to leverage the various programming options (Macro, SAS Code)

Some of these are easier than others to develop. Most of these skills are hard won, so if you are trying to train someone up on anything but the most basic Procs, my advice would be to hire an experienced person and have them train new staff. Someone who is new to SAS, in my opinion, needs someone close by to answer questions.

Wednesday, October 3, 2007

The Analytic Value Chain - Extract, Transform, and Load the data

Related posts
The Analytic Value Chain introduced
Defining the problem
Determine Data Requirements
Locate Relevant Data

Extract, Transform, and Load (or ETL) is database administrator talk for getting data ready to analyze. For those who want a more formal definition, check out Wikipedia. I'll talk a little bit about each, but most of the action is in Transform.

From the data analyst perspective, there is not much to say on extraction; you need to get the data out of your systems and you need people who have the requisite skills. Think SQL. If you are hiring data analysts make sure they can write SQL and can architect a simple database. It will be a big help when they need to merge data sets or give requirements to your IT staff.

As I said above, transformation is much more interesting, involving both cleaning the data and creating new variables for analysis. Lets talk about data cleansing first. By data cleansing, I mean how your team handles missing values and outliers.

Though unglamorous, data cleansing is a critical task. Without having a consistent method for data cleaning, everyone will invent their own method. Once people start sharing data sets, a lack of consistency leads to bad analysis. Further, you want to ensure that your vendors and your data warehouse play by the same rules as your analysts. You might think it is obvious that missing values should be coded as missing, but your DBAs are not analysts or statisticians and they have different priorities.

True story on handling of missing data: We had one data set that tracked values over of something over time, call it Revenue per Week. If Revenue per Week for Time2 was missing, the DBA pulled the values from Time1 and plugged it into Time2. In our case, we found data from t1 was being used in t71. As a result, the data set was unnaturally stable over a very long period of time. Because of our data quality checking efforts, we found the problem and now missing data is coded as missing.

The second part of transformation involves creating variables. As I have said before, I like a focused approach to data analysis. If you start transforming variables for analysis (taking the square and cubes, etc.) you add variables. Seems to run counter to taking a focused approach, I know. I am not suggesting creating every possible transformation, but you are going to need a few to help you capture non-linear effects or normalize your data.

In my experience, if you take take the square, the log (base 10), and the inverse of your variables you are going to get most of the value out of your transformations. You are going to need to use some common sense (what is the square of a categorical variable?) but in general, you are not going to need to go crazy creating new variables. However, your data analysts are going to want to create every possible transformation that they can think of; it is easy and they might need them later. Serious emphasis on might. In my opinion, the time would be better spent thinking about what transformations are actually useful. Also, each of those transformations creates another column of data that has to be processed, slowing down your analyses. My recommendation is to use the big three and if you can logically think through other variables that may need to be transformed, tackle those on an one-off basis.

Another type of transformation is creating interaction terms. I had written a long and involved descriptions of how interactions work and why you should care, but I deleted it; I think interactions deserve deserves its own post. Interactions capture additional impact of two variables combines that is not captured by considering each variable separately. The short version is that you can (and should) create interaction variables by multiplying variables together. The challenge: it is difficult to know which interactions to create. You could guess by now that I am not in favor of creating every possible interaction terms, you would create an enormous data set that would not be useful. I am a fan of creating interactions that, due to your understanding of your business, you think exist and making them a regular part of your transformation process.

Last piece, Load. I have nothing to say here. In most analytic platforms, once you have extracted the data, it is loaded. Really, the process for data analyst should be called ELT.

Take aways? You want to be thoughtful about your transformation process. People often create tons of variables as a proxy for thinking deeply about the problem they are trying to solve. Creating big, dumb, datasets have real costs associated with them. Not least of which is that they are hard to analyze. Also, create a standard method for data cleaning and make sure everyone knows the standard.

OOO

Sorry for the radio silence. I have been traveling for the last 10 days and have little access to an internet connection. Today, I am speaking at the "Optimization Summit" in San Francisco, traveling back to DC tomorrow, and back in the harness on Friday. I have a new "Value Chain" post in the hopper and I'll post it in a couple of days.

Monday, September 24, 2007

The Analytic Value Chain - Locate relevant data

Related posts
The Analytic Value Chain introduced
Defining the problem
Determine Data Requirements

This post overlaps with the previous Value Chain post on determining the data requirements, but is meant to be a bit more tactical. If you have followed my advice from the previous post, you have a good sense of what data you are going to need. Now you need to find it. In most organizations I have worked with, the data to solve any given business problem exists. The challenge is that the data often exists in a place that is not accessible to most of the folks in the organization. The data may not be in a production environment (it is sitting on a server under someone's desk) and if it is in production, the data warehouse might be so large that no one really knows what is in there (this is not an uncommon problem in real warehouses. I had an expert in physical warehousing once tell me that a really good warehouse knows where a specific pallet can be found 80% of the time). I once had both problems at the same time. I found two data sets that answered a critical business problem when combined, but were running on different desktops in two different parts of my client's organization. I found the data by luck, but wound up doing a very impactful analysis. My value was in carrying the data sets, on floppy disks, back to my PC. Obviously, you can't analyze data that you can't find.

So, what to do? I don't have a ton of advice here, but I have found somethings work pretty well in identifying the data that is out there and making it accessible. First, treasure your staff that really know your data infrastructure because they have hard fought knowledge (for those in my organization, you know who you are. And you know how much I value your contributions. You also know that I am understating.) There is no replacement for just having experience in your data infrastructure. Having said that we rely on people power, metadata helps. And even the best staff are not going to be able to find useful data if you don't have data dictionaries for every dataset.

Second, create tables that aggregate your most useful data. We do this and can get our hands on useful data sets pretty quickly; in minutes. We evaluate the variables in those tables about once a year (or after any major strategy shift) to ensure that the data set maintains its usefulness. This data set has an additional value in that it can be shared with your entire organization and forms a common "fact base" for the organization.

Third, try to think ahead and ask your team to be on the lookout for certain types of data. There are business questions that I know I am going to want to take a look at and by communicating to the team my topics of interest, they can make serendipitous discoveries.

Currently using Google Presentation and Speadsheet

What pieces of crap. The flaws are too numerous to list. I have confidence that they'll get better over time, but for now, unusable. Back to Office.

It is too bad, this is the first time I have not liked a Google product.

Wednesday, September 19, 2007

The Analytic Value Chain - Determine Data Requirements

Related posts
The Analytic Value Chain introduced
Defining the problem

So, once you have a good understanding, you should probably grab every piece of data you can find, transform the variables every way you can think of and start analyzing to figure out the variables that impact your dependant variable, right?

This way lies madness and spurious results. And is an approach that a naive data analyst would take (experts in data-mining might have a different take, but I was trained as a research scientist and I don't like the kitchen sink approach.) Before you start grabbing data, you should spend some time thinking about the behavior you are trying to predict and what phenomena might drive that behavior. Let me go a step further, I recommend that you generate hypotheses on what relationships you expect to find. Write 'em down. I go so far as to create something I call a conceptual map that graphically shows how I think the all of the variables, including the dependant, impact each other. Once I see the relationships, I can quickly generate hypotheses. Maybe you are thinking, what does this have to do with determining your data requirements? Hang on, I am getting to it.

Once you have the hypotheses, then you can start thinking about what data you need to test those hypotheses. In my world, it is important to have a good idea of what kinds of analyses we are going to be doing because the most time we have access to many more variables than we can possibly look at. We can't go on fishing expedition. Some other reasons to think hard about your data requirements:

  • In rapidly changing businesses, you'll spend more time finding data than analyzing the data. So, parsimony is key, you don't want to spend time getting data that is not going to be useful.
  • Most large enterprises have real restrictions on who can access specific types of data. So, even if you find the data, you'll have to figure out who owns the data.
  • Merging and analyzing larger datasets takes processor time. The smaller the datasets, the less time the analysts spend waiting for results.
  • The world is complicated. The more variables you have, the greater the chance of finding a spurious correlation. Also, if you have a good set of hypotheses and a conceptual map, you'll have a better sense if a relationship makes sense.

In sum, planning for your data needs, though it takes time up front, saves you time in the end.

Friday, September 14, 2007

SAS or SPSS

Someone asked me the other day if I prefered SAS or SPSS. As with so many other things, it depends. Here is a very good discussion on the differences between the two packages. From my perspective, you only go with SAS if you need their functionality, typically required if you are conducting very advanced statistics or handling very large data sets. Also, their support is very good (which you need because their product is complicated) and a very complete product suite (think end-to-end.) SPSS is much more user friendly and produces better looking output that can be easily imported into Office applications. But their product suite is a little more limited. They also have solutions, but they are more of a research focused company. So, for my everyday use, I like SPSS, but like to have access to SAS for when I need the big guns. If you would like to talk about the differences, feel free to drop me a line at ken.rona@gmail.com. I am happy to offer guidance.

Thursday, September 13, 2007

Principles of Innovation

Over the past couple of years, I have been thinking about innovation. Specifically, what (and should ) I do with my teams to encourage innovations. Someone brought to my attention a video from Google's Vice President, Search Products & User Experience, Marisa Mayer.



Marisa talks about nine guiding principles that she believes are key enablers of innovation in her business. I would not argue that we all should adopt her principles, just that it was an interesting discussion and that having such a set of guiding principles is probably something worth thinking about. Mine are (somewhat redundantly with Marisa's):

1. Transparency is your friend.
2. Don't be afraid to iterate and experiment. Always test. Without it, you will not learn.
3. Noble failure is ok. The corollary is that stupid failure is not ok.
5. Share the credit. There is plenty to go around and you will likely have more than 1 good idea in your life.
6. Try to work around people you would be honored to hang with - and recruit someone when you find them! I am fortunate in this regard.
7. Don't have an opinion. Have facts.
8. You can build a business on users, but you need to run it on money. So, while you are innovative, think about how you are going to make money with the idea.
9. Be brave (but not disrespectful) in the presence of senior executives. That is how most of them got to be in their position.
10. Encourage dissent in your organization.
11. Be dissatisfied, but don't be a jerk. It makes you want to make things better.