Author Archives: Steve Todd

About Steve Todd

Steve Todd is a retired Dell Technologies Fellow and former EMC Distinguished Engineer who spent nearly four decades building high-tech products for the information storage industry. A prolific inventor named on over 400 U.S. patent applications, his innovations have generated billions of dollars in revenue. He served as Vice President of Data Innovation and Strategy in the Office of the CTO and is the author of two books on corporate innovation, Innovate with Influence and Innovate with Global Influence. He holds B.S. and M.S. degrees in Computer Science from the University of New Hampshire and writes about technology at his blog, Information Playground.

Shanghai Software

GUEST POST from Steve Todd

How can I enable innovation from thousands of miles away?

After twenty-five years of delivering IT products for EMC, I’ve spent the last fourteen months as the Director of EMC’s Global Innovation Network. The role has been a whirlwind of new responsibilities; my background as a software engineer was certainly helpful, but my day-to-day work contained little to no overlap with developing software.

During the month of April 2012 my role was expanded. I received an interesting new assignment: guide the efforts of several dozen software engineers. This new duty is unique in many ways, not the least of which is that the entire team is located in China.

Perhaps the most significant aspect of the new role, however, is that the software engineers are all part of EMC Labs China. They are not tied to any business unit. They are members of EMC’s Office of the CTO, and therefore free to explore and discover new directions for the company.

I spent the first few months learning their efforts, skillsets, and personalities. I encountered a fairly staggering array of work that represents five years worth of maturity (EMC’s research presence in China dates back to 2007). Externally, the team is active in industry standards, university research, industry initiatives, and they attend tier-1 conferences. They’ve published papers and won awards (the latest one being the May 2012 “Best Demo” award for the PEACOD project at DASFAA).

With very little research background myself, I’m looking forward to having them teach me a thing or two about how it’s done.

Their internal work, however, allows me to climb back into the software saddle. They are currently driving a large number of software projects that will come to represent new EMC product offerings. As they transfer knowledge and prototypes into EMC’s business units, I have the unique opportunity to create a pipeline of new work that represents next steps for our company and our industry.

I have never managed a team of researchers (in fact I have never actually managed any software engineers in my entire career). I also have zero experience managing employees in a different geography. Within a few weeks I was struggling to formulate a management approach. Fortunately, the EMC Innovation Network has a straightforward vision: spark the creation and delivery of high-value ideas. The tasks at hand are actually very easy to articulate: communicate CTO vision to the team, work together to create a compelling pipeline of new work that flows to EMC’s business units, and provide career growth to the team.

Easy to say, hard to do.

I decided to work with Human Resources to kick start my relationship via the NMAP exercise: New Manager Assimilation Program. Every researcher in China provided anonymous feedback describing what they think about me as well as what I needed to know about them. My next step was to meet them face-to-face. I set off to China with four pages worth of feedback.

I touched down in Shanghai during the second week in July, where I would spend half my trip. The remainder of the journey would take me to Beijing (the team is distributed between the two cities).

I’m typing this post on my way back to Boston. I have many stories to tell. In fact, I have a series of stories that will likely take several weeks to blog my way through.

I made a pretty interesting discovery during my trip.

It’s not about the software. It’s about the relationships.

More to follow,

Steve

image credit: innovationplayground.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 6 Innovation Analytics: Operationalize

GUEST POST from Steve Todd

This is the last in a series of posts describing a methodology (EMC’s Data Analytics LifeCycle) for using analytics to measure innovation at a multi-national corporation. This lifecycle is taught at the Data Science and Big Data Analytics course created by EMC, and I’ve blogged my way through each phase of the lifecycle and have arrived at the end (Phase 6).

As a review, here is a graphical view of the lifecycle, followed by a summary of all the posts written thus far:

Phase 5

Global Knowledge Flight Patterns

Phase 4

Voice Measurement

Boundary Spanner Validation

Finding Boundary Spanners

Phase 3

Longitudinal Studies

Hypothesis Exploration

Model Planning

Phase 2

Data Quality

Exploring the Data

ELT

Data Preparation

Phase 1

Creating The Analytics Plan

Hypothesis Generation

Introduction to Phase 1

Introduction to Innovation Analytics

Given the foundation of the first five phases, let’s finish with the final phase.

Phase 6 is called “Operationalize”. My team and I have not yet reached this phase. My understanding of Phase 6, however, is influencing our journey through the steps. The journey that my team has undergone so far can be summarized as follows:

Running analytics against a sandbox filled with notes, minutes, and presentations from innovation activities has yielded great insights into EMC’s innovation culture.

Phase 6 moves the analytic models out of the sandbox and into production. The course advises a production “pilot” be run first (as opposed to deploying the model on a wide-scale). This approach minimizes risk. Smaller-scale deployment allows the team to learn about the performance and make adjustments before a full deployment.

Phase 6 may require a new team of people to join the initiative (the people that are responsible for running the production environment). These people will help feed data sets into the production model. During the execution of the model in the production environment, it is important to detect anomalies on inputs before they are fed into the model. This may not be 100% possible. Consider doing a logistic regression on a training set of the data if possible.

What does this specifically mean for the project I’ve been running? The points mentioned below are key aspects to remember for any company wishing to run innovation analytics:

  • We need more data, which means we need a marketing initiative to convince people to submit (or inform) the global community on their innovation/research activities.
  • This data is sensitive and some thought needs to go into “who” can run the model and “who” see the results
  • In addition to running models, a parallel initiative will likely be to access the repository for search (people want to search for innovation/research initiatives). This may impact the performance of the analytics.
  • We need a mechanism to continually re-evaluate the model after deployment. Assessing the benefits is one of the main goals of this stage, as well as defining a process to retrain the model as needed.

This last point represents a challenging and often overlooked aspect of Phase 6. The team needs to assess whether the model is meeting goals and expectations, and if desired changes are actually occurring. The data may change over time, or live data may morph to the point where the model needs to be updated or retrained.

As I reach the end of this series of blog posts, I’d like to thank Dave Dietrich, who has proofread nearly all of my posts for accuracy!

As the efforts of the data scientists come to a close, Dave has one final piece of advice:

Hold a post-mortem with the analytic team to discuss what would change in the process or project if you had the chance to do it over again.

image credit:melansonconsult.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 2 Innovation Analytics: Data Quality

GUEST POST from Steve Todd

I am blogging my way through the Data Analytics Lifecycle as taught in EMC’s Data Science and Big Data Analytics course.  I am running a Data Analytics project that employs a team of volunteer data scientists from around the world, and I have communicated an analytic plan to them (along with a set of hypotheses). The entry into Phase 2 of the process (typically the longest phase) has resulted in preparing the data, loading it (without transform), and exploring it.

At this point it is worth mentioning a quote that I heard during the course:

If you do not have data of sufficient quality or cannot get good data, you will not be able to perform the subsequent steps in the lifecycle process.

How Clean is the Data?

One of the data scientists on my team was evaluating a tool called Tableau, which can be used for data exploration (among other things). They began to use the tool and explore the data that had been previously loaded into the analytics sandbox. They sent me the following screenshot (I zoomed in and circled my name):

I am showing up twice in the database because some entries have a space before my first name. This is a classic problem (and not always easy to fix).  Addressing this problem within the sandbox is clearly a much easier proposition than doing so in the production database. However, it could take a long time to get it right (another reason why phase 2 takes so long).

Who typically does this work?  Is it a database admin (DBA)? A data engineer? Both typically play a role in Phase 2. The “Database Administrator” provisions and configures the database environment to support the analytical needs of the working team. The “Data Engineer” tends to have deep technical skills to assist with tuning SQL queries for data management and extraction. They also support data ingest to the analytic sandbox.  These people can be one in the same, but many times the data engineer is an expert onqueries and data manipulation (and not necessarily analytics as such).  The DBA may be good at this too, but many times they may simply be someone who is primarily skilled at setting up and deploying a large database schema, or product, or stack.

In addition to mis-spelled names, the data scientists exploring the data are starting to uncover missing data that will help them prove the hypotheses. For example, consider one of the hypotheses generated in phase 1:

H5: Knowledge transfer activity can identify research-specific boundary spanners in disparate regions.

The association of boundary spanners to the geographic location where they work requires that any names found in the sandbox (e.g. “Todd,Steve”) have an associated location (e.g. Hopkinton, MA).

Our data scientists found that this data was nowhere to be found within the sandbox.

In addition to DBAs and Data Engineers, IT often plays a large role in Phase 2.  For our project, once the names were “cleansed”, we had to bring in IT resources to help generate geographic associations via our employee database. In our particular case we were fortunate: not only did IT grant access to our request, but the IT resource had Data Engineering skills and cleansed the data for us! In general, bringing in additional data from the IT realm is no easy task. Access to these types of assets is typically a very tough, time-consuming part of Phase 2.

I could write paragraph upon paragraph describing issues that we’ve come across (and solved) for Phase 2. It may be more useful, however, to summarize some of the lecture material that describes common problems:

  • Consistency of data types (e.g. confirm that all numeric types contain numeric fields)
  • Data feeds can often change over time (e.g. someone removes a column without telling anyone)
  • Fields that contain calculations (e.g. interest charges) may change over time (if interest rates change over time)
  • What are the legal ranges of data and are there any values that are out of bounds?
  • Is the data standardized/normalized? If so, what is the scale?
  • Are geospatial data sets consistent (e.g. metric versus english units, two-letter state abbreviations versus full-names)?

During this phase the data scientist may discern what to keep and what to discard. They had probably formed an opinion of what model they will use. Data exploration and cleansing has either validated their assumptions or caused them to select a different model. Data cleansing is a big job, so the objective should be to determine “what is enough?”. What is clean enough data? What is sufficient quality for the operating context? What will properly enable the analysis?  These questions give people boundaries for the data cleaning, which is quite intensive.

Phase 3 is about model planning. How does one know when they are ready to leave phase 2 and move on to phase 3 (keep in mind that a return to Phase 2 is highly likely!)?

In general, Phase 3 begins when the data quality is “good enough” to start building the model. In my case, once we had cleaned up erroneous names and associated the names with geographies, the team had sufficient reason to enter Phase 3.

I will relate our team experience with Phase 3 in future posts.


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 2 Innovation Analytics – Exploring the Data

GUEST POST from Steve Todd

At this point in the series of articles on the Data Analytic Lifecycle, raw data has been identified and imported into a Data Analytics sandbox. The data (a mix of structured and unstructured data) is depicted below. Contained within the sandbox now lies a large amount of data representing research and innovation activities occurring globally throughout my corporation (EMC).

At this point in the lifecycle it is recommended that the data scientists and engineers begin to “get used to the data”. This means that they can use any number of tools to inspect the format, structure, and quality of the data. They have already immersed themselves into the analytic plan generated in Phase 1, and they are likely looking for data sources to validate the hypotheses in the analytic plan.

This process will almost certainly identify “gaps” in the data that will prevent the data scientists from being able to prove or disprove the hypotheses.

At my corporation we have created a volunteer team of data scientists that ensure that our approach scales globally.  One of these data scientists is Vladimir Suvorov of Saint Petersburg, Russia. Vladimir described the tools and approach he used to explore the data:

R-studio provides the easiest way for initially examining the data placed within a sandbox. R-studio provides a simple connection string and SQL query on one side and a powerful statistics and graphing package on the other. Besides that, R is open-source software, so it can present additional value for the community.  I recommend R for fast prototyping and quick summaries of the data.

Vladimir created a chart that focuses solely on university activities. This chart uses color coding techniques to categorize the types of university activities that have occurred in previous months. What is actually happening here is broadly considered “data exploration”.  In this phase, people like Vladimir explore the data, assess data quality, examine basic information about the data itself (relationships in the data, trends, etc.), and they try to understand what kind of data they actually have.  This lets them form ideas of how they can test in future phases, and what kinds of insights they may be able to drive towards.

Vladimir’s simple visual chart shows the pace of visits to universities (yellow), as well the emergence of a dark blue color, which represents meetings of the Research Advisory Team (RAT) for the purpose of discussing university research funding for 2012. These meetings continue throughout the first quarter of the current year (2012-03).

The relatively small amount of employee lectures at universities (red) and professor visits to EMC (aqua) is a potential indicator that programs could be put in place to accelerate these types of exchanges.

This chart falls under the category of “descriptive statistics”.

It’s an important activity in Phase 2 which, as just explained, can already lead to actionable conclusions.

During the Data Science and Big Data course, I learned that bar charts like this, while helpful for the data exploration phase, may not be the best choice in future phases. They can actually be improved via a few simple tips. Stacked bar charts are good for showing a few aggregated data points and exploring the data points.  When looking at data over time (as above), it is generally better to display line charts, to make it easy to see trends (people generally have a more difficult time examining trends in stacked bar charts).  Likewise, best practice is to use very soft colors and strong, emphasis colors as a way of highlighting key points; red is also a signal for things like “danger” or problem areas, especially in certain world countries.  In summary, a graphic like this gives us good initial information at this stage in the process. For more polished visualizations (used later in the lifecycle), Data Scientists need to address these kinds of aesthetic considerations.

During the data exploration exercise the data scientists and engineers begin to notice that (a) certain data needs conditioning and/or normalization, and (b) there are data sets that cannot be found anywhere which are critical to proving analytic hypotheses.

These activities will be further profiled in the next post. Once again, thanks to David Dietrich, who not only taught the course I attended, but continues to oversee this series of posts.

image credit: silverdane.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 2 Innovation Analytics: ELT

GUEST POST from Steve Todd

This series of articles describes an analytic lifecycle being used to gain insight into the innovation and research practices of a multi-national corporation (EMC). After creating an analytic plan in Phase 1, a previous post described the Data Preparation phase of the lifecycle. This phase involves the creation of an enormous sandbox (e.g. ten times the size of a data warehouse configuration).  Data scientists and engineers are encouraged to extract data from many sources and load it into the sandbox unchanged. This approach may seem a bit revolutionary (most processes transform the data first). This lifecycle, however, is geared towards the data scientist. The possession of the raw data allows for more robust analysis. The diagram below provides an overview of this approach.

As I mentioned in the last post, there are two types of data that will allow a data scientist to analyze innovation. The first type, depicted on the left, is a structured SQL database containing thousands of innovation ideas submitted by employees over a five year period. The second type of data consists of minutes and notes from global innovation and research activities. This content is highly unstructured.

It’s worth taking a moment to discuss how the global team of users and data scientists came up with relevant innovation activities (depicted below). EMC’s product line consists of high-tech products and services that have been introduced into the marketplace. Tracing the lineage of these products and services usually results from an idea that happened long ago during a specific activity. The team came up with a candidate list of activities that is often associated with innovation and research:

Visiting universities, creating publications, attending conferences, visiting customers and partners, holding internal knowledge sessions, holding idea contests, and creating intellectual property are all activities commonly associated with innovation, and therefore an effort was made to gather six months worth of these activities from data sources worldwide.

After resisting the urge to transform these documents before loading them, a next logical step in Phase 2 is to explore the data. I will step through an example of this process in my next post.

image credit: users.rowan.edu.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 2 Innovation Analytics – Data Preparation

GUEST POST from Steve Todd

There is a massive amount of innovation and research data globally distributed amongst the 50,000+ employees at my corporation (EMC). In this current series of blog posts I’ve been theorizing that analytics can allow my Innovation Team to unlock key insights from the data and accelerate innovation world-wide. The problem is large but I had the good fortune to attend the first offering of my company’s Data Science and Big Data Analytics course. The course described a logical set of steps for testing business theories as part of a Data Analytics Lifecycle. I’ve been following these steps and after a number of posts describing the critical first step (the generation of an analytic plan), it is a good time to move on to Step 2: Data Prep.

As the arrows on the diagram indicate, these steps are iterative in the process. Proceeding to Phase 2 is often a matter of whether or not you are comfortable sharing the analytic plan with your peers. If so, then the data preparation phase can begin.

The analytic plan assists the data scientist in identifying the business problem, a set of hypotheses, the data set, and a preliminary plan for the creation of algorithms that can prove or disprove the hypotheses.  Once the analytic plan has been delivered and socialized, the next step is all about the data. In particular, the next step is all about conditioning the data.

The data must be in the right shape, structure, and quality to enable the subsequent analysis.

Building an Analytic Sandbox

In my last post I mentioned that the data set in question falls into two categories: (a) a production “idea submission” server (essentially a large-scale database containing structured data), and (b) a globally-distributed set of unstructured documents representing knowledge expansion within the corporation in the form of minutes and notes about innovation/research activities.

These data sets cannot be analyzed in their current production formats. In addition, it is possible that the data is not of sufficient quality. Furthermore, the data is likely inconsistent. All of these possibilities add up to the fact that a separate analytic sandbox must be created to run experiments. Industry practice states that on average the size of this sandbox should be roughly ten times the size of the data in question (e.g. the current size of your enterprise data warehouse). Keep these things in mind when creating the sandbox:

  • You are going to need strong bandwidth and network connections to your sandbox.
  • Collect as much data as you can, including summary data, structured/unstructured, raw data feeds, call logs, web logs, etc. This is why the sandbox needs to be large.
  • Determine the type of transformations you will need to assess data quality and derive statistically useful measures.
  • Transform the data after it is in the sandbox (ELT: Extract, Load, Transform, as opposed to ETL). This allows analysts to choose to (a) transform the data or (b) use the data in its raw form. It’s worth pointing out that this method is the opposite of best practice for some Data Warehousing use cases. While ETL is a widely accepted practice, the sandbox approach prefers ELT.
  • Acquire the right set of tools for the transformation. Good examples would be Hadoop for analysis, Alpine Miner for creating analytic workflows, and R for many general purpose transformations.

Sandbox creation typically requires assistance from IT, a DBA, or the person that controls the enterprise data warehouse.

Once the sandbox is created, there are three key activities that allow a data scientist to conclude whether or not the data is “good enough”.

  1. Familiarize yourself with the data thoroughly. List out all the data sources and determine whether key data is available or more information is needed. This can be done by referring back to the analytic plan to determine if you have what’s needed, or if more data must be loaded into the sandbox.
  2. Perform data conditioning. Clean and normalize the data. During this process discern what to keep versus what to discard.
  3. Survey & Visualize the data. Create overviews, zoom and filter, get details, and begin to create descriptive statistics and evaluate data quality.

I learned in the course that this part of the process is expected to take at least 50% of the time spent on the entire data analytics lifecycle (and 80% is not uncommon)!  Indeed, as our team went through this process for innovation analytics we had to work through quite a few issues before being able to work on implementing a model.

I will describe these issues in the next post.

Thanks again to David Dietrich for his research on the data lifecycle and ongoing support throughout this series of blog posts.

image credit: stevetodd.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 1 Innovation Analytics – Creating the Plan

GUEST POST from Steve Todd

The finishing touch for Phase 1 of the Data Analytics lifecycle is the creation of an Analytic Plan. In the same way that requirements drive all phases of a software project, the analytic plan lays the foundation for all of the work in an analytics project.

I’ve mentioned in previous posts that this part is not easy. Analytic Plans are new to me. Before starting I need to give credit where credit is due. David Deitrich has been a driving force behind our Data Science and Big Data Analytics curriculum, and a regular contributor to this series of articles.

There are four initial components of an Analytic Plan:

1. Framing of the Business Problem

In my case I am trying to accelerate innovation within my corporation (EMC). Three problems faced by the corporation are (a) the tracking of knowledge growth throughout our global employee base, (b) ensuring that this knowledge is effectively transferred within the corporation, and (c) that this knowledge is most effectively converted into corporate assets. Executing on these three elements more effectively should accelerate innovation, which is the lifeblood of our company.

2. Initial Hypothesis

In my last post I described eight different initial hypotheses theorizing how analytics can assist in solving the business problem. These eight hypotheses were boiled down to one high-level hypothesis statement:

An increase in geographic knowledge transfer improves the speed of idea delivery.

This hypothesis paves the way for what data we will need and what type of analytic methods we will likely use.

3. Data

The data that the project will rely on fall into two categories.

  1. The first category represents five years’ worth of idea submissions into EMC’s Innovation Showcase process. The Showcase process is a formal, organic innovation process whereby employee ideas from around the globe are submitted, vetted, judged, and incubated. The data is a mix of structured (idea counts, submission dates, inventor names) and unstructured (the ideas themselves) content.
  2. The second category encompasses minutes and notes representing innovation and research activity from around the world. This data is also a mix of structured and unstructured. The structured data, once again, includes items such as dates, names, and geographic location. The unstructured documents contain the “who, what, when, and where” information that represents rich data about knowledge growth and transfer within the company. This type of information, however, is often stored in business silos that have little to no visibility across disparate research teams.

The first repository (the idea submissions) is centralized. The second data set (centralized research and innovation minutes/notes) will be gathered from throughout the corporation and contain 6 months worth of global data.

4. Model Planning – Analytic Technique

Model Planning represents the conversion of the business problem into a data definition and a potential analytic approach. In other words the rubber is beginning to hit the road in terms of creating algorithms. A model contains the initial ideas on how to frame the business problem as an analytic challenge that can be solved quantitatively. There is a strong link between the hypotheses and the analytic techniques that will eventually be chosen. Described below are a few algorithms and approaches that make sense given the hypotheses. They do not represent a complete list but give the reader a sense for this activity within the analytic plan.

Keep in mind that model selection is an “art form”. Some people are better at it than others. It requires iteration and overlap with phase 2 (Data Prep). Multiple types of models are applicable to the same business problem.  Selection of methods can vary depending on the experience of the Data Scientist’s comfort zone. In other cases model selection is more strongly dictated by the problem set.

  • Use Map/Reduce for extracting knowledge from unstructured documents. At the highest level, Map/Reduce imposes a structure on unstructured information by transforming the content into a series of key/value pairs. Map/Reduce can also be used to establish relationships between innovators/researchers discussing the knowledge.
  • Natural language processing (NLP) can extract “features” from documents, such as strategic research themes, and can store them into vectors.
  • After vectorization, several other techniques would be appropriate:
    • Clustering (e.g. k-means clustering) can find “clouds” within the data (e.g. create ‘k’ types of themes from a set of documents).
    • Classification can be used to place documents into different categories (e.g. university visits, idea submission, internal design meeting).
    • Regression analysis can focus on the relationship between an outcome and its input variables. What happens when an independent variable changes?  It can help in predicting outcomes. This could suggest where to apply resources for a given set of ideas.
    • Graph theory (e.g. Social Network Analysis) will be an important way to establish relationships between employees who are submitting ideas and/or collaborating on research.

At this point I have generated some hypotheses, described potential data sets, and chosen some potential models for proving or dis-proving the hypotheses.  During this process I have been sharing my thoughts in bits and pieces with my peers, and I feel confident that I have enough data to draft a high-level analytic plan and submit it for formal review. I’ve attached a template slide below.

The last two rows in the Analytic Plan overview (Results & Key Findings, Business Impact) are a reminder to me that I am working toward Step 5 of the Analytic Lifecycle: Communicate the Results. As the business user I participate most heavily in the beginning and the end of the Lifecycle.

I’ve spent a lot of time on this first step. Any analytic project lead should do the same. With the Analytic Plan as the foundation, it’s time to move on to Step 2: Data Prep.

image credit: stevetodd.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 1 Innovation Analytics – Hypothesis Generation

GUEST POST from Steve Todd

In this series I have introduced the concept of a Data Analytics lifecycle and began to explain how it guides the analysis of innovation at my corporation (EMC).

As with any lifecycle, the first phase (Discovery) lays the foundation for the rest of the steps. Two of the key activities in the Discovery phase are (a) the creation of hypotheses, and (b) the creation of an analytic plan. In this post I will introduce relevant hypotheses; my next post will dive into with the analytic plan.

I found it useful to ask myself the following question:

Given a global repository representing employee ideas, discussions, minutes, and notes about innovation and research activities, what can I measure?  What measurements can accelerate corporate innovation world-wide?

I found the exercise below to be the most difficult part of the curriculum. It is clearly the most relevant, however. Business theories about data must turn into statements that can be proved or disproved.

In this phase of the analytics lifecycle, the curriculum encourages business users to generate as many relevant hypotheses as possible. All of these can help guide data scientists in further phases. Each statement below starts with a “stream of consciousness” on the business problem, and ends with a specific hypothesis that data scientists can either prove or disprove.

Hypothesis #1: Local Measurement of Innovation Activity

I believe that innovation can be measured for a given geography. This measurement can take a number of forms, including number of participants, percentage of time dedicated to innovation, local idea-to-implementation timeframes, and geographic reach of innovators (how far outside the workplace does innovation activity extend)?  As the Director of a Global Innovation Network, I’d like to know how these activities map to corporate strategy.

IH1: Innovation activity in different geographic regions can be mapped to corporate strategic directions.

Hypothesis #2: Geographic Innovators

In every locale world-wide there are typically a set of people who pursue innovation with passion and consistency.  Their contributions may be hidden from the corporate eye. They also may be focused on particular activities, such as idea contests, visits to customers/partners, frequent visitors to a university, or the generation of intellectual property.  If they have not yet taken the initiative to extend their visibility to the global stage, I believe that they can be found via analytics and connected to relevant knowledge sources.  I believe their skill in the delivery of ideas would be improved. For this topic there are two hypotheses to prove:

IH2a: The length of time it takes to deliver ideas decreases when global knowledge transfer occurs as part of the idea delivery process.

IH2b: Innovators that participate in global knowledge transfer deliver ideas more quickly than those that do not.

Hypothesis #3: Effective Ideators

When it comes to idea generation, some employees have an advanced ability to suggest ideas that resonate with the decision makers. Rarely are their ideas dismissed outright. They often (but not always) have a track record of idea delivery as well.  I believe that analytics can help me find these people. I also believe that the form of their ideas can be analyzed to reveal clues as to why their ideas are likely to be funded. In turn, the format of any idea submission can be evaluated for its value.

IH3: An idea submission can be analyzed and evaluated for the likelihood of receiving funding.

Hypothesis #4: Geographic Knowledge

Very often certain geographies have a reputation for excellence in a certain area of knowledge.  I believe that analytics will reveal that this knowledge can be found in other locales as well (or conversely it may not be found where it was assumed to be).  In general, different geographies will likely reveal themselves to be hubs of expertise in any number of areas. Knowing this fact would not only facilitate the matching of problems to local innovators in a certain region, it also may provide the opportunity to join different locales together for problem-solving exercises.

IH4: Knowledge discovery and growth for a particular topic can be measured and compared across geographic regions.

Hypothesis #5: Knowledge transfer facilitation via boundary spanners

There are certain employees that have arisen within a geography and made connections with other geographies for the purpose of collaboration. They may not have high visibility within a corporation aside from the direct connections that they have made on their own.  I believe that not only can analytics identify these people, but it can also classify the type of knowledge that these individuals are transferring. These “boundary spanners” can be targeted and trained as “innovation facilitators” and united at a corporate level.

IH5: Knowledge transfer activity can identify research-specific boundary spanners in disparate regions.

Hypothesis #6 Corporate research gaps and assignments

A corporate research roadmap needs a portfolio of initiatives to go along with it.  I believe analytics can enable a dashboard view of particular strategic initiatives (e.g. cloud computing) and determine how much research activity (if any) is occurring across the corporation. This view can also be extended to profile funding activity on particular themes. I also believe that analytics can recommend the best place to perform research as new items are added to corporate research roadmaps.

IH6: Strategic corporate themes can be mapped to geographic regions.

Hypothesis #7 Incubation Lineage and Asset Generation

I believe that the path that knowledge takes, from a local innovator, to a corporate boundary spanner, to an implementation team, to a delivered asset, can be traced and measured. I also believe that this measurement, once studied, can reveal ways to accelerate innovation and point out areas of knowledge that are yet to be converted. I’ve long been a fan of provenance, and I love the concept of “idea lineage”. The lineage can be studied to reduce asset delivery time.

IH7a: Frequent knowledge expansion and transfer events reduce the amount of time it takes to generate a corporate asset from an idea.

IH7b: Lineage maps can reveal when knowledge expansion and transfer did not (or has not) result(ed) in a corporate asset.

Hypothesis #8 New areas of innovation and research

Finally, I believe that predictive analytics can reveal areas of focus for future innovation, research, and investment. What knowledge should be expanded? Who should collaborate on that theme? What kind of assets could result?

IH8: Emerging research topics can be classified and mapped to specific ideators, innovators, boundary spanners and assets.

If I were to sum up the list above into one hypothesis, it would look something like this:

An increase in geographic knowledge transfer improves the speed of idea delivery.

In my next post I will describe how this hypothesis can be integrated into an analytic plan.

image credit: stevetodd.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

Phase 1 Innovation Analytics

GUEST POST from Steve Todd

How do you analyze innovation at a large multi-national corporation? I’ve been using the Data Analytics Life Cycle (depicted below):

These steps have helped me to internally construct a strategy for analyzing global innovation processes and methodologies at my own corporation (EMC). I’ve published some of the results of this effort in previous posts. The analytic lifecycle was something I learned while attending the new Data Science and Big Data Analytics course.  Who came up with these steps? They are essentially an overview of industry best practices and experiences as summarized by EMC Education Services (for a deeper dive into the course content, register and take a look here). The business driver that caused me to attend the course was my belief that analytics could help me discern promising new opportunities in my role as Director of the EMC Innovation Network. The first thing I learned at the course is that I am not a data scientist. Any successful analytics project has a set of key players. In addition to data scientists, key roles include project sponsors and managers, DBAs, data engineers, and business intelligence analysts. My role is essentially that of a business user: someone who consults and advises on how to operationalize the end result of the analytic exercise. The Discovery Phase Step number 1 in the analytics lifecycle is all about the business domain. Here are some of the key activities that are critical in this phase:

  • Frame the business problem as an analytic challenge that can be solved in phases.
  • Understand what’s been done in the past.
  • Assess the resources supporting the project (people, technology, time, and data).
  • Form initial hypotheses.
  • Determine readiness to move to the next phase.

Framing the Business Problem My company has 50,000+ globally distributed employees, many of whom innovate on a daily basis. The main business problem, from an innovation standpoint, is to ensure that we have an innovation pipeline that continually introduces new revenue sources and cost improvements. Innovation is the lifeblood of the company, especially in the field of high-tech. The most important word in EMC’s innovation lexicon is knowledge. The mission of our EMC Innovation Network is to (a) expand knowledge locally, (b) transfer it globally, and (c) leverage it strategically.  The continual introduction of new revenue sources and cost improvements comes down to leveraging new knowledge that has been transferred and shared across our global employee base. Our company needs to analyze the expansion, transfer, and leverage of knowledge. The insight gained from this process will improve the innovation pipeline (one of the hypotheses that I will expand upon in future posts). Understanding What’s Been Done in the Past When it comes to a repository for innovation data, EMC has a five-year history of global ideas submitted by employees. The repository amounts to roughly 6000 ideas. In addition to the idea repository, each business unit has their own repository and collaboration site describing innovation and research activities specific to their business. As I mentioned in a previous post, the idea of running analytics across this type of data is a fairly novel approach. We’ve already analyzed year-over-year idea submission totals, with an emphasis of the geographic location of the submitter. Other than that, there has been no previous attempt to analyze knowledge expansion and transfer on a global scale. This realization surfaces a clear pain point: any future analysis would require some sort of centralization of global knowledge activities related to innovation. This meant that our project would likely be phased. The team would start with the idea repository and focus on measuring research and innovation activities in a later phase. Assessing the Resources The resources required to run this analytics initiative is a good news/bad news situation. The good news is that I work in the CTO Office and my team has an excellent lab with excellent lab managers. Compute/network/storage resources are not an issue. The bad news is that I need two teams of people to help with this project and none of them report into my organization:

  1. I need sponsoring organizations, managers, and high-level DBAs to help with the initial phases.
  2. I need data engineers and scientists to execute the low-level work.

I’ve solved this problem by forming two global, volunteer teams within my own company.  I met with MIT Professor Peter Gloor to ask his advice on how to motivate globally distributed teams outside of my functional organization. He gave me great advice that worked. Every other Tuesday morning I meet with Team #1: the steering committee for the project.  On alternating Tuesdays I meet with Team #2: data scientists-in-training! These teams are distributed throughout the U.S., China, India, Israel, Egypt, Russia, and Europe. Initial Hypotheses and The Analytic Plan This next sentence is essentially the jewel of the course: The hypotheses and analytic plan form the foundation of everything that comes after it. Before taking the course, my hypotheses fell into two buckets:

  1. Descriptive analytics of what is currently happening in my organization will spark further creativity, collaboration, and asset generation.
  2. Predictive analytics will advise executive management of where it should be investing next.

After taking the course, I realized that I need to spend much more time a more comprehensive set of hypotheses and a formalized analytic plan. I will be diving into each one of these in future posts. image credit: stevetodd.com


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.

A Strategy for Innovation Analytics

GUEST POST from Steve Todd

Using analytics to improve innovation processes is, well, innovative in and of itself.

I spoke with Tom Davenport last week about using analytics to understand innovation. While my own company (EMC) and other corporations use analytics across many different parts of the business,Tom and I agreed that it is unique to use data scientists to gain insight into innovation processes.

So how exactly is it done?

I have received requests for more detail on the approach that we are using, especially from managers that are unfamiliar with the field of analytics.

In response to these requests I’ve decided to write a series of posts that describe the evolution of my own experience applying analytics to innovation data. These posts will roughly fall into three categories:

1. A description of the data that my team collects internally.  Most of my posts have focused exclusively on idea submissions from employees, but the larger repository contains innovation and research data beyond just idea submissions.

2. The training that I underwent to manage my own personal team of Big Data Scientists.  This training contains a six-step process that I undertake to gain innovation insight into the data we have collected over the years.

3. My ongoing use of social media to accelerate the collection of innovation and research data world-wide, and the specifics of the approach that I take to motivate teams of “volunteer” data scientists at worldwide locations.

If you are interested in doing some pre-reading along these lines, I recommend taking a look at the posts that I have already written on the topic. I also recommend taking a deeper look at the training that I undertook to increase my own expertise in this area. The diagram below depicts the approach that I use for gaining analytic insight.

Please leave me a note in the comments section if there are specific topics you’d like to see explored along the way.

As Bill Schmarzo put it, be prepared to embark on a most excellent journey.

imagecredit: theosgroup


Subscribe to Human-Centered Change & Innovation WeeklySign up here to get Human-Centered Change & Innovation Weekly delivered to your inbox every week.