Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Tuesday, May 6, 2008

Quality Data helps us go GREEN!

Yesterday was another day of coming home from work late and having to force the door open to get past all the junk mail inside. After picking up, taking into the kitchen and spending 10 minutes going through I was yet again presented with another fine example of poor data quality (i.e. the majority of organizations really don’t have a grip on their customer data let alone the ability to household).

3 copies of a news letter from the same software company (no names mentioned!), the exact same letter from a State Insurance agency for both my wife and I, and then two copies of the Crate & Barrel latest summer catalog addressed to me (how on earth I became registered on their list I’ll never know!).

I wonder what the impact to the environment would be if organizations simply got a better understanding of their customer data and improved their marketing functions alone?

So once I finished my nightly chore of “shredding” I did some quick research to see what sort of impact to the environment today junk mail has. Check out the following facts listed by New America Dream:
  • More than 100 million trees’ worth of bulk mail arrive in American mail boxes each year – that’s the equivalent of deforesting the entire Rocky Mountain National Park every four months. (New American Dream calculation from Conservatree and U.S. Forest Service statistics)

  • In 2005, 5.8 million tons of catalogs and other direct mailings ended up in the U.S. municipal solid waste stream – enough to fill over 450,000 garbage trucks. Parked bumper to bumper these garbage trucks would extend from Atlanta to Albuquerque. Less than 36% of this ad mail was recycled. (U.S. Environmental Protection Agency)

  • The production and disposal of direct mail consumes more energy than 3 million cars. (New American Dream calculation from U.S. Department of Energy and the Paper Task Force statistics)

  • Citizens and local governments spend hundreds of millions of dollars per year to collect and dispose of all the bulk mail that doesn’t get recycled. (New American Dream estimate from EPA statistics)

  • California's state and local governments spend $500,000 each year collecting and disposing of AOL’s direct mail disks alone. (California State Assembly)

With companies trying to put on a more “Green” face you would think this would be a nice eco friendly place to start. Imagine the impact of cutting bulk/junk mail in half by just knowing who your customer is and the fact that you may have multiple that live at the same address?

Even though the challenges surrounding customer data are not new, more is being spoken in the industry around Customer Data Integration. Check out Tony Fisher’s article on TDWI for an introduction on Data Quality and the Emergence of Customer Data Integration as well as go directly to such vendor sites as DataFlux and Trillium Software for innovative solutions that work to address data quality challenges, deduplication and relationship identification.

Lastly while not being one to solicit an audience, if you do have any interest in helping the environment and stopping all that junk mail look at GreenDimes. I signed up last night… I’ll let you know how it works out!

Monday, April 28, 2008

What Gets in the Way of Good Analytics?

Today at Bank Systems and Technology, there’s an article on the increasing importance of analytics to the banking industry. The story is fairly typical in the genre – “we used to manage by gut, but better information about our customers can help us in so many ways!”

What caught my attention was that quite a few of the contributed quotes came from places on the org chart that just don't exist at most organizations – the “Director of Statistics and Modeling” and the “Department of Insight and Innovation” to name two. These references were threaded alongside a frequent comparison of “mature” analytics areas, such as credit card predictive modeling and “growing” areas, such as customer attrition modeling. This might suggest that organizations who create a dedicated function related to analytics and related disciplines are more successful at spreading the competency internally than those organizations that leave it to chance. This is certainly the position put forth by Thomas Davenport in Competing on Analytics, and is certainly intuitive in some respects.

It’s easy to envision a success story for such a group – evangelizing the power of analytics, introducing new skills to functions without a historical strength in analysis, etc. But what are the likely barriers and points of failure? How can an organization considering such an investment get ahead of the curve and mitigate the risk?

I’d speculate there are a handful of key reasons for struggle or failure:

  1. Lack of a starting point / quick win “pilot” - Perhaps it is difficult for a Center of Excellence-type structure to get off the ground without one demonstrated benefit within the first year or so
  2. Insufficient data trail - For businesses or domains without a solid trail of transactional information, it might be tougher to get started (there goes my idea for a chain of cash-only restaurants with no POS system)
  3. Lack of data architecture / infrastructure investment - If a new analytics team’s first report includes a request for $5 million just to organize the data, rough roads may be ahead
  4. Active resistance to the scientific approach - If a CEO is commonly heard to say “you guys think too much,” is that an organization likely to be hospitable to analytics?

What do you think is the biggest barrier? One I didn’t identify? What are the keys to success in building an organization's overall competency in analytics?

Friday, April 25, 2008

The Data Integration Challenges and BI (Part Two)

In Part One of this topic Brian introduced some of the key data integration challenges for a typical BI engagement and left off by highlighting some of the specific data integration challenges that included:

(i) Transformation of data that does not meet expected rules (contents of data elements and the validation of referential integrity relationships for example)

(ii) Mapping of data elements to some standard or common value

(iii) Cleansing of data to improve the data content (for example to cleanse and standardize name and address data) that extends the data transformation process a step further

(iv) Determining what action to take when those integration rules fail

(v) Ensuring proper ownership of the data quality process
In this second part of the article he takes a little deeper into several of these components.

Data transformations may be as simple as replacing one attribute value with another or validating that a piece of reference data exists. The extent of this data validation effort is dependent on the extent of the data quality issues and may require a detailed data quality initiative to understand exactly what data quality issues exist. At a minimum the data model that supports the data integration effort should be designed to enforce data integrity across the data model components and to enforce data quality on any component of that model that contains important business content. The solution must have a process in place to determine what actions to take when a data integration issue is encountered and should provide a method for the communication and ultimate resolution of those issues (typically enforced by implementing a solid technical solution that meets each of these requirements).

As Organizations grow via mergers and/or acquisitions, so too does the number of data sources and eventually lack of insight into overall corporate performance. Integration of these systems upstream may not be feasible and so the BI application may be tasked with this integration dilemma. A typical example is the integration of financial data from what used to be multiple Organizations or the integration of data from different geographical systems.

This integration is a challenge. It must consider (i) the number of sources to be integrated, (ii) commonality and differences across the different sources, (iii) requirements to conform attributes [such as accounts] to a common value but retain visibility to the original data values and (iv) how to model this information to support future integration efforts as well as downstream applications. This task is indeed a challenging one. All attributes of all sources must be analyzed to determine what is needed and what can be thrown away. Common attribute domains must be understood and translated to common values. Transformation rules and templates must be developed and maintained. The data usage must be clearly understood especially if the transformation of data is expected to lose visibility into any data that is transformed (for example if translating financial data to common charts of accounts).

Making Information Accessible to Downstream Applications

With this data integration effort in place, it is important to understand the eventual usage for this information (downstream applications and data marts) and to ensure that downstream applications can extract data efficiently. The data integration process should be designed to support the requirements for integrating data, that is to support the data acquisition and data validation/data quality processes (validation, reporting, recycling, etc), to be flexible to support future data integration requirements and to support historical data changes (regardless of any reporting expectations that may require a subset of this functionality requirement). The data integration process should also be designed to support the push or pull of data in addition. With that in mind the data integration model should provided metadata that can assist downstream processes (timestamps for example that indicate when data elements are added or modified), partition large data sets (to enable efficient extraction of data), reliable effective dating of model entities (to allow simple point in time identification) and be designed consistently.

The data integration process may at first seem a daunting process. But by breaking the BI architecture into it’s core components (data acquisition, data integration, information access), developing a consistent data model to support the data integration effort, establishing a robust exception handling and data quality initiative and finally implementing processes to manage the data transformation and integration rules, the goal of creating a solid foundation for data integration can be met.

Thursday, April 24, 2008

Boston Globe and Central Intelligence Agency to Speak on Data Architecture at DIG 2008

I am pleased to announce that the Boston Globe and Central Intelligence Agency will be speaking on the topic of Creating One Version of the Truth at DIG 2008. In addition to these two organizations, Dan Power, an industry guru, will be presenting on master data management.

Dennis Newman will present how the Boston Globe has established an Enterprise Information Management (EIM) initiative to address data integrity challenges to support an enhanced customer reporting platform. The Globe established a set of common definitions for customer-centric metrics to deliver sales and marketing analytics.

David Roberts will discuss the Central Intelligence Agencies approach to maximize value from enterprise data assets. The CIA is highly dependent on quality data and information to drive decisions. David will present the CIA’s enterprise data architecture and the value that the intelligence community has gained by having a robust data platform.


Dan Power from Hub Solutions design will present on the importance of establishing a master data management initiative and platform. Dan has over 20 years in enterprise technology with a specialization in master data management (MDM), customer data integration (CDI) and enterprise data architecture.

Friday, April 18, 2008

The Data Integration Challenge and BI (Part One)


This week I've asked a collegue of mine, Brian Swarbrick, to provide insight into some of the typical Data Integration challanges faced when developing Business Intelligence solutions. Brian, an expert in large scale data warehouse and data integration initiatives was so enthusiastic on the subject that we have decided to split his blog into two parts!! Next week I will publish Part Two. Thank you Brian!


The Data Integration Challenge and BI
The goal of any BI solution should be to provide accurate and timely information to the User organization. The User must be shielded from any complexities related to data sourcing and data integration. It is up to the development team to ensure that they deliver a robust architecture that meets these expectations.

The most important aspect of any BI solution is the design of the overall BI framework that encompasses data acquisition, data integration and information access. There are challenges in designing each of these components correctly but often data integration is the one that is the most complex yet important component of the BI solution that must be developed. A solid architecture is required to support the data integration effort (see Claudia Imhoff’s article on why a Data Integration Architecture is needed).

So what are some of key the challenges and considerations that should be addressed when thinking about data integration?

First, unless your project is tasked with building "one off’’ or departmental type solutions, it is important to separate the integration component of the architecture from the analytical component (this is the point where some readers may disagree, but separation of these components allows for a more flexible and scaleable architecture over time – a must for any Enterprise solution today). With this rule in place, the data integration team can focus on what they do best (data integration) and the analytical team can focus on what they do best (designing for reporting and analytics).

With this structure in place the data integration team has some tough challenges ahead of them that must be addressed:

(i) Identifying the correct data sources of information

(ii) Identifying and addressing data quality and integration challenges

(iii) Making information accessible to downstream applications

Identifying the Correct Data Sources of Information
Before data can be integrated it must be identified and sourced. As simple as this sounds it in not unusual for an Organization to have multiple sources of the same data. It is important to identify the data source that is the true ‘system of record’ for that information, contains the elements that support current information requirements and can extend to support future information requirements. Choose the data source that makes the most sense and not the one that is the easiest to get to.

Once the appropriate sources of information have been identified, the integration team must then determine how best to access that information. The team must identify how often the data needs to be extracted (once a day, week, etc) and how the data will be extracted (push or pull, direct or indirect). The frequency should be based on future as well as current requirements for information. It is easier to build based on what is required for today than for what may be planned or needed tomorrow. Data volumes should be a consideration when determining the optimum acquisition method and often a more frequent data sourcing process may be beneficial irrespective of the final reporting expectations (this is a good example of where separation of integration and analytics has merit since the data integration layer can be designed for optimum integration without impact to the requirements of the analytic environment).

Getting at the data itself is often more politically challenging that technically challenging. Source data may exist in internally developed as well as packaged and externally supported applications.

Pull paradigms are good when:

(a) Tools are available that can connect directly to the source systems (that’s a given) and when needed provide options for change data capture mechanisms

(b) Access to the systems is allowed; just because you can connect to a source system does not mean that the IT organization will allow that to happen – these solutions can be invasive and direct access may not be welcomed or allowed (so make sure you consider this)

(c) Source volumes are small and all data is being extracted in full or there is a means to identify new or changed records. The latter is a definite consideration when data volumes are large but there must be a means to identify these changes and it must be reliable and efficient else source invasiveness becomes a concern (especially if the source system must perform tuning to support these downstream processes)

Push paradigms (even when enterprise tools for pulling data are available) are good
options when:

(a) Data with the desired granularity, frequency and content is readily available in a different format and can be leveraged

(b) Direct access to source systems is not an option and/or the IT prefers to source the data that is needed. In this scenario a solution for change data capture may need
to be developed

(c) It is easier for IT to identify the data to be pulled and provide it instead of downstream applications pulling the data directly

Before determining the best choice for your project you also need to consider the limitations of the tools available within your environment

Identifying and Addressing Data Integration Challenges
Once the method for data acquisition has been addressed, data must be cleansed, transformed and integrated to support downstream applications such as data marts. So what does this mean and what are the potential challenges?

The size of the data integration effort is dependent on several factors: (i) the number of data sources being integrated and the number of source systems from which data is provided (ii) quality of data within each of those systems, (iii) quality of data and integration across those source systems, (iv) the Organization’s priority for improving data quality in general. When integrating data the Organization has the choice of enforcing data quality during the integration process or ignoring it.

So what are some of the key challenges for a typical data integration effort? These typically include:


(i) Transformation of data that does not meet expected rules (contents of data elements and the validation of referential integrity relationships for example)

(ii) Mapping of data elements to some standard or common value

(iii) Cleansing of data to improve the data content (for example to cleanse and standardize name and address data) that extends the data transformation process a step further

(iv) Determining what action to take when those integration rules fail

(v) Ensuring proper ownership of the data quality process

So what are some of the challenges and considerations within each of these areas? Tune in to Part Two of this article when we will address some of these considerations as well as addressing the need for making information easily accessible downstream of the integration process.

Thursday, April 17, 2008

Politics: There's No "I" in "DIG"

What do sports, politics and DIG have in common? Well, of course, it’s prediction markets. There’s Protrade, and Tradesports and the Iowa Electronic Markets and, well, Las Vegas itself, kind of. But thinking across to the other themes of the conference, the similarities disappear. Much has been said and written about the use of data and analytics in sports (Moneyball, footballoutsiders.com, 82games.com), but the closest most politicos get to analysis is focus groups, commissioned polls and a cornucopia (or is it hodgepodge?) of cognitive biases (“we need to focus on ‘soccer moms!’”).

In the last few years, some individuals and organizations have begun to make a dent in this space; notably among them Get Out the Vote: How to Increase Voter Turnout by a couple of Yale professors who base their recommendations on actual research. More recently, Brendan Nyhan at Duke reports on his blog the founding of “The Analyst Institute,” which states as its mission “for all voter contact to be informed by evidence-based best practices. To ensure that the progressive community becomes more effective with every election, we facilitate and support organizations in building evaluation into their election plans.”

It’s not as if there isn’t incentive to win, and it’s not as if there’s a lack of interested funding. So why is politics behind the curve on data and analytics? Is there a rational (or irrational) belief that politics need to be managed by gut? Or are there structural reasons? Or am I mistaken in thinking politics is late to the game, and that McCain is hiding the next Billy Beane somewhere on the Straight Talk Express?

Wednesday, April 9, 2008

In the Mood

How are you getting your engines revved for Vegas the DIG conference?

Consider this an open thread to share book and article recommendations related to data, analytics or enterprise 2.0. The poster with the most compelling suggestion will...be treated to their choice of a soft drink or adult beverage at the Green Valley Ranch by legendary DIG conference chair Pete “Memphis Ruined My Week” Graham.

Monday, April 7, 2008

Shifting Mindsets on BI

Pete Graham recently wrote a post on Using Business Intelligence in E2.0 that challenged each of us to bring business intelligence (BI) into the business conversation (verses creating a business conversation around BI). It was a prickly role reversal for those of us who like to look at the information value chain in a linear fashion beginning with data: data -> information -> knowledge (picture below of basic analytical information systems strategy). However, he provided a gentle but persuasive reminder that our mental mindsets and diagrams need to shift.

Let me explain. The idea of information and its use within business is an old idea, but its mastery reigns rather elusive. There are three core competencies that need to be achieved: Data IN, Information OUT, & Knowledge AROUND.

Data IN
Every time something happens within a business, there exists the opportunity for us to capture a piece of “data” that records its occurrence. For instance, when someone walks into a retail outlet, their visit can be recorded with a date stamp and time stamp. When the visitor buys a greeting card, the transaction is stored, inventory is marked down, and cash can be credited. If the person happens to pay by credit card, the purchase is tagged with the person’s card number. If the customer scanned their loyalty card, the transaction is immediately tagged with their profile information - and on and on. We could go on to name thousands of activities that are tracked within our organizations. These transactions let us know that something has happened!

This is not surprising. We live in a digital world where many of our actions are recorded. The challenge for businesses is to store this point-in-time data in a timely fashion and in such a way that it can be accessed quickly and easily in the future. I call this exercise, the “Data IN” process. This is the opportunity for our organizations to capture all of the happenings within our business ecosystem. Unfortunately, this raw data is unwieldly to the average business person.

Information OUT
Therefore, an organization is tasked with putting this data into context so that users can see an evolving narrative about their business. This narrative helps us to understand the what, when, and how of our businesses and their performance within the marketplace. We get to see the single occurrence (or piece of data) with the context of the business story. This process of transforming data into “information” is invaluable and gives us the digestible analytics to manage, measure, and improve our businesses.

Getting “Information OUT” is achieved by answering both traditional and current business questions with information about the past or with forecasts about the future.

Knowledge AROUND
The last piece of the information value chain is to seize the Aha! moments and business insights and push them out to the organization. For instance, a store manager who sees a declining trend in her customer base may realize that a profound shift is taking place in her market. With the combination of some analytical reporting and some field observation, she may notice that a local competitor has cut deeply into her customer base. This “knowledge” needs to be shared with her organization so that other store managers can prevent a similar decline and so functional groups within the organization can support or assist with planning a response (or change to the business). Our companies have a need to easily and quickly share insights throughout the organization, or broadcast “Knowledge AROUND”.

Today’s E2.0 tools have brought renewed energy to the business conversation represented by the Knowledge AROUND piece of the value chain. Tools like blogging, microblogging, wikis, prediction markets, etc… are democratizing the voice of the market facing parts of our organizations! This is exciting because it allows the conversation that is happening out in the field – between the people in the field and the market (customers, vendors, etc… ) to more effectively influence the information value chain. To Pete’s point, at the beginning of this post, our organizations need to bring BI into the business conversation. If we do, we have the opportunity to consistently adapt to fulfill the needs of our changing markets.

Let’s keep thinking about the paradigm shifts required to bring BI to E2.o. What do you think? What topics should we be discussing?

Sunday, April 6, 2008

Information Quality & Master Data Management?

Master Data Management is the process used to create and maintain a “system of record” for core sets of data elements and their associated dimensions, hierarchies and properties which typically span business units and IT systems.

Master Data, often referred to as “Reference Data”, may in your organization take the form of Charter of Accounts, Product Catalogue, Stores Organization, Suppliers and Vendor Lists but to name a few.

In his article “Demystifying Master Data Management”, Tony Fischer uses Customer as an example of Master data and how, if not understood and managed appropriately, can cause all sort of headaches for a company, in this case the CEO himself!

“Years ago, a global manufacturing company lost a key distribution plant to a fire. The CEO, eager to maintain profitable relationships with customers, decided to send a letter to key distributors letting them know why their shipments were delayed—and when service would return to normal.

He wrote the letter and asked his executive team to "make it happen." So, they went to their CRM, ERP, billing and logistics systems to find a list of customers. The result? Each application returned a different list, and no single system held a true view of the customer. The CEO learned of this confusion and was understandably irate. What kind of company doesn't understand who its customers
are?”

So what are the typical barriers that hinder organizations from addressing their master data management problem? My colleagues and I typically encounter four primary barriers:

Multiple Sources and Targets: Reference data is created, stored and updated in multiple transactional and analytic systems causing inaccuracies. Synchronization challenges between disparate systems

Ability to Standardize: Most organizations cannot agree on a standardized view of master data. There are a lack of audit policies that comply with federal regulations

Organizational Ownership: Disagreement within the organization as to who takes ownership of master data management, business or IT. Assignment of accountability with cross-functional processes is difficult

Centralization of Master Data: Organizational resistance to centralizing master data since there is a sense that control will be lost. Challenges to find a technology solution that supports existing systems and the lifecycle of master data management


Organizations that are addressing such barriers typically have a successful master data management process in place that contains the following components:

Data Quality: Focus on the accuracy, correctness, completeness and relevance of dataIncorporate validation processes and checkpoints. Effort is highest in the beginning of a MDM initiative to correct quality issues.

Governance: Cross functional team formed to establish organizational standards for MDM related to ownership, change control, validation and audit policies. Focus includes establishing a standard meeting process to discuss standards, large changes and organizational issues.

Stewardship: Assignment of ongoing ownership of MDM stewardship. Typically MDM stewards are business users. Accountable for the implementation of standards established through MDM governance

Technology: Create an architectural foundation that aligns with the other three components. Implement a technology that centralizes reference data. Align processes with the technology solution to synchronize master data across source and analytic systems


As we can see, master data management is not a one-time initiative but rather a long-term program that runs continuously within the organization. To be successful organizations need to instill an iterative approach that helps develop a program that continuously monitors, evaluates, validates and creates master data in a consistent, meaningful and well communicated way.

What is your organization doing about Master Data Management? Have you had success in establishing a Data Governance program? Who own the process in your organization, IT or the business?

Thursday, March 27, 2008

What’s all the hype around unstructured data?

Check out the DM Review article by Michael GonzalesComprehensive Insight: Structured and Unstructured Analysisfor an introduction into the topic of unstructured data and releasing its potential.

Michael provides some insight into the challenges organizations face in dealing with unstructured data vs. structured and how technology has been evolving to help better leverage such information assets.

With an estimate that “more than 85 percent of all business information exists as unstructured data” it is no wonder that technology vendors are putting more focus on how to extend their products to make this information more accessible and usable.

Although the article gives some interesting insight into the evolution of the technologies it doesn’t provide any insight into how to actually integrate and store this data in the traditional data warehouse. How does one integrate such unstructured data in the form of documents, images, video content, and other multimedia formats? Is such data actually relevant to data warehouses and CPM processes? Perhaps not the actual content but perhaps the metadata associated with the content (e.g. x number of documents types, average occurrence of y in videos of type z, number of emails on subject w, etc).

Vendors that are beginning to address the storage and integration of such unstructured data into existing solutions are primarily the large database vendors. Certainly “Big Blue” (IBM) boasts support for analysis of unstructured data with its DB2 Warehouse 9.5 product offering and Microsoft SQL Server 2008 is touted by Microsoft to “provide a flexible solution for storing and searching unstructured data”.

Although the advances in vendor technologies are providing a means of storing such information in a manner that makes it accessible, is the typical organization yet ready to focus its resources on doing so? When so many have yet to fully realize the benefits of provisioning to the business traditional structured data, e.g. Financial, Operational, Customer, etc, you have to beg the question as to whether this should yet be a high priority?

What is your organization doing? Have you implemented any creative solutions? What are the demands from the business?

Wednesday, March 19, 2008

Real-time Data Usage

This past weekend was a major weekend for me, one that would determine my core happiness for the rest of the year!

You see I’m Rugby mad. I’ve played it, I’ve coached it, and now I watch it, incessantly.

This past weekend was the concluding weekend of the European 6 Nations Rugby Championship, an annual international tournament between the home nations of Europe (England, Ireland, Scotland, Wales, France and Italy). A tournament that started in 1871 and every year since has been an excuse for all to bring out their nationalistic pride and cheer for their ancestral team! For me it’s Wales, the land of my forefathers, the land of daffodils and of course sheep. Wales had the opportunity to win the tournament in style, beat France at home and raise the championship trophy undefeated, Grand Slam winners. And did they do it? They sure did!

Rugby, just like most American sports is now a professional sport and with it many changes have come. Dragged out of the traditions of amateur pastimes where the local butcher was your star player, teams today are forced to continually explore all possible avenues in an attempt to better themselves and obtain competitive advantage over their rivals. No longer is it good enough to just employ the best players and coaching staff, teams are looking elsewhere.

One interesting field of innovation that we were given a brief insight into during one of the games was the usage of statistical information real-time by the Welsh coaches that allowed them to then make real-time adjustments to how their team and players were approaching the game. Using the interesting technology Sportstec the Welsh team was actively making adjustments that helped provide a competitive advantage over their opposing team.

Around the field “spotters” were employed to feed information into a database application on specific events happening. Number of times a player passed to his left vs. his right, how many carries an individual had with the ball, number of times the ball was kicked from a certain place vs. passed, number of missed tackles by each player, success rate of a particular move, etc... By providing such detailed information on actions performed by their team as well as the opposition, the coaches were then able to react and make tactical real-time changes, for example adjustments to the team’s strategy on the field, instruction to specifically focus and improve in certain aspects of the game, as well as instruction to target identified weaknesses in the opposing team.

Did having this level of information access have a direct result on the outcome of the game? Who knows, but one thing for sure Wales beat Italy in this game 47-8, when on average the other teams who beat Italy did so by only 6 points! The other thing, did I mention, Wales won the Championship, undefeated!

So where next? If teams are able to get hold of and make use of such real-time data I just wonder how much further they could go.

What if players began to wear RFIDs in their shirts so we know where they are at anyone time. The ability to understand in real-time how much time they have spent in one location, how much time it took them to reposition, the average distance they make while running with the ball, efficiency of path travelled by each player when covering a kick-off. What about collecting information from body skins that can sense applied pressure? Could we measure the level of impact endured in a tackle thus giving the ability to predict level of fatigue vs. amount where performance begins to degrade, an opportune time for a tactical substitution perhaps?

For more detailed commentary on how the Welsh team and others are finding innovative ways to capture and use information check out the videos @ http://www.sportstec.com/videos.htm

Wednesday, March 12, 2008

Top 5 mistakes in Data Warehousing

Top 5 reasons why many data warehouse managers fail to deliver successful data warehouse initiatives:

Data Quality: Quality of source system data that is to be integrated into the data warehouse is “overrated” and thus time to resolve is “underestimated”

  • Bad information in means bad information out. The CPM applications that will source data from the warehouse will suffer diminishing adoption if not addressed upstream
  • Data integration strategy must include methodology to address erroneous data
  • Significant level of involvement from business and IT to help resolve (decision and execution of) challenges

Data Integration: Lack of robust data integration design results in incomplete and erroneous data and unacceptable load times

  • What happens when you are the process of loading data and you start receiving exceptions to what is expected? Is data rejected and you are now faced with the dilemma of partial data loads? How do you avoid manual intervention?
  • What checks and balances do you have in place that ensure what you are extracting from source systems is being populated into the target? Can you audit your data movement processes to ensure completeness as well as satisfy regulatory obligations?
  • Your processes can handle the data volumes you are dealing with today but can they handle the data volumes of tomorrow? How easy is it to reuse existing processes when adding additional source systems/subject areas to your Warehouse?

Data Architecture: Creating a solution that is not able to scale after an initial success will result in a redesign of the architecture

  • After the first success the business will quickly want to extend the usage of the solution to a greater number of users, will the performance continue to live up to expectations?
  • As users mature and adoption improves so will the complexity of information usage, i.e. more advanced queries, can the design continue to perform as expected?
  • Increased usage and maturity results in the demand to integrate into the solution additional data sources/subject areas. Is the architecture easily extensible?

Data Governance & Stewardship: With no controls established around data usage, its management and adherence to definitions, data silos and erroneous reporting begin to reappear

  • Stakeholders must be identified and give decision rights to help improve the quality and accuracy of your common data
  • Practices around the managing of standard definitions of common data and business rules applied must be established
  • Understand who is responsible for the data and hold them accountable

Change Management: Not preparing an organization to utilize what is being built results in the investment in data warehouse not being fully realized and thus deemed a failure due to low user adoption

  • “Build it and they will come”; providing information access does not necessarily equate to information usage.
  • Helping the business understand how they can leverage these newly available data often results in changes to the way that they work. “Day in the life of” today vs. “day in the life of” tomorrow
  • Education and training programs are required
  • Integrated project teams (business and IT) are essential to the success of data warehouse initiatives, with individuals becoming champions within the organization for change and adoption

Saturday, March 8, 2008

Importance of having accurate data to drive decisions

I recently read an article in the Boston Globe discussing the unexpected rise in costs related to the state's universal healthcare plan. The program requires that every single state resident have healthcare coverage (a soon-to-be national topic based on the outcome of the presidential elections). What the article highlights is that the anticipated costs of the program could double and that the state has not budgeted for the increase in costs. The primary driver of the increased budget is that the state underestimated the number of state residents that do not have healthcare coverage. The legislature had two numbers to use to drive the budget model, one source being the state's estimate of 460,000 and the second source being the US Census Bureau estimate of 748,000. Unfortunately for the state, they used a number somewhere in between and they are now realizing that their budget will be short.

Beyond the political nature of the article, what I find interesting about this situation is you can clearly see the importance of having accurate data when making a decision. More appropriately said, the state had already made a decision to provide universal healthcare for every state resident, but because of the poor quality of data, how they allocated resources (i.e. money) is being significantly impacted.

Projected Enrollment by Year

Thursday, March 6, 2008

Building an enterprise semantic layer

I recently read a blog post on FastForward by Paula Thorton mentioning a Reuters technology infrastructure called Calais. The purpose of Calais, to put it in simple terms, is to provide a service to automatically put context to unstructured data. The unstructured data could be in the form of news articles, blog postings, or any other text based content. The Calais web service would then identify the entities, facts and events based on the natural language descriptions in the text. The service would then return a descriptive model of the unstructured data.

The reason why this is important or at least is generating excitement is that many believe Web 3.0 will be based on this type of inference of content. Web 2.0 is highly dependent on many individuals applying context themselves through concepts like tagging, social sharing and general broader collaboration.

Now, the majority of the excitement and focus is on the user base that is currently driving the Web 2.0 movement. As someone who has done a majority of my professional work inside the four walls of an enterprise in areas such as data architecture, data integration, business intelligence and corporate performance management, I see incredible opportunity in something like Calais. A majority of the effort associated with building internal measurement systems like dashboards and management reporting applications is in developing a single semantic layer of metadata. Aside from the effort to develop the semantic layer, it is typically inconsistent because the organization lacks the ability to agree on a common business taxonomy that describes the enterprise.

When we think about bringing Web 2.0 technologies (social networks, wikis, prediction markets, blogs) to the enterprise, the critical first step is building out a metadata layer that puts descriptors on data to create information and starts to establish context. There is plenty of unstructured data floating around organizations in the form of documents, internal web sites, email communications and traditional knowledge centers. In addition, there is an inordinate amount of structured data, which I would argue, can have very little value when it lacks context. It requires speaking to an info-worker who can explain the report, dashboard or data set. These sets of structured data are typically living/trapped within silos of different functions such as finance, marketing, sales and operations. If you look at the collective structured and unstructured data/information that an organization captures across disparate groups/functions, being able to "infer" the entities, facts, and events will start to build the context needed to make better informed decisions. Add the ability to link people through the social network of an organization to share and disseminate information, and you are starting to see the value in implementing a Web 2.0 platform across the enterprise.

Tuesday, March 4, 2008

Welcome to the Data Theme

How often do you sit back and spend a moment to actually think about how much data we actually generate and use on a day to day basis? Whether it be in the office sending an email to a colleague, or updating a financial spreadsheet for your CFO we are using and generating data. Do you ever think that when you simply stop on the way home to fill up your car with gas, or scroll through the TV guide looking for your favorite show to TIVO, you are generating and using data?

With advances in processor speeds, data storage capabilities, and application technologies our capacity to generate and capture data is forever increasing. Many organizations are leveraging this data, turning it into usable, sustainable information that can be used as an asset to help gain competitive advantage. The majority of organizations however are simply overwhelmed as to what to do and where to start.

Over the next couple of months I welcome you to join me as we explore the theme of Data. We will look to discuss a variety of topics that relate to the challenges faced by organizations who are working to develop an "information architecture for the 21st century". Not only will we discuss the traditional well heeled topics such as "Data Strategy", "Data Quality" and "Data Integration", but also others that often influence the success of organizational initatives due to lack of prioritization or simple underestimation, such topics may include "Data Governance", "Master & Meta Data Management", and "Managing Unstructured vs Structured Data" to name a few. Of course if there is a topic of interest that you wish to discuss I fully encourage the suggestion.

***

Glyn D. Heatley bio - I'm a Director and Leader in the Information Strategy and Architecture Practice at The Palladium Group. I bring over 13 years of experience in delivering large scale Corporate Performance Management solutions with a primary focus in Data Management and Business Intelligence. Over the years I've gained experience in all areas of the Data Warehouse Life Cycle including Requirements Gathering, Solutions Architecture, Data Architecture, Data Integration, Business Intelligence and Project Management.