This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Tuesday, November 29, 2011

Navigating eDiscovery in the Cloud Shouldn't Be That Difficult

In a follow up to my Blog post titled, "eDiscovery in the Cloud: The Sky Is Not Falling", this Blog post is dedicated to the premise that successfully navigating eDiscovery in the cloud is not as complicated as many are indicating it should be or as complicated as many are making it.

Successfully navigating the brave new world of eDiscovery in the cloud is really just a matter of education and a willingness to move beyond the status quo.  There is no doubt that if you don't pay attention, you and your team will perish on the rocks. However, don't pass on taking the eDiscovery in the cloud journey because it is too dangerous or give up before you at least make an attempt to learn how to save your ship.

First of all, in case anyone missed the memo, the cloud train has left the station.  As an example, independent research firm Forrester Research predicted in a research report published earlier this year titled, “Sizing the Cloud” that the global cloud computing market would reach $241 billion in 2020 compared to $40.7 in 2010.  So, more than likely, whether you want your data in the cloud or not, it is moving quicker than you think.  And, as an end-user, unless you have some kind of cloud storage phobia, it really shouldn't matter that much.  The real debate doesn't start until you couch the question(s) about cloud computing in terms of what happens when your have to perform the delicate and often times messy operation of eDiscovery in the cloud.  If you are a glutton for punishment and like to dwell on all of the negative things that could possible happen in the life then I encourage you to read "The Promise of the Cloud Meets the Obligations of E-Discovery", published on the Law.com website on October 12, 2011 by Brendan M. Schulman and Samantha V. Ettari.  This article does a great job of indicating that the sky is falling and that we are all doomed.  However, as I indicated in the my response to this piece, "cloud computing has already made it and most of us are just fine, eDiscovery in the cloud and all!!"  But, the devil is always in the details and therefore what does this mean in practical terms?

Further, please note that if you are currently doing a bad job of eDiscovery in general, you had better read the Schulman and Ettari article as the sky is going to fall if you attempt to perform eDiscovery in the cloud under your current practices. Once you have completed reading that article and if you still want a road map for successful implementation of eDiscovery in the cloud, come back and finish reading this blog post.

What is eDiscovery in the Cloud?To properly perform eDiscovery in the cloud,  you first have to understand what it is and, probably more importantly, what it is not.  The current crop of litigation technology vendors have done a great job of confusing the market in regards to eDiscovery in the cloud.  However, I believe that over the next 12-18 months, the market will become much more educated and some amount of consensus will begin to form regarding a more realistic and concise definition of eDiscovery in the cloud.

eDiscovery in the cloud is NOT uploading all of your potentially responsive ESI to a litigation service provider's data center and then accessing that ESI via the Internet to perform searches and document review.  That may be Early Case Assessment (ECA) or document review delivered under a Software-as-a-Service (SaaS) model.  But, it is not eDiscovery in the cloud.

Likewise, eDiscovery in the cloud is NOT manually collecting big chunks (that's a technical term) of potentially responsive ESI from your cloud provider and the performing eDiscovery with that ESI the same way you process ESI from your corporate network or from unconnected desktops and laptops (BTW - I am in the process of investigating the nightmare of collecting ESI from your cloud provider and plan to author a Blog post of my findings before the end of the year.  So, if anyone has any input, send it to me and I will consider including it in my post).

eDiscovery in the cloud ultimately means having a virtual eDiscovery process that actually runs in the cloud right alongside of your cloud storage and allows you to perform, Early Case Assessment (ECA) including First Pass Review, possibly preservation and legal hold management, definitely forensically sound collection and the generation of an industry standard load file and/or full on document review and production.  In addition, eDiscovery in the cloud also means that you can operate these processes remotely through an Internet based user interface and don't have to have operational bodies physically inside the cloud data center(s) to perform any of the normal magic that is currently required by many of the legacy hosted eDiscovery platforms. 

Further, eDiscovery in the cloud should also include what I am going to call (for lack of a better term at this point) federated eDiscovery to enable an organization to "perform eDiscovery" on data no matter where it resides.  Currently, users that are supported by competent IT organizations, don't have to worry about where ESI is physically located.  Therefore, eDiscovery professionals shouldn't have to worry either.  This would include ESI behind the corporate firewall, housed with different cloud service providers or housed with the same cloud service providers in different data centers potentially in different countries (don't get me started on the debate regarding the legal issues with moving ESI in and out of countries as that is the topic of a future Blog post). Please note that I am not oblivious to the challenges of moving large amounts of data around.  However, we all might be surprised to learn that class 5 rapids have been successfully navigated in other industries.

Is this Definition Realistic
This definition of eDiscovery in the cloud may sound like something that only Scotty, the engineer from the Star Trek Enterprise, could cobble together with technology from the next century and a good amount of duct tape.  However, the technology exists today and is ready to be utilized with little or no duct tape required.  Therefore, the only real speed bumps on this journey will be convincing the cloud service providers to install the appropriate eDiscovery technology as a standard part of their technology stack, enlisting a new generation of eDiscovery consultants to support the development of best practices for eDiscovery in the cloud and finally to show the market that eDiscovery is no longer a reason to NOT move your data to cloud.  I realize that these are not insignificant roadblocks.  However, providing eDiscovery as a standard part of it's technology stack is a homerun for cloud service providers and the associated services represents a blue water/green field market opportunity for eDiscovery consultants and possibly service provides. Therefore, resistance should be minimal and buy-in should be quick.

What's Next?
In the coming weeks I will be releasing my initial list of eDiscovery technology vendors that can support my vision of eDiscovery in the cloud along with an initial overview of the best practices.  If anyone has any input that you believe should be included in these upcoming Blog posts, send them to me and I will consider including them.

In the mean time, if you are concerned with moving your data to the cloud and are hesitant because you are concerned about eDiscovery or if you are currently faced with the daunting task of extracting your ESI from a cloud service provider, contact me as I can help you successfully navigate your way through this paradigm shift.

Labels: , , , , , , ,

Monday, November 1, 2010

Building ROI for an eDiscovery Cloud Computing Model

Cloud Computing is an important step in the evolution of Enterprise Information Technology Systems.  And, without a doubt represents that next big paradigm shift in eDiscovery Information Technology. 

As such, I plan to dedicate my blog for the remainder of the year to investigating and reporting on Enterprise eDiscovery in the Cloud.  As an adjunct to this blog, I have also started a LinkedIn Group called “Enterprise eDiscovery in the Cloud”.  Click Here to join this LinkedIn Group.

As a place to start the discussion, I want to present and then begin to analyze an absolutely outstanding whitepaper titled, “ Building Return on Investment from Cloud Computing”, by Mark Skilton, Director, Capgemini and other members of the Cloud Business Artifacts Project, that I found on the OpenGroup Website.

Cloud Computing is not a sliver bullet; is not going to be easy or inexpensive to implement; and, may not necessarily be the correct infrastructure and/or application delivery mechanism for either vendors or end-users within the eDiscovery market.  However, done properly with an appropriate amount of planning and with the support of the “right” partners.  I predict that Cloud Computing will be the “next big thing” and will propel eDiscovery technology capabilities to a level not readily available today.
The full text of the whitepaper by Mark Skilton and other members of the Cloud Business Artifacts Project is as follows;

Introduction

Cloud Computing has been described as a technological change brought about by the convergence of a number of new and existing technologies.
The promise of Cloud Computing is primarily the following key technical characteristics (see Above the Clouds [4]):
  • The ability to create the illusion of infinite capacity; the performance is the same if scaled for one, to a hundred, or a thousand users with consistent service-level characteristics.
  • Abstraction of the infrastructure so applications are not locked into devices or locations.
  • Pay-as-you-go usage of the IT service; you only pay for what you use and with no or minimal up-front investment costs. You typically just use the service through a connection and device.
  • The service is on-demand; able to scale up and scale down with near instant availability. Typically, no forward planning forecast is required.
  • Access to applications and information from any access point.
But this is only half the story. These technical characteristics can also be found in non-disruptive technology solutions. The rate of change and magnitude of cost reduction and specific technical performance impact of Cloud Computing are not just incremental, but can give a five to ten times order of magnitude improvement.

The Capacity-Utilization Curve

The famous graph used by Amazon Web Services illustrating the capacity versus utilization curve has become an icon in Cloud Computing. The model illustrates the central idea around Cloud-based services enabled through an on-demand business provisioning model to meet actual usage.


The Capacity versus Utilization Curve
Why this matters to business is that one of the core precepts of Cloud Computing is to avoid the cost impact of over-provisioning and under-provisioning. This is in addition to the opportunity for cost, revenue, and margin advantages of business services enabled by rapid deployment of Cloud services with low entry cost, and the potential to enter and exploit new markets.
We contend that in years from now, when Cloud Computing is seen in a historical context, the capacity versus utilization curve will be seen as an iconic model that had the same effect as previous well known business models.


Iconic Business Models
Examples of these include:
  • The Moore’s Law model that establishes the concept of exponential growth in computational power but has subsequently been seen in other technology areas, including storage and network
  • The technology hype cycle that established the emergence of innovation lifecycles, and is developed in publications by Charles H. Fine [2] and Clayton M. Christensen [3]
  • The Boston Consulting Group Growth-Share Matrix that can be used to show how key industrial markets and products and services undergo transitions as the maturity lifecycles emerge, grow, and recede
But what does this potential icon mean for business? Matching capacity and actual utilization on demand improves operational efficiency, but is that all there is to it? Capacity and utilization are Key Performance Indicators (KPIs). They measure how much or how little something is being used. But is this aligned and being used to generate Return on Investment (ROI)?

Race to the Bottom versus Quality of Service (QoS)

The positioning of Cloud Computing, while initially seen as a disruptive technology influence on both buyer and seller prospects, is now evolving into a trade-off between low-cost arbitrage and added value Quality of Service (QoS).


Race to the Bottom
The terms “race to the bottom” or similarly the “prisoner’s dilemma” refer to the competing drive between participants in a market driven by the need to make the greatest cost savings. The term is often seen in a negative context, as the lower costs and margins are seen as a detriment to the participants. Massively scalable services from Cloud Computing providers have the effect of driving down costs and prices, as the dynamics of competition are shifted by the presence of potentially rapid cost reductions and huge data center investments.
The counter-balance to this is the Quality of Service (QoS), and the associated Cost of that Service (CoS) that characterizes the value of the cost per unit of performance provisioned (see the discussion on the Financial Value Perspective of Moving from CAPEX to OPEX and Pay-as-you-go).
The differentiator of Cloud Computing is not just the utility infrastructure computing services, but includes all the higher-level services that enhance and build business service value. We see this as the influence and scope of the movement from IT-centric to business-centric services across a wider services continuum, with utility services for infrastructure at one end, and with business-centric software and business processes delivered as a service from the Cloud at the other.
In addition, the need to provide adequate security should be considered. People are willing to pay a little more for a service if they are assured that there will be good security measures in placed.
This issue is highlighted in this White Paper as it has a direct bearing on the Cloud Computing ROI debate and how it is measured:
  • Pricing and costing of Cloud services
  • Funding approaches to Cloud services
  • Return on Investment (ROI)
  • Key Performance Indicators (KPIs)
  • Total cost of ownership (TCO)
  • Risk management
  • Decisions and choices evaluation processes for Cloud services
The discussion of the market dynamics of Cloud Computing is not developed further in this White Paper, but is a recommended area of research going forward as more products and services become Cloud-enabled.

Traditional IT Compared to Cloud Computing

The iconic capacity versus utilization curve of Figure 1 provides a yardstick of current thinking in Cloud Computing provisioning. It is shown here for increasing demand, but the same model can be applied to both growth and decline of capacity demand in a periodic pattern typical of many businesses.
The following shows some of the characteristics of traditional IT compared to Cloud-Computing.
Traditional IT
  • Hardware is hosted on the premises of the organization and/or manage hosted.
  • Hardware and software is provisioned for peak demand.
  • Service management monitoring is used to generate forecasts of demand usage and current SLA performance.
  • Chargebacks and compensations are used to adjust usage and payments.
  • Under-provisioning and over-provisioning of capacity can result from unforeseen demand changes.
  • Business invests in ownership of assets that can be enhanced and extended through IT programs and development.
  • Changes to IT involve migration and divestment/investment issues and programs.
Cloud Computing
  • Hardware and/or software is hosted off-premise (public or hybrid) or on-premise as a private Cloud service.
  • Services are provisioned and used based on actual demand, providing this elasticity as a managed service.
  • Services are typically focused on short-term “burst” demand to gain cost savings over provisioning and owning the assets.
  • Statistical automated scaling is used to optimize the shared virtual assets.
  • Risk is transferred from the buyer to the seller/provider of the Cloud service.
  • Cloud sellers and providers seek to grow amortized economies of scale through increasing the numbers of users of the shared resources.
  • The IT infrastructure and operation is masked from the service user. Cloud is more than just SaaS.

New Technology Adoption Lessons from Other Industries

Understanding the characteristics of Cloud Computing can be assisted by observations from other industries that are in the process of transformation. Lessons from alternative power sources such as solar energy and wind power contain examples of issues that resonate with familiarity in the Cloud Computing context.
When comparing Cloud Computing to solar and wind energy there are similar adoption issues:
  • Potential unlimited energy resource
  • Challenges to distributing efficiently
  • Demonstrating its value over traditional/other alternatives

Solar Energy and Wind Power
Just taking a look at the solar energy and wind power characteristics:

New Technology Adoption from a Buyer’s and Seller’s Perspective

Examining the issues of effective wind power or solar energy compared to contemporary energy sources draws parallels with the challenges we also see in defining new technology adoption.
Sellers of resources and services characteristically focus on their operation and technical development, and how they can enable effective business models for existing and potential new customers and markets.
Buyers are typically not concerned about how the sources of energy, resources, or services were generated and delivered. They seek to understand whether their businesses can be supported by the products or services, and whether these can be reliable and cost-effective. Buyers want to know the choices on offer and how they may be able to enhance or swap resources and services for improved business performance.
The following shows some of the concerns of buyers and sellers of new technology.
Buyers
  • Don’t care where the service comes from or what medium was used to generate it.
  • Does the service address the business requirement?
  • Is the QoS reliable?
  • What are the switching costs from one energy provider to another?
  • Is the service cost effective?
  • There are advocates and dissenters.
  • Can I use the service when and where I need it?
Sellers
  • Efficiencies of production compared to existing alternatives.
  • Storage of the service and use in service to meet on–demand needs at point of use.
  • The cost of (architecture) to deploy and distribute the service.
  • Purchasing incentives and direct governance investment.
  • There are advocates and dissenters.
  • Location, security, and access?
In Cloud Computing the common themes in engaging sellers and buyers in the new technology provisioning model includes three key questions:
  • How does Cloud Computing compare to traditional IT? This principally relates to the comparison of service-level performance and license costs.
  • What can I not put in the Cloud? Answers typically include UNIX systems, mainframes, and very high I/O applications, but pretty much anything can be co-located or hosted in an elastic virtual container environment. Beyond the technical definitions there are the business processes and provisioning models that set Cloud Computing apart from its predecessors of utility computing and virtualization.
  • How does Cloud Computing impact revenue and budget lines? This issue involves the cost/performance enabled by virtualization and economies of scale, and the lowered need for up-front investment. Movement of revenue to Cloud providers may need to be balanced by sale of added-value services.

Above and Beyond the Clouds

There are many definitions and viewpoints provided by the sellers of what is now termed “Cloud Computing”. Much of the vocabulary used is defined from the perspective of IT performance and capacity, and the impact of cost savings of asset ownership and variable seller service costs. Yet all these have direct cost-benefit impact on the business consumers of the end services and how they compete and deliver products and services in their industry.
Many business IT departments have addressed emerging trends through actions to drive cost reduction and leverage IT service providers’ adoption of Cloud style services. Many industry organizations and leading IT suppliers of software, hardware, and services, seeking to address their customer needs, have vigorously evaluated and followed a Cloud-style strategy. The challenges and issues are in the transition from the current traditional IT to the new potential capabilities of Cloud Computing. They must be expressed in a language that business end users can understand, and relate to investment, cost improvements, or business performance.
Closer to home in understanding the issues of Cloud Computing, we can draw upon the University of California, Berkeley RAD Lab. White Paper Above the Clouds [4]. This highly informative paper identifies the technical issues for Cloud adoption, and its potential for business benefits and technical challenges. The paper also sheds light on the value that these technical scenarios can provide to business:
  • Avoid missed business opportunities from under-provisioning and over-provisioning (Page 1)
  • Responsive Service to variable demand (Page 2)
  • Pursue emergent and explorative new business market opportunities hitherto unforecast or predicted (Page 2)
  • Perform cost associative tasks fast and lower cost (Page 2)
  • Decouple utility services and brokering from business front end (the “fab-less” chip foundries example) (Page 3)
  • Make more money from amortizing economies of scale (Page 4)
  • Leverage existing investments through hybrid means (Page 4)
  • “Anywhere” services, “border-less” delivery (Page 4)
This interpretation of the Above the Clouds White Paper in a business-issue context illustrates some interesting aspects of Cloud Computing potential.
The decoupling of resources and provisioning (what can be termed “back end”) from the “front end” business use through intermediation follows much the same argument as Why Buy the Cow [5] from Ivar Subrah on how on-demand powers the economy and The Big Switch [6] by Nicolas Carr that takes the analogy even further with distributed industrial IT services.
The business of IT becomes that of leveraging the products and services through competing platforms and channels to defend, attack, and build customers and market share (see the discussion on the Importance of a Business Perspective of the Cloud).
However, a secondary issue from this, described as a “race to the bottom” by many industry observers, arises where commodity charging lowers the cost, removing margin benefits. The counter-balance to this is the Cost of Service (CoS) and Quality of Service (QoS) charges and added value on top of the products and services. We explore this in the next section.
Even after the passing of time from the date of publication of Above the Clouds [4], the technical challenges are still evident, but they are becoming less so with the evolution Cloud services and technology.
The Above the Clouds paper famously asserts that Private Cloud is not Cloud Computing (Page1) as it is not open to the general public, which is part of the original definition. This illustrates that definitions of Cloud are still evolving, and perhaps the definition commonly assumed in the marketplace today does not necessarily require openness, though it still retains the central concepts of elastic capacity and provisioning.

Building Return on Investment from the Cloud

The central theme of this White Paper is how to go beyond the initial capacity and utilization benefits described in Cloud Computing.
The view of capacity and utilization is a technology provider/seller viewpoint which is essentially based on key performance indicators (KPIs) rather than business benefit metrics.
  • IT capacity, as measured by storage, CPU cycles, network bandwidth, or workload memory capacity is an indicator of performance.
  • IT utilization, as measured by uptime availability and volume of usage is an indicator of activity and usability.
But effective cost/performance ratios and levels of usage activity do not necessarily imply proportional business benefits. They are just indicators of business activity that are not in themselves more valuable than lower operating cost. There are, however, business metrics that translate the indicators of the capacity-utilization curve to direct and indirect benefits to the business, as illustrated in Figure 5.


Business Metrics Derived From Capacity/Utilization
These metrics are described in the following sections.

Speed of Cost Reduction – Cost of Adoption/De-Adoption

The introduction of Cloud Computing as an option transforms cost of ownership and changes the dynamics of the provisioning cycle in a number of fundamental ways.
The speed and rate of change of cost reduction can be much faster using Cloud Computing than traditional investment and divestment of IT assets. In Cloud Computing the buyer can move from a CAPEX to an OPEX model through purchasing the use of the service rather than having to own and manage the assets of that service. This responsibility is transferred to the service provider.


Speed of Cost Reduction, Cost of Change
The use of Cloud Computing to the user also potentially means a movement to a pay-as-you-go style billing model which can have different tariffs and contractual obligations compared to traditional IT ownership. These can include minimum usage periods and flexible pricing per usage profiling (see the discussion on the Financial Value Perspective of Moving from CAPEX to OPEX and Pay-as-you-go).
The key issue is the ability to adopt and remove the service either at the point of use (to scale up and down) or to make choices to use new services or change service provider.
Migration between Cloud services is still a challenge. There are portability and interoperability issues.. And hosting corporate and personal data and knowledge on the Cloud can make customers of Cloud services dependent on the providers.
There is a trade-off between the benefits of speed, cost, and Quality of Service (QoS) from a particular Cloud service provider and their ecosystem of services versus the flexibility and choice of alternative services and Cloud solutions.
The cost of change in an ROI business case is less in Cloud Computing as the choice of selected Cloud services is more stable and more cost-effective than traditional ownership.

Optimizing Ownership Use

The use of IT has become an enduring feature in all organizations today. The investment in data, knowledge, and infrastructure assets and software code now represent many lifeblood operations for businesses.
But many of the issues of cost of ownership are often decoupled from choices made during selection of new IT, and the impact on the long-term running and maintaining of these IT services and subsequent business usage is not properly considered:
  • Technology design choices and purchasing are often done by strategic or tactical contractual purchasing based on project requirements, with little consideration for optimizing running and maintainance over the whole system lifecycle, but
  • The cost of maintenance and modifications often represent a significant part of the asset lifecycle beyond initial provisioning.
The ability to “design and provision for run”, so that the choices of IT procurement are aligned with the best options and performance for long-term operation, has long been an ideal goal of business and IT. But, while technical trends such as OO, SOA, and Web 2.0 have brought functional improvements, the improvements in runtime infrastructure support have always been illusive.
A key aspect of moving to Cloud Computing is the ability to select hardware, software, and services from defined design configurations to run in production. Cloud Computing in effect seeks to bridge the design-time and run-time divide and optimize service performance. Patches and upgrades or new technology are in theory invisible to the end user of the service as they are included as part of the automatic asset management features.


Optimizing Ownership Use
Cloud Computing can help an enterprise achieve the goal of a more cost-effective asset–management lifecycle process for the IT portfolio, to optimize both design and run-time performance.
While the capacity-utilization curve can reflect overall usage, this can be broken down to identify which assets need to be supported in this way and to rationalize, consolidate, and optimize the assets that need to perform for business goals.
The key benefits to Cloud Computing ROI from a business case perspective are in the optimization of the total asset portfolio.

Rapid Provisioning

Elastic provisioning to scale up and down to actual demand creates a new way for enterprises to scale their IT to enable business to expand.
The provisioning time compression from a week to hours, for example, demonstrated by Cloud Computing sellers/providers is a means to rapid provisioning that is not just about saving time but is also defining a new business operating model.
Organizations can review and develop business plans and then deploy infrastructure and services in a more rapid and proactive way.
Customization and development, testing, and support can also been seen in a new light with the provision of IT services in a dynamic fashion targeted at business needs.
Buyers and sellers can view rapid provisioning as a marketplace of services. Sellers can offer rapid provisioning services that sustain buyer needs for existing IT services, offer choices for innovation, and enable rapid introduction of new technology.


Optimizing Time to Deliver/Execution
The impact of rapid provisioning on ROI business cases can be profound. Examples in the government/federal sector as well as financial services and consumer goods are already evident and pointing the way to the emergence of online Cloud-based marketplaces as de facto standards for current and future trading between suppliers and buyers of services.

Increase Margin (Make More Money)

One of the core precepts of Cloud Computing is to avoid over-provisioning and under-provisioning. This is in addition to the opportunity for cost, revenue, and margin advantages of business services enabled by rapid deployment of Cloud services with low entry cost, and the potential to enter and exploit new markets.
What makes Cloud Computing exciting is that the potential for business is not just the incremental change improvement but the disruptive transformational effect Cloud Computing can have from the possibilities of new business operating models.
Cloud Computing enables business to pursue new and existing markets by rapid entry and exit of the products and services. It enables enterprises to “land and expand” in markets with an infrastructure and service capacity that can grow with the business.


Optimizing Margin
Cloud Computing removes the need for additional infrastructure to test and enter markets for business (a key benefit feature particularly for small to medium size organizations).
Cloud Computing has an impact on the margin through cost reduction and through economies of scale to make more use of the same resources.
The impact on a ROI business case is that an enterprise can make more money or better use of existing investments through Cloud Computing.
There are many cases of companies large and small that can enter and develop service offerings through a “long tail” (refer to the publication by Chris Anderson [7]) approach enabled by Cloud Computing infrastructure. Existing and new markets can be attacked and entered through speculative and well-timed interventions to exploit and grow business performance.

Dynamic Usage – Elastic Provisioning and Service Management

The focus on capacity and utilization can be taken further through arrangements with Cloud Computing service providers/sellers that enable dynamic usage provisioning.
Traditional licensing associated with ownership, number of users, support, and maintenance costs and services are being challenged by the pay-as-you-go model found in on-demand Cloud Computing.
Cloud Computing is more than restructuring software and hardware and support licenses into a kind of periodic rented or lease license. It is targeting the end usage of the services at the point of real business need of the number and scope of users of the IT service (see the discussion on the Importance of a Business Perspective of the Cloud).
With either fixed usages volumes or variable functional usage, new innovative consumption models enabled by Cloud Computing allow businesses to consider using IT in a flexible and agile way.
This can range from the “freemium”, “contractless” service that you use or pay by credit card and advertizing revenues or specific pre-allocated bands of services and software functionality over a defined period of use (see the discussion on the Financial Value Perspective of Moving from CAPEX to OPEX and Pay-as-you-go).


Elastic Provisioning
Cloud Computing can change the ownership process from buyer to seller in the sense that IT becomes a commodity purchase, and buyers focus on outcome-based performance and choices.
The impact of dynamic provisioning on the Cloud Computing ROI business case is that the façade of service management becomes more “digital”. With the emergence of Internet services, it is now commonplace in residential markets for services and products to be viewed and provisioned online.
This expectation is now translated via Cloud Computing into the business world, where online service catalogs, self service, and automated services are increasingly part of the consumer experience.

Risk and Compliance Improvement

The green sustainability issue is equally valid in Cloud Computing and seen by a number of industry observers as an argument that moving into a Cloud environment will help organizations improve their carbon footprint.
This perhaps shifts the problem to the Cloud service providers as their industry potentially becomes a huge sink for electrical emissions from massive scale computing data center services.
It is expected that this challenge will be met as technological and design improvements address the energy consumption growth patterns. The benefit to the economic and emission footprint from the use of shared services is expected to have an improved impact compared to leverage of existing assets.


Green Costs of Cloud – Sustainability
A secondary effect, however, is the growth of more Cloud service users as the Cloud Computing paradigm takes off. As the cost and emission footprint per Cloud service falls, more services per cost can be consumed. As more advanced services emerge (the term “multiplicity” – see [8] – is now starting to emerge as large workload and cost-sensitive processing is moved to the Cloud and multiple processing made possible) then so can “usage creep” occur as the consumption rates further increase the usage of Cloud Computing services.
The alignment of compliance is a wider issue that includes green and legislative issues facing organizations and specific industry sector policies.
The impact on the ROI business case from using Cloud Computing services is directly relevant to sovereignty, security, and management of services risk containment. Cloud Computing domains cut across these complex issues and are directly affected by decision processes to adoption off-premise services.

 

Discussion: Financial Value Perspective of Moving from CAPEX to OPEX and Pay-as-you-go

Software as a Service (SaaS), utility computing, and Cloud Computing are recent themes in IT that seek to change the provisioning and utilization of IT.
Key to this is the change in cash flow and cost of capital investment.

Cash Flow

Moving to a pay-as-you-go model means the cashflow of your business is changing. Sources of revenue and outgoing cash expenditure are on a usage basis based on a unit such as time, volume, or component. Cash flow – Cash Flow after Taxes – is a financial measure of a business ability to generate cash flow through its operations. Moving from a CAPEX to an OPEX model develops the use of operational expenses rather than capital assets and the treatment of operating statements rather than balance sheet management. Cash flow describes revenue, cash, and working capital changes that flow within part of the operating expenses liquidity and available usage of funds. Adopting the Cloud Computing paradigm seeks to make more money (increase revenues) while driving capital costs down through greater efficiencies of working capital and OPEX changes. Calculations of Net Present Value (NPV) of investments often need to consider the discounted cash flows of the cost of capital (WACC) to assess the value of the investment return. Cloud Computing seeks to minimize or zero upfront investment and to drive improved asset usage ratios, Average Revenue Per Unit, Average Margin Per User, and cost of asset recovery.

Cost of Capital

Moving from CAPEX to OPEX is a change in the basis of capital investment usage as upfront and ongoing costs are changed by the Cloud Computing business model. The focus is on the ability to maximize the leverage of that capital to acquire IT and business services while minimizing the risk to the business in capital used for initial investment and ongoing maintenance charges. While moving away from investments in long-term assets may be seen as context of Cloud Computing, this implies a move towards long-term OPEX-style service where QoS and costs are still equally relevant regardless of asset ownership. The common factor is the business performance and SLA requirements.
A company with a high cost of capital (WACC) and which would benefit from bringing in their tax shield (high CFAT), is a candidate for shifting CAPEX to OPEX – but other aspects of the business context may contradict that candidacy such as availability of appropriate solutions and security constraints on using shared services. If CAPEX to OPEX is desired, then the company should be considering and evaluating outsourcing solutions, including public Cloud solutions, hybrid Cloud, and Private Cloud solutions.
Cash flow can be an important indicator if CAPEX to OPEX is the focus. Pay-as-you-go can be seen as easier on cash flow than pay-upfront. But both cash flow considerations may not necessarily exist in the same business scenario. For example, a business may want to improve cash flow through moving to a direct usage model but still retain investment in CAPEX for differentiated private business processes.

OPEX

Using an OPEX model can potentially remove and release capital that would otherwise be used for initial investment and ownership of IT assets. Alternatively, investment in a Cloud Computing platform may require capital investment and changes to the payment and funding of the service as it is amortized over a wider shared service model for economies of scale.
The cost of capital from sources of equity and cost of debt point of view can change for private and public/federal industries that have stock market/shareholders or government sources of funding.
If the overall goal is to maximize the use of capital by best use of the debt and equity funds, in Cloud Computing the use of OPEX moves the funding towards optimizing capital investment leverage and risk management of those sources of funds.

Pay-as-you-go and Pay-by-the-drink

There are other ways of getting the equivalent to pay-as-you-go besides outsourcing/public Cloud. Financing and leasing are both forms of pay-as-you-go, as is a monthly software “rental fee” – or any other form of software licensing which shifts payments into the future. A close cousin to “pay-as-you-go” is “pay-by-the-drink” – usage-based billing. This type of billing can be construed to help with cash flow, but arguably, usage-based billing is only beneficial (to the subscriber) if bill amounts are predictable and controllable. If not, then neither the subscriber or the provider can budget effectively, and consequently the subscriber pays a premium for bursting capacity, and/or the provider (and thus the subscriber) oversubscribes the resources and runs the risk of a capacity shortage (“brownout”).
If the billing basis is not tied to business activity or business outcome metrics, then most commercial utility service buyers typically opt for a monthly or annual baseline fixed rate. In other words, of the billing is tied to metrics which the business can predict and control (business metrics) then the preference is for usage-based billing: but if the billing is based on IT infrastructure and/or application metrics which the business cannot readily correlate to the business activity enabled, then fixed rate billing is preferred. Likewise, residential buyers of utility services such as cell phone service are being offered fixed rate monthly billing to ease budgeting.
The following section examines some of the metrics and performance indicators that drive business towards the Cloud Computing value model.

Cloud Computing Key Performance Indicators and Metrics

Cloud Computing introduces an expanded context for service-oriented business and IT.
Developing ROI models that show how Cloud Computing adoption can benefit both business and IT consumers and providers involves examining the key technology features and business operating model changes.
This section gives an overview of ROI models to support Cloud Computing assessments and business cases in two aspects:
  • Key Performance Indicator ratios that target Cloud Computing adoption, comparing specific metrics of traditional IT with Cloud Computing solutions. These have been classified as cost, time, quality, and profitability indicators relating to Cloud Computing characteristics.
  • Key Return on Investment savings models that demonstrate cost, time, quality, compliance, revenue, and profitability improvement by comparing traditional IT with Cloud Computing solutions.
The overview of Cloud Computing ROI models considers both indicators and ROI viewpoints.
Figure 12 shows an overview of Cloud Computing ROI models and KPIs.

Cloud Computing ROI Models and KPIs

Cloud ROI Cost Indicator Ratios

Figure 13 shows the cost indicator ratios, and outline explanations are given below.

 
Cloud Computing ROI Models – Cost Indicator Ratios
Availability versus recovery SLA:
  • Indicator of availability performance compared to current service levels
Workload – predictable costs:
  • Indicator of CAPEX cost on-premise ownership versus Cloud
Workload – variable costs:
  • Indicator of OPEX cost for on-premise ownership versus Cloud; indicator of burst cost
CAPEX versus OPEX costs:
  • Indicator of on-premise physical asset TCO versus Cloud TCO
Workload versus utilization %:
  • Indicator of cost-effective Cloud workload utilization
Workload type allocations:
  • Workload size versus memory/processor distribution; indicator of % IT asset workloads using Cloud
Instance to asset ratio:
  • Indicator of % and cost of rationalization/consolidation of IT assets; degree of complexity reduction
Ecosystem – optionality:
  • Indicator of number of commodity assets, APIs, catalog items, self service

Cloud ROI Time Indicator Ratios

Figure 14 shows the time indicator ratios, and outline explanations are given below.

Cloud Computing ROI Models – Time Indicator Ratios
Timeliness:
  • The degree of service responsiveness
  • An indicator of the type of service choice determination
Throughput:
  • The latency of transactions
  • The volume per unit of time throughput
  • An indicator of the workload efficiency
Periodicity:
  • The frequency of demand and supply activity
  • The amplitude of the demand and supply activity
Temporal:
  • The event frequency to real-time action and outcome result

Cloud ROI Quality Indicator Ratios

Figure 15 shows the quality indicator ratios, and outline explanations are given below.

Cloud ROI Quality Indicator Ratios
Experiential:
  • The quality of perceived user experience
  • The quality of User Interface (UI) design and interaction – ease-of-use
SLAresponse error rate:
  • Frequency of defective responses
Intelligent automation:
  • The level of automation response (agent)

Cloud ROI Profitability Indicator Ratios

Figure 16 shows the profitability indicator ratios, and outline explanations are given below.

Cloud ROI Profitability Indicator Ratios
Revenue efficiencies:
  • Ability to generate margin increase/budget efficiency per margin
  • Rate of annuity revenue
Market disruption rate:
  • Rate of revenue growth
  • Rate of new market acquisition

Cloud ROI Savings Models

Figure 17 shows the savings models, and outline explanations are given below.
 
Cloud Computing ROI Savings Models
Speed of time reduction:
  • Compression of time reduction by Cloud adoption
  • Rate of change of TCO reduction by Cloud adoption
Optimizing time to deliver/execution:
  • Increase in provisioning speed
  • Speed of multi-sourcing
Speed of cost reduction:
  • Compression of cost reduction by Cloud adoption
  • Rate of change of TCO reduction by Cloud adoption
Optimizing cost of capacity:
  • Aligning cost with usage, CAPEX to OPEX utilization pay-as-you-go savings from Cloud adoption
  • Elastic scaling cost improvements
Optimizing ownership use:
  • Portfolio TCO , license cost reduction from Cloud adoption
  • Open Source adoption
  • SOA re-use adoption
Green costs of Cloud:
  • Green sustainability
Optimizing time to deliver/execution:
  • Increase in provisioning speed
  • Reduced supplycchain costs
  • Speed of multi-sourcing
  • Flexibility/choice
Optimizing margin:
  • Increase in revenue/profit margin from Cloud adoption

Discussion: The Importance of a Business Perspective of the Cloud

From a business perspective, the way an organization operates differentiating business processes and their Quality of Service (QoS) is key to business operating success. Identifying competitive business processes as well as standard commodity operations will improve the focus of innovative market growth and cost of service optimization activities made possible by business models based on Cloud Computing opportunities.
Just focusing on infrastructure improvements may result in cost rationalization but may miss the impact and value of applications and business processes to the end customer. QoS is an essential ingredient in evaluating the business effectiveness. The elements of QoS are made up of infrastructure, resources, activities, and services spanning the whole lifecycle of business.

Amortization of Economies of Scale

In Cloud Computing the operating challenges experienced from one customer can be proactively fixed for all the other customers of the Cloud service by using a shared platform. Amortization of problems is just one example of how a Cloud solution can achieve more favorable QoS levels. So, value can be leveraged from amortizing economic economies of scale across the collective membership potential of a service ecosystem created by the Cloud.

Business Portfolio Focus

Just looking at Cloud Computing from a technical infrastructure point of view is potentially missing the wider picture of the impact of technology on the business.
Overall, what matters is defining the value to business. Value can be defined in many ways. It does not just mean the financial values of Total Cost of Ownership (TCO) and Return on Investment (ROI), but can also mean customer value, seller provider value, broker value, market brand value, corporate value, as well as technical value of the investment.
Your business is a portfolio of business processes. Using portfolio management techniques, group your business processes into three domains where the processes in each domain have common IT enablement solution selection criteria (for example, differentiating based on IT, differentiating not based on IT, and not differentiating), and apply the solution selection criteria.
The business perspective also includes consideration of whether using Cloud services can help facilitate interactions with business partners or partner organizations – for example, by using SOA or EDI through the Cloud – and whether using Cloud services may endanger any existing interactions, where suppliers of data impose particular conditions for handling confidential data.
The work of the Cloud Business Artifacts (CBA) Project in The Open Group Cloud Computing Work Group is seeking to identify the key Cloud buyer questions and in a language business can understand and use to target solutions to meet real business requirements.

Conclusion

Cloud Computing is an important stage in the development of IT systems, comparable with the emergence of the mainframe, the minicomputer, the microprocessor, and the Internet.
Cloud Computing can provide many advantages over conventional approaches to IT provisioning, which can translate into significant improvements in ROI. But what makes it particularly exciting is that its potential effect on business is not just incremental improvement, but disruptive transformation through new operating models.
This White Paper provides an analysis of how to build and measure ROI that will help businesses to reap the benefits of Cloud Computing, and take advantage of its potential for incremental improvement and disruptive transformation of business processes.
Our understanding of Cloud Computing is currently at an early stage. This is an initial analysis. ROI models will evolve as the technology matures. This evolution will be reflected, and key indicator ratios will be described in more detail, in future deliverables of The Open Group Cloud Computing Work Group and its Cloud Business Artifacts and Cloud Business Use-Cases projects.

Labels: , , , , , ,

Thursday, May 20, 2010

When to Bring eDiscovery In-house or Move to Cloud Computing?

Over the past 18 months, the “big buzz” in the eDiscovery market has been that the Global 2000 are bringing eDiscovery in house.  The value proposition is that outsourcing to legal vendors, service providers and/or outside counsel is just too expensive.  And, ultimately, ensuring that it (eDiscovery) is done properly (i.e. in a legally defensible manner, etc.) is the responsibility of the General Counsel (GC) and other C level executives (i.e. CIO, CFO, CEO) anyway.

However, what about members of the Global 2000 that don’t have enough litigation to justify bringing it (eDiscovery) in-house?  And, what about all the rest of the enterprises outside of the Global 2000 (the majority of the enterprises worldwide) that never have to deal with the issues and costs of eDiscovery?

First of all, with the inevitable convergence of eDiscovery and Governance, Risk and Compliance (See “ Government Intervention and Oversight Driving the Convergence of eDiscovery and Governance, Risk and Compliance (GRC)”) all enterprises worldwide are going to have to deal with the issues of information management and reporting as it relates to eDiscovery and GRC.
So, just because you never have to respond to requests to produce information due to litigation, it is negligent to not  be prepared (e.g. have a data retention policy and a plan for eDiscovery).  And now, with the accelerating increase in government intervention and oversight, Governance, Risk and Compliance (GRC) are almost certainly going to affect all enterprises worldwide in some way shape or form.

So, given the cost of “outsourcing” along with this inevitable new playing field, when do you bring eDiscovery in-house? Or, given the convergence of eDiscovery and GRC, maybe a better question is when should the enterprise have an  in-house plan and associated capabilities to support whatever legal and/or GRC requests come along?

The answer for some may be to keep outsourcing.  However, those that do had better brush up on their responsibilities (outsourcing or not) as the courts are no longer accepting ignorance as a defense.

Another alternative maybe to investigate Cloud Computing as a cost effective way of moving eDiscovery / GRC in-house without having to invest in all of the IT infrastructure.  Obviously, archiving data (e.g. email) “in the Cloud” has matured to the point where it is almost a “no brainer” for most enterprises to consider (no hate mail from the anti Cloud contingency please as your security concerns are beginning to wear a bit thin).  And, most if not all of the applications that it takes to support eDiscovery / GRC processing such as Early Case Assessment (ECA), Legal Holds, Search and Analytics and Document Review are now available as Infrastructure-as-a-Service (IaaS), Platform-as-a-Service (PaaS) and Software-as-a-Service (SaaS).  Moving large amounts of data around and maintaining an appropriate chain-of-custody are still issues that need to be watched closely in this environment.

Enterprises worldwide, whether in or outside the Global 2000, are having to deal with the paradigm shift and subsequent realities of eDiscovery and Governance, Risk and Compliance (GRC) reporting.   And, the pressures of the worldwide financial crisis along with the accelerating increase in government intervention and oversight have unfortunately added another layer of complexity.

However, Cloud Computing is now a very viable alternative when considering moving your eDiscovery and GRC operations in-house.

Labels: , , , , , , , , , ,

Tuesday, May 11, 2010

eDiscovery Data Mapping Should be a Top Priority for General Counsel and CIOs within the Global 2000

A key aspect and legal requirement of eDiscovery is the creation of a data map to determine precisely what information is available within an organization and where it resides. This is a process that should begin long before a company ever finds itself in court. Surprisingly, over the past 2 years, I haven’t found more than a hand full of General Counsel and CIOs at some of the largest companies in the world that truly understand the importance of eDiscovery Data Mapping, the risks involved in not having a working eDiscovery Data Map nor the fact that they should be leading the charge to develop and manage an enterprise wide eDiscovery Data Map.

Ganesh Vednere, a manager at Capgemini, wrote an excellent overview of the key aspects of eDiscovery Data Mapping titled, “The Quest for eDiscovery: Creating a Data Map”. that appeared on the November / December 2009 Informatics site. In this overview, Ganesh indicated that, at a minimum, the enterprise should consider completing the following tasks as a general practice to start the eDiscovery Data Mapping process:

Get a list of all systems – and be prepared for a few surprises
Begin the process by creating a list of all systems that exist in the company. This is easier said than done, as in many cases, IT does not even have a full list of all systems. Sure, they usually have a list of systems, but don’t take that as the final list! Due diligence involves talking to business process owners, employees, and contractors, which often brings to light hidden systems, utilities, and home-grown applications that were unbeknownst to IT. Ensure that all types of systems are covered, e.g. physical servers, virtual servers, networks, externally hosted systems, backups (including tapes), archival systems, and desktops, etc. Pay special attention to emails, instant messaging, core business systems, collaboration software, and file shares, etc.

Document system information
After the list of all systems is known, gather as much information about each as possible. This exercise can be performed with the help of system infrastructure teams, application support teams, development teams, and business teams. Here are some types of information that can be gathered: system name, description, owner, platform type, location; is it a home grown-package, and does it store both structured and unstructured data; system dependencies (i.e., what systems are dependent on it and what systems does it depend on); business processes supported, business criticality of the system, security and access controls, format of data stored, format of data produced, reporting capabilities, how/ where the system is hosted; backup process and schedule, archival process and schedule, whether data is purged or not; if purged, how often and what data gets purged; how many users, is there external access allowed (outside of the company firewall), are retention policies applied, what are the audit-trail capabilities, what is the nature of data stored, e.g. confidential data, nonpublic personal information, or still others.

Get a list of business processes
Inventory the list of business processes and map it to the system list obtained in the step above to ensure that all the various types of ESI are documented. The list of business processes is also useful during the discovery process, when one can leverage the list to hone in on a particular type of ESI and obtain information about how it was generated, who owned the data, how the data was processed, how it was stored, and so on. A list of business processes can also be useful when assessing information flows.

Develop a list of roles, groups, and users (custodians)
Obtain the organizational chart and determine the roles and groups across the business and the business processes. Document the process custodians and map out who had privileges to do what. Understand the human actors in the information lifecycle flow.

Document the information flow across the entire organization
Determine where critical pieces of information got initiated, how the information was/is manipulated, what systems touch the information, who processes the information, what systems depend on the information, and so on. Understanding the flow of information is key to the data mapping/discovery process.

Determine how email is stored, processed, and consumed
Given the large percentage of business information and business records that reside in email, special attention needs to be placed on email ESI. Typically email is the first thing that opposing counsel go after, so determining whether email retention and disposition policies are consistently enforced will be key to proving good faith. There are a number of automated tools that will enable you to create email maps, link threads of conversation, heuristically perform relevancy search, extract underlying metadata, and so on. Before deciding to buy the best-of-breed solution, however, perform due diligence on existing email processes. Understand how employees are using email. Are they creating local archives (.PST files), are they storing emails on a network or a repository, are they disposing of them at the end of retention periods, are they using personal emails to conduct official business, and so on. Identify deficiencies and violations in email policies before the opposing counsel does.

Identify use of collaboration tools
SharePoint will have the lion’s share of the collaboration space in many organizations, but even then you must ensure that all other tools – whether they are social networking tools, Web-based tools, or home-grown tools – are included in the data-mapping process. You need to carefully document the types of information being stored on each of these tools. Sometimes company information has a nasty habit of being found in the most unlikely of places. Wherever possible work with compliance, information management, or records management groups to establish usage policies to prevent runaway viral growth of these tools. If the organization already has thousands of unmanaged SharePoint sites, work with IT and business to institute governance controls to prevent further runaway growth.

Don’t forget offsite storage
After inventorying and mapping all systems, one would think the job is done. Alas, there is more work ahead. Offsite storage is an often under-appreciated aspect of the discovery process. It is quite reasonable to assume that there might be substantial evidence stored offsite which might become incriminating at a later date. Offsite storage may contain boxes or tapes full of records whose existence was somehow never properly documented, with the result that they cannot be located unless someone opens the box or attempts to recover the tape data. These records continue to live well past their onsite cousins. This means the organization continues to have the record in backup tapes (or paper) and other formats that it purportedly claimed to have destroyed. The search for records in offsite storage is made more complicated if the offsite storage process did not create detailed indices about the contents. If there are tapes labeled “2007 Backup Y: Drive,” then it may become quite an arduous task to determine what information is really contained in those tapes. Nevertheless the journey must be started. It could involve anything from a full-scale review of all tapes, followed by reclassifying and re-filing the tapes, to perhaps a review of just the offsite storage manifests. It could also involve a search for critical information or a clean-up of the last three years’ worth of tapes, and so on.

eDiscovery Data Mapping Platform
This is an impressive beginning best practice. However, I would add the requirement of seriously considering investing in an eDiscovery Data Mapping platform to help guide you through the eDiscovery Data Mapping process and then manage all of the information that you discovery. After all, just creating your eDiscovery Data Map is just the beginning of the process. The real value of creating the eDiscovery Data Map will be seen when your enterprise uses the eDiscovery Data Map to support your first legal matter and enables you to more fully meet the legal requirements of the the Federal Rules of Civil Procedure (FRCP) and the associated state and local rules.

Genome from Exterro, is an excellent example of a new generation of data mapping solutions that are dynamic, shifting to reflect your company's information universe as it evolves.Through intelligent workflows and automated processes, Genome enables IT departments and legal teams to quickly visualize and analyze data source information to proactively scope case parameters. Over the next couple of weeks, I will be reviewing the solutions that are currently available along with some recommendations for which platforms will provide the best return on your investment (ROI).

The full text of Ganesh Vednere’s overview is as follows:

A key aspect of ediscovery is the creation of a data map to determine precisely what information is available within an organization and where it resides. This is a process that should begin long before a company ever finds itself in court.

The phone rings. It is the general counsel. The organization may be sued over patent infringement. Counsel knows that this could be “The Big One.” All sorts of data, documents, metadata, emails, and other forms of information may be required. Counsel asks IT: Do you have, or can you get together, a list of all systems and the data they contain?” There is a long, silent pause on the phone. Then the IT manager says “Well, we do have a list of systems. Let me send it your way.” Counsel gets the list. It is nothing close to the data map it needs. Instead it is a list of servers, their IP addresses, platform configuration, and their physical rack location in the data center. Good information for disaster recovery purposes, but not particularly helpful in court.

“Well, this is the best I’ve got,” comes the retort from IT. “We do not have a data map nor would we know how to create one – and, by the way, do you really think we have the bandwidth to work on this now?”

Why You? The Challenge of Data Mapping
So who gets stuck with the job? You do. You might argue that “IT manages all the infrastructure and stuff, why couldn’t they just run an inventory on their systems?” And IT will reply that “Well, we do manage the infrastructure, but we know very little about the inputs, outputs, documents, records, and other information on these applications. Go talk to the business side.” And you go to business, and business will tell you that “I just use the system and click these buttons on the screen. The system is a black box to me. I have no idea about all of the underlying data, metadata, and data structures. I suggest you talk to the operational folks.” And you talk to the operational folks, and they say, “What are you talking about? We just execute business processes. Don’t ask us about data and metadata. Go talk to the analyst who worked on the system design.” And you look for the analyst, and you eventually learn that …“Oh, she was a consultant and she left the project three years ago.”
The challenges are many but the data map must be created. And the job is yours. So where do you start? By going back to the beginning … the very beginning.

How Did We Get in Such a Mess?
Let’s take a look at a typical mid-size organization. It has several thousand employees and contractors with offices in the U.S. and E.U. The sheer volume of information that resides in just one division is mind-boggling. New information sources keep popping up, employees keep creating new SharePoint sites on their own, and there is use (or misuse) of social collaboration tools, to say nothing of several hundred IT systems in play at any point in time. Data is moved and migrated from one place to another without proper documentation or communication, more and more tape backups are being created, and some employees are making copies of data on thumb drives or worse, emailing them to their personal email addresses.

How did things ever become such a mess? There are manifold reasons: IT is traditionally kept at arm’s length on compliance and uninvolved with information management and governance during systems design and development. Records management departments, on the other hand, often institute sound policies and retention schedules but have a tough time putting these into practice and getting people to adhere to them. On the legal side, general counsels often work against themselves: becoming increasingly exasperated over the large amount of money spent on searching, processing, and producing electronically stored information (ESI), they often push hard to cut costs, thereby shortcircuiting the process.

Must an Organization Have a Data Map?

It may seem surprising that even today, many successful organizations do not have a data map, or at best, a superficial one. It is not that organizations are lacking in will, however, but that the process seems too daunting. Consider the “typical organization” above. If there are several hundred IT systems and other home-grown business applications, one must not only know what these systems are, and where they are located, but also the types of information (documents, records, other content) that are produced from these systems and additional information such as data format, data location, whether the data is updated by other systems, or transformed into other formats, etc.

To add to the complexity, a determination also needs to be made as to whether a piece of ESI can be extracted and presented using reasonable and customary means. For example, if an IT system was retired and the data backed-up on tape, it is reasonable to assume that extracting the tape, processing the information, and presenting it in a readable format may not be easy since the underlying version of the software no longer exists. Counsel, however, must be able to assess whether this is indeed the case. Having a data map eases some of these tasks and makes it easier for counsel to relate the information as needed.

Data Mapping Considerations

In the current economic environment, companies are bracing themselves for an uptick in the number of lawsuits. Whether the matter is related to regulators, customers, consumers, employees, or business partners, companies are often required to provide ESI in court. While this should be sine qua non for most organizations, many are simply too overwhelmed to be able to react fast enough and are thus placing themselves at a much greater risk. If that’s the case in your enterprise, here are some initial steps that will help you move forward.
  1. Understand the prevailing legal environment. Organizations are not created equally, and not all have the same set of applicable legal requirements. It is therefore important to analyze the type of environment that the organization operates in, the jurisdiction it is under, and the various federal and state laws, regulations, and common industry standards that apply to it, with regard to ESI. While the contents of a data map by itself do not directly correlate to a particular law or regulation, it is useful to know what checks and controls need to be established during the datamapping process and ensure that there are no “show-stopper” questions in court around how the data map was created or what the process was.
  2. Use a partnership model and obtain buy-in from senior management. It is important that each entity within an organization have a vested stake in the success of any data-mapping project. This means that management in each of these organizational fiefdoms must understand what a data map is, how it will be used, and what the process of creating one is. Getting buy-in from these senior managers is a crucial first step and must be completed prior to the start of the process. Additionally, it is important that people of the appropriate rank are selected to work on the project. Folks who are deep in the weeds will generally have a lot more information about data flows and how processes and people work together versus the senior executive who operates in more of a decision-making capacity.
  3. There is little point in pursuing a “big-bang” approach for the data map. Instead, work towards a phased approach. Prioritize which divisions or lines of business to focus on first and then address the remaining ones later. Work with line managers to determine what, if any, information has been collected on systems and processes within their particular areas. Standard industry lists may be employed as a starting point, e.g. HR, Accounting, Communications and Marketing, etc. Begin the first phase of the process here and then iteratively build upon what’s already available.
  4. Use the right technology. As more capital is allocated towards automating ediscovery, vendors will naturally gravitate towards building specialized software for this mission. Time, cost, and relevancy of results will drive the success of vendor products. While some organizations have attempted to build custom tools, more and more prefer choosing established products or service offerings to guide them through the rediscovery and data-mapping process. Already many vendors have begun mapping their offerings to the electronic discovery reference model (EDRM) and other industry standards. This market is still maturing and organizations should not go out and immediately purchase a top-rated vendor’s software without due consideration of the organization’s unique circumstances.
Creating the Data Map
Once you’ve worked your way through each of considerations above and taken action as needed, you’re ready to start the actual data-mapping process. It is lengthy but well-defined and can be broken down into each of the following steps:

The Data Mapping Process

  1. Get a list of all systems – and be prepared for a few surprises. Begin the process by creating a list of all systems that exist in the company. This is easier said than done, as in many cases, IT does not even have a full list of all systems. Sure, they usually have a list of systems, but don’t take that as the final list! Due diligence involves talking to business process owners, employees, and contractors, which often brings to light hidden systems, utilities, and home-grown applications that were unbeknownst to IT. Ensure that all types of systems are covered, e.g. physical servers, virtual servers, networks, externally hosted systems, backups (including tapes), archival systems, and desktops, etc. Pay special attention to emails, instant messaging, core business systems, collaboration software, and file shares, etc.
  2. Document system information. After the list of all systems is known, gather as much information about each as possible. This exercise can be performed with the help of system infrastructure teams, application support teams, development teams, and business teams. Here are some types of information that can be gathered: system name, description, owner, platform type, location; is it a home grown-package, and does it store both structured and unstructured data; system dependencies (i.e., what systems are dependent on it and what systems does it depend on); business processes supported, business criticality of the system, security and access controls, format of data stored, format of data produced, reporting capabilities, how/ where the system is hosted; backup process and schedule, archival process and schedule, whether data is purged or not; if purged, how often and what data gets purged; how many users, is there external access allowed (outside of the company firewall), are retention policies applied, what are the audit-trail capabilities, what is the nature of data stored, e.g. confidential data, nonpublic personal information, or still others.
  3. Get a list of business processes. Inventory the list of business processes and map it to the system list obtained in the step above to ensure that all the various types of ESI are documented. The list of business processes is also useful during the discovery process, when one can leverage the list to hone in on a particular type of ESI and obtain information about how it was generated, who owned the data, how the data was processed, how it was stored, and so on. A list of business processes can also be useful when assessing information flows.
  4. Develop a list of roles, groups, and users (custodians). Obtain the organizational chart and determine the roles and groups across the business and the business processes. Document the process custodians and map out who had privileges to do what. Understand the human actors in the information lifecycle flow.
  5. Document the information flow across the entire organization. Determine where critical pieces of information got initiated, how the information was/is manipulated, what systems touch the information, who processes the information, what systems depend on the information, and so on. Understanding the flow of information is key to the data mapping/discovery process.
  6. Determine how email is stored, processed, and consumed. Given the large percentage of business information and business records that reside in email, special attention needs to be placed on email ESI. Typically email is the first thing that opposing counsel go after, so determining whether email retention and disposition policies are consistently enforced will be key to proving good faith. There are a number of automated tools that will enable you to create email maps, link threads of conversation, heuristically perform relevancy search, extract underlying metadata, and so on. Before deciding to buy the best-of-breed solution, however, perform due diligence on existing email processes. Understand how employees are using email. Are they creating local archives (.PST files), are they storing emails on a network or a repository, are they disposing of them at the end of retention periods, are they using personal emails to conduct official business, and so on. Identify deficiencies and violations in email policies before the opposing counsel does.
  7. Identify use of collaboration tools. SharePoint will have the lion’s share of the collaboration space in many organizations, but even then you must ensure that all other tools – whether they are social networking tools, Web-based tools, or home-grown tools – are included in the data-mapping process. You need to carefully document the types of information being stored on each of these tools. Sometimes company information has a nasty habit of being found in the most unlikely of places. Wherever possible work with compliance, information management, or records management groups to establish usage policies to prevent runaway viral growth of these tools. If the organization already has thousands of unmanaged SharePoint sites, work with IT and business to institute governance controls to prevent further runaway growth.
  8. Don’t forget offsite storage. After inventorying and mapping all systems, one would think the job is done. Alas, there is more work ahead. Offsite storage is an often under-appreciated aspect of the discovery process. It is quite reasonable to assume that there might be substantial evidence stored offsite which might become incriminating at a later date. Offsite storage may contain boxes or tapes full of records whose existence was somehow never properly documented, with the result that they cannot be located unless someone opens the box or attempts to recover the tape data. These records continue to live well past their onsite cousins. This means the organization continues to have the record in backup tapes (or paper) and other formats that it purportedly claimed to have destroyed. The search for records in offsite storage is made more complicated if the offsite storage process did not create detailed indices about the contents. If there are tapes labeled “2007 Backup Y: Drive,” then it may become quite an arduous task to determine what information is really contained in those tapes. Nevertheless the journey must be started. It could involve anything from a full-scale review of all tapes, followed by reclassifying and re-filing the tapes, to perhaps a review of just the offsite storage manifests. It could also involve a search for critical information or a clean-up of the last three years’ worth of tapes, and so on.
Conclusion
In today’s highly litigious world, creating a data map is one of the primary steps in responding to litigation requests. It is vital that organizations get a solid foundation by focusing time, energy and resources in doing it right – and creating it long before it’s needed.

Labels: , , , , , , , , , , ,

Friday, October 9, 2009

Standards Need to Emerge for Collecting and Processing Electronically Stored Evidence (ESE)

Most litigators and their litigation support staff that have been practicing over the past 5-10 years could probably teach a class on the process of preservation, collection, processing, review and production of paper evidence. Or, at least they could stand at a whiteboard and draw a basic workflow diagram of the basic steps.

However, with the dramatic and accellerating increase in the amount of Electronically Stored Information (ESI) which I like to call Electronically Stored Evidence (ESE), the subsequent technical issues and the associated changes to the Federal Rules of Civil Procedures (FRCP), very few, if any of the same litigators and their staff, can now even describe the most basic workflow to to get ESE for a trial. Therefore, although many are talking of their importance (myself included), eDiscovery standards of any substance, are a long way off.

This is certainly not the fault of the lawyers as they have never been required to have much of true understanding of the technology of processing evidence in order to be successful litigators. However, the bar has now literally been raised and litigators can't even provide adequate representation without an indepth understanding of these new issues.

Maybe we should consider requiring a license or some type of ceritification to practice law when eDiscovery is involved? Or, has ESE become so intertwined in our matters that there isn't a case without eDiscovery and therfore every lawyer that want to litigate anything should have to be certified?

As a place to start this discussion / debate, we need to start identifying the basic components of ESE and how it is stored, how to preserve it, how to extract it (the new word for collection), how to process it, how to review it, and how to produce it.

Wouldn't it great if 5 years from now, litigators could stand at a whiteboard and diagram and explain the basic "standard" components of the workflow for processing Electronically Stored Evidence (ESE)?

Eric P. Blank addresses these issues in an excellent article titled,"The Need for E-Discovery Standards: A Call From the Trenches", posted on October 5, 2009 on the EDD Update Blog.

Eric P. Blank is the founder and managing attorney of Blank Law + Technology PS. His practice focuses on electronic discovery counseling, e-security response planning and implementation, investigations and computer forensics. Mr. Blank has conducted more than 300 investigations into computer and software-related torts and employee misconduct since 2001 and has frequently been a court-appointed special master or neutral in e-discovery matters.

The full text of Mr. Blank's post is as follows:

Most discussion about standards in electronic discovery focuses on the big-picture issues of scope, cost and cost shifting.

These are important questions eloquently argued in the courts. However, they overlook the mundane, pick-and-shovel e-discovery concerns that affect every case. I’m talking about the elementary technical issues of preservation, extraction, processing, review and production.

I’m talking about extracting data from electronic storage media, processing the data and its metadata into a document review software application platform, supporting the review and producing the data as discovery or evidence.

Outside the e-discovery world, the first stage of this process is known as Extract, Transform, Load (ETL). Identifying and overcoming the challenges of ETL have occupied computer scientists for decades. Principal obstacles to effective ETL include widely diverse and poorly documented storage repositories, asynchronous multimedia platforms, constantly evolving software, hardware and software anomalies, and human error, usually with respect to initial planning.

E-discovery vendors on the ground face those obstacles and more. Consider, as just a few of many examples, the following:

Mobile phones and PDAs: In some models, data can be extracted through forensic imaging. In others, such as many of those without SIM cards, data can only be pulled through live file extraction. Click here and here to read my earlier blog posts about the difference between forensic imaging and live file extraction. In any case, the question is this: Should data extraction scope be defined by current technical capabilities, or should there be a single common standard – such as live files only – for those instances when mobile phones and PDAs are subject to e-discovery?

A multitude of file types: Extraction and processing applications address dozens, sometimes hundreds, of file types. These file types are usually associated with, and identified by, a particular file extension, such as .doc or .xls. However, custom extensions are easy to apply – documents I create might have a .epb file extension, for example – and it is also simple to apply a nonstandard extension to a particular file type (e.g., a .doc extension to a PDF file). These are often missed, or improperly processed, by extraction and processing software.

Computer forensics software in the hands of an experienced technician can reveal documents by file type without relying on extension format and such, but doing so is costly and time consuming. What checks should be done for mislabeled or unusual file extensions? When are such checks required?

Metadata: Most of us think of metadata in basic terms such as the putative author, creation date, modification date, last-access date and so forth. However, metadata varies widely across data types. Microsoft Office documents, for example, have more than 100 metadata fields. It is also possible to create custom fields with many document types. Nearly all of these, such as the ubiquitous P-size and L-size, are nearly never important in civil litigation.

“Nearly never” is not, however, the same as “never.” Such data can be extracted, but it is not, as a rule, supported by processing software, which renders it unavailable at the attorney review level. Is it possible to agree on which metadata fields should be preserved and processed? When they should be processed? Which fields are important forensically? When all fields should be preserved?

Rapid technological change: Software is updated all the time. This affects how metadata is produced and the appearance of electronic documents. Processing software hasn’t kept up. It’s also inconsistent. For example, the last-access date on a Word 2007 document running in Windows Vista is affected differently than an Office XP Word document running on Vista. Both documents, however, are processed the same, as if the metadata means the same, when it does not. How should inconsistencies like this be addressed? What should the typical approach be?

Webmail: Screenshots of Web-based email services such as Hotmail are a common and inexpensive workaround to downloading actual Hotmail files. Which method is preferred? Is either method not preferred? As third-party cloud data repositories multiply, what constitutes best practices with regard to extraction methods will become a critical question.

Capture rates: What percentage capture rate is acceptable for processing software? Many files are often not processed by even the best technology, and must be laboriously hand processed. In a million-item processing job, a 1 percent miss rate equals 10,000 documents not processed and available for review. Is 99 percent acceptable? Is 98 percent? Note: If you think that the processing rate for your document review software is 100 percent, you’re kidding yourself.

Searching: Keyword searching, including keyword searches supported by “fuzzy” search techniques, are giving way to conceptual searching, which is the future of document search and review. Conceptual searching, however, involves proprietary algorithms and processes with a wide range of accuracy. What standards must conceptual searching meet to be accepted? How are these standards applied? When, if ever, is conceptual searching disallowed?

File format: In e-discovery today, most documents are produced in .Tiff format. Putting aside the larger question of whether .Tiff should be the standard for producing electronic documents, what about documents such as spreadsheets that don’t translate well into .Tiff files? In what format should presentation-type documents be produced? As slide shows? As workbook copies with notes and presenters’ comments? How are native files to be tracked and authenticated as a best practice?

Today, e-discovery consultants decide many of these questions on their own or after consulting with litigation counsel. In essence, a consultant decides when it is and isn’t practical to extract files from a system, whether to image a particular hard drive and whether to put aside as unreadable a back-up tape from a set of tapes that must be searched.

Much of the time, the consultant makes the “right” decision, as subsequently decided by the court, the client or the opposing party. It’s a rare consultant, however, who won’t admit that adopting e-discovery standards would bring enormous benefit to the practical challenges of data extraction, processing and production.

I'll be discussing these and other issues in the future. Any of the problems mentioned above could be an entire article. I look forward to working with the legal and technical community to address these “technical” standards – as opposed to the widely discussed “strategic” standards which may ultimately be addressed by changes in the Federal Rules of Civil Procedure.

Labels: , , , , , ,