This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Wednesday, November 3, 2010

eDiscovery from Database Management Systems

I have spent the majority of my career building software and services companies that develop and sell enterprise class applications running on large SQL databases for the Fortune 2000. Therefore, when I read Jason Kruse’s article on the Law Technology News Blog titled, “Database Discovery Is Dubious, but Unavoidable”, I had to smile and agree.

Jason basically states that if the legal community thought that harvesting electronic evidence from email systems was difficult, they haven’t seen anything yet as extracting information from databases (where most enterprise electronic information is stored) is going to prove to be much more challenging. Further, there is going to have to be much more cooperation between the legal department and the Information Technology (IT) department.

As an enterprise application development expert, I have spent many hours trying to figure out how to get data into databases. As an eDiscovery technologist, I am looking forward to the challenge of getting it back out again. The next few years are going to be fun.

The full text of Jason’s article is as follows:

Structuring data into databases has long been a solution to store complex data that can be retrieved and reported in variable ways. That data solution, however, has a legal problem in the e-discovery context.

It took years for many litigators and judges to become comfortable with discovery of e-mail and other electronic records. But as more forms of electronic records enter into discovery disputes, lawyers are back on unfamiliar ground. "There are still types of evidence that lawyers prefer to ignore and hope will go away, the way e-mail discovery was ten years ago," says Rob Brunner, who leads the Financial and Enterprise Data Analytics practice at FTI Consulting. "And I hate to say it, but e-mail was an easy problem compared to what's next."

Brunner is speaking specifically of structured data, especially electronic evidence from databases. However, structured data includes a broad swath of content types, including common sources most might consider a document, such as e-mail. Any time a software system, whether a large, enterprise database or an e-mail server, pulls information from a number of different files and merges them into a single view, it is functioning like a database. Unstructured data is commonly defined as data that is not stored in a database or in a semantically tagged document.

Structured data is discoverable in litigation, but lawyers are finding that there is little guidance for handling it. "This is an issue that's only going to demand more attention, because we're swimming in this kind of information," says Brunner. "In trying to explain why discovery of this information is important, I like to point out that 500 out of 500 Fortune 500 companies have structured data. It's something you won't be able to avoid in many cases."

With structured data, information is in the form of separate files that are linked so that information can be pulled from different sources, analyzed, and compiled. For example, a typical corporate human resources system contains information about employees that can be viewed as individual employee records or compiled into statistics about the entire work force.

Like the mass of ice below an iceberg's waterline, the amount of structured data is often the bulk of corporate data, but is rarely seen. According to the Data Warehousing Institute, a technology research firm, approximately 47 percent of corporate data is structured in nature, compared to 31 percent of unstructured data. (The remaining 22 percent was described as semi-structured data.)

IGNORED NO LONGER
The Sedona Conference, a nonprofit legal think tank largely concerned with preservation and production of electronically stored information in civil litigation, has announced that it will publish a commentary on the discovery of information from databases after December 2010.

This will be one of the first such efforts to provide guidance for the discovery of structured data. "The commentary focuses on what is the basis of relevance in a database," says Conrad Jacoby, the founder of efficientEDD.com and editor of the forthcoming document. "Structured data is so difficult to define that we have to start by building the most basic groundwork for discovery of this information."

Unfortunately, the drafting committee found the issue of discovery of structured data so problematic that it had to scale back its ambitions. Initially, the organizers had hoped to address structured data in many forms, including the emerging problem of structured data that is accessed over the web. But in the end, the commentary only addresses database evidence, and ignored all other structured sources. "This is a complicated conversation to have and sometimes the sides talked past one another," says Jacoby. "It's a matter of vocabulary. You can have really excellent lawyers and technical people, but their vocabulary and logic are not the same. The same words can have different meanings and you can go around and around and around."

For example, Jacoby says defining a word as simple as "search" created a headache for his group. In many cases, a database only logs the first words of the text field, meaning a lot of data is not easily retrievable. "You might look at a database and assume that a database query would search all records in the database," he says. "But it turns out that some information is not indexed or searchable. So then what are you searching? Is it even possible to get all relevant information out of a database cost-effectively?"

Structured data is often important for litigation, especially for establishing damages and issues of liability. Unfortunately, it is often ephemeral and endlessly changing. Jocoby points out that automated transaction logs for many businesses are continually deleted and overwritten from point of sale systems. A cash register often keeps a record until cash out, and then the record is uploaded to a regional, then a national database, then it is often overwritten when a credit card transaction clears. "Even finding a record is hard," he says. "A record could be in multiple places or none of the expected places."

A CREDIBILITY PROBLEM
The Sedona Conference and many court systems still struggle with very basic questions, such as how to identify discoverable structured data for litigation. Databases are different in almost every organization. Even the common systems are typically customized for each customer. But even more problematic, many databases are purpose-built and understood in depth by only a few people. "The nastiest issues arise with proprietary systems," says Craig Carpenter, vice president of marketing with Recommind. "You tend find them in large multinationals and the odds are that only 15 people on the planet know how to access some of these systems."

Discovery of databases can become expensive, but for different reasons than the discovery of e-mail and other records. In e-mail, much of the cost is in human review to protect privilege when producing a collection of records. With databases, cost overruns are more likely to arise when you are trying to get information out of a system. "In many cases, pulling a single record is impossible without preserving the larger data set," says Jacoby. "That's when people say, 'fine, then give me the whole database,' and the producing party resists, and now you have a fight on your hands."

Though database files are discoverable in electronic form, courts have been reluctant to grant plaintiffs broad access to them for litigation. Courts struggle with information that is not contained in discrete documents, unlike the static artifact traditionally considered to be a document. There is little case law regarding authenticating structured data, but there are procedures that can be used to try to verify information pulled from such sources is complete.

For example, experts recommend that the most basic step to take in database discovery is to review the regular checks a database makes of the records it produces, which make sure that the results of a database search query and the production information match. Most databases are designed with complex reporting and data mining tools and lawyers can take advantage of these functions to obtain detailed reports of information being produced. "There is no industry checklist, but you can do some verification to make sure data is at least not corrupted," says Brunner.

Because information stored in a database is constantly changing, both the producing and requesting parties can manipulate data and present it in any light they choose. "Just because it comes out of a database doesn't mean it is accurate," says Jacoby. "I think courts make that mistake, and it's important to make sure that a judge understands a database record can be wrong."

Database evidence is often incomplete and misleading, and validating the integrity of data does not validate its meaning. Unlike written records, which can be read and interpreted based on their literal meanings, database records are often stripped of context when produced for litigation. "A pharmaceutical company may have 10,000 adverse records in a database for a particular drug, but how many of those are legitimate complaints?" Jacoby says. "You can slice and dice data all kinds of ways that may be statistically valid but misleading."

To head off this problem, lawyers need to be conversant in the language of structured data and be able to explain issue to the court so that records are presented accurately. Experts say that the problems of database discovery are so complicated and technical that sometimes the only way to communicate the issues is to find simple analogies to technology that laymen understand. "I testified a year and a half ago in a $250 million suit, and I struggled to explain to the judge how things work," says Brunner. "I described the system as a big calculator, as in 'you put the data into the computer and add it up to create a record.' It was an oversimplification, but it worked."

Unfortunately, e-discovery vendors have been slow to respond to this issue. Brunner specializes in this kind of discovery, but he says the industry has yet to provide a reliable, reusable road map for structured data in e-discovery as it has for other data types. "Services have developed to address the low hanging fruit like e-mail and other file types that are relatively easy to build a solution for," says Brunner. "We're just beginning to tackle the question of how you create a repeatable process for structured data."

Labels: , , ,

Thursday, August 28, 2008

Oracle vs. SQLServer Argument Hits eDiscovery Market

Having spent the last 20 years in the enterprise class software market I am very familiar with the age old argument about whether Oracle's DBMS or Microsoft's SQLServer is better. The Oracle bigots will tell you that SQLServer is slow, doesn't scale and just wasn't architected for true enterprise class use. SQLServer Biggots will tell you that Oracle is way too expensive. I will get in to some of the specifics later in this post. However, first I wanted to point out that the Oracle vs. SQLServer argument has now hit the eDiscovery market with CaseCentral's announcement that they have Migrated their On-Demand eDiscovery Platform to Oracle® Real Application Clusters.

Given the fact that Tom Thimot, CaseCentral President and CEO served as the vice president of central U.S. sales at Oracle, where his team grew license revenues from $50 million to more than $250 million in a two-year period and was also a key leader in the worldwide Oracle applications vertical organization that grew revenues by more than $100 million, it is no big surprise that he is an Oracle Biggot.

The Full Text of the Press Release is as follows:
REDWOOD SHORES, Calif., Aug. 20 /PRNewswire-FirstCall/ --

-- CaseCentral, the leading secure SaaS platform for corporations looking to take control of eDiscovery, has migrated its Software-as-a-Service (SaaS) platform to Oracle Database and Oracle Real Application Clusters to deliver better performance, scalability and availability, Oracle announced today.
-- CaseCentral's platform allows corporations to apply disciplined business process to litigation and regulatory matters, reducing risk and business disruption, boosting productivity, and controlling costs. CaseCentral has delivered its proven, SaaS platform to over 1,100 customers and more than 7,250 registered users.
-- CaseCentral actively manages several hundred terabytes of evidence -- 95 percent of which is unstructured data such as emails, office documents, and images -- and has hosted over 25,000 individual litigation matters.
-- San Francisco-based CaseCentral initially deployed its clustered database environment in 2007. Its purpose-built Java-based platform is deployed on a multi-node cluster of HP BladeSystem servers running Linux.
-- CaseCentral utilizes Oracle Enterprise Manager to help provide the monitoring and management necessary to meet the mission-critical needs of their clients. CaseCentral will migrate customers currently supported by Microsoft SQL Server to the clustered Oracle Database environment.

Supporting Quote
"CaseCentral offers a highly scalable, on demand software platform that allows companies to own the eDiscovery process and their law firms to own the execution," said Ted Sergott, Chief Technology Officer, CaseCentral. "At a moment's notice, a new or existing client can send us millions of documents and terabytes of data to support an eDiscovery matter and thanks to our architecture, we are fully prepared to accommodate that volume of mission-critical data. Running on Oracle Database and Oracle Real Application Clusters, CaseCentral offers customers a scaleable, secure and available platform to meet their eDiscovery needs."

Supporting Resources
About Oracle Database: http://www.oracle.com/database
About Oracle Real Application Clusters: http://www.oracle.com/clusters
About Oracle Enterprise Manager: http://www.oracle.com/enterprise_manager/index.html
To download free, evaluation versions of Oracle software, go to: http://www.oracle.com/technology/software/index.html

About CaseCentral
CaseCentral is the leader in on-demand eDiscovery software for corporations looking to take control of eDiscovery. CaseCentral enables companies to efficiently and defensibly respond to today's legal and compliance challenges, consistently, accurately and faster, while delivering overall savings of 30-60 percent and increasing earnings per share (EPS) by up to 1.1 percent. The company's on-demand eDiscovery software platform, with its secure, multi-party architecture and configurable litigation workflow engine, includes best-practice solution templates that enable companies to be operational within hours. CaseCentral is the first to provide eDiscovery business intelligence dashboards giving customers real-time insight into review rates, quality rates and costs per document by case, firm or user. CaseCentral is used by more than 25 of the Fortune 100 and 81 of the AmLaw 100. Founded in 1994, CaseCentral is consistently chosen to handle many of the most complex and highly visible litigation projects in the nation. For more information, call 1.800.714.2727 or visit http://www.casecentral.com/.

About Oracle
Oracle (ORCL) is the world's largest enterprise software company. For more information about Oracle, please visit our Web site at http://www.oracle.com/.

Oracle vs. SQLServer
According to a recent whitepaper available on the Microsoft site title "Leaping Forward: SQL Server 2008 Compared to Oracle Database 11g", Microsoft SQL Server has steadily gained ground on other database systems and now surpasses the competition in terms of performance, scalability, security, developer productivity, business intelligence (BI), and compatibility with the 2007 Microsoft Office System. It achieves this at a considerably lower cost than does Oracle Database 11g.

So, Microsoft contends that Microsoft® SQL Server® 2008 outperforms Oracle in the areas that matter to your business. The following summarizes some of the mission-critical areas in which SQL Server 2008 excels:

Performance and Scalability: SQL Server scales to some of the world’s largest workloads, evidenced by strong industry standard benchmark results. Customers such as Unilever, Citi, Barclays Capital, and Mediterranean Shipping Company support their most mission-critical applications on SQL Server. Customers running SQL Server 2008, including large ISVs such as Siemens and RedPrairie, report excellent experiences with the latest scalability enhancements. SQL Server is recognized as Best Seller and Top Growth Best Seller by CRN Magazine.

Security: The National Vulnerability Database (NIST) reports over 330 critical security vulnerabilities in Oracle database products over the last four years. During that same period, SQL Server 2005 experienced ZERO vulnerabilities. This result comes from secure engineering processes as part of the Trustworthy Computing Initiative, comprehensive security features, and a robust Microsoft Update infrastructure. This winning combination reduces both security risks and patching downtime for customers. According to one expert, Oracle is five years behind Microsoft in patch management. Computerworld reports that two-thirds of Oracle DBAs do not apply security patches.

Developer Productivity: SQL Server works with Microsoft Visual Studio® to help provide an integrated development experience, allowing developers to work in one environment across the client, mid-tier, and data-tier. SQL Server 2008 takes a step further with new development features. In contrast, Oracle’s array of tools and SDKs, assembled via acquisition, require developers to learn and work across numerous interfaces. In fact, IDC reports that Microsoft is the number one application technology platform of choice.

Business Intelligence: SQL Server is part of the Microsoft integrated Business Intelligence platform, which spans data warehousing, analytics and reporting, score carding, planning, and budgeting. SQL Server is in the Leader’s quadrant in both Gartner’s Magic Quadrant for BI Platforms and Magic Quadrant for Data Warehousing. SQL Server 2008 introduces more innovation with new data warehousing and business intelligence features. According to Oracle’s latest price list, the company currently charges up to an additional 800% or more on top of their base database fees for similar functionalities.

Microsoft Office System Integration: SQL Server helps customers gain better business insight and make faster decisions through the product's tight integration with the familiar Microsoft Office System user interface. For example, add-ins such as Data Mining for Excel uses both SQL Server and Microsoft Office to provide insight into customer data. IDC recognizes Microsoft as the fastest growing BI tool vendor. Oracle has Microsoft Office Plug In, which includes subset of the functionalities that SQL Server provides, but charges an additional $30,000 per processor.

Total Cost of Ownership: SQL Server has a simple tiered SKU licensing model. Oracle, on the other hand, has a complex array of options and add-ins that are required to develop, deploy, and manage most large-scale applications. The SQL Server integrated development environment and easy-to-use development tools lead to improved Time to Solution and Time to Value for applications and business insight. SQL Server is highly successful in the areas of self-tuning and automated administration, resulting in a much simpler deployment and management profile than Oracle Database 11g. SQL Server is designed to work seamlessly with the rest of the Microsoft software stack, which can help provide smoother development and deployment experience and higher performance than Oracle.

Summary
As I started this post, I grew up in the enterprise software market and therefore learned to tow the Oracle line and bash SQLServer as a nice little departmental database that would never have what it takes to play in the big leagues. However, I believe that Microsoft has come a long way and now that I am "playing" in the legal market, I would recommend that any technologies sitting on the SQLServer platform will have to be considered just as strong as the technologies sitting on Oracle.

And, I hope that someone out there in the litigation market challenges this position so that we can rekindle some of the debate that was so much fun for so many years in the general enterprise markets. And, it is arguments / debates such as these that will help to usher the litigation market into the world of leading edge technology.

Labels: , , , ,