This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Friday, June 25, 2010

The Conceptual Search Game is Finally On!!

I have been writing (probably preaching to some) about advanced search technology, the differences between conceptual search and keyword search and the importance of advanced search technology in both the Early Case Assessment (ECA) and document review phases of eDiscovery for the past 3 years.

Concept Search Cash Law Emerging http://ediscoveryconsulting.blogspot.com/2008/06/concept-search-case-law-emerging.html
Concept Search vs. Keyword Search http://ediscoveryconsulting.blogspot.com/2008/12/concept-search-vs-keyword-search-in.html
Litigators Need ESI Analytics – Not Boolean Search Tools http://ediscoveryconsulting.blogspot.com/2010/05/litigators-need-esi-analytics-not.html

However, I have been somewhat disappointed in regards to the level of adoption of true conceptual search technology by the leading Litigation Technology vendors.  That appears to be changing.  As an example, in a June 17, 2010 post by StoredIQ titled, “Email Search: Nowhere to Hide”, the author provides an overview of StoredIQ’s search technology with specific focus on Natural Language Processing (NLP).  The reason that I am pointing this out is that 12 months ago, StoredIQ would not have been spending marketing dollars or Blog space on Natural Language Processing (NLP) because the market didn’t know what it was and didn’t care.

In this Blog post, StoredIQ now contends, “Probably of greatest interest to litigators during the discovery process is StoredIQ’s ability to perform natural language processing (NLP), which is the ability to extract linguistically derived natural language concepts from within email and user files including people, places and things. Legal teams can immediately search using over 250 out-of-the-box concepts and attributes including credit card accounts, social security numbers and stock symbols. NLP identifies word usage based upon context within a sentence. For example, NLP can identify if the word ‘will’ is used to identify a person’s name, a legal document or an auxiliary verb showing intent. StoredIQ has proprietary technology for adaptive sentence boundary disambiguation (ASBD) which substantially increases the precision of Natural Language Processing to address common grammatical deficiencies that are present in many business documents. No other information management technologies have this capability. NLP is a critical capability necessary to accurately perform eDiscovery, records management or risk management as full text indexes alone cannot provide the required level of precision.

Interestingly enough,  in my discussions with General Counsel and their litigation support teams from the Information Technology departments over the past 6 months, I have found a new awareness and appreciation for true conceptual search or semantic search or NLP.    So, StoredIQ is on the right track with their current product  offerings and I would bet that they have a product roadmap with more of the same.

So, I guess the conceptual search game is on and the other litigation technology vendors had better take notice of what their clients are saying in regards to what search technology they need.

The full text of the  StoredIQ Blog post is as follows:

A recent article by Jacob Goldstein, 23 Things Not To Write In An Email, illustrates the type of granularity as well as breadth of keywords that can be used by litigators during the legal discovery process to search for relevant information. He points out some keywords that may raise a legal red flag and should be used carefully when constructing emails. However, today’s technology search capabilities provide such precise, complete and accurate results, that there just isn’t anywhere to hide.
For instance, StoredIQ’s advanced search capabilities can look within compressed files, email archives and email attachments, in addition to the text contained in the email message itself. It can also search non-printable text within a document or email and can search through comments and revisions. In addition to search using keywords, StoredIQ supports many advanced search capabilities including:
  • Single term search
  • Multiple term search
  • Concept-based search
  • Boolean operators
  • Logical grouping of terms
  • Wildcards within search terms or Boolean expressions
  • Proximity searches
  • Natural language entities
  • Regular expressions
  • Macro-based searches
  • Object level attributes
  • By hash value (digital signatures)
Probably of greatest interest to litigators during the discovery process is StoredIQ’s ability to perform natural language processing (NLP), which is the ability to extract linguistically derived natural language concepts from within email and user files including people, places and things. Legal teams can immediately search using over 250 out-of-the-box concepts and attributes including credit card accounts, social security numbers and stock symbols. NLP identifies word usage based upon context within a sentence. For example, NLP can identify if the word ‘will’ is used to identify a person’s name, a legal document or an auxiliary verb showing intent. StoredIQ has proprietary technology for adaptive sentence boundary disambiguation (ASBD) which substantially increases the precision of Natural Language Processing to address common grammatical deficiencies that are present in many business documents. No other information management technologies have this capability. NLP is a critical capability necessary to accurately perform eDiscovery, records management or risk management as full text indexes alone cannot provide the required level of precision.
I know a lot of these terms can be a mouthful, but the underlying take away is that legal teams have the technology to precisely and accurately search electronic data, including email, making it much easier for litigators to discover data that was at one time hidden from them.

Labels: , , , , , , , , , ,

Monday, May 24, 2010

Litigators Need ESI Analytics – Not Boolean Search Tools

This morning as I was enjoying my Monday morning coffee and reviewing the latest “Blog postings and press releases” in eDiscovery, I came across an article highlighting a research study published by AIIM on March 29, 2010, titled, “Users Need Content Research Tools, Not Basic Search Tools” that indicated users are more interested in finding good content analysis tools as opposed to just search tool. It is interesting that this study was not specifically written about eDiscovery. However, I would suspect that if AIIM had polled just eDiscovery professionals that they would have gotten an even louder call (more than 70%) for good eDiscovery analytics.

eDiscovery technology vendors are making tremendous strides in regards to enabling users to conduct some analytics during the Early Case Assessment (ECA) phase and even during the Document Review phase of the EDRM. However, I contend that most searches that are done today are keyword searches based upon a list of keywords that were produced by outside counsel (probably Associates and paralegals) with little or nor real knowledge of the Electronically Stored Information (ESI).

This archaic and potentially dangerous practice is no doubt a huge step forward from sticky notes, yellow pads and Excel spreadsheets. However, with the advanced analytical tools that have been on the market for general business analysis for years, there is no excuse for eDiscovery professionals to not be using advanced analytic search technology such as conceptual search for eDiscovery.

I have written extensively about this topic on this Blog. Following are links to some of those posts:

Become eDiscovery Superheroes: http://ediscoveryconsulting.blogspot.com/2010/03/become-ediscovery-superhero-with.html

Concept Search vs. Keyword Search in eDiscovery: http://ediscoveryconsulting.blogspot.com/2008/12/concept-search-vs-keyword-search-in.html

The Fog is Lifting on Conceptual Search in eDiscovery:
http://ediscoveryconsulting.blogspot.com/2008/12/fog-is-lifting-on-concept-search-in.html

Conceptual Search Case Law Emerging: http://ediscoveryconsulting.blogspot.com/2008/06/concept-search-case-law-emerging.html

The New Generation of eDiscovery Search: http://ediscoveryconsulting.blogspot.com/2009/02/new-generation-of-ediscovery-search.html
Web 3.0 in eDiscovery: http://ediscoveryconsulting.blogspot.com/2009/12/web-30-in-ediscovery.html

The full text of the AIIM article is as follows:

Silver Spring, MD – March 29, 2010 - According to a recent survey report by content management association AIIM, organizations could derive much higher business value from content analytics tools than from simple search-engines. Sophisticated content reporting across text documents and rich media file-types has created the opportunity to report and research across unstructured content, bringing the same capabilities of strategic insight and improved decision-making as Business Intelligence (BI) reporting brings to structured content.

Over 70% of respondents in AIIM’s survey would find advanced content analysis functions “Extremely useful” or “Very useful.” They rate their current ability to “research” content for business insight, or to monitor desirable or undesirable activity as 3 to 6 times less than their ability to simply “search” across different content types. Relatively new as a recognized toolset, content analytics tools provide trend analysis, content assessment, pattern recognition and exception detection. Applications include fraud detection in claims or loan applications, pattern detection in inspection reports, detecting unauthorized use of copyright material, analysis of healthcare records against other citizen databases, automatic redaction (blanking out) of sensitive information, and sentiment analysis in customer correspondence or social media sites.

According to Doug Miles, Director of AIIM research activities, “In much the same way that BI tools opened up structured corporate data in finance and ERP systems to give managers true insight into business operations, content analytics can leverage the investments in content management systems to measure subtle trends and sentiments in assessment reports, correspondence, emails, and social media sites. Meanwhile, analysis tools for rich media file types, such as video and audio, are providing much better management of these valuable assets, as well as improving the ability to detect fraud and crime.”

One particularly useful application of analytics is that of measuring the relevance and likely duplication of stored documents and records, with a view to reducing the size of content stores in order to save storage space, particularly during system migration or company merger activities. Only 15% of respondents had any automated tools for this kind of content assessment.

The AIIM report projects a considerable increase in spend on content analytics technologies over the next two years, as well as increases for Digital Asset Management (DAM) and enterprise search applications.

Based on over 500 responses, the AIIM research report is entitled “Content Analytics – research tools for unstructured content and rich media.” Part of the AIIM Industry Watch series, the full report is free to download from the AIIM website. It is underwritten by Allyis, IBM and Media Beacon.

About the research
The survey was taken by 527 individual members of the AIIM community between February 9th and February26th using a Web-based tool. Invitations to take the survey were sent via e-mail to a selection of the AIIM worldwide community members.

About AIIM
AIIM (http://www.aiim.org/) is the community that provides education, research, and best practices to help organizations find, control, and optimize their information. For over 60 years, AIIM has been the leading non-profit organization focused on helping users to understand the challenges associated with managing documents, content, records, and business processes. The AIIM community includes over 65,000 ECM users and professionals.

About Allyis
Allyis develops and supports technologies that help businesses operate, share information, and communicate more effectively. Whether developing an employee intranet to connect a dispersed workforce, designing a knowledge management strategy to surface talent and expertise, or providing content management support, Allyis leverages people and technology to make business more efficient and effective. http://www.allyis.com/

About IBM
As a content, process and compliance software market leader, IBM ECM delivers a broad set of mission-critical solutions that help solve today’s most difficult business challenges: managing unstructured content, optimizing business processes and helping satisfy complex compliance requirements. More than 13,000 global organizations and governments rely on IBM ECM to improve performance and remain competitive through innovation. http://www.ibm.com/.

About MediaBeacon
MediaBeacon, Inc. is the leading provider of Digital Asset Management, Enterprise Search and secure role-based media distribution portals of digital content technology. With some of the largest DAM deployments known to date and hundreds of global enterprise customers, MediaBeacon is a proven leader in the industry. For more information, visit http://www.mediabeacon.com/

Labels: , ,

Monday, March 29, 2010

Become an eDiscovery Superhero with Conceptual Search and Categorization Technology

I grew up watching Batman and Robin on a really small black and white TV and reading Superman comic books. And, I have really enjoyed the resurgence of the Batman movie series. Given this fondness for superheros, I wanted to post a recent paper entitled, "Become an ESI Superhero" by Herbert L. Roitblat, Ph.D. of OrcaTec LLC and eDSG.

Dr. Roitblat takes a really clever approach to explaining eDiscovery categorization technology in terms that we can all understand. He contends that its time for us all to put on our tights and join the Justice League of Superheros. I'm not sure that the world is ready to see some of us in tights. But, Dr. Roitblat's analogy is very helpful in understanding how these new technologies can be used to appear as a superhero.

He contends that, "Categorization can help to make the review much more reliable by using the machine to learn the decision patterns of an expert and then using those decisions as recommendations for more detailed review. Think of it as a form of Vulcan mind meld that transfers the expertise of the best expert available to the rest of the review staff. "

He states that, "Semantic clusters help to gain a quick overview of what the collection is about and also suggest key terms and phrases that can be used to identify responsive documents. They provide a quick method of determining which documents merit further immediate review and which can be safely set aside. After review, if some documents in a cluster are marked responsive and others are not, it may be useful to examine why not."

And, he concludes that. "Concept selection gives you x-ray vision into the meaning of your collection. It lets you identify what the words mean in this particular context and to identify documents based on their meaning. It helps to highlight the documents that are most about a specific concept from the specific point of view of the context. For example, among the Enron emails, the word "osprey" is not used to refer to the bird, or the aircraft but to one of the off-books partnerships that got Enron into so much trouble."

If all of this really works as advertised, it will make us all look like superheros. Whether or not you decide to wear the tights is up to you.

The full text of Dr. Roitblat's paper is as follows:

Super heroes have super powers. Their powers enable them to do things far beyond the capabilities of ordinary people. Superman had his X-ray vision. Wonder Woman had a lariat that compelled complete honesty. The Flash had super speed. Batman had his superior intellect and technology.

These members of the Justice League each had powers that would come in handy managing today's eDiscovery. If these powers were available to lawyers today, would you use them?

How would or should lawyers go about making decisions about using their super powers? There are legal issues, certainly. But as far as I can see, the really critical question is whether these powers provide capabilities beyond those of traditional eDiscovery processes.

The eDiscovery powers I'm thinking about revolve around technology. We cannot afford to wait for a lightning bolt to strike. They include categorization, clustering, and concept searching. Tools like these have the power to help you see inside the case materials, derive the honest information from them with super speed, and amplify your superior intellect.

The goal for the eDiscovery process is to identify the ESI that is potentially relevant and to separate it from the ESI that is not. Ten years ago, we could argue about paper versus plastic, about whether it was more or less efficient to review documents on a computer screen or on paper. At the time, I worked on a case involving 13 million pages, which the producing party was determined to produce on paper. That would have been about 65 tons of paper, the weight of an old-fashioned steam locomotive. There was simply no way that the receiving party could go through all of that paper in a timely way. Just sorting the pages into date order was a daunting task.

Since that time, the volume of ESI that must be considered has only continued to leap over tall buildings. It is no longer practical to have the managing partner on a case read through all of the available documents. Few attorneys actually have the super power, the Flash's speed, for example, to read through millions of pages in a short time, so they resort to other means to help them get through it.

Categorization
One of the oldest approaches is to hire an army of temporary attorneys to read the documents. When Verizon was acquiring MCI in 2005, they responded to a DOJ second request by hiring 225 attorneys for four months, working 16 hours per day, 7 days per week. And all this effort was needed to review just 1.3 terabytes or 1.6 million documents. The review, alone, cost over $13 million or about $8.50 per document. As the volume of ESI continues to grow into multi-terabyte collections, this approach is simply not sustainable.

There is also evidence to suggest that this approach is not as accurate as it might be. What level of attention can a reviewer sustain while reading documents 16 hours a day, seven days a week, week after week? We had one client, in fact, in an unrelated matter, who asked us to use search technology to filter out the jokes because the reviewers were spending way too much time on them, rather than plowing through the potentially relevant material.

We (Roitblat, Kershaw, & Oot, 2010) recently published a paper in the Journal of the American Society for Information Science and Technology (JASIST) analyzing the performance of human reviewers in comparison to two computer-assisted review systems. None of the authors of that paper has any financial relationship with the reviewers or the companies providing the computer categorizers.

We set out to examine the idea that these computer systems could yield results that are comparable to those that would be obtained with a human review, and we found that they were. If we somehow knew which documents were truly responsive, then we could compare the judgments made during the first review with these true judgments. Unfortunately, no such oracle was available, so instead, we had to settle for a comparative method. We assessed the level of agreement between a new traditional review by new professional reviewers with the original review and between the computer systems and the original review. By comparing a new human review with the original review, we get an assessment of how well the traditional approach captures reliable aspects of the document collection. If the computer systems perform no worse, then it may be reasonable to use systems like this, rather than to spend the time and money needed to hire humans to do the work. I'll return to this assertion in a bit, after discussing some of the results.

After being trained on the issues involved, two teams of experienced professional reviewers were given a random sample of 5,000 documents to review for responsiveness. We could now assess the level of agreement of each of the teams with the original review and with each other. Team A agreed with the responsiveness decisions of the original review on about 76% the documents. Team B agreed with the original review on about 72% of the documents. They agreed with each other on about 70% of the documents.

You might think that the reason for such low agreement between the teams and the original review was because of the extreme conditions under which the original review was conducted. I have little doubt that the original review could have been done more reliably with more time (and more expense), but these results do not support that conclusion because the two re-review teams agreed with each other, even less often than they agreed with the original review, and were under practically no time pressure. Instead, it would seem that the low level of agreement is more likely explained by the unreliability of human responsiveness judgments in general. People are just not very reliable at identifying responsive documents, even in a small collection of a few thousand documents.

The two computer systems, on the other hand, agreed with the original review on about 83 % of the documents, certainly no worse than the level of agreement to be expected based on the human review . Replacing the army of human reviewers with a computer did not decrease the level of agreement, but it would very likely save a substantial amount of money.

Alan Turing, considered by many to be the father of computing, developed a test for assessing machine intelligence. His argument was that intelligence is a function. If the outputs of a machine are indistinguishable from the outputs of a human under a specific set of circumstances, then the computer could be said to implement the same function as the human—in this case the intelligence function.

In Turing's test all communication is done through a written medium, say a computer keyboard. The tester is supposed to have a conversation with a partner in another room. If the tester cannot tell whether she is communicating with a person or a machine, then the machine is said to be producing equivalent results and to be executing the same function. The computer could be said to be genuinely intelligent.

We could apply this same methodology in assessing computer versus human judgments of responsiveness. If we cannot tell the difference between the two kinds of systems, then we can conclude that they perform the same function. Based on the data reported in JASIST, that seems to be the case. If anything, the computers were a bit more reliable than the humans. If Batman were an attorney, he would use this kind of technology to gain an advantage over his adversaries. Think of it as a kind of BatReview.

Not every attorney is convinced by these results. For some, this skepticism reflects a romantic notion that there must be something special about having real live humans read every document. Their claim is that humans will find responsive documents that the computer will miss.

Although it may be true that humans will find documents that the computer might miss, it is at least equally true that one human will find documents that another human might miss and conversely, miss documents that another human reviewer might find responsive. It is also true, that the computer is likely to find responsive documents that the human might miss. The agreement between two groups of humans was no higher than the agreement between the computer systems and the original review. The available evidence does not support the claim that humans will find more responsive documents or that they will retrieve fewer nonresponsive documents than either of these categorization systems will.

Another claim is that the character of the documents missed by the computer will somehow be different from the character of the documents missed by the humans. This claim is more difficult to assess, because it is not obvious how to measure the character of the documents that are missed by one system or another. In order for this notion to be valid, however, and still allow the machines to achieve levels of agreement that are equal to or higher than those achieved by humans it would require that the humans miss an equal number of documents that the computer finds responsive. It implies that there is some systematic difference in the documents that the people find and the computer does not and some systematic difference in the documents that the computer finds and the people do not. It's not impossible that both systematic differences exist, but its practical significance is, at best, elusive.

How well any system, whether human or machine, performs is a matter of measurement. It seems unreasonable to claim that humans are somehow better than computers at distinguishing responsive from nonresponsive documents without measuring their performance along dimensions that matter. For example, the low level of agreement between human reviewers often comes as a shock to attorneys. Every review should include quality measures, whatever technology is used to perform it. Superman was not a superhero just because he flew around, that would only make him super. He was a superhero because he was successful at thwarting bad guys. His effectiveness could be measured by the number of evilness of the villians he defeated. The effectiveness of review should be measured by the ability to identify responsive documents and eliminate the nonresponsive.

Practically every process can be improved once it is measured. The steps taken to improve the quality of the review should be determined by the degree of risk in the case, in other words, they should be based on judgments of reasonableness.

The two computer systems used in the JASIST study did not make up their categorization by themselves. They don't actually decide what is responsive and what is not responsive. They form their categories on the basis of input from people. They implement the judgment of their "trainers" rather than make up their own. One system used in this test learns how to distinguish between responsive and nonresponsive documents from example judgments made by reviewers. The system is given a set of documents that the reviewers determined were responsive and a set that the reviewers determined were nonresponsive. From its analysis of these two sets, it derives a set of computational rules that distinguish between the two and applies these rules systematically to the remaining documents. This is the most common form of machine learning and automated categorization.

The other system was trained by linguists and attorneys to distinguish between responsive and nonresponsive documents. The trainers read the request and the training information provided to the original reviewers. They then adjusted the system's algorithms until it distinguished between documents in the way that these people determined was accurate. In both cases, the computer simply systematically implemented the decisions made by its team of trainers without getting bored, distracted, tired, or needing a vacation.

Keyword selection
Over the last few years, many attorneys have grown comfortable with one weak form of machine classification. Many of them have been using keywords to select or cull documents for further review. The attorneys pick a set of keywords or Boolean queries to use to select documents. These terms may be created by one side or negotiated between the two sides. Any document that contains one of these keywords or matches the query is selected for further processing, the others are simply ignored. For the most part, if a document does not have one of the keywords in it, it is never looked at again, so its information is effectively lost.

Keyword searching, although important, is the weakest form of machine classification available, hardly up to superhero standards, kind of like the Green Lantern with yellow things. The success of keyword searching depends critically on the ability of the attorneys to pick the right words that identify the responsive documents and do not overwhelm with nonresponsive ones. For more than 20 years we have known that attorneys are only about 20% successful at guessing the right words to search for (Blair and Maron, 1985). In the 2008 TREC Legal Track, they had two sides of an issue create search terms. The "defendant's" search terms retrieved just 4% of the responsive documents. The "plaintiff's" search terms retrieved 43% of the responsive documents, but 77% of the documents returned were nonresponsive, much higher than the 59% nonresponsive rate for the "defendant."

As you would expect, when they negotiated the search terms, as many of the thought-leaders in eDiscovery suggest, the results were intermediate, but still not stellar. Negotiated search retrieved only 24% of the responsive documents (http://trec.nist.gov/pubs/trec17/papers/LEGAL.OVERVIEW08.pdf, p. 10). In the 2007 TREC, it was even lower. Of the documents that were retrieved, 72% were found to be nonresponsive. Keyword culling serves to reduce the volume of documents that will be considered for review, but it does not do a very good job of identifying responsive documents. Cooperation may be the better policy, but it's no lariat of truth, unless further steps are taken.

One reason for the poor performance of these search terms might be that they were based on the attorney's expectations for the words that ought to be in the collection and ought to distinguish between responsive and nonresponsive documents, rather than for the words that actually were there. For example, people often misspell words in emails. A name like "Brian," for example, might be spelled "Brain," or "Bryan." "Believe" might be misspelled as "beleive." Document authors, especially email authors, may use nonstandard abbreviations. They can be very creative in the words they choose to use. There are over 200 synonyms, for example, for the "think." Did they "buy," "purchase," "acquire," or just "get" a new car, for example? Groups all have specialized jargon. Lawyers may have one way to talk about issues in the case, but the document authors rarely think or write like lawyers.

Identifying the right terms to search for can be helped by looking at the words that are actually in the ESI. One easy way to do this is to simply print out or display an index of the collection. This list is sometimes called a word wheel. Then, the parties can select terms from this list, rather than trying to make up a list from their imagination.

Clustering
A more powerful technique is to use semantic clustering to group documents with similar content. Using one of a number of techniques, the computer groups together similar documents and finds a word or a phrase that describes that group. A quick examination of these groups is often enough to determine whether the documents they contain are likely to be responsive or not. The labels can be used as potential search terms. Some of these clusters may be obviously responsive, some obviously nonresponsive and others may require looking at a few of the documents to figure out.

There is an added benefit if a given document can be in more than one cluster because documents can be about more than one topic. An email might say something like, "I'm bringing pizza to the party on Saturday, and by the way, the money that we stole is now in our Swiss bank account." If that email were clustered into only the pizza party cluster, no one would ever see it again and that information would be lost.

The clusters can also be used to select or cull. You can scan down a list of clusters and determine whether the documents in that cluster are likely to be responsive or not. Any document in a potentially responsive document could then be reviewed further, but documents that never appear in a responsive cluster can be set aside.

Here are a few of the cluster labels derived from a subset of the Enron data:

- ibm
- 2q
- bal month
- transactions
- software
- luis gasparini
- coaches
- generators
- lon ect
- hou ees
- el paso
- ebner daniel
- officials
- epmi long term northwest
- bonus
- game

If the issue involves financial information, then documents in 2q, epmi long term, transactions, and bonus clusters may be relevant. Terms like these would make good search terms, and you know that they are actually used in the documents and that there are enough of them to make up a cluster. Documents in the coaches, software, and game clusters are unlikely to be relevant. Because the same document can appear in more than one cluster, it will still be selected for review if it appears in at least one responsive cluster, no matter how many nonresponsive clusters it appears in.

Concept Selection
Concept searching is another tool that can help to amplify your powers. Concept searching identifies the meanings of words using any of a number of different technologies, including ontologies, thesauri, latent semantic indexing (LSI), and language modeling. Ontologies and thesauri are usually created by experts, who program word relations into the system. These knowledge engineers identify that the word "car" is related to the word "vehicle," for example. Systems based on one of these approaches contain only those relationships that have been explicitly programmed into them.

LSI and language modeling derive the meanings of words automatically from the context in which those words are used. These systems reflect the actual patterns of word use and are capable of discovering unanticipated relationships.

Concept searching works by expanding the query that the user submits to include the original term and additional, related terms. Concept searching identifies documents as relevant, even if the exact query term happens not to be in the document, because it searches not just for the original term submitted by the user, but for it and related terms. This approach tends to push the most relevant documents, the ones with the query term and the most context, to the top of the results list and to add a small number of related documents, which do not happen to have the query term, to the tail end. Concept search tends to mitigate the difficulty of guessing the right terms to search for, because it learns what terms are related to which. It helps to find documents based on their meaning, rather than solely on the presence of specific words.

Conclusion: Using These Superpowers
The eDiscovery superpowers described above can really help to reduce the cost and burden of eDiscovery. Categorization can help to make the review much more reliable by using the machine to learn the decision patterns of an expert and then using those decisions as recommendations for more detailed review. Think of it as a form of Vulcan mind meld that transfers the expertise of the best expert available to the rest of the review staff.

Documents identified by the categorizer can be reviewed first. It prioritizes processing so that the most likely documents can be completed first. It allows the reviewers to quickly gain exposure to the documents most likely to be relevant and learn from the examples what makes them relevant.

After the review, a sample of the decisions made by the human reviewers can be fed back to the categorizer and it can be retrained. If the reviewers were consistent in their review, then there should be few discrepancies between the categories assigned by the computer and the categories assigned by the reviewers. These discrepancies can then be examined to resolve those inconsistencies.

By examining the actual words used in the documents, the quality of keyword selection can be greatly improved. It then becomes easier to negotiate sensibly about which keywords to use. It becomes easier to demonstrate that you have met the requirement of the Federal Rules to conduct a reasonable search of the ESI. Most importantly, you increase your chances of finding the responsive documents without over-burdening the collection with irrelevant ones.

Semantic clusters help to gain a quick overview of what the collection is about and also suggest key terms and phrases that can be used to identify responsive documents. They provide a quick method of determining which documents merit further immediate review and which can be safely set aside. After review, if some documents in a cluster are marked responsive and others are not, it may be useful to examine why not.

Concept selection gives you x-ray vision into the meaning of your collection. It lets you identify what the words mean in this particular context and to identify documents based on their meaning. It helps to highlight the documents that are most about a specific concept from the specific point of view of the context. For example, among the Enron emails, the word "osprey" is not used to refer to the bird, or the aircraft but to one of the off-books partnerships that got Enron into so much trouble.

Armed with these superpowers you are now ready to combat your adversaries in the world of eDiscovery. Go put on your tights and join the Justice League.

Labels: , , , , , ,

Tuesday, June 23, 2009

Early Case Assessment (ECA): From Key Word Search to Automating Document Relevance

Just before the Christmas holidays in 2007 I was teaching a CLE on the “Changes to the Federal Rules of Civil Procedure (FRCP)”. There had to have been 150 litigators, litigation service providers and consultants in the room listening to me preach about this unchartered new world called eDiscovery. My initial morning session was an overview of the current status of paper based litigation processing, how the accelerating increase in the volume of Electronically Stored Information (ESI) was steaming down the eDiscovery tracks like an out of control locomotive, how the changes to the FRCP would effect us all and what they all needed to know that day to get ready. One of topics that I touched on was the need to become more familiar with key word search methodology and the new technologies that would enable the legal industry to better manage the process of finding relevant data and assessing cases early on in the eDiscovery lifecycle. The reason that I am recalling this CLE seminar was because I will always remember the comments from an older litigator that packed up his belongings during the first break, walked up to the front and announced that, “all of this fancy new technology was never going to replace good old fashion legal hard work, understanding the relevant facts and key information for case wasn’t as complicated as I was making it sound and therefore he wasn’t going to stick around for the rest of my session.”

Well, the remainder of 149 attendees stayed and I hope that they all walked away with a better understanding of the importance of understanding key word search methodology and the new technologies that would enable the legal industry to better manage the process of finding relevant data and assessing cases early on in the eDiscovery lifecycle. Over the subsequent 18 months, key word became a big issue in the eDiscovery lifecycle and spawned a whole new technology arena called Early Case Assessment (ECA) that focused on reducing the amount of “relevant” data that had to ultimately be reviewed by a lawyer. And, having just completed another CLE this past week on ECA, I think that “most” eDiscovery professionals understand the keyword search methodology and the new technologies that enable them to better manage the process of finding relevant data and assessing cases early on in the eDiscovery lifecycle.

However, just went we thought we had it had it all figured out, a study from the Text Retrieval Conference (TREC) indicated that the keyword method tends to miss most of the relevant documents, while yielding mainly irrelevant documents. As a result, only a fraction of the relevant documents make it to the detailed review stage, while most of the documents that are submitted to review are in fact not relevant.

Moreover, the study also found that the keyword approach is typically binary, meaning that documents are either included or not. There is no graduated scale of relevance. This rigid approach does not allow for relative ranking of documents, making it extremely difficult to manage and prioritize document review.

However, as with any market that is going through a paradigm shift looking for its center and trying to normalize on some standards, the litigation technology vendors have been all over the keyword search issue and are starting to release their new solutions into production. One of the very first players in the industry to address the issue of document relevance is Equivio, a leading provider of near de-duping and email thread management technology. They have just launched Equivio>Relevance™, an expert-guided system that enhances the eDiscovery process through automated document prioritization.

Getting back to the comments of my lawyer friend that walked out of my CLE back in 2007, fancy new technology may never replace good old legal hard work. However, in today’s new world of Terabytes of potential evidence in even some of the small matters, we need all the technical help that we can get. And, it appears that Equivio is stepping up and offering us all at least a fighting chance to find the documents that we need successfully mange our cases.

The Full Text of Equivio’s Press Release is as follows:

Kensington, MD, June 22, 2009 – Equivio, a provider of software for managing data redundancy, announced today that it has launched Equivio> Relevance™, an expert-guided system that enhances the eDiscovery process through automated document prioritization.
Traditionally, attorneys use keywords to pre-filter documents prior to detailed review. According to the TREC (Text Retrieval Conference) studies, the keywords method tends to miss most of the relevant documents, while yielding mainly irrelevant documents. As a result, only a fraction of the relevant documents make it to the detailed review stage, while most of the documents that are submitted to review are in fact not relevant.

Moreover, the keywords approach is typically binary, meaning that documents are either included or not. There is no graduated scale of relevance. This rigid approach does not allow for relative ranking of documents, making it extremely difficult to manage and prioritize document review.

Equivio>Relevance™ is designed to address these limitations, introducing a higher level of flexibility, control and accuracy into the eDiscovery process. Based on initial input from a lead attorney, Equivio>Relevance uses statistical and self-learning techniques to calculate graduated relevance scores for each document in the data collection. Equivio>Relevance also uses a statistical model to calculate the precision and recall achieved by the software. This statistical model is used to provide a new level of measurability and control in the eDiscovery arena, while also helping to ensure the defensibility and transparency of the process.

Equivio>Relevance drives value throughout the eDiscovery flow through:

  • Early case assessment: Equivio>Relevance facilitates rapid assessment of the key issues and concepts in a case.
  • Culling: Equivio>Relevance achieves high levels of recall and precision, helping overcome the challenges of over and under-inclusion that characterize traditional keyword methods.
  • Review prioritization: By organizing the review set according to relevance rankings, Equivio enables prioritization of document review. This allows attorneys to immediately focus on the most relevant documents.
  • Review quality assurance: By identifying discrepancies in the responsiveness designations of Equivio>Relevance vis-à-vis the human review team, the application helps find responsive documents missed in the detail review. Similarly, the discrepancies can be used to locate documents incorrectly marked by the human review team as responsive.
Equivio>Relevance also generates a list of keywords that characterize relevant documents in the collection. These automatically-generated keywords can be used to supplement and enhance the manual list of keywords built by the legal team.

"Equivio is committed to developing and delivering innovative technologies that will help litigators improve the quality and consistency of their eDiscovery processes," said Amir Milo, CEO of Equivio. "Equivio>Relevance™ enables attorneys to review fewer and more relevant documents, lowering review costs and reducing the risk of missing key data."

About Equivio

Equivio enables the management of data redundancy in content-centric business processes. Equivio's technology zooms in on unique data, allowing you to read less, think more, win big™. With products for grouping near-duplicates, capturing email threads and determining document relevance, Equivio powers a broad range of business applications, including eDiscovery, records management, email archiving, data retention and intelligence. To learn more about winning with Equivio, visit http://www.equivio.com/.

Labels: , , , , ,

Thursday, February 12, 2009

The New Generation of eDiscovery Search

Train Leaving the Station The New Generation of eDiscovery Search technology train is getting ready to leave the station. However, after walking the tradeshow floors and attending many of the breakout sessions at last weeks LegalTech in New York, it is obvious that there is a tremendous amount of confusion regarding the definition and scope of the New Generation of eDiscovery Search technology and more importantly, how the courts view the use of such technology.

With the accelerating volume of Electronically Stored Information (ESI) or what I like to call Electronically Stored Evidence (ESE), the current legacy search technologies built into the current legacy eDiscovery tools and the associated best practices for document review are beginning to have a hard time "keeping up". Further, there is a tremendous amount of confusion and trepidation among litigators in regards to potential malpractice claims, sanctions and adherence to Rule 702 and Daubert challenges associated with employing the New Generation of eDiscovery Search technology. Finally, litigation technology vendors, whether purposely or not, have confused the market with fancy new marketing terms like "conceptual search", "transparent search", "linguistic search" and "clustering" without any real explanation of how they work and how to correctly employ them. Therefore, I thought that it was time to restart the campaign to both educate and lobby the eDiscovery industry in regards to the New Generation of eDiscovery Search.

First of all, I want to start with education regarding the pertinent issues. Without a doubt, one of the best article posted over the past 12 months on the legals issues surrounding the topic of what I am calling the New Generation of eDiscovery Search, was written by By Wayne C. Matus and John E. Davis in the New York Law Journal on October 31, 2008, titled, "Do Your Searches Pass Judicial Scrutiny?"

Since it is my impression that many of these very important issues are still either unknown to most in the eDiscovery "business" or are being ignored, I contacted Mr. Matus this week to get permission to repost his article.

Following is the full text of of "Do Your Searches Pass Judicial Scrutiny?":
Electronically stored information is increasing exponentially, and bills from law firms and discovery vendors to deal with this vast sea of data escalate significantly each year. Jason Baron, the director of litigation at the National Archives and Records Administration, believes that ESI is growing so fast that even with unlimited funds and human resources it will soon be impossible for humans to review these large document populations.[FOOTNOTE 1] Still, lawyers faced with potential malpractice claims and sanctions are loath to try new methods for handling the problem. It is time for change.

The traditional means used by litigators to address ESI is the application of keywords and Boolean search terms to identify relevant and non-privileged materials.[FOOTNOTE 2] While acknowledging that this method is unquestionably deficient, a recent article published in this publication concluded that "the available evidence suggests that keyword and Boolean searches remain the state of the art and the most appropriate search technology for most cases."[FOOTNOTE 3] We agree that, in a perfect world, if the parties can nevertheless meet and confer, and agree upon keywords to reduce the population to manageable proportions, the traditional judgmental method can be made to work. However, this is an imperfect world where plaintiffs and defendants do not always agree, and are not always equally motivated, to reduce costs. In fact, it is often quite the opposite. Moreover, even where the sides use judgmental sampling to agree upon keywords, the costs nevertheless usually remain too high.

THE JUDGMENTAL APPROACH

The judgmental approach to keywords ultimately fails because of "recall" and "precision." "Recall" measures how completely a process captures target data. "Precision" measures efficiency - the amount of irrelevant data captured along with the target data. Keywords, as judgmentally used by lawyers, recall too little, while capturing much that is irrelevant. An early landmark empirical study by David Blair and M.E. Maron[FOOTNOTE 4] showed that while lawyers thought they were retrieving about 75 percent of the relevant data, the true results were more like 20 percent. A subsequent study, conducted by the Text Retrieval Conference,[FOOTNOTE 5] confirmed this result, finding that only 22 percent of relevant documents were recalled using keyword search techniques, as opposed to approximately 78 percent found by other search techniques.[FOOTNOTE 6] Many lawyers will also tell you that it is common for reviewers to find only 10 to 40 percent of the recalled documents to be relevant, meaning lawyers are reading mostly junk.
We advocate two different approaches to yield better and more efficient results. First, we suggest that keywords are best used coupled with statistical, rather than judgmental, sampling. Second, we suggest that experienced counsel and vendors working with a combination of advanced conceptual search techniques can more efficiently and effectively deal with large amounts of ESI, resulting in a narrowed and enriched review set with a concomitant reduction in lawyer hours.

KEYWORDS DONE RIGHT

In Victor Stanley v. Creative Pipe, 250 FRD 251 (D. Md. 2008), Chief Magistrate Judge Paul W. Grimm of the U.S. District Court for the District of Maryland found counsel had waived the attorney-client privilege as to 165 inadvertently produced documents -- despite the use of 70 separate keyword searches in conducting their privilege screen -- because, among other reasons, counsel had failed to conduct "quality assurance testing." Clearly, judgmental sampling did not pass judicial scrutiny, while statistical sampling would likely have.

Counsel seeking to conduct a proper keyword search should instead consider the following steps:

• Sample the data. Counsel should isolate a random and statistically significant sample of the relevant datasets and then conduct a manual review of such data for relevance and privilege.[FOOTNOTE 7] This will educate counsel as to what to expect from the larger population and help in formulating keywords.
• Analyze and rank keywords. Counsel should then create and run search terms against the sample set, and (based on the information derived from the sample review) analyze their effectiveness by "recall" and "precision." This preliminary knowledge of the contents and richness of particular datasets will provide the basis to predict retrieval and review costs, and whether, for example, counsel should conduct any or just a limited review of such data.[FOOTNOTE 8]
• Review and repeat until satisfied that the search plan is defensible. This approach is plainly iterative in nature; it is the rare set of searches that achieves acceptable returns without adjustment. Successive application and fine-tuning of terms should permit counsel to achieve defensible levels of recall with superior precision rates. Practitioners should take note: One of the main factors cited by the court in Victor Stanley to determine if a party has conducted a reasonable search is if the party has reviewed a sample of the results to "assess its reliability, appropriateness for the task, and the quality of its implementation."[FOOTNOTE 9]
There is no consensus as to what percentage of recall will pass muster. Instead, counsel must be able to explain to the court the "reasonableness" under the circumstances of each step of the process, including the point at which a party was satisfied with the effectiveness of its search terms.[FOOTNOTE 10]

KEYWORDS: TO DISCLOSE OR NOT

Search terms created by counsel are generally protected, at least initially, by the attorney work-product doctrine, as their "mental impressions, conclusions, opinions, or legal theories ... concerning the litigation."[FOOTNOTE 11] But how does one show "reasonableness" of the search methodology without disclosing the keywords? Three recent cases, Victor Stanley, O'Keefe and Equity Analytics,[FOOTNOTE 12] have required that attorneys be able to explain and defend to the court, at its request, the methodology used to employ the search terms. One court has indicated it might find a waiver of privilege and require disclosure of the search terms.[FOOTNOTE 13]

OTHER FILTERING TECHNOLOGIES

Magistrate Judge John M. Facciola of the U.S. District Court for the District of Washington, D.C., recently pointed to authority that "concept searching" applications -- which use statistical and linguistic models to search for ideas as well as words and impose order upon disparate documents -- are "more efficient and more likely to produce comprehensive results" than keyword or Boolean searches.[FOOTNOTE 14] For example, the TREC 2007 Legal Track study found that 78 percent of relevant documents in a dataset were not found by Boolean keyword searches, but only by alternative search techniques.[FOOTNOTE 15] It is little wonder that Magistrate Judge Grimm has expressed optimism that concept-based searches studied by TREC would supplant keywords as the preferred method "for a variety of ESI discovery tasks."[FOOTNOTE 16]

These findings appear to have been confirmed. Earlier this year, the eDiscovery Institute disclosed its preliminary assessment of the study it conducted on the performance of computerized document review against human review. The study was conducted against a dataset drawn from the Verizon-MCI merger consisting of 1.3 terabytes and over two million documents. They concluded that computer systems allowed a comparable level of performance to be achieved with fewer people, less time and lower cost. While actual cost of traditional review was over $13.5 million, computer-assisted review was projected to cost just a fraction of that amount.[FOOTNOTE 17]

Concept search methodologies fall into three basic (and sometimes overlapping) categories:

• Probabilistic. This technique relies upon probabilistic search models such as "Bayesian classifiers," which evaluate and classify documents based on the interrelationships, proximity and frequency of usage of words found therein. The model may be given a "head start" by a sample set of relevant documents developed by attorneys at the outset of the process, which the computer analyzes and applies to the remaining documents. This technique can order groups and documents based on perceived potential importance to assist in the review process.
• Rule-based (or "clustering"). This statistically driven process analyzes the prevalence of words in documents and, based on such analysis, groups together documents interpreted as featuring like concepts. This technique can order documents by perceived potential importance as well.
• Linguistic. Sometimes referenced as "fuzzy search models," this technique seeks documents containing all forms of a target word or its synonyms in a general and/or case-specific thesaurus. Linguistic approaches may also rely upon statistics to analyze documents for terms along the same subject lines -- or sometimes to identify documents that use different ways to make the same point.[FOOTNOTE 18]
Differing tools often produce differing results, but some combination of each of these approaches (as well as Boolean keyword searches) -- using a transparent, iterative and measured process as described above -- can be used to best effect. As successive waves of ESI are received (as is often the case), moreover, certain of the concept-searching applications "learn" and become better at identifying correlations, associating documents with particular attributes with concepts of interest to counsel and minimizing false positives. Further, the statistics generated by this process permit counsel to draw educated lines as to where review should proceed and, sometimes more importantly, when it is reasonable to stop. The advantages of these powerful, computerized techniques become even more apparent where ESI reaches the terabyte range and the steep recall/precision tradeoff exhibited by keyword analyses may reach unacceptable levels. Two jurists have indicated in opinions an interest in hearing from experts as to such new approaches.[FOOTNOTE 19]

CONCLUSION

Given escalating volumes of ESI, with no end in sight, and the general impatience of courts with e-discovery mistakes, counsel and their clients soon may have no choice but to adopt discovery tools that are more efficient and precise than traditional Boolean search techniques. Courts have already put practitioners on notice of this emerging obligation. Combining measured approaches to search methodologies with advanced techniques can greatly assist in the organization of ESI and the cost-effective conduct of litigations and investigations. The future is now for these state-of-the-art search techniques.
Wayne C. Matus is a litigation partner in the New York office of Pillsbury Winthrop Shaw Pittman and one of two national leaders of the firm's e-discovery practice. John E. Davis is a senior associate in the firm's New York office specializing in e-discovery. Sandra Barragan, an associate at the firm, assisted in the preparation of this article.

::::FOOTNOTES::::

FN1 "EDD Showcase: Discovery Overload," Law Technology News, January 2008.
FN2 Although keyword searches and Boolean term searches are undeniably distinct, for purposes of this article we will refer to them interchangeably.
FN3 See "Assessing Alternative Search Methodologies," H. Christopher Boehning and Daniel J. Toal (NYLJ, April 22, 2008).
FN4 "An Evaluation of Retrieval Effectiveness for a Full-Text Document Retrieval System," Communications of the Association for Computing Machinery at 289-99, March 1985.
FN5 TREC is sponsored by the National Institute of Standards and Technology (NIST) and the Advanced Research and Development Activity of the Department of Defense.
FN6 These were the results of TREC 2007, the second year of the Legal Track study. See Jason R. Baron, Douglas W. Oard, Paul Thompson & Stephen Tomlinson, Overview of the TREC-2007 Legal Track, at §6 (linked at http://trec-legal.umiacs.umd.edu/). In a prior TREC-6 Ad Hoc Task study, for keywords to achieve just 50 percent recall, the architects had to accept a dismal 20 percent precision rate (whereby four of every five documents selected by keywords were nonresponsive). See H5 White Paper, Concept Search: Perceived Security, Actual Risk, at 2, citing Voorhees, Ellen M., and Harman, Donna, Overview of the Sixth Text REtrieval Conference (TREC-6), in NIST Special Publication 500-240: The Sixth Text REtrieval Conference (TREC 6), ed. E.M. Voorhees and D.K. Harman, 1-24 (Gaithersburg, MD: NIST 1997), and Voorhees, Ellen M., and Harman, Donna, Overview of the Seventh Text REtrieval Conference (TREC 7), ed. E.M. Voorhees and D. K. Harman, 1-24 (Gaithersburg, MD: NIST 1998).
FN7 E.g., Treppel v. Biovail Corp., 233 FRD 363, 374 (SDNY 2006).
FN8 See McPeek v. Ashcroft, 212 FRD 33, 35 (D. D.C. 2003) (ordering sampling of backup tapes to determine whether they contained relevant documents); Wiginton v. DB Richard Ellis Inc., 229 FRD 568, 570 (N.D. Ill. 2004) (ordering sampling of archived material based on keywords to determine whether it should be restored); see also Victor Stanley, 250 FRD at 261, citing The Sedona Conference Best Practices Commentary on the Use of Search & Information Retrieval Methods in E-Discovery, 8 Sedona Conf. J. 189 (2007) [hereinafter, "The Sedona Best Practices"]. For example, the sampling process may reveal that certain data sources (such as local hard drives) or file types (such as Microsoft Access files) have such low yield that the collection and review effort is not worthwhile. While opposing counsel may not agree to such decision, the data provided by the sample review will provide the evidentiary support needed to defend the reasonableness of such steps to the court.
FN9 Victor Stanley, 250 FRD at 256.
FN10 See, e.g., Security Financial Life Insurance Company v. Dept. of Treasury, 2005 WL 839543, *4 (D. D.C. April 12, 2005) ("In deciding whether an agency's document search is adequate, the issue is not whether other responsive records might possibly exist, but whether the search was adequate, judged by a reasonableness standard.") (internal citations omitted); see also Victor Stanley, 250 FRD at 261 n.10 ("the cost-benefit balancing factors of [FRCP] 26(b)(2)(c) apply to all aspects of discovery").
FN11 Fed. R. Civ. P. 26(b)(3); see Lockheed Martin Corp. v. L-3 Comm. Corp., 2007 WL 2209250 (M.D. Fl. July 29, 2007) ("documents containing instructions about how to conduct the [ESI] search and what specifically to search for are opinion work product" and therefore protected as attorney work product privileged material); see also Gibson v. Ford Motor Co., 2007 WL 41954, at *6 (N.D. Ga. Jan. 4, 2007) (document retention notice that included a list of search terms reflected attorney mental impressions and so constituted protected work product).
FN12 Victor Stanley, 250 FRD at 256, United States v. O'Keefe, 537 F.Supp.2d 14 (D. D.C. 2008), and Equity Analytics, LLC v. Lundin, 248 FRD 331 (D. D.C. 2008).
FN13 Counsel, early in the process, should consider disclosure of keywords and other aspects of the search protocol to the adversary and the court, and invite their comment and approval, as a means of managing discovery costs and risk. Magistrate Judge Grimm in Victor Stanley Inc., 250 FRD at 256, found, among other things, that defense counsel's failure to disclose the keywords used to screen for privileged documents in defending its search methodology justified a finding of waiver as to the inadvertently produced documents. The court provided a checklist for attorneys to follow when preparing the methodology to be used to gather and produce ESI: Attorneys should consider the reasonableness of "the keywords used; the rationale for their selection; the qualifications of the [creators of the search] to design an effective and reliable search and information retrieval method; whether the search [is] a simple keyword search, or a more sophisticated one, such as one employing Boolean proximity operators ... ." Id.
FN14 Disability Rights Council v. Washington Metropolitan Transit Authority, 242 FRD 139 (D. D.C. 2007), citing George L. Paul & Jason R. Baron, "Information Inflation: Can the Legal System Adapt?" 13 Rich. J.L. & Tech. 10 (2007).
FN15 TREC 2007 Legal Track.
FN16 Victor Stanley, 250 FRD at 261 n.10.
FN17 http://www.ediscoveryinstitute.org/research/index.html.
FN18 Such tools are described in further detail in The Sedona Best Practices at 191-216 & Appendix. While we are unaware of a court that has expressly endorsed this approach, parties have used these search techniques in conducting, among other things, internal investigations.
FN19 Magistrate Judge Grimm in Victor Stanley, 250 FRD at 260, and Magistrate Judge Facciola in O'Keefe, 537 F.Supp.2d at 24, and Equity Analytics, 248 FRD at 333, have indicated that conducting and defending e-discovery may at times require experts. Indeed, Magistrate Judge Facciola stated that search methodologies in e-discovery may be scrutinized under Rule of Evidence 702.


Labels: , , , , , , , , , ,

Tuesday, December 9, 2008

Concept Search vs. Keyword Search in eDiscovery

Having grown up in the enterprise class solutions world with relational databases, I am very comfortable with SQL query based searching. And, over the past several years, with my focus on the litigation market and litigation technology, I have now become very familiar with keyword searching against scanned and OCR'd document files. However, with with my passion for "leading edge" technology and/or solutions that can meet the demands of a market going through a paradigm shift, I have not been overly excited or impressed with the state of search technology in eDiscovery.

However, that has all changed with the emergence of conceptual search technology. As such, I have spent a tremendous amount of time researching conceptual search and how it compares from both a technology standpoint and from as business value standpoint. And, although I have come to the early conclusion that there is room and a need for all three, I have also determined that there is still a tremendous amount of confusion in regards to concept search vs. keyword search technology and the best use of both.

Therefore, in an effect to keep the followers of my blog informed, I have been following a series of excellent posts on the ediscovery 2.0 blog discussing conceptual search. Following is the full text of the latest post titled "Concept Search Versus Keyword Search in Electronic Discovery" by Will Uppington:

In my last post, I started a discussion on the myths surrounding concept search. The first myth I dispelled was the “concept search is concept search” myth. The myth is that there is an agreed upon definition of concept search. In actuality, when people in e-discovery use the term concept search, they don’t always mean the same thing. Frequently they are not actually talking about concept search technology at all and are actually talking about concept or content categorization technology, which is very different. The second myth that needs dispelling is that concept search is better than keyword search.

The thinking behind this myth goes something like this:

Keyword search has a lot of problems. It is prone to being over-inclusive, i.e., finding some non-relevant documents, and under-inclusive, i.e., not finding some relevant documents. Concept search technologies are new and interesting and using these technologies you can find documents that keyword search can’t find. Therefore, concept search must be better than keyword search.


Let’s examine this thinking. The first two statements are accurate. Keyword search is not perfect and can produce over- and under-inclusive results. And concept search and content categorization technologies can both help identify documents that keyword search technologies might not find. However, the conclusion that concept search is better than keyword search is not valid and doesn’t follow from these two statements. Why?

In order to answer this question, we first need to go back to the difference between concept search and content categorization. Because these are different technologies, we really need to separately compare concept search versus keyword search and content categorization versus keyword search. Let’s start with content categorization and keyword search.

The issue with this comparison is that keyword search and content categorization do different things. Keyword search can be used in many ways in e-discovery. The two most common are: (1) analysis or case assessment: finding the hot documents and understanding the matter by determining who knew what, when, how and why, etc., and (2) culling: removing non-responsive documents and/or identifying potentially privileged documents in order to reduce a large, starting set of documents to a smaller set before review.
Content categorization, on the other hand, has historically been used within the review phase of e-discovery. Categorization can help reviewers to better understand the documents they are reviewing and thus potentially increase the speed of review. Practitioners with whom I have worked also find that categorization can be useful during analysis by helping to understand a matter and identify potentially important keywords.

However, content categorization has not been used as part of culling. First, culling needs to be transparent. You need to be able to get agreement with or at least explain to the opposing side and the court exactly how you have culled the data set. If you cull based on categories of documents that have been generated by a proprietary, black-box algorithm, it’s going to be difficult to gain agreement on or explain your culling methodology. This is why the typical method of culling is still to use keyword search and either agree on the set of search terms with the opposing side or to use e-discovery search best practices to perform keyword searches on your own.

Second, content categorization has its own issues when it comes to being over- and under-inclusive. There is no guarantee that your group of documents that have been categorized as being related to, for example, a company’s hiring policies include all of the documents in your matter related to hiring policies or that they do not include some documents that may not really be related to hiring policies. Content categorization, like keyword search and virtually every information retrieval technology, is not perfect.

So what about concept search technology? Surely, concept search technology is better than old, boring keyword search. Well, actually it’s not that clear-cut. The problem with concept search technology is that while it might find more relevant documents than plain keyword search, it will also likely find more false positives. Imagine searching for documents containing “terminate” in an employment matter and your concept search technology automatically searching for “fire”, “dismiss”, etc. as well. You’ll find more documents related to the termination of employees, but you’ll also find a lot more non-relevant documents concerning house fires, the fire department, etc.

So concept search can help address the under-inclusive problem with keyword search, (though it won’t solve it) and can be helpful during analysis. But it can often increase the over-inclusive problem. In addition, today’s concept search technologies share the transparency problem with concept categorization. These technologies have largely been designed as “black boxes”, which as I have discussed in the past, makes sense for Enterprise search but not for e-discovery search, and, as a result, could also be potentially difficult to explain and defend. For these reasons, concept search technology isn’t used very much in e-discovery today. In order for its use to become widespread, it will need to become more transparent. But that’s a topic for another day.

The bottom line here is that despite all the hype, concept search and content categorization technologies do not solve all the challenges of e-discovery search. Both of these technologies can be very useful and the technology behind them is always improving. However, as most of the experienced practitioners I work with already know, these technologies are generally better thought of as supplements to keyword search, not replacements. The important question is not whether to use one technology over the other but which technology is best suited to your objectives and how best to use all the available technologies to achieve the desired goal.

Labels: , , , , ,

Saturday, June 14, 2008

eDiscovery Search Predictions for 2008

As I continued my education eDiscovery search platforms over the past couple of weeks and subsequent update to my eDiscovery Paradigm Shift Blog, I came across a really interesting update by Stephen E. Arnold on the current state of the search market in general titled "Search Rumor Round Up, Summer 2008". And, although it is not specific to search in the eDiscovery space, his overview of search is outstanding and very applicable to what we have to look forward to in eDiscovery.

As I pointed out in my post titled "eDiscovery Search Case Law Emerging", the courts are starting to catch up in regards to the value and impact of search in the eDiscovery process along with the subtle nuances of search technology. And, as I talk to law firms and the legal departments of many of the Fortune 500 about their eDiscovery issues, I am finding that the topic of search and more recently conceptual search is being raised more often in the context of culling, de-duping, finding potentially responsive and privileged docs and gaining a better understanding of their data in general. However, I have also found a complete lack of understanding of current search technology and it applicability to the real needs of the litigators and their litigation services consultants.

Further, I am curious to understand the foundation for this weeks acquisition of Attenex by FTI. Was this a fire sale because the conceptual search market has not matured and therefore Attenex has not been able to reach its full potential as a conceptual search based review platform? Or, was this a brilliant move by FTI to add yet another leading edge technology to its roster based upon accelerating market demand? Just for the record, I happen to think that it is the latter. But, there are some intriguing arguments for the former.

The observations and comments in Mr. Arnold's article that I believe are most applicable to eDiscovery are as follows:

Rumor 1: More Consolidation in Search
As eDiscovery technology matures and the obvoius winners in technology and the appropriate strategy/formula for success begins to emerge, we are seeing the same basic consolidation in eDiscovery in general and will continue to see even more eDiscovery technology consolidation over the remainder of 2008 and in to 2009.

Rumor 5: Search Will Become a Commodity
I believe this prediction to be true at the desk top / SaaS based eDiscovery platform level where the user is wanting to search several 100,000 docs in a short period of time. However, at the service center production level where the document pool is terabytes of data, I believe that there is still room for several search technology leaders that can figure out how to get these massive search project done in hours as apposed to days.

Rumor 6: Search Is a Component of Other Enterprise Software
I believe that there is no doubt that user are quickly going to expect sophisticated search to become a seamlessly integrated part of any eDiscovery platform.

Rumor 9: Key Word Search Is Dead
This is my favorite prediction for eDiscovery as I can visualize all the current eDiscovery vendors cringing.

Rumor 10: A Hardware Maker Will Put Search on a Chip
Coming from a background of integrating the appropriate and ripe software technologies into firmware / hardware solutions, I absolutely agree with this prediction. And, it fits really nicely with the "behind the firewall appliance" direction that some of the archiving vendors are going.

All of this being said, the full text of Mr. Arnold's article is as follows:

I am fortunate to receive a flow of information, often completely wacky and erroneous, in my redoubt in rural Kentucky. The last six months have been a particularly rich period. Compared to 2007, 2008 has been quite exciting.

I’m not going to assure you that these rumors have any significant foundation. What I propose to do is highlight several of the more interesting ones and offer a broader observation about each. My goal is to provide some context for the ripples that are shaking the fabric of search, content processing, and information retrieval.

The analogy to keep in mind is that we are standing on top of a jello dessert like this one.

The substance itself has a certain firmness. Try to pick it it up or chop off a hunk, and you have a slippery job on your hands. Now, the rumors:

Rumor 1: More Consolidation in Search
I think this is easy to say, but it is tough to pull off in the present economic environment. Some companies have either investors who have pumped millions into a search and content processing company. These kind souls want their money back. If the search vendor is publicly traded, the set up of the company or its valuation may be a sticky wicket. There have been some stunning buy outs so far in 2008. The most remarkable was Microsoft’s purchase of Fast Search & Transfer. SAS snapped up the little-known Teragram. But the wave of buy outs across the more than 300 companies in the search and content processing sector has not materialized.

Rumor 2: Oracle Will Make a Play in Enterprise Search
I receive a phone call or two a month asking me about Oracle SES10g. (When you access the Oracle Web site, be patient. The system was sluggish for me on June 14, 2008.)The drift of these calls boils down to one key point, “What’s Oracle’s share of the enterprise search market?” The answer is that its share can be whatever Oracle’s accountants want it to be. You see Oracle SES10g is linked to the Oracle relational database and other bits and pieces of the Oracle framework. Oracle’s acquisitions in search and retrieval from Artificial Linguistics more than a decade ago to Triple Hop in more recent times has given Oracle capability. As a superplatform, Oracle is a player in search. So far this year, Oracle has been moving forward slowly. An experiment with Bitext here and a deployment with Siderean Software there. Financial mavens want Oracle to start acquiring search and content processing companies. There are rumors, but so far no action, and I don’t expect significant changes in the short term.

Rumor 3: Microsoft Will Tidy Up Its Search Operations
This rumor suggests that Microsoft, a giant company with many barons and dukes controlling fiefdoms, can deploy one search solution. I don’t think that will happen quickly. The Certified Gold Partners who make better search systems than those available from Microsoft can rest easy for the foreseeable future. Search is too complicated in general and within Microsoft for a one-size-fits-all solution. I anticipate more search options, not fewer. Coveo, Exalead, ISYS Search Software, and others will benefit from the Microsoft approach to search for months, if not years.

Rumor 4: Semantic Search Will Unseat Google
Semantic technology is now within reach of almost any search and content processing vendor. The technology is relatively well known and the processing power is available at a reasonable cost. By itself, semantic search will not be enough to shift the market share that Google is amassing in the consumer search and enterprise markets. Google’s been chugging along for a decade, and it has yet to meet significant competition other than itself. Semantic technology is a component, not a Google killer in the hands of a competitor at this time.

Rumor 5: Search Will Become a Commodity
No, as I described in my Web log post on May 12, 2008, about the “search elephant”, search has too many different meanings for one solution to sweep the board. Each unit of a company has many different search and content processing needs. It is, therefore, difficult to convince the legal department to use the open source Lucene tool for eDiscovery. The legal eagles will want to use a service from Brainware or Stratify. Down the hall, the chemical engineers need to find chemical structure. Search consists of niches, and these will bump heads, overlap, and been quite confusing to sort out. In that confusion, consultants and different vendors thrive.

Rumor 6: Search Is a Component of Other Enterprise Software
This is a rumor related to “search will become a commodity”. True, enterprise software vendors will include more robust search and content processing systems in their software, market the heck out of the enhancements, and bundle it with whatever the client wants to buy. But enterprise applications open the door to point solutions that meet specific needs. So search certainly will become ubiquitous and the ecosystem will spawn new species of information access. Nope, search is going to be with us for a long, long time.

Rumor 7: The Google Search Appliance Doesn’t Work
False. The GOOG has more than 10,000 licensees, a fleet of partners, and the OneBox API that can make the Google Search Appliance work like Roy Roger’s prescient horse, Trigger. The GOOG has had an impact on the enterprise search market. It’s easy to complain about the Google Search Appliance. It’s harder to explain how a company with such an interesting approach to sales can sell such a large number of units. Obviously a certain sector of the market wants these Google boxes.

Rumor 8: Social Search Will Revolutionize Enterprise Search
Nope. Social functions can be useful, but in regulated industries, there are some challenges associated with social search. Social search is the equivalent of a restaurant’s weekly special. If the customers gobble enough of the dish, the special could be promoted to a specialty. It’s early days for social search in an enterprise, but it’s not too soon for law enforcement and military intelligence people to embrace the concept. Social search is quite useful in certain areas, but one needs to have lots of social data to crunch to see the technology in full flower.

Rumor 9: Key Word Search Is Dead
Key word search is hard for many people. Alternatives and options are needed. But key word search is too useful in certain types of research to go the way of the dodo. Investors like to think that a whizzy interface without a search box is the next big thing in search. Interfaces are becoming more important by the hour. But an interface without a way to look for words and phrases won’t carry the day.

Rumor 10: A Hardware Maker Will Put Search on a Chip
What’s happening is research and investigation. The Exegy appliance could be boiled down to a smaller gizmo. At some point in the future, any search appliance could be reduced to firmware. I think the likelihood of search on a chip is high, but it’s not something you will be able to buy in 2008. Intel invested in Endeca for a reason. Intel had a brief love affair and then a messy divorce with search vendor Convera years ago. Other chip-centric outfits are poking around in this area as well. On the horizon, yes, but appliances will be about as close to search in a single package that we will have in 2008.

Observations
Most people don’t realize that search is like a giant jello dessert. There is a shimmery, attractive quality to the whole thing. When you start to pick it apart, the substance becomes slippery and tough to pin down. It’s easy to be fooled by surface changes like semantic search and social search, which are like a squirt of whipped topping on the jello. Do you have candidates for rumors you think I should have included in my round up. If so, use the comments section of this Web log to post your favorites. To avoid legal hassles, I may have to edit some of your inputs. Ah, life in the modern world is so rewarding.

Labels: , , , , , , ,