This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Thursday, February 12, 2009

The New Generation of eDiscovery Search

Train Leaving the Station The New Generation of eDiscovery Search technology train is getting ready to leave the station. However, after walking the tradeshow floors and attending many of the breakout sessions at last weeks LegalTech in New York, it is obvious that there is a tremendous amount of confusion regarding the definition and scope of the New Generation of eDiscovery Search technology and more importantly, how the courts view the use of such technology.

With the accelerating volume of Electronically Stored Information (ESI) or what I like to call Electronically Stored Evidence (ESE), the current legacy search technologies built into the current legacy eDiscovery tools and the associated best practices for document review are beginning to have a hard time "keeping up". Further, there is a tremendous amount of confusion and trepidation among litigators in regards to potential malpractice claims, sanctions and adherence to Rule 702 and Daubert challenges associated with employing the New Generation of eDiscovery Search technology. Finally, litigation technology vendors, whether purposely or not, have confused the market with fancy new marketing terms like "conceptual search", "transparent search", "linguistic search" and "clustering" without any real explanation of how they work and how to correctly employ them. Therefore, I thought that it was time to restart the campaign to both educate and lobby the eDiscovery industry in regards to the New Generation of eDiscovery Search.

First of all, I want to start with education regarding the pertinent issues. Without a doubt, one of the best article posted over the past 12 months on the legals issues surrounding the topic of what I am calling the New Generation of eDiscovery Search, was written by By Wayne C. Matus and John E. Davis in the New York Law Journal on October 31, 2008, titled, "Do Your Searches Pass Judicial Scrutiny?"

Since it is my impression that many of these very important issues are still either unknown to most in the eDiscovery "business" or are being ignored, I contacted Mr. Matus this week to get permission to repost his article.

Following is the full text of of "Do Your Searches Pass Judicial Scrutiny?":
Electronically stored information is increasing exponentially, and bills from law firms and discovery vendors to deal with this vast sea of data escalate significantly each year. Jason Baron, the director of litigation at the National Archives and Records Administration, believes that ESI is growing so fast that even with unlimited funds and human resources it will soon be impossible for humans to review these large document populations.[FOOTNOTE 1] Still, lawyers faced with potential malpractice claims and sanctions are loath to try new methods for handling the problem. It is time for change.

The traditional means used by litigators to address ESI is the application of keywords and Boolean search terms to identify relevant and non-privileged materials.[FOOTNOTE 2] While acknowledging that this method is unquestionably deficient, a recent article published in this publication concluded that "the available evidence suggests that keyword and Boolean searches remain the state of the art and the most appropriate search technology for most cases."[FOOTNOTE 3] We agree that, in a perfect world, if the parties can nevertheless meet and confer, and agree upon keywords to reduce the population to manageable proportions, the traditional judgmental method can be made to work. However, this is an imperfect world where plaintiffs and defendants do not always agree, and are not always equally motivated, to reduce costs. In fact, it is often quite the opposite. Moreover, even where the sides use judgmental sampling to agree upon keywords, the costs nevertheless usually remain too high.

THE JUDGMENTAL APPROACH

The judgmental approach to keywords ultimately fails because of "recall" and "precision." "Recall" measures how completely a process captures target data. "Precision" measures efficiency - the amount of irrelevant data captured along with the target data. Keywords, as judgmentally used by lawyers, recall too little, while capturing much that is irrelevant. An early landmark empirical study by David Blair and M.E. Maron[FOOTNOTE 4] showed that while lawyers thought they were retrieving about 75 percent of the relevant data, the true results were more like 20 percent. A subsequent study, conducted by the Text Retrieval Conference,[FOOTNOTE 5] confirmed this result, finding that only 22 percent of relevant documents were recalled using keyword search techniques, as opposed to approximately 78 percent found by other search techniques.[FOOTNOTE 6] Many lawyers will also tell you that it is common for reviewers to find only 10 to 40 percent of the recalled documents to be relevant, meaning lawyers are reading mostly junk.
We advocate two different approaches to yield better and more efficient results. First, we suggest that keywords are best used coupled with statistical, rather than judgmental, sampling. Second, we suggest that experienced counsel and vendors working with a combination of advanced conceptual search techniques can more efficiently and effectively deal with large amounts of ESI, resulting in a narrowed and enriched review set with a concomitant reduction in lawyer hours.

KEYWORDS DONE RIGHT

In Victor Stanley v. Creative Pipe, 250 FRD 251 (D. Md. 2008), Chief Magistrate Judge Paul W. Grimm of the U.S. District Court for the District of Maryland found counsel had waived the attorney-client privilege as to 165 inadvertently produced documents -- despite the use of 70 separate keyword searches in conducting their privilege screen -- because, among other reasons, counsel had failed to conduct "quality assurance testing." Clearly, judgmental sampling did not pass judicial scrutiny, while statistical sampling would likely have.

Counsel seeking to conduct a proper keyword search should instead consider the following steps:

• Sample the data. Counsel should isolate a random and statistically significant sample of the relevant datasets and then conduct a manual review of such data for relevance and privilege.[FOOTNOTE 7] This will educate counsel as to what to expect from the larger population and help in formulating keywords.
• Analyze and rank keywords. Counsel should then create and run search terms against the sample set, and (based on the information derived from the sample review) analyze their effectiveness by "recall" and "precision." This preliminary knowledge of the contents and richness of particular datasets will provide the basis to predict retrieval and review costs, and whether, for example, counsel should conduct any or just a limited review of such data.[FOOTNOTE 8]
• Review and repeat until satisfied that the search plan is defensible. This approach is plainly iterative in nature; it is the rare set of searches that achieves acceptable returns without adjustment. Successive application and fine-tuning of terms should permit counsel to achieve defensible levels of recall with superior precision rates. Practitioners should take note: One of the main factors cited by the court in Victor Stanley to determine if a party has conducted a reasonable search is if the party has reviewed a sample of the results to "assess its reliability, appropriateness for the task, and the quality of its implementation."[FOOTNOTE 9]
There is no consensus as to what percentage of recall will pass muster. Instead, counsel must be able to explain to the court the "reasonableness" under the circumstances of each step of the process, including the point at which a party was satisfied with the effectiveness of its search terms.[FOOTNOTE 10]

KEYWORDS: TO DISCLOSE OR NOT

Search terms created by counsel are generally protected, at least initially, by the attorney work-product doctrine, as their "mental impressions, conclusions, opinions, or legal theories ... concerning the litigation."[FOOTNOTE 11] But how does one show "reasonableness" of the search methodology without disclosing the keywords? Three recent cases, Victor Stanley, O'Keefe and Equity Analytics,[FOOTNOTE 12] have required that attorneys be able to explain and defend to the court, at its request, the methodology used to employ the search terms. One court has indicated it might find a waiver of privilege and require disclosure of the search terms.[FOOTNOTE 13]

OTHER FILTERING TECHNOLOGIES

Magistrate Judge John M. Facciola of the U.S. District Court for the District of Washington, D.C., recently pointed to authority that "concept searching" applications -- which use statistical and linguistic models to search for ideas as well as words and impose order upon disparate documents -- are "more efficient and more likely to produce comprehensive results" than keyword or Boolean searches.[FOOTNOTE 14] For example, the TREC 2007 Legal Track study found that 78 percent of relevant documents in a dataset were not found by Boolean keyword searches, but only by alternative search techniques.[FOOTNOTE 15] It is little wonder that Magistrate Judge Grimm has expressed optimism that concept-based searches studied by TREC would supplant keywords as the preferred method "for a variety of ESI discovery tasks."[FOOTNOTE 16]

These findings appear to have been confirmed. Earlier this year, the eDiscovery Institute disclosed its preliminary assessment of the study it conducted on the performance of computerized document review against human review. The study was conducted against a dataset drawn from the Verizon-MCI merger consisting of 1.3 terabytes and over two million documents. They concluded that computer systems allowed a comparable level of performance to be achieved with fewer people, less time and lower cost. While actual cost of traditional review was over $13.5 million, computer-assisted review was projected to cost just a fraction of that amount.[FOOTNOTE 17]

Concept search methodologies fall into three basic (and sometimes overlapping) categories:

• Probabilistic. This technique relies upon probabilistic search models such as "Bayesian classifiers," which evaluate and classify documents based on the interrelationships, proximity and frequency of usage of words found therein. The model may be given a "head start" by a sample set of relevant documents developed by attorneys at the outset of the process, which the computer analyzes and applies to the remaining documents. This technique can order groups and documents based on perceived potential importance to assist in the review process.
• Rule-based (or "clustering"). This statistically driven process analyzes the prevalence of words in documents and, based on such analysis, groups together documents interpreted as featuring like concepts. This technique can order documents by perceived potential importance as well.
• Linguistic. Sometimes referenced as "fuzzy search models," this technique seeks documents containing all forms of a target word or its synonyms in a general and/or case-specific thesaurus. Linguistic approaches may also rely upon statistics to analyze documents for terms along the same subject lines -- or sometimes to identify documents that use different ways to make the same point.[FOOTNOTE 18]
Differing tools often produce differing results, but some combination of each of these approaches (as well as Boolean keyword searches) -- using a transparent, iterative and measured process as described above -- can be used to best effect. As successive waves of ESI are received (as is often the case), moreover, certain of the concept-searching applications "learn" and become better at identifying correlations, associating documents with particular attributes with concepts of interest to counsel and minimizing false positives. Further, the statistics generated by this process permit counsel to draw educated lines as to where review should proceed and, sometimes more importantly, when it is reasonable to stop. The advantages of these powerful, computerized techniques become even more apparent where ESI reaches the terabyte range and the steep recall/precision tradeoff exhibited by keyword analyses may reach unacceptable levels. Two jurists have indicated in opinions an interest in hearing from experts as to such new approaches.[FOOTNOTE 19]

CONCLUSION

Given escalating volumes of ESI, with no end in sight, and the general impatience of courts with e-discovery mistakes, counsel and their clients soon may have no choice but to adopt discovery tools that are more efficient and precise than traditional Boolean search techniques. Courts have already put practitioners on notice of this emerging obligation. Combining measured approaches to search methodologies with advanced techniques can greatly assist in the organization of ESI and the cost-effective conduct of litigations and investigations. The future is now for these state-of-the-art search techniques.
Wayne C. Matus is a litigation partner in the New York office of Pillsbury Winthrop Shaw Pittman and one of two national leaders of the firm's e-discovery practice. John E. Davis is a senior associate in the firm's New York office specializing in e-discovery. Sandra Barragan, an associate at the firm, assisted in the preparation of this article.

::::FOOTNOTES::::

FN1 "EDD Showcase: Discovery Overload," Law Technology News, January 2008.
FN2 Although keyword searches and Boolean term searches are undeniably distinct, for purposes of this article we will refer to them interchangeably.
FN3 See "Assessing Alternative Search Methodologies," H. Christopher Boehning and Daniel J. Toal (NYLJ, April 22, 2008).
FN4 "An Evaluation of Retrieval Effectiveness for a Full-Text Document Retrieval System," Communications of the Association for Computing Machinery at 289-99, March 1985.
FN5 TREC is sponsored by the National Institute of Standards and Technology (NIST) and the Advanced Research and Development Activity of the Department of Defense.
FN6 These were the results of TREC 2007, the second year of the Legal Track study. See Jason R. Baron, Douglas W. Oard, Paul Thompson & Stephen Tomlinson, Overview of the TREC-2007 Legal Track, at §6 (linked at http://trec-legal.umiacs.umd.edu/). In a prior TREC-6 Ad Hoc Task study, for keywords to achieve just 50 percent recall, the architects had to accept a dismal 20 percent precision rate (whereby four of every five documents selected by keywords were nonresponsive). See H5 White Paper, Concept Search: Perceived Security, Actual Risk, at 2, citing Voorhees, Ellen M., and Harman, Donna, Overview of the Sixth Text REtrieval Conference (TREC-6), in NIST Special Publication 500-240: The Sixth Text REtrieval Conference (TREC 6), ed. E.M. Voorhees and D.K. Harman, 1-24 (Gaithersburg, MD: NIST 1997), and Voorhees, Ellen M., and Harman, Donna, Overview of the Seventh Text REtrieval Conference (TREC 7), ed. E.M. Voorhees and D. K. Harman, 1-24 (Gaithersburg, MD: NIST 1998).
FN7 E.g., Treppel v. Biovail Corp., 233 FRD 363, 374 (SDNY 2006).
FN8 See McPeek v. Ashcroft, 212 FRD 33, 35 (D. D.C. 2003) (ordering sampling of backup tapes to determine whether they contained relevant documents); Wiginton v. DB Richard Ellis Inc., 229 FRD 568, 570 (N.D. Ill. 2004) (ordering sampling of archived material based on keywords to determine whether it should be restored); see also Victor Stanley, 250 FRD at 261, citing The Sedona Conference Best Practices Commentary on the Use of Search & Information Retrieval Methods in E-Discovery, 8 Sedona Conf. J. 189 (2007) [hereinafter, "The Sedona Best Practices"]. For example, the sampling process may reveal that certain data sources (such as local hard drives) or file types (such as Microsoft Access files) have such low yield that the collection and review effort is not worthwhile. While opposing counsel may not agree to such decision, the data provided by the sample review will provide the evidentiary support needed to defend the reasonableness of such steps to the court.
FN9 Victor Stanley, 250 FRD at 256.
FN10 See, e.g., Security Financial Life Insurance Company v. Dept. of Treasury, 2005 WL 839543, *4 (D. D.C. April 12, 2005) ("In deciding whether an agency's document search is adequate, the issue is not whether other responsive records might possibly exist, but whether the search was adequate, judged by a reasonableness standard.") (internal citations omitted); see also Victor Stanley, 250 FRD at 261 n.10 ("the cost-benefit balancing factors of [FRCP] 26(b)(2)(c) apply to all aspects of discovery").
FN11 Fed. R. Civ. P. 26(b)(3); see Lockheed Martin Corp. v. L-3 Comm. Corp., 2007 WL 2209250 (M.D. Fl. July 29, 2007) ("documents containing instructions about how to conduct the [ESI] search and what specifically to search for are opinion work product" and therefore protected as attorney work product privileged material); see also Gibson v. Ford Motor Co., 2007 WL 41954, at *6 (N.D. Ga. Jan. 4, 2007) (document retention notice that included a list of search terms reflected attorney mental impressions and so constituted protected work product).
FN12 Victor Stanley, 250 FRD at 256, United States v. O'Keefe, 537 F.Supp.2d 14 (D. D.C. 2008), and Equity Analytics, LLC v. Lundin, 248 FRD 331 (D. D.C. 2008).
FN13 Counsel, early in the process, should consider disclosure of keywords and other aspects of the search protocol to the adversary and the court, and invite their comment and approval, as a means of managing discovery costs and risk. Magistrate Judge Grimm in Victor Stanley Inc., 250 FRD at 256, found, among other things, that defense counsel's failure to disclose the keywords used to screen for privileged documents in defending its search methodology justified a finding of waiver as to the inadvertently produced documents. The court provided a checklist for attorneys to follow when preparing the methodology to be used to gather and produce ESI: Attorneys should consider the reasonableness of "the keywords used; the rationale for their selection; the qualifications of the [creators of the search] to design an effective and reliable search and information retrieval method; whether the search [is] a simple keyword search, or a more sophisticated one, such as one employing Boolean proximity operators ... ." Id.
FN14 Disability Rights Council v. Washington Metropolitan Transit Authority, 242 FRD 139 (D. D.C. 2007), citing George L. Paul & Jason R. Baron, "Information Inflation: Can the Legal System Adapt?" 13 Rich. J.L. & Tech. 10 (2007).
FN15 TREC 2007 Legal Track.
FN16 Victor Stanley, 250 FRD at 261 n.10.
FN17 http://www.ediscoveryinstitute.org/research/index.html.
FN18 Such tools are described in further detail in The Sedona Best Practices at 191-216 & Appendix. While we are unaware of a court that has expressly endorsed this approach, parties have used these search techniques in conducting, among other things, internal investigations.
FN19 Magistrate Judge Grimm in Victor Stanley, 250 FRD at 260, and Magistrate Judge Facciola in O'Keefe, 537 F.Supp.2d at 24, and Equity Analytics, 248 FRD at 333, have indicated that conducting and defending e-discovery may at times require experts. Indeed, Magistrate Judge Facciola stated that search methodologies in e-discovery may be scrutinized under Rule of Evidence 702.


Labels: , , , , , , , , , ,

Saturday, May 10, 2008

In Search of Integrated Conceptual eDiscovery Search Technology

Over the past 6 months I have been investigating cost effective, integrated conceptual eDiscovery search technology delivered under a SaaS model. The basis for this investigation is to identify a way to extend the current capabilities of eDiscovery search through a forward thinking search technology that can be tightly integrated on the same Microsoft stack based eDiscovery platform with email archiving and other proactive data retention technology, Electronic Data Discovery (EDD) software and an Online Review Tool (ORT). My finding are that the current state of forward thinking search technology is such that it requires the support of a separate and proprietary database and therefore does not lend itself to integration with EDD and ORT platforms that sit on standard SQLServer solutions.

Where this current state of the market leaves the user is with a choice of either moving large amounts of data or least large amounts of index files and associated data between platforms or investing in a completing propriety eDiscovery solution.

In the process of this investigation, I have found several outstanding articles that touch on the various topics incumbent in this discussion. The first article, found on Law.com, titled "In Search of Better E-Discovery Methods" by H. Christopher Boehning and Daniel J. Toal, does an excellent job of discussing some of the standard criteria for new search technology and whether or not it surpasses currently available keyword and Boolean search technology.

The second article is actual a Blog posting by Cher Devey, titled "Alternative Search Technologies - Too Good to be True" on her "eDiscovery Myth or Reality?" Blog. Ms. Devey discusses the concept and viability of human intervention into the search process. (Please note that the full text of Ms. Devey's Blog Post can be found at the bottom of this posting).

The full text of Mr. Boehning's and Mr. Toal's article is as follows:

As the burdens of e-discovery continue to mount, the search for a technological solution has only intensified. The holy grail here is a search methodology that will enable litigants to identify potentially relevant electronic documents reliably and efficiently.

In an effort to achieve these often competing objectives, litigants most commonly search repositories of electronic data for documents containing any number of defined search terms (keyword searches) or search terms appearing in a specified relation to one another (Boolean searches). These search technologies have been in use for years, both in litigation and elsewhere, and accordingly are well understood and widely accepted by courts and practitioners.

But keyword and Boolean searches are far from perfect solutions; they are blunt instruments. Such searches will identify only those electronic documents containing the precise terms specified. These methodologies therefore will not catch documents using words that are close, but not identical, to the specified search terms, such as abbreviations, synonyms, nicknames, initials and misspelled words.

On the other hand, using more search terms may reduce the risk that an electronic search will miss a relevant document, but only at the price of increasing -- often quite dramatically -- the number of irrelevant documents found in the search. This is a serious problem because counsel must manually review whatever documents the searches yield in order to sift out non responsive materials, make privilege determinations and designate confidential documents. Keyword and Boolean searches thus require a careful balance to be struck: Unduly restrictive searches may miss too many responsive documents while over broad searches threaten stratospheric discovery costs.

Against this backdrop, courts and litigants understandably have been intrigued by the claims of those promoting alternative search technologies, such as "concept searching." The vendors of such technologies suggest their search strategies are able to identify the overwhelming majority of responsive documents while virtually eliminating the need for lawyer involvement in the review process.

Such claims strike many in the legal community as too good to be true. And their skepticism is appropriately heightened because the precise methodologies that such vendors use often are shrouded in mystery, owing to their stated desire to safeguard their proprietary processes and techniques. But this also means their tantalizing claims cannot readily be subjected to independent scrutiny. The question thus posed -- and still largely unexplored -- is whether these alternative search technologies have anything to offer and, if so, how best to evaluate the competing technologies and the often sensational claims of their promoters.

To evaluate whether an alternative search technology might be helpfully employed in any particular case, it is first essential to understand how it works. Some of the principal alternative search technologies, which fall under the broad heading of "concept searching" methodologies, are as follows:

Clustering. Whereas keyword and Boolean searches mechanically apply certain logical rules to identify potentially relevant documents, clustering relies on statistical relationships, which results in documents containing similar words being clustered together in relevant categories. The clustering tool compares each document in a pool to "seed" documents, which have already been designated as relevant. The more words a document has in common with a seed document, the more likely it is to be about the same subject and therefore to be responsive. Moreover, clustering tools generally rank documents based on their statistical similarity to the seed documents.

Taxonomies and ontologies. A taxonomy tool is used to categorize documents containing words that are subsets of the topics relevant to a litigation. For example, if one of the topics of interest is "dogs," a taxonomy tool would capture documents that mention "golden retrievers," "poodles" and "chihuahuas." Ontology tools perform similar searches, but are not confined to identifying subset relationships. Building on the last example, an ontology tool would capture documents that mention "kennels" or "veterinarians."

Bayesian Classifiers. Bayesian search systems use probability theory to make educated inferences about the relevance of documents based on the system's prior experience in identifying relevant documents in the particular litigation. The search results then would be ranked based on the predicted likelihood of their relevance to the litigation.

HOW APPROACHES COMPARE
These alternative search technologies may sound promising in concept, and the claims about their efficiency and accuracy likely add to their allure, but the question remains whether these approaches outperform the standard search approach.

Keyword searching (including with the use of Boolean connectors), its acknowledged limitations notwithstanding, has secured such widespread acceptance for a reason. As an initial matter, the technology and search methodology is well understood and familiar to anyone who has used Westlaw, Lexis or similar search engines. It therefore can be easily discussed with both opposing counsel and judges. The simplicity of keyword searching also doubtlessly promotes negotiated resolution of discovery disputes because the parties have less reason to fear that ignorance about the technology will lead them to strike a bad bargain.

But the simplicity of keyword searching is also its principal weakness. Keyword searches capture only documents containing the precise terms designated, which virtually assures that such a search will miss relevant documents. And, on the other side of the equation, keyword searches will mechanically capture every document -- whether relevant or not -- containing any search term. This means keyword searches may be both substantially under- and over-inclusive. Concept searching systems, by contrast, are not dependent on a particular term appearing in a document and therefore may locate documents a Boolean search would not. But they may suffer from other infirmities.

So how does concept searching stack up? The best evidence to date comes from the Text REtrieval Conference, which in 2006 designed an independent research project to compare the efficacy of various search methods. In view of the prevalence of keyword and Boolean searches in litigation today, TREC was particularly interested in determining whether the alternative search methodologies outlined above were better than Boolean.

As its starting point, the TREC study used a test set of 7 million documents that had been made available to the public pursuant to a Master Settlement Agreement between tobacco companies and several state attorneys general. Attorneys assisting in the study then drafted five test complaints and 43 sample document requests (referred to as topics). The topic creator and a TREC coordinator then took on the roles of the requesting and responding counsel and negotiated over the form of a Boolean search to be run for each document request.

In addition to the Boolean searches, computer scientists from academia and other institutions attempted to locate responsive documents for each topic utilizing 31 different automated search methodologies, including concept searching. The results were striking. On average, across all the topics, the negotiated Boolean searches located 57 percent of the known relevant documents.

But none of the alternative search methodologies reliably performed any better. That is to say, for each topic, the Boolean search did about as well as the best alternative search methodology.

Interestingly, although the Boolean searches generally outperformed the alternative search protocols, the methods did not necessarily retrieve the same responsive documents. In fact, when all of the responsive documents found by the 31 alternative runs were combined, TREC discovered that the alternative search runs collectively had located, on average, an additional 32 percent of the responsive documents in each topic.

As a result, while the Boolean search generally equaled or outperformed any of the individual alternative search methods, those searches also captured at least some responsive documents that the Boolean search had missed.

COST ANALYSIS
This suggests that even if alternative search methodologies have not yet been shown to beat Boolean searches, their use to supplement Boolean searches might increase the number of responsive documents located. But at what cost? The potential benefits of locating any additional documents through use of an alternative search methodology would still have to be weighed against the cost, both in money and resources, required to locate them.

The relevant cost here is not just the price of using the alternative search technology, but also the number of false positives identified by the approach (i.e. documents retrieved by the search, but turn out not to be responsive). Any automated search method -- whether a keyword or concept search -- will yield false positives, which counsel must review and filter out prior to production, which can be a costly process. It therefore is far from clear that use of an alternative search methodology in addition to a keyword or Boolean search will be appropriate in any particular case, a question the TREC study does not attempt to address.

For now, the available evidence suggests that keyword and Boolean searches remain the state-of-the-art and the most appropriate search technology for most cases. This seems particularly true when keyword or Boolean searches are used in an iterative manner, where litigants: (i) negotiate search terms and Boolean operators, (ii) run the agreed-upon searches, (iii) review the preliminary results, and (iv) adjust the searches through a series of meet-and-confers. This type of "virtuous cycle of iterative feedback" has been endorsed by courts and commentators alike.

The intuition of the legal community that an iterative approach to electronic discovery promotes reliability and efficiency finds empirical support in the TREC study. As part of its study, TREC employed an expert tobacco document searcher who used an "interactive" search methodology.

TREC found that the expert searcher located, on average, an additional 11 percent of the relevant documents beyond those that had been located by the initial Boolean searches, which means that an interactive Boolean approach ultimately located 68 percent of the relevant documents -- far better than any of the alternative search methodologies.

CONCLUSION
It may be that alternative search methodologies eventually will surpass the performance of keyword and Boolean searches, but that day does not yet seem to have arrived.

The independent research conducted to date suggests that, for the time being at least, nothing beats Boolean, particularly when used as part of an iterative process.

That does not necessarily mean that alternative search technologies are not worth considering, either independently or along with Boolean or keyword searches. But practitioners would be well advised to carefully scrutinize the marketing claims of the purveyors of such technologies and to factor in often substantial direct and indirect costs of such approaches.

H. Christopher Boehning and Daniel J. Toal are litigation partners at Paul, Weiss, Rifkind, Wharton & Garrison. Associate Jason D. Jones and Aaron Gardner, the firm's discovery process manager, assisted in the preparation of this article.

The Full Text of Ms. Devey's Blog Posting is as follows:

It seems that alternative search technologies (alternative to the familiar Keyword and Boolean searches) touted by Vendors are considered as ‘too good to be true’. Check it out yourself at In Search of Better E-Discovery Methods By H. Christopher Boehning and Daniel J. Toal, New York Law Journal April 23, 2008

The above legal article also mentioned the Text Retrieval Conference (TREC) 2006 study which was also examined by Will Uppington in the article, Better Search for E-Discovery, March 11th, 2008

What I find interesting in Will Uppington’s article is the finding; ‘One of the best ways to get better search queries is to commit human resources to improving them, by putting a “human-in-the-loop” while performing searches’.

Reading in between these two ‘search themed’ titles, one from the legal side and the other from a technical perspective, highlighted the contrasting findings and interpretation on the TREC 2006 study

What else can we say/talk about the ‘human-in-the loop’, the ‘virtuous cycle of iterative feedback’ & “interactive” search methodology?

Well such phrases/concepts are not new. What is new is that the ‘human actions’ aspects are creeping (awareness?) into the ediscovery space. Other knowledge researchers outside the ediscovery domain have been busily coming up with phrases/concepts such as the ‘concept searching’ methodologies. Reality (or inertia adoption) testing of such newer technologies are clearly not well understood (too good to be true?) by the courts and practitioners.

On human actions and computer programs, a beautiful quote comes from my friend, Roger C: “While computer programs can write other computer programs, they can’t write the first program”.

Labels: , , , , , ,