This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Monday, July 12, 2010

Is Autonomy the Undisputed Heavy Weight Champion?

I spent the first segment of my career (please note that I didn’t use the word half as it is just too depressing) in the enterprise software market competing against Oracle with a variety of startups and early stage organizations and fielded a few consulting organizations that relied on non-Oracle database systems (e.g. Microsoft SQL).  I am citing this history as experience to draw an analogy between the early days of “dealing with” the market perception of  Oracle to the current phenomena of the market considering Autonomy as the undisputed heavy weight champion of  the enterprise search and content management world for eDiscovery and Governance, Risk and Compliance (GRC).

And, as with the crowing of any champion, there are supporters and detractors on both sides spewing innuendo, have truths, unsubstantiated case studies, etc.  Its almost as exciting as a national election or the World Cup (well maybe not the World Cup).
As an example, in a July 12, 2010 Blog post by Stephen Arnold titled, “Autonomy: A Real Success. CMSWatch: Maybe Another Real Miss?”,  Mr.  Arnold takes exception to a Blog post from the CMSWatch Blog by Tony Byrne.  Both of these posts are excellent as they “stir the Autonomy pot” and therefore I will include them at the end of my post.  However, I think that, given Mr. Arnold’s agnostic nature, it is interesting that he seems to an advocate for Autonomy.

Mr. Arnold states that Autonomy is on track to hit $1.0 billion by the end of calendar 2010. The company has a proven track record of improving the performance of the companies it acquires. Autonomy’s management has demonstrated its ability to integrate quickly its acquired products with IDOL (the firm’s integrated data operating layer). The result is Autonomy’s knack of transforming the acquired companies’ position in their markets.
He goes on to state that there are other data that shed light on Autonomy’s track record, which I have documented Autonomy’s technology in my writings such as Beyond Search (Gilbane, 2009), the Enterprise Search Report (CMSWatch.com, 2004-2006), and Successful Enterprise Search Management (Galatea, 2009). Here are three points that must not be overlooked:
  1. Autonomy has 20,000 plus customers plus around 1,000 licensees of its technologies for use in other enterprise software and systems
  2. Autonomy has made intelligent acquisitions that has given the firm a strong presence in eDiscovery, rich media, and fraud detection. Autonomy has recently pushed into online marketing using capabilities from Ineterwoven and its IDOL framework. My research reveals that Autonomy has acquired companies to bring its technology to new markets so more content can be understood.
  3. Autonomy has grown its revenues and generated a profit, making it possible for other UK based technology companies to ride the Autonomy horse in the race for government and venture funding.
I am not going to dispute any of Mr. Arnold’s statements as that is not the point of my post (please note that later on I will dispute the fact that all of Autonomy’s clients are happy).  I am just using these posts as proof that the same arguments that I heard about Oracle are now beginning to surface about Autonomy.  Let’s list a few of them:
  1. Oracle/Autonomy has so many clients and licenses that why would a prospect even consider another vendor?
  2. Oracle/Autonomy has acquired all of the best new technologies on the market so why would a prospect even consider another vendor?
  3. Oracle / Autonomy management is so obviously smarter than anyone else in the market and therefore why would a prospect even consider another vendor?
The market bought into this “line of thinking” and the Oracle sales teams perpetuated the myths and used it to their advantage.  In fact, I can remember prospects telling me that they really liked my Oracle “knock off” (their words not mine) and understood the technological superiority of my product and understood and appreciated the TOC and ROI arguments that I had developed.  But, in the end,  and as communicated to them by the Oracle sales executive, even if Oracle turned out to be the wrong decision for all the reasons that I had sighted, they wouldn’t have to worry about losing their job if they went with Oracle and therefore that is what they were going to do.    At an even more frustrating level, I can sight prospect after prospect that spent the better part of a great lunch meeting railing against Oracle and then after the check came and was paid, indicated that they were unfortunately going to buy another round of Oracle licenses.

Don’t take my stories as “sour grapes” as I did in fact have my fair share of  success selling against Oracle.  And, don’t forget that at one point, the standard Oracle line in the industry was “What in the heck does Microsoft know about building databases?”  We all know the answer to that question (well, maybe not everyone).  The point being that technology vendors like Oracle and now maybe Autonomy think that they can reach a point of being too big to fail and so big and stable that it would be crazy to even consider alternative solutions.

Well,  the technology highway is riddled with the wreckage of companies and technologies that thought that they owned a marketplace and then one day a couple of “guys in a garage” proved them wrong.   Please note that I am not underestimating what Autonomy has done, the insightful intelligence of their founders,  their market share,  their momentum , their marketing budget or the their massive sales force. Nor am I overestimating the value of new ‘garage based” and superior technology.  I am merely stating that, just like we all saw with competition infringing upon the coronation of Oracle, we may also have the opportunity to witness some amount of legitimate competition taking market share away from Autonomy.

Now as promised, some thoughts on the “happiness” of  Autonomy users.  Mr. Arnold states in another of his pieces on Autonomy that based on research by IDC, “If the data compiled for the report are accurate, Autonomy has a big footprint and happy customers. Among the thousands of Autonomy licensees are AOL, BAE Systems, BBC, Bloomberg, Boeing, Citigroup, Coca Cola, Daimler AG, Deutsche Bank, DLA Piper, Ericsson, FedEx, Ford, GlaxoSmithKline, Lloyds TSB, NASA, Nestle, the New York Stock Exchange, Reuters, Shell, Tesco, T-Mobile, the U.S. Department of Energy, the U.S. Department of Homeland Security and the U.S. Securities and Exchange Commission”.

Once again, I am not going to dispute what Mr. Arnold and IDC are saying.  However, I am going to offer an opinion and state some unscientific anecdotes that I have.  First of all, as with Oracle, many of the Fortune 500 have a license or two of just about everything available on the market and therefore it would not seem unusual for any of the organizations stated to have a copy of Autonomy software.

Second,  enterprise buyers are very politically savvy by nature and therefore one they make a decision to purchase something like Oracle or Autonomy, it would not be within their “political nature” to indicate publically they had made a poor decision.

Now, for the unscientific anecdotes.  I spent a good portion of just about every “working” day of my career talking with Fortune 2000 business line, IT and Legal Users about their needs (pain), past successes, past failure and plans for future purchases and “nirvana“ solutions.  Obviously, Autonomy has come up a time or two. And, without fail, the main themes that I have heard from Autonomy users are; (1) they don’t much like the proprietary nature of IDOL; (2) they think that it is too expensive; (3) Autonomy over promised under performed; (4) the integration of the acquired technologies didn’t seem to be working very well,  and;  (5) they were concerned that Autonomy had gotten so big and unwieldy that it (Autonomy) would have a very hard time being innovative in the future.  And, with the goal of full disclosure, I would have to say that, just as was the case with many Oracle clients that I talked to, most of the Autonomy users had plans to purchase additional licenses.

In summary, I am actually very impressed with what Mike Lynch has done and would not argue that they in fact be the undisputed heavy weight champion of the search and content management market.  However, I would also caution both current and potential users to not overestimate Autonomy’s ability to keep pace and no underestimate the power of the free market and its ability to nurture new search and content management technology support healthy competition.

After all , who would have ever thought that we would be using our iPhones or iTouches for eDiscovery (http://www.marketwire.com/press-release/iClearwell-Sets-Standard-as-First-E-Discovery-Companion-Application-iPhone-iPad-1288443.htm).

The full text of Mr. Arnold excellent Blog post is a follows:

In Harrod’s Creek, I can easily spot the real squirrel hunters. They have food. Mostly laconic, these hunters have a big pile of dead squirrels as proof of their competence. There is also the smell of fresh burgoo wafting from their log cabins. I can smell ability from my goose pond.


Lousy hunters have empty gun belts and squirrels shot when snacking on store bought food used to lure the critters. That’s a real danger — cheap tricks or just shooting wildly, often putting bird shot in an innocent’s backsides or the face like the 2006 incident between Vice President Dick Cheney and Texas lawyer Harry Whittington. Some faux hunters have just shot themselves in the foot. Ouch!

Azure chip consultants is a synonym for “bad hunter” in my opinion. Source: http://api.ning.com/files/LCP2NCaWo-ptCqGncB3hGsX8vuh8dnDzSJ0iLnkibas_/18holeinhandG.jpg

One of my two or three readers sent me a link to a write up called “Don’t Ogle Search If You Really Want Content Management”. In my opinion, the write up relies on insinuation, not facts. (I think that some folks are immune to facts, but I find facts useful.) In the article’s headline, the word “ogle”, for example, is one I don’t associate with information retrieval. (The publisher of this “ogle” opinion piece caught my attention in July 2008 with its similar assault on Attivio. My response to that misleading article is here.)

Yet another example of factless criticism of a vendor appears in this segment of the “ogle” write up about Autonomy, one of a very small number of search and content processing vendors with a consistent track record of technical breadth, sales, revenue, and profit:

From an initial focus on enterprise search tools, Autonomy has become a roll-up vendor after acquiring a variety of other information management suppliers such as Interwoven. As a financial strategy this can be successful, and investors seem to cotton to Autonomy. As a technology strategy, vendor roll-ups are problematic. Autonomy’s technology strategy is to rip legacy search subsystems from acquired products, replace them with some pieces from its own IDOL toolset, and then promote its particular approach to search as a distinct advantage for you. Specifically, Autonomy will try to sell you on the value of “meaning-based computing.” Even if you can get your mind around what meaning-based means, you should remain skeptical that Autonomy has technically spectacular or original services here. More importantly, you risk getting sidetracked from your original goal of, say, creating a user-friendly repository for your 50,000 Office documents.

These statements are presented without verifiable foundation to support the allegations in my opinion.

Autonomy is on track to hit $1.0 billion by the end of calendar 2010. The company has a proven track record of improving the performance of the companies it acquires. Autonomy’s management has demonstrated its ability to integrate quickly its acquired products with IDOL (the firm’s integrated data operating layer). The result is Autonomy’s knack of transforming the acquired companies’ position in their markets.

But there are other data that shed light on Autonomy’s track record, which I have documented Autonomy’s technology in my writings such as Beyond Search (Gilbane, 2009), the Enterprise Search Report (CMSWatch.com, 2004-2006), and Successful Enterprise Search Management (Galatea, 2009). Here are three points that must not be overlooked:

1.Autonomy has 20,000 plus customers plus around 1,000 licensees of its technologies for use in other enterprise software and systems
2.Autonomy has made intelligent acquisitions that has given the firm a strong presence in eDiscovery, rich media, and fraud detection. Autonomy has recently pushed into online marketing using capabilities from Ineterwoven and its IDOL framework. My research reveals that Autonomy has acquired companies to bring its technology to new markets so more content can be understood.
3.Autonomy has grown its revenues and generated a profit, making it possible for other UK based technology companies to ride the Autonomy horse in the race for government and venture funding.

In December a year or so ago, at the International Online Conference, in my for-fee, end note debate, I challenged Andrew Kanter (Autonomy), Charlie Hull (Lemur Consulting), and Dr. Charles Oppenheim (Loughborough University) about their views of search, content processing, and related fields. In front of an audience of about 300 search professionals, I pointed out that key word search was dead. I pointed out that most search systems did not understand the meaning of processed information. Autonomy’s Andrew Kanter strongly and politely disagreed with me. As I recall, he said to the audience and me:

Autonomy IDOL is the only product in the market that can understand the meaning and concepts of all information in any language, including audio and video. This has big implications for the content management market as no other vendor can do this.

I demanded some concrete examples to support his position. Mr. Kanter without missing a beat gave me four concrete examples drawn from Autonomy’s work in intelligence, search enabled applications, fraud detection, and rich media.

What did I do?

I listened, considered the evidence, and I conceded defeat. Facts convinced me.

Despite the proliferation of marketing baloney, facts about search and content processing are more important than insinuations and unsubstantiated generalizations. There is no excuse for any one to venture into a technical jungle without adequate preparation and forethought. Take a short cut at Booz, Allen & Hamilton, and you would have bene fired when Dr. William P. Sommers ran the the firm’s Technology Management Group. Furthermore, unsubstantiated assertions are the method of some high school journalists and third tier consultants in my opinion.

As readers of this blog know, I am no fan boy of any search and content processing vendor. Spend 15 minutes with me, and you will learn that I can and will point out fact-based strengths and weakness of the companies I monitor. (For a list of the firms I track, navigate to this link.) MBA double-talk and half-baked arguments create confusion. Verbal wild firing brings little benefit to those trying to understand the complexities of digital information.

Autonomy is on track to hit $1.0 billion by the end of calendar 2010. The firm has been able to acquire firms that add to Autonomy’s customer base and extend Autonomy’s meaning-based technology to new markets. Its Zantaz acquisition expanded Autonomy’s footprint in eDiscovery, boosted Autonomy’s position in cloud computing, and was a spark that set off a wave of activity in the eDiscovery and online archiving sector. That’s how you kill real squirrels. Take aim. Bang. New revenue.

In the search and content processing market, Autonomy has been a leader in marketing and on target acquisitions. In my opinion, OpenText (the Canadian company competing with Autonomy in some enterprise markets) has had to follow Autonomy in a “me-too” fashion in an effort to try to keep pace.

For me, Autonomy continues to grow in a landscape littered with failures. The more important question is, “Why isn’t Autonomy like Delphes, Entopia, Fast Search & Transfer, InQuire, Siderean, STAIRS III, and many other search vendors which have run aground?

There will be no answers in an analysis which lacks facts and insight.

And there is another important question, “Why shoot wildly at Autonomy?”

I don’t know. Grandstanding like squirrel hunters with an automatic rifle and a bag of walnuts? Fun? Frustration with life, business, or traffic to the company Web site? A need to differentiate one’s business from a farrow of third-tier consultants?

My hunch is that some “experts” see an opportunity to make money by appointing themselves mavens or satraps in search and content processing. Grabbing for a brass ring from a carousel pony seems harmless enough. The logic could run like this: I use Google. Google makes search easy. Therefore, search is easy.

What better way to get sales leads than to make unsubstantiated claims about a company that in terms of customers, financial performances, and scope is arguably one of the world’s leading vendors in information retrieval and processing?

Anyone looking at search market facts should be able to figure out that unsubstantiated assertions have zero impact on a billion dollar enterprise. Wild shots and tricks call attention to the hunter, not the prey. And what is the point of the “ogle” write up in my opinion: Marketing hoo-hah or a need for attention?

I don’t know the difference between slander and libel. I will leave that to legal eagles.

I do know that Autonomy has happy customers. Autonomy has hundreds of OEM deals that continue from year to year. Sue Feldman, IDC’s search expert, recent analysis of Autonomy was fact-based and positive. Also, I know that Autonomy is growing. What some search dilettantes do not appreciate is that search start ups in the UK have a better shot at getting funding due to the uplift from Autonomy’s success.

And there is the matter of profits.

Outfits like Goldman Sachs and other market makers pay attention to companies that make money. With your job on the line in a search procurement, would you go with a struggling vendor or with an outfit that was a market leader?

Does Autonomy have weaknesses?

What software vendor doesn’t? My Overflight service had a glitch on July 7, 2010. My goslings had to troubleshoot a mistake I made in a script. Software is difficult to make perfect. You can ask Google how the Buzz code is working out. Why not chase down Microsoft and ask about the Kin? If you have a moment, get Larry Ellison on the line and ask, “How is SES11g working out for you?”

To sum up: Constructive criticism based on solid technical understanding and demonstrable evidence delivers real value. Here in Harrod’s Creek, when I want burgoo I seek out the real hunters who hit their target without resorting to unsportsmanlike tricks. You may want to avoid the faux hunters who take short cuts, sport Travel Smith vests, and glittering generalities.

Labels: , , , , , ,

Friday, June 25, 2010

The Conceptual Search Game is Finally On!!

I have been writing (probably preaching to some) about advanced search technology, the differences between conceptual search and keyword search and the importance of advanced search technology in both the Early Case Assessment (ECA) and document review phases of eDiscovery for the past 3 years.

Concept Search Cash Law Emerging http://ediscoveryconsulting.blogspot.com/2008/06/concept-search-case-law-emerging.html
Concept Search vs. Keyword Search http://ediscoveryconsulting.blogspot.com/2008/12/concept-search-vs-keyword-search-in.html
Litigators Need ESI Analytics – Not Boolean Search Tools http://ediscoveryconsulting.blogspot.com/2010/05/litigators-need-esi-analytics-not.html

However, I have been somewhat disappointed in regards to the level of adoption of true conceptual search technology by the leading Litigation Technology vendors.  That appears to be changing.  As an example, in a June 17, 2010 post by StoredIQ titled, “Email Search: Nowhere to Hide”, the author provides an overview of StoredIQ’s search technology with specific focus on Natural Language Processing (NLP).  The reason that I am pointing this out is that 12 months ago, StoredIQ would not have been spending marketing dollars or Blog space on Natural Language Processing (NLP) because the market didn’t know what it was and didn’t care.

In this Blog post, StoredIQ now contends, “Probably of greatest interest to litigators during the discovery process is StoredIQ’s ability to perform natural language processing (NLP), which is the ability to extract linguistically derived natural language concepts from within email and user files including people, places and things. Legal teams can immediately search using over 250 out-of-the-box concepts and attributes including credit card accounts, social security numbers and stock symbols. NLP identifies word usage based upon context within a sentence. For example, NLP can identify if the word ‘will’ is used to identify a person’s name, a legal document or an auxiliary verb showing intent. StoredIQ has proprietary technology for adaptive sentence boundary disambiguation (ASBD) which substantially increases the precision of Natural Language Processing to address common grammatical deficiencies that are present in many business documents. No other information management technologies have this capability. NLP is a critical capability necessary to accurately perform eDiscovery, records management or risk management as full text indexes alone cannot provide the required level of precision.

Interestingly enough,  in my discussions with General Counsel and their litigation support teams from the Information Technology departments over the past 6 months, I have found a new awareness and appreciation for true conceptual search or semantic search or NLP.    So, StoredIQ is on the right track with their current product  offerings and I would bet that they have a product roadmap with more of the same.

So, I guess the conceptual search game is on and the other litigation technology vendors had better take notice of what their clients are saying in regards to what search technology they need.

The full text of the  StoredIQ Blog post is as follows:

A recent article by Jacob Goldstein, 23 Things Not To Write In An Email, illustrates the type of granularity as well as breadth of keywords that can be used by litigators during the legal discovery process to search for relevant information. He points out some keywords that may raise a legal red flag and should be used carefully when constructing emails. However, today’s technology search capabilities provide such precise, complete and accurate results, that there just isn’t anywhere to hide.
For instance, StoredIQ’s advanced search capabilities can look within compressed files, email archives and email attachments, in addition to the text contained in the email message itself. It can also search non-printable text within a document or email and can search through comments and revisions. In addition to search using keywords, StoredIQ supports many advanced search capabilities including:
  • Single term search
  • Multiple term search
  • Concept-based search
  • Boolean operators
  • Logical grouping of terms
  • Wildcards within search terms or Boolean expressions
  • Proximity searches
  • Natural language entities
  • Regular expressions
  • Macro-based searches
  • Object level attributes
  • By hash value (digital signatures)
Probably of greatest interest to litigators during the discovery process is StoredIQ’s ability to perform natural language processing (NLP), which is the ability to extract linguistically derived natural language concepts from within email and user files including people, places and things. Legal teams can immediately search using over 250 out-of-the-box concepts and attributes including credit card accounts, social security numbers and stock symbols. NLP identifies word usage based upon context within a sentence. For example, NLP can identify if the word ‘will’ is used to identify a person’s name, a legal document or an auxiliary verb showing intent. StoredIQ has proprietary technology for adaptive sentence boundary disambiguation (ASBD) which substantially increases the precision of Natural Language Processing to address common grammatical deficiencies that are present in many business documents. No other information management technologies have this capability. NLP is a critical capability necessary to accurately perform eDiscovery, records management or risk management as full text indexes alone cannot provide the required level of precision.
I know a lot of these terms can be a mouthful, but the underlying take away is that legal teams have the technology to precisely and accurately search electronic data, including email, making it much easier for litigators to discover data that was at one time hidden from them.

Labels: , , , , , , , , , ,

Thursday, June 12, 2008

Concept Search Case Law Emerging

In a followup to my post regarding my investigaiton of a SaaS based conceptual search technology for the eDiscovery market: http://ediscoveryconsulting.blogspot.com/2008/05/in-search-of-integrated-conceptual.html, I came accross an interesting article on the Law.com site by Leonard DeutchmanPennsylvania Law Weekly titled "When E-Discovery Is Put to the Test, Will federal rules on expert testimony govern admission of search engine results?".

This outstanding article discusses the issues of how the courts view search technology in light of Disability Rights Council v. Washington Metropolitan Transit Authority, 242 F.R.D. 139 (D.D.C. 2007) and United States v. O'Keefe, 537 F. Supp. 2d 14 (D.D.C. 2008) and Equity Analytics v. Lundin, 2008 U.S. Dist. LEXIS 17407 (D.D.C. Mar. 7, 2008).

Please note that I am still evaluating search technology and will post my finding as soon as I have completed my investigations. However, I felt as though this article raised so many pertinent and timely issues that I wanted to post it before my findings were complete.

The fulll article is as follows:

An influential federal district judge whose opinions on e-discovery are well respected may have set e-discovery on a path toward its most searching scrutiny yet.

In Disability Rights Council v. Washington Metropolitan Transit Authority, 242 F.R.D. 139 (D.D.C. 2007), Judge John M. Facciola recommended "concept searching," -- the use of complex search engines that make use of linguistic or statistical patterning to locate responsive e-mails and electronic -documents, in order for a tardy producer of discovery to wade through voluminous electronically stored information quickly. Interestingly, Facciola made no mention of whether the use of concept searching tools should be subject to Federal Rule of Evidence 702, which governs the admission of scientific or expert testimony.

Recently, however, in United States v. O'Keefe, 537 F. Supp. 2d 14 (D.D.C. 2008) and Equity Analytics v. Lundin, 2008 U.S. Dist. LEXIS 17407 (D.D.C. Mar. 7, 2008), Facciola held that any challenges to or defenses of search methodology in producing e-discovery must be scrutinized under Rule 702, and so ordered hearings under Daubert v. Merrill Dow Pharmaceuticals, 509 U.S. 579 (1993).

These rulings give rise to the question of what a Daubert hearing for an e-discovery search engine would look like.

HOW AND WHERE TO SEARCH
The first issue the court would address is how search engines search. The most direct approach is keyword searching, which take three basic forms:
Direct searching for keywords, e.g., "Locate all files with 'Jones.'"
Boolean searching, e.g., "Locate 'Jones' or 'Smith'," "Locate 'Jones' but not 'Smith,'" and other combinations.

Proximity searching, e.g., "Locate 'Jones' within 25 words of 'Smith.'" Often such searching is restricted by date range, e.g., "Locate all e-mails with 'Jones' created after January 1 but before July 1, 2007 only."

Concept searching, as has already been briefly discussed, takes a different approach. It targets information relating to a concept even if specific keywords are not present (e.g., a series of e-mails mentioning the words "Clinton," "McCain" and "Obama" would likely concern the 2008 U.S. presidential election, even if the phrase "presidential election" does not appear).

Some concept searching tools use "taxonomies" or "ontologies," that is, compilations of both commercially available data and data supplied by the client pertinent to the case collected from the lawyers and key players. Some concept searching uses linguistic analysis examining how the communicants discuss matters, while other approaches, such as "clustering" and "latent semantic indexing," use mathematical probabilities to determine whether a given file is related to a given concept. For an excellent discussion of concept searching, see "The Sedona Conference Best Practices Commentary on the Use of Search and Informational Retrieval Methods in E-Discovery."

A DIFFICULT HEARING
Regardless of which approach the search engine takes, the actual Daubert hearing will prove difficult for two practical reasons, both stemming from the fact that search engine applications are proprietary. First, it simply will be hard to get the designer to appear at the hearing to testify as to how the engine works. Second, the designer will fight giving the best evidence of the efficacy of the engine, that is, the engine's source code, because that code is proprietary. Should the code be revealed, the design would lose its value, as anyone could use that code without having to obtain a license (i.e., a copy of the application) from the designer.

Proprietary applications can be validated without their source code revealed, but only under certain circumstances. The easiest is where a specific positive finding needs to be corroborated. For example, if a proprietary forensic search tool such as Guidance Software's EnCase reports that a file is found at a particular location on a hard drive, an examiner can use an "open source" tool, i.e., a tool whose source code is known and which has been validated, to confirm the finding. Such corroboration, however, does not validate the search tool, only the result of the use of that tool at a particular time. To validate a proprietary tool generally using open source tools requires months of work, thousands of hours by highly experienced analysts, such as the FBI put in when validating EnCase. Of course, each time a new version of a tool comes out, more hours of validation are needed. Thus, while this means may be reliable, it is hardly practical.

A second means would be to use another proprietary tool -- say, Access Data's FTK -- to run the same search as EnCase performed and compare results. This method, however, is not truly scientific, since identical search results are just as likely to confirm that the two engines are identically flawed as they are reliable.

A third means to validate a proprietary search engine without revealing its code would be to search test data sets with known test results and which contain the types of data that the engine would search when regularly deployed. Comparing the results of the searches by the proprietary search engine to the known results should validate or invalidate the search engine. Again, however, such testing is extremely time-consuming and expensive. The designer would have to engage in such testing and publish its results; one could hardly expect the typical user of the search engine to engage in such studies.

As previously stated, the second Daubert hearing issue is where the searching was done. Specifically, the issue would be whether only the files actively stored on a hard drive, for example, were searched, or whether deleted files, temporary files or file fragments in the "unallocated space" of a hard drive were also searched. When ESI is gathered, unless bit stream, forensic images (i.e., exact copies of every 1 and 0 on a piece of media) are made, the deleted files, etc., will not even be present to search. To search for such ESI, forensic tools must be used. Thus, in United States v. O'Keefe, for example, the defendant challenged the government's search results for potentially exculpatory evidence in its possession by arguing that by not looking "everywhere" on the drive for deleted files or file fragments, the government had not fully discharged its duty to search everywhere.

The problem with searching "everywhere," however, is not so much a Rule 702 problem as a practical one: forensic searches of every possible file fragment take impossibly long, and if many hard drives and servers are involved, the impossible becomes unthinkable. O'Keefe, however, raises another issue, one far more interesting and conceptually difficult: for search engines, passing the Daubert test may depend upon whether one is trying to prove that something is there or that something is not there.

EVIDENCE AND ITS ABSENCE
Anyone who remembers examining scientific method when taking high school or college science classes will recall the question whether the absence of evidence that "x" is present means that "x" truly is not present or whether the test for finding "x" was simply insufficient. For example, while a PET Scan's positive finding for cancer is conclusive, a failure to detect cancer may mean the absence of cancer or that the PET Scan failed to detect cancer that was present.

Thus, the acceptance of a test as scientific proof under Daubert and Rule 702 is more likely when the test is to prove that something is present than absent. In Sanders v. Texas, 191 S.W.3d 272 (Ct. App. 2006), for example, the Texas Court of Appeals had no trouble affirming the trial court's findings that the expert's use of EnCase to create a bit stream, forensic image of the defendant's hard drive and his search of the drive to uncover child pornography -- both positive findings -- was scientifically valid. Since Encase's findings can be corroborated by a tool other than the proprietary one used, the validity of the imaging and search is much easier to establish.
In O'Keefe, Equity Analytics v. Lundin and the prototypical e-discovery matter, the typical challenge is the opposite of the typical challenge in a criminal matter: the requesting party's typical challenge to e-discovery production is not that it is inauthentic but that it is incomplete. The Daubert challenge in e-discovery cases is to prove that the search results yielded "everything."

If the search engine in question were an open-source tool, the challenge could be more easily met: the search engine's methodology would be open for all to test, and it would either work when searching test sets with known results or produce anomalies or mistakes. However, if the search engine is proprietary, proving the negative (it did not miss anything) by proving the positive (this is how it searches) is not available to the tool's proponent. The "third means" discussed above -- subjecting the search tool to known test data to see whether it missed any "hits" -- could work, but that means is extremely time-consuming, expensive and beyond the capability of the typical user.

THE MYTH OF PERFECTION
The Sedona Conference commentary provides an interesting method of "corroboration." It cites a study in which review attorneys, doing a "manual" review of discovery, were asked how much responsive data they were able to find. The attorneys guessed 75 percent, but a detailed analysis revealed that they had found only 20 percent. Using that study to illustrate what the Sedona Conference's commentary refers to as the "myth of perfection," i.e. that review attorneys slogging through e-documents and e-mails will catch responsive ESI that concept search engines will miss, the commentary makes the scientifically questionable but legally valid point that the validity of concept search tools must be determined by measuring concept search results against the actual results of review attorneys, not against results of a "perfect" search. If concept searching improves upon review practice as it now stands, it is a valid litigation tool.
In making its point, the Sedona Conference commentary returns to a touchstone of discovery practice: that when producing discovery, a "perfect review ... of information is not possible ... . The governing legal principles and best practices do not require perfection in making disclosures or in responding to discovery requests."

The Daubert challenge raised by Facciola, then, may be met not by judging the scientific validity of a search engine in an absolute way, but by judging how valid it is to suit the purposes of e-discovery production, an undertaking which involves many factors, such as the costs in time, money and energy to the producing party and their marginal benefit to the requesting party and the litigation, that have no bearing on the scientific validity of the search engine. In other words, the ultimate acceptance of an e-discovery search tool may be informed by its relative perfection but will ultimately depend, like so many other things in the law, upon the totality of circumstances.

Labels: , , , , ,

Tuesday, March 25, 2008

Corporations are Evidence Machines

In my never ending quest to find the best technologies in the industry, I have recently discovered Humanizing Technology (HT), a emerging player in the sophisticated search technology market.

The Problem
As I have been "preaching"on this Blog, after Enron and the changes to Federal Rules of Civil Procedure (FRCP), it has become increasingly risky for companies to rely solely upon reactive or remedial measures with regard to electronic compliance. Through (FRCP) directives, courts continue to impose more stringent eDiscovery requirements, increasingly mandating that companies better understand and manage electronic communications and records. These mandates, combined with the sheer volume of electronic communications and the proliferation of communications technologies, pose a serious risk to companies.

An October 1, 2007 Forbes magazine article titled, The Data Explosion, noted: "Corporations are evidence machines, generating terabytes of electronic documents, e-mails and digitally recorded phone calls each year." Non-compliance problems, many of which are perpetrated in electronic communications, can be and frequently are image-damaging publicity events for companies as well as a basis for significant financial and legal risk. Managing e-compliance in the midst of the data explosion, as opposed to having it manage you, is a key challenge to company boards, executives, and managers as well as their professional advisors.

HT History
HT began as a technology development company in early 2000 and soon began focusing upon applications that required the mining and extraction of hard-to-find information as a core competency. As such, HT developed a unique concept search technology and deployed it as a utility within its news search applications and patent search product. In these products, the technology proved its ability to deliver unique concept search capabilities and so HT embarked upon efforts to leverage its concept search technology by applying it to other markets.
In mid-2007, HT de-coupled its concept search technology from the news and patent search products to deploy it within various text-search applications. In just a few short months, HT proved the technology’s uniqueness within the following applications:
  1. A significant mid-size manufacturing business searching various data stores that, when complete, will involve up to 2 tera-bytes of data searched.
  2. Two leading Universities for searching intellectual property, research expertise, knowledge base information, and library archives.
  3. A law firm searching its database of legal documents.
  4. A medical information application searching various data stores for relevant patents, patient record data, and clinical trial results.
  5. An Indiana law enforcement agency searching data stores of evidence.

Based on the success of the technology, HT sought to identify various business and/or legal problems that were complex and costly to organizations but could be solved by deploying superior technology that searches, finds, and retrieves critical information. Initially, HT was advised to deploy its technology toward electronic discovery. While this is an attractive market and involves complex business and legal issues, HT believes that it makes more sense to help companies avoid trouble (i.e., be proactive) rather than simply helping them get out of trouble (i.e., be reactive). Thus, the emerging area of electronic compliance became the focus of HT’s concept search. HT is bringing together expertise and state-of-the-art tools for the implementation of Best Practices in electronic compliance for its customers.

The HT Solution
HT’s Audit Quality Search (AQS) technology seeks to assist corporations in risk management by carrying out various aspects of a compliance audit program including routine internal monitoring, more thorough periodic internal audits, and very thorough external text audits. HT has the ability to assist executives, officers, board members, high level managers, audit committee members, and professional advisors in assessing regulatory compliance gaps, identifying and managing risk and, in general, carrying out a broad set of corporate governance and oversight responsibilities. HT AQS can be used in the following ways:

  1. As an investigative tool to research specific suspected wrong-doing.
  2. As a gap analysis tool to carry out proactive but general compliance assessments.
  3. As a compliance audit tool to provide a basis for sampling and summarizing a company’s overall state of compliance (similar to the function of financial audits).
  4. As a records management tool to analyze electronic record data stores and provide a basis for making retention/deletion decisions.
  5. As a due diligence tool to analyze electronic record data stores for completed and/or prospective acquisitions.

By utilizing AQS technology as part of a comprehensive compliance audit program, company executives, managers, board member, audit committee members and professional advisors can reduce financial risk and legal exposure by implementing Best Practices and "reasonable state-of-the-art methods" for assuring compliance.

I would encourage anyone who reads this Blog to contact Michael Mulcahy, the VP of Business Development at HT to get more information about this interesting and very promsing new search solution.

Labels: , , , , , , , , ,