This Page

has been moved to new address

The eDiscovery Paradigm Shift

Sorry for inconvenience...

Redirection provided by Blogger to WordPress Migration Service
----------------------------------------------------- Blogger Template Style Name: Snapshot: Madder Designer: Dave Shea URL: mezzoblue.com / brightcreative.com Date: 27 Feb 2004 ------------------------------------------------------ */ /* -- basic html elements -- */ body {padding: 0; margin: 0; font: 75% Helvetica, Arial, sans-serif; color: #474B4E; background: #fff; text-align: center;} a {color: #DD6599; font-weight: bold; text-decoration: none;} a:visited {color: #D6A0B6;} a:hover {text-decoration: underline; color: #FD0570;} h1 {margin: 0; color: #7B8186; font-size: 1.5em; text-transform: lowercase;} h1 a {color: #7B8186;} h2, #comments h4 {font-size: 1em; margin: 2em 0 0 0; color: #7B8186; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px;} @media all { h3 { font-size: 1em; margin: 2em 0 0 0; background: transparent url(http://www.blogblog.com/snapshot/bg-header1.gif) bottom right no-repeat; padding-bottom: 2px; } } @media handheld { h3 { background:none; } } h4, h5 {font-size: 0.9em; text-transform: lowercase; letter-spacing: 2px;} h5 {color: #7B8186;} h6 {font-size: 0.8em; text-transform: uppercase; letter-spacing: 2px;} p {margin: 0 0 1em 0;} img, form {border: 0; margin: 0;} /* -- layout -- */ @media all { #content { width: 700px; margin: 0 auto; text-align: left; background: #fff url(http://www.blogblog.com/snapshot/bg-body.gif) 0 0 repeat-y;} } #header { background: #D8DADC url(http://www.blogblog.com/snapshot/bg-headerdiv.gif) 0 0 repeat-y; } #header div { background: transparent url(http://www.blogblog.com/snapshot/header-01.gif) bottom left no-repeat; } #main { line-height: 1.4; float: left; padding: 10px 12px; border-top: solid 1px #fff; width: 428px; /* Tantek hack - http://www.tantek.com/CSS/Examples/boxmodelhack.html */ voice-family: "\"}\""; voice-family: inherit; width: 404px; } } @media handheld { #content { width: 90%; } #header { background: #D8DADC; } #header div { background: none; } #main { float: none; width: 100%; } } /* IE5 hack */ #main {} @media all { #sidebar { margin-left: 428px; border-top: solid 1px #fff; padding: 4px 0 0 7px; background: #fff url(http://www.blogblog.com/snapshot/bg-sidebar.gif) 1px 0 no-repeat; } #footer { clear: both; background: #E9EAEB url(http://www.blogblog.com/snapshot/bg-footer.gif) bottom left no-repeat; border-top: solid 1px #fff; } } @media handheld { #sidebar { margin: 0 0 0 0; background: #fff; } #footer { background: #E9EAEB; } } /* -- header style -- */ #header h1 {padding: 12px 0 92px 4px; width: 557px; line-height: 1;} /* -- content area style -- */ #main {line-height: 1.4;} h3.post-title {font-size: 1.2em; margin-bottom: 0;} h3.post-title a {color: #C4663B;} .post {clear: both; margin-bottom: 4em;} .post-footer em {color: #B4BABE; font-style: normal; float: left;} .post-footer .comment-link {float: right;} #main img {border: solid 1px #E3E4E4; padding: 2px; background: #fff;} .deleted-comment {font-style:italic;color:gray;} /* -- sidebar style -- */ @media all { #sidebar #description { border: solid 1px #F3B89D; padding: 10px 17px; color: #C4663B; background: #FFD1BC url(http://www.blogblog.com/snapshot/bg-profile.gif); font-size: 1.2em; font-weight: bold; line-height: 0.9; margin: 0 0 0 -6px; } } @media handheld { #sidebar #description { background: #FFD1BC; } } #sidebar h2 {font-size: 1.3em; margin: 1.3em 0 0.5em 0;} #sidebar dl {margin: 0 0 10px 0;} #sidebar ul {list-style: none; margin: 0; padding: 0;} #sidebar li {padding-bottom: 5px; line-height: 0.9;} #profile-container {color: #7B8186;} #profile-container img {border: solid 1px #7C78B5; padding: 4px 4px 8px 4px; margin: 0 10px 1em 0; float: left;} .archive-list {margin-bottom: 2em;} #powered-by {margin: 10px auto 20px auto;} /* -- sidebar style -- */ #footer p {margin: 0; padding: 12px 8px; font-size: 0.9em;} #footer hr {display: none;} /* Feeds ----------------------------------------------- */ #blogfeeds { } #postfeeds { }

Thursday, August 2, 2012

StoredIQ Reinvents Itself in a Big Data Way


Over the past five (5) years StoredIQ has had more than its fair share of ups and downs.  Founded in 2001, venture backed StoredIQ began to establish itself as a "next generation" player in the eDiscovery software market around 2005.  However, after being overlooked for a large consolidation move in 2009, StoredIQ seemed to loose its way and couldn't figure out if they were in the Information Governance market competing with Autonomy, IBM and Symantec or in the eDiscovery Early Case Assessment (ECA) market competing with Clearwell Systems.

2010 became a pivotal year as they brought on Phil Myers as the new CEO.  With 29 years of experience in the technology industry and having managed three successful start-up companies, Phil made adjustments in personal, mission and strategy and got StoredIQ back in the game.

In 2011, Phil hired Tom Bishop as The new Chief Technology Officer (CTO).  Bishop was the former chief technology officer of IBM Tivoli. After Tivoli, Bishop served as CTO of VIEO, Inc., where he was named “Chief Technology Officer of the Year” by InfoWorld magazine vice president and CTO at BMC Software where he was responsible for product vision and direction, including advancing Atrium, the company’s innovative open-architected foundation for Business Service Management solutions.  Tom was the right technology leader at the right time to figure out what the market wanted StoredIQ to be and how to get them there technically.

Throughout 2011, StoredIQ executives met with customers, prospects and other industry thought leaders to try and establish their corporate identity.  More importantly, they tried to figure out if they were going to build product to compete in the Information Governance or eDiscovery markets.  Where they ended up may surprise some of you.

Named by Gartner as a 2012 "Cool Vendor" in Risk Management, Privacy and Compliance, StoredIQ ended up in the middle of "Big Data" with its new mission to enable organizations to actively manage their vast and ever-increasing amounts of unstructured data.  So, with a slight twist on the approach and who they are now selling to, StoredIQ actually ended up in both Information Governance and eDiscovery.  You see, at the root of any Information Governance or eDiscovery project or process is the ability to identify, collect, index and analyze Big Data.  And, that's what StoredIQ is now doing.
I had the pleasure of spending an hour today with Phil Myers, StoredIQ's CEO and Amir Jaibaji, Vice President of Product Management for StoredIQ.  They walked me through their "new strategy" and gave me a quick demo of DataIQ, their recently announced data analytics module that provides users with an exceptionally unique visual overview and approach to analyze unstructured data.  It’s very visual, fast and provides an abundance of information that you probably didn’t even know that you had about your data.  Whether you are an analyst in the Information Technology (IT) department managing storage utilization, a risk manager looking for “open shares” in SharePoint or a General Counsel trying to forecast the cost of pending litigation, DataIQ is just what you have been hoping for. It was impressive to say the least and  if it is any indication of where Myers and Bishop have taken StoredIQ, they have not only reinvented themselves, they had established themselves as a formidable player in the Big Data analytics market.

Over the next couple of weeks, I plan  to spend more time with StoredIQ and will report on what I find.  My expectations are very high.

Labels: , ,

Friday, July 13, 2012

Five Initial Steps to Meet the Governance, Risk and Compliance Obligations Brought on by Today's Big Data File Stores

The accelerating increase in the amount of unstructured Electronically Stored Information (ESI) is leaving IT organizations struggling with how to store and manage all of this new information. Aside from just providing the underlying storage infrastructure to host this amount of data, companies are also faced with the task of properly managing their Big Data file stores to meet existing governance, risk and compliance obligations. To do so, there are five steps they can take now to position their organization to meet them.


According to a 2010
report by IDC, the amount of information created, captured or replicated has exceeded available storage for the first time since 2007. The size of the digital universe this year will be tenfold what it was just five years earlier. According to this same IDC report, the volume of unstructured ESI is expected to grow at over 60% CAGR (Compounded Annual Growth Rate).

According to Forrester Research and as
reported in an article that appeared on Forbes website last week:
  • The average organization will grow their data by 50 percent in the coming year
  • Overall corporate data will grow by a staggering 94 percent
  • Database systems will grow by 97 percent
  • Server backups for disaster recovery and continuity will expand by 89 percent
Overseeing the expansion of storage space and ensuring that the data is protected has become a minor part of the overall task of Big Data file storage and management. Business stakeholders and the Information Technology (IT) organizations from enterprises of all sizes and across all industries must now face a list of Governance, Risk and Compliance (GRC) regulations to which they have to legally comply or face potentially fatal financial penalties to the enterprise. 

The most obvious laws to which they are subject include:
  • Sarbanes-Oxley (SOX)
  • Health Insurance Portability and Accountability Act (HIPAA)
  • Gramm-Leach-Bliley (GLBA)
  • Federal Information Security Management Act (FISMA)
  • Consumer Information Protection Laws
  • Federal Rules of Civil Procedure (FRCP)

Further, the list of new regulations is growing. The passage of The Patient Protection and Affordable Care Act (PPACA) will result in the US Government adding 159 new agencies, programs, and bureaucracies to assist with the compliance of over 12,000 pages of new regulations. Over the past ten years, in response to the threat of international terrorism, the US Department of Homeland Security (DHS) has added hundreds of new regulations. Finally, cyber terrorism, including acts of deliberate, large-scale disruption of enterprise computer networks, is now a reality that all businesses must face.

In the face of this, Big Data file storage and management vendors, along with the associated industry consultants, have developed a list of hardware and software requirements and associated value propositions to help enterprise buyers decide which Big Data file storage and management platforms to purchase.

But before they buy, there are five steps that buyers should take first to ensure they are prepared to meet the governance, risk and compliance obligations brought on by today's Big Data file stores:
  • Internal Collaboration: File management and Governance, Risk and Compliance (GRC) requirements affect business stakeholders from the boardroom to IT to the manufacturing floor and loading dock to the accounting office. The development of cross functional workgroups and the promotion of internal collaboration between functional experts is the key to successfully identifying, understanding and addressing all of the requirements and issues involved in Big Data file management across the entire enterprise.
  • Network Architecture Planning:  Over the past 25 years, enterprise architectures grew with little or no planning resulting in wasteful redundancy and little or no access to all the enterprise data as may be required to comply with today’s GRC requirements. The advent of the Internet and now cloud computing has brought this decades of poorly planned networks to light resulting in them become more of an enterprise liability than an asset. The time is now for IT to hit the restart button and explore new options such as virtualization, hybrid cloud architectures and the use of cloud service providers (CSPs) that enable them to better leverage, manage and optimize their existing infrastructure..
  • Security:  The introduction and proliferation of portable storage devices, Wireless Internet, mobile computing devices, enterprise Software-as-as-Service (SaaS) applications, cloud storage, blogs and social media such as Facebook, LinkedIn and Twitter, data theft and cyber attacks are a real issue for which many (and arguably most) companies do not have a good answer. Now is the time for IT to take a serious look at their internal file access policies and move as quickly as possible to address any existing shortcomings.
  • Data Retention Policy Development and Implementation: Sarbanes-Oxley (SOX), the Health Insurance Portability and Accountability Act (HIPAA) and the Federal Rules of Civil Procedure (FRCP) all have very specific data retention guidelines for what types of ESI data an enterprise has to keep and how long to keep it.  Enterprises must investigate and document these requirements, development data retention policies and acquire the appropriate software to ensure compliance.
  • Technology Vendors and Consulting Partners: Business stakeholders and IT management may be overwhelmed with the task of addressing the issues of successfully meeting the GRC obligations of big file storage and management. If this is the case, reach out to the hardware and software vendor community and askhow their solutions support these issues. If required, engage the services of vendor independent consulting partners to act as trusted advisors to assist in the successful navigation of the required cultural transitions and the acquisition of the best technology platforms.

The accelerating increase in the amount of unstructured Electronically Stored Information (ESI) is putting IT organizations on the defensive as they struggle to figure out how to store and manage all of this new information. However, overseeing the expansion of storage space and ensuring that appropriate backups are completed has become a minor part of the overall task of big file storage and management.

Rather business stakeholders and IT staff need to act now to first bring their infrastructure under control so they can get in front of the growing list of GRC regulations to which they are subject. By following the five steps outlined above, enterprises will be in a position so that when they purchase a product, they will have a good grasp of what their true enterprise challenges are and have a high probability of bringing in a product that addresses them.

Labels: , , , , , , , , , , , ,

Tuesday, September 6, 2011

eDiscovery and Big Data Analytics

If cloud computing in general is the next challenge facing information governance and eDiscovery.  Then, Big Data Analytics is one of the specific issues that information governance and eDiscovery technologist are going to have to conquer.

Cloud computing is an IT infrastructure choice, storage architecture and application delivery mechanism.  And, combined with mobile computing, a choice that will results in more Electronically Stored Information (ESI) or Electronically Stored Evidence (ESE), if you are a lawyer, than the total information from all previous generations.

And, whereas the eDiscovery industry has been struggling to collect, process and analyze terabytes of data in a reasonable amount of time for a reasonable cost, this new paradigm of cloud computing and its associated federated data stores is already producing peta and exabytes of data.  Hence, the term Big Data (its actually all a matter of perspective).  Of even more concern is the fact that as the eDiscovery market has been struggling to appropriately and accurately analyze structured data, the new paradigm of Big Data in the cloud is largely unstructured data and therefore largely left out of the eDiscovery equation.

Today, whether right or wrong and due largely to a lack of understanding and associated technological and financial restraints, very little unstructured data is even considered during 26(f) strategies and is therefore left out of most litigation. From an eDiscovery standpoint, there is no doubt that there are "smoking guns" hiding in some Big Data store as unstructured data and as such brings a whole new meaning to the phrase of "looking for a needle in a haystack".

However, I believe that there is hope as Big Data analytics do in fact exist and the technology is evolving. As Srinivasan Sundara Rajan from HP points out in September 6, 2011 article on SOA World Magazine Site titled, "Traditional vs Big Data Analytics," "Big data analytics provide new ways for businesses and government to analyze unstructured data which so far have been rejected by the data cleansing routines in a typical enterprise data warehouse scenario."  This same technology will be useful for information governance and eDiscovery.  And, from a requirements standpoint,  may prove to be a very interesting and financially rewarding vertical for the technologists to address.

The full text of the Srinivasan Sundara Rajan's article is as follows:

Big Data Analytics Convergence Among the Major IT Companies
Major IT companies acquiring analytics software and application providers has been the order of the day. We have seen the words ‘Big Data Analytics' being used in many solutions for the enterprise.
‘Big Data' is the general term used to represent massive amounts of unstructured data that are not traditionally stored in a Relational form in enterprise databases. The following are the general characteristics of Big Data.
  • Data storage defined in order of PETA BYTES, EXA BYTES and much higher in volume to the current storage limits in enterprises which TERA BYTES.
  • Generally it is considered as Unstructured data and not really falling the under the relational database design which the enterprises have been used to
  • Data Generated using unconventional methods outside of data entry like, RFID, Sensor networks etc...
  • Data is time sensitive and consists of data collected with relevance to the time zones
In the past, the term ‘Analytics' has been used in the business intelligence world to provide tools and intelligence to gain insight into the data through fast, consistent, interactive access to a wide variety of possible views of information.

Very close to the concept of analytics, data mining has been used in enterprises to keep pace with the critical monitoring and analysis of mountains of data. The biggest challenge is how to unearth all the hidden information through the vast amount of data.

Traditional DW Analytics vs Big Data Analytics
The analytics of enterprise data toward meaningful insights into the information that exists over a period of
time in that context is why Big Data Analytics makes it different from traditional data warehouse analytics.


Traditional Data warehouse AnalyticsBig Data Analytics
Traditional Analytics  analyzes on the known data terrain that too the data   that is well understood.  Most of the data warehouses have a elaborate ETL processes and database constraints, which means the data that is loaded inside a data warehouse is well under stood, cleansed and in line with the business metadata.The biggest advantages of the Big Data  is it is targeted at unstructured data outside of traditional means of capturing the data. Which means there is no guarantee that the incoming data is well formed and clean and  devoid of any errors.  This makes it more challenging but at the same time it gives a   scope for much more insight into the data.
Traditional Analytics is built on top of the relational data model,  relationships between the subjects of interests have been created  inside the system and the  analysis is done based on them.In typical world, it is very difficult to establish  relationship between all the information in a formal way, and  hence unstructured data in the form  images, videos, Mobile generated information, RFID etc... have to be considered in big data analytics. Most of the big data analytics database are based out  Columnar databases.
Traditional  analytics is batch oriented  and  we need to wait for nightly ETL and transformation jobs to complete before the required insight is obtained.Big Data Analytics is aimed at  near real time analysis of the data using the  support of the software meant for it
Parallelism in  a traditional analytics system is achieved  through  costly hardware like MPP (Massively Parallel Processing) systems   and / or  SMP systems.While there are appliances in the market for the Big Data Analytics,  this can also be achieved  through commodity hardware and new generation of analytical software like Hadoop or other Analytical databases.
Traditional Data warehouse AnalyticsBig Data Analytics
Traditional Analytics  analyzes on the known data terrain that too the data   that is well understood.  Most of the data warehouses have a elaborate ETL processes and database constraints, which means the data that is loaded inside a data warehouse is well under stood, cleansed and in line with the business metadata.The biggest advantages of the Big Data  is it is targeted at unstructured data outside of traditional means of capturing the data. Which means there is no guarantee that the incoming data is well formed and clean and  devoid of any errors.  This makes it more challenging but at the same time it gives a   scope for much more insight into the data.
Traditional Analytics is built on top of the relational data model,  relationships between the subjects of interests have been created  inside the system and the  analysis is done based on them.In typical world, it is very difficult to establish  relationship between all the information in a formal way, and  hence unstructured data in the form  images, videos, Mobile generated information, RFID etc... have to be considered in big data analytics. Most of the big data analytics database are based out  Columnar databases.
Traditional  analytics is batch oriented  and  we need to wait for nightly ETL and transformation jobs to complete before the required insight is obtained.Big Data Analytics is aimed at  near real time analysis of the data using the  support of the software meant for it
Parallelism in  a traditional analytics system is achieved  through  costly hardware like MPP (Massively Parallel Processing) systems   and / or  SMP systems.While there are appliances in the market for the Big Data Analytics,  this can also be achieved  through commodity hardware and new generation of analytical software like Hadoop or other Analytical databases.


Use Cases for Big Data Analytics
Enterprises can understand the value of Big Data Analytics based on the use cases and how the traditional problems can be solved with the help of Big Data Analytics. The following are some of the usages.

Customer Satisfaction and Warranty Analysis: Probably this is the one big area that most product-based enterprises are worried about. As of today, there is not a clear way of gauging the issues with the products and the associated customer satisfaction, unless they come in a formal way in an electronic form.
  • Information regarding quality is collected through various external channels and most of the times the data is not clean
  • As the data is unstructured there is no way to relate the associated issues, so that the long-term fix can be given to customer.
  • Classification and grouping of problem statements are missing , resulting enterprises not able to group the issues
From the above discussion, utilizing the Big Data Analytics for customer satisfaction and Warranty analysis will help enterprises gain insight into the much-needed customer mind set and solve their problems effectively and to avoid them in their new product lines.

Competitor Market Penetration Analysis: In today's economy where the competition is high, we need to gauge the areas where the competitors are strong and their pain points through an analysis within the legal means. This information is available in a variety of web sites, social media sites and other public domains. Big data analytics on this data can provide an organization with much needed information about Strength, Weakness, Opportunities and Threats for their product lines.

Healthcare / Epidemic Research & Control: Epidemics and seasonal diseases like influenza start with certain patterns among the people and they spread to a larger section if they are not detected early and controlled. This is one of the biggest challenges for growing as well as developed nations. The current issue most of the times the symptoms vary between the people and various health care providers treat them differently. There is also not a common classification of symptoms across people. Adopting Big Data Analytics on this typically unstructured data will help the local governments to effectively tackle the outbreak situations.

Product Feature and Usage Analysis: Most product companies, especially consumer products, keep adding lot of features to their product line, however it may happen that some of the features are not really used by the consumers and some are used more and effective analysis of this data captured by various mobile devices and other RFID based inputs can provide valuable insights to the product companies.

Future Direction Analysis: The trends in each business are analyzed by research groups and this information is available through industry specific portals or even common web blogs. Constant analysis of this futuristic data will help enterprises to look forward to future and bring them to their product lines.

Summary
Big data analytics provide new ways for businesses and government to analyze unstructured data which so far have been rejected by the data cleansing routines in a typical enterprise data warehouse scenario. However as evident from the use cases above, these analyses will go a long way in improving the operations of the organizations. We will see more convergence of the products and appliances in this space in the coming days.

Labels: , , , ,