Click for full size

The Redgrave LLP team just released a fabulous case study, “Generative AI for Complex Document Review”. To summarize, they put Relativity aiR for Review up against Relativity Active Learning (RAL) using a 45k collection to test both workflows for defensibility under caselaw and the Sedona Conference TAR 1 Reference Model parameters.  The paper is a dense, academic read that will consume a goodly chunk of billable time, though their Law.com article is a much easier read and surfaces good points. I encourage peers designing and supporting relevance workflows under reasonableness and proportionality standards to find the time to digest the full paper. This Part 1 blog will highlight some items and potential implications. In Part 2 I will use the study metrics in one of my ROI cost models for a different perspective on the “98% fewer attorney review hours” headline.

“Greg’s second-pass analysis is a useful extension of what we set out to measure.  Our study focused on first-pass responsiveness accuracy, where the data showed generative AI can find more responsive documents and miss fewer than active learning.  How a team structures the review that follows, and what it costs, depends heavily on the assumptions a given matter brings to it.” — Robert Keeling, Redgrave LLP

**eDJ disclaimer- Although the Redgrave LLP study is based on Relativity functionality, my commentary tries to take a broader view on validating and justifying AI workflows independent of platform. Too few in our profession rigorously test our tools and measure the impact of our workflow approaches. So please do not take my comments as a criticism of the Redgrave team’s good work or Relativity’s excellent platform. Instead, I am trying to extract important implications and offer alternative perspectives on the results based on the study as it was conducted.**

eDJ Take-aways (your mileage may vary):

aiR for Review approach – Single partner used a reviewed 250 sample set to establish 6% richness and train aiR for Review prompt over 43 iterations with research searches taking a total of 18 hours (3 working days). I have done this exercise recently with much fewer iterations and time. Maybe I just had a richer set to feed the Prompt Starter, but Jeremy Pickens believes that prompt iterations are taking 20-32 hours when time is documented in studies. Redgrave team made it clear that prompt development does not require partner rates. The paper calls out the low richness and offers reasonable suggestions on how the recommended 20-30% richness would improve precision. My take is that if you have a very low richness collection, spend some time filtering and find a higher richness training set within it.

Exclude errors and unsuitable items – Redgrave conservatively counted the 12 AI errored docs as relevance positives. They did not specify filtering out items with too little/much extracted text or formats unsuitable for text analysis.  While this treatment of AI errors is logical for an academic case study, I find it misleading because practical aiR workflows should include a separate handling stream for errors and items not suitable for AI. I generally assume 5%+ will have to be manually reviewed or handled by family rules based on filters. 

Overall, the AI approach was not Precise(getting just the relevant).

It did well at Recall(does not miss many relevant).

It was better than humans at limiting Elusion (missing relevant docs).

The Active Learning approach – Traditional Cimplifi contractor team of 24 over 7 days running 1,123 hours of RAL prioritized review (~40 docs/hour with QC/training). RAL’s “prioritization” barely prioritized anything — 77% of the collection (34,409 of 45,004 docs) had to be reviewed before the queue fell below the relevance cutoff, even after richness only rose from 6% to 11% in the queue. The ‘post cutoff’ review continued and only found 12 responsive documents (0.1%). Peer reviewers commented that they have never seen TAR 2 results this poor in real cases. Validation 7% richness indicates ~3,150 relevant documents. The goal of Active Learning or AI prediction is to minimize the volume of non-responsive documents reviewed while making a reasonable, defensible effort for complete responsive production. If a prioritized review was producing such low richness in the first day most of us would have brainstormed ways to enrich the collection. See my 2nd pass argument below.

 The Validation test – A partner reviewed 1,000 random docs, initially finding 73 responsive. This 1,000 doc sample was compared with the aiR and RAL predictions to understand how the different approaches performed against a Subject Matter Expert (SME) for ‘ground truth’. The ground moved beneath the SME’s feet after re-reviewing the positive conflicts with aiR’s predictions.

The informed re-review has a circularity problem. The SME changed 10 out of 151 conflicting positive calls to “Responsive” after seeing aiR’s own rationale, though they computed the study metrics BEFORE they did the re-review.  That’s not independent confirmation of aiR’s output, it’s the grader being shown the answer key’s reasoning. RAL’s conflicts never got the same second look because RAL does not provide ‘reasoning’ to consider. The Wilson Confidence Intervals are wide (aiR 78.2–93.4%)

The elusion-vs-population-recall distinction – CAL training treats a reviewer’s wrong “not responsive” call as ground truth for the next training round, and elusion sampling can’t see that. Redgrave’s §5.1 explains RAL’s tool reported 100% recall on its own elusion sample, but blind population recall was only 64%. This is the sleeper finding of the whole study. That’s a bigger defensibility concern than any precision numbers. Validated recall depends entirely on what you sampled from. Redgrave and Relativity deserve credit for plainly calling this issue out and publishing the results. Too many of my historical case study engagements for providers immediately converted to internal ‘eyes only’ when I delivered independent findings like these.

Performance comparison – While aiR for Review yielded 87% recall, RAL lagged with 64% recall. There is a lot more context worth digesting. I cannot help but think that a higher richness collection would have generated higher recall. This is a strong case for running this kind of validation analysis on a closed matter.

Redgrave §4.2 is worth quoting. “The agreement analysis reveals a consistent asymmetry: when aiR for Review and RAL disagreed, aiR for Review overwhelmingly erred toward over-inclusion—flagging documents as responsive that the expert deemed not responsive. RAL’s disagreements ran in the opposite direction—it tended to discard documents that turned out to be responsive.”

The Richness-Precision relationship – Redgrave’s §5.2 provides a clear explanation and visualization of how the precision of both models would improve if the relative richness were increased. We generally improve richness with better scoping/collection/filtering criteria up front. Redgrave’s study inherited a public production, so you must give them credit for playing the hand they were dealt. They also called out the impact of review topic complexity on human review as context for interpreting their results.

The case for 2nd Review – The more that I test these systems the more I understand the value of the traditional 1st/2nd pass review workflow. Modern legal holds preserve-in-place millions of potentially relevant items. For over 25 years we have leveraged collection and processing scopes to increase richness to an acceptable 15-30% that has proven to deliver high recall (>80%) with very low elusion (<1-3%). First pass contract review excludes the 70-85% not relevant so that counsel can put eyes on every doc to be produced AND surface the few hot docs that may lead to case resolution. In the final predictive production analysis, aiR for Review would have produced twice the 2nd pass volume compared to RAL. However, aiR’s elusion rate was 34% that of RAL per the study. One could only hope that using the recommended minimum richness would curb aiR’s over-inclusiveness to control the cost of the probable 2nd pass review in high-stakes matters.

All attorney hours are not equal – Sneak Peek at Part 2 – Applying generic* Amlaw 200 and global eDiscovery AI/hourly rates to the study hours yields a fascinating counterpoint to the “98% fewer attorney review hours” attention grabber. 18 partner hours is indeed 98% less than 1,123 Cimplifi hours. However, factoring in 20 hours for validation sampling brings that down a touch. **The validation sampling applies to both workflows. The real kicker was putting it through my ROI time-cost model (blog Part 2) to see that the 1st pass Study workflows are almost identical much closer in cost because of the incredible range in billable rates. **I agree with Redgrave and peer feedback that the validation pass should be done on RAL, but sticking with the study parameters for this comparison.

RAL AI, prompt + validation
Hours 1,123 18.5
Cost $57,405 $33,451

Forgive me for not showing my financial homework yet, but this analysis has already consumed too many unbillable hours and needs to get to Redgrave LLP and Relativity for feedback before being published. See Part 2 for full breakdown.

Summary:

Well-constructed academic case studies on AI assisted review are rare, valuable and appreciated. I believe the study effectively demonstrates AI and Active Learning workflows can produce reasonable, defensible relevance reviews with each approach presenting different advantages and challenges. I am not convinced that either workflow negates the value and need for some kind of 2nd pass in high-stakes matters. The study does not contemplate the complexities of modern ESI types, familial relationships, PII/CBI or required request classification/organization. It provides clear and detailed information for peers to ingest and interpret for their own unique discovery profile. So, block out some time and do your homework. Share what I got right or wrong. Look forward to publishing Part 2 next where I use the study metrics to project larger time/cost comparisons with more ‘in the trenches’ models. I will be making the model workbook available for download.

I am looking forward to RelFest and seeing roadmap meet reality. I am still booking briefings and peer conversations, so grab a time slot if you are joining me in Chicago!

Greg Buckles wants your feedback, questions or project inquiries at Greg@eDJGroupInc.com.  Reach out for a free 15 minute ‘Good Karma’ call if he has availability. He solves problems and creates eDiscovery solutions for enterprise and law firm clients.

Greg’s blog perspectives are personal opinions and should not be interpreted as a professional judgment or advice. Greg is no longer an investigative journalist and all perspectives are based on best public information. Blog content is neither approved nor reviewed by any providers prior to being published. Do you want to share your own perspective? Greg is looking for practical, professional informative perspectives free of marketing fluff, hidden agendas or personal/product bias. Outside blogs will clearly indicate the author, company and any relevant affiliations. 

Greg’s latest nature, art and diving photographs on Instagram.

0 0 votes
Article Rating