Legal Cyber Academy
All insights

Technology-Assisted Review: What Courts Have Actually Approved

By the Legal Cyber Academy editorial team ·

What courts have actually approved

Technology-assisted review is judicially approved as a method a producing party may choose for searching electronically stored information, beginning with Magistrate Judge Andrew Peck's opinion in Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182 (S.D.N.Y. 2012). Courts permit TAR when the producing party elects it, have refused requesting parties' motions to force a producing party to use it, and have refused to let a requesting party block it. Those decisions endorse no vendor and no particular tool, and none sets a recall threshold.

What Da Silva Moore held, and the four things it expressly disclaimed

The operative sentence is narrow: "This judicial opinion now recognizes that computer-assisted review is an acceptable way to search for relevant ESI in appropriate cases." 287 F.R.D. at 183.

The closing paragraph, at 193, does most of the work that vendor marketing ignores. It states that computer-assisted review need not be used in all cases; that the exact ESI protocol approved there will not be appropriate in all future cases; that the opinion endorses no vendor; and that it endorses no particular computer-assisted review tool. What the Bar should take from it, the court wrote, is that computer-assisted review "is an available tool and should be seriously considered for use" in large-data-volume cases — with counsel still obliged to design an appropriate process "including use of available technology, with appropriate quality control testing," while "adhering to Rule 1 and Rule 26(b)(2)(C) proportionality."

Two qualifiers travel with the holding, and both matter to how the case can be cited. First, the parties had already agreed to use predictive coding; what they disputed was implementation — custodians, data sources, a proposed cutoff at the top 40,000 documents, seed-set transparency. The court said so directly: "The decision to allow computer-assisted review in this case was relatively easy — the parties agreed to its use (although disagreed about how best to implement such review)." Id. at 189. Da Silva Moore is not authority that a court will impose TAR over objection. Second, this was a magistrate judge's discovery ruling; Rio Tinto records that it was affirmed. 2012 WL 1446534 (S.D.N.Y. Apr. 26, 2012).

Two doctrinal anchors carry the reasoning. The court held that "[w]hile this Court recognizes that computer-assisted review is not perfect, the Federal Rules of Civil Procedure do not require perfection," citing Pension Committee of the University of Montreal Pension Plan v. Banc of America Securities, 685 F. Supp. 2d 456, 461 (S.D.N.Y. 2010). 287 F.R.D. at 191. And it grounded the analysis in Rule 1 — the rules must "be construed, administered, and employed by the court and the parties to secure the just, speedy, and inexpensive determination of every action and proceeding" — and in proportionality.

Check the vintage of that second anchor before you quote it. Da Silva Moore reproduced the proportionality factors from Rule 26(b)(2)(C)(iii) as that subsection then read. The 2015 amendments moved those factors into Rule 26(b)(1), which now defines the scope of discovery as nonprivileged matter "relevant to any party's claim or defense and proportional to the needs of the case," listing the importance of the issues, the amount in controversy, the parties' relative access to relevant information, the parties' resources, the importance of the discovery in resolving the issues, and whether the burden or expense outweighs the likely benefit. Today's Rule 26(b)(2)(C)(iii) says something different: the court must limit discovery that is "outside the scope permitted by Rule 26(b)(1)." A brief that block-quotes the 2012 opinion's rule text as current law is quoting a superseded rule.

So when a brief says TAR is "court-approved," ask what proposition that is offered for. The practice is approved. The product is not, and no opinion in this line says otherwise.

You probably cannot compel TAR on the other side, and you almost certainly cannot block it

Three years after Da Silva Moore, Peck wrote in Rio Tinto PLC v. Vale S.A., 306 F.R.D. 125, 127 (S.D.N.Y. 2015), that "the case law has developed to the point that it is now black letter law that where the producing party wants to utilize TAR for document review, courts will permit it."

The footnote to that sentence is the part practitioners need. Rio Tinto records that where the requesting party sought to force the producing party to use TAR, the courts refused — citing the Biomet M2a Magnum hip implant litigation and Kleen Products v. Packaging Corp. of America. Peck added a caveat worth preserving: in those cases the producing parties had spent over $1 million on keyword search (in Kleen) or on keyword culling followed by TAR (in Biomet), "so it is not clear what a court might do if the issue were raised before the producing party had spent any money on document review." Id. at 127 n.1.

Do not stretch that into a rule that no court will ever push a party toward TAR. Rio Tinto's own footnotes record the counter-example. In EORHB, Inc. v. HOA Holdings LLC — the Hooters case — Vice Chancellor Laster told the parties at an October 2012 hearing, on his own motion, that "I would like you all, if you do not want to use predictive coding, to show cause why this is not a case where predictive coding is the way to go." The parties then agreed that the defendant would use TAR and the plaintiff would not, and he approved both approaches. Id. at 128 n.4. The proposition that actually holds is narrower than "courts never order TAR": it is that a requesting party's motion to force TAR on a producing party has failed.

The Tax Court reached the same place from the other direction in Dynamo Holdings Ltd. Partnership v. Commissioner, 143 T.C. 183 (2014). The IRS opposed predictive coding as an "unproven technology." The court disagreed — "we understand that the technology industry now considers predictive coding to be widely accepted for limiting e-discovery to relevant documents and effecting discovery of ESI without an undue burden" — and held "that petitioners must respond to respondent's discovery request but that they may use predictive coding in doing so." Its reasoning was institutional temperament rather than technology: "the Court is not normally in the business of dictating to parties the process that they should use when responding to discovery," any more than it would dictate whether a paralegal, a junior attorney or a senior attorney performs a paper review. The remedy for a suspected shortfall is a later motion to compel, not advance methodology control. Note the forum before borrowing the language: Dynamo applies the Tax Court's Rules 70 and 72 — which that court observed are "similar to corresponding provisions found in the Federal Rules of Civil Procedure" — and its own Rule 1(d), not the FRCP directly.

The most useful recent decision runs the other way on the same principle. In Berger v. Graf Acquisition, LLC, C.A. No. 2023-0873-LWW (Del. Ch. Oct. 21, 2024), the defendants offered to review the full universe captured by the plaintiff's broad search terms — roughly 125,000 documents — if they could use a TAR protocol; the plaintiff rejected the offer and demanded manual review. Vice Chancellor Lori W. Will held that whether to employ TAR "is a matter for the producing party to decide," that "[i]t is not up to the requesting party to block TAR if the producing party prefers it," and that the defendants may use TAR to reduce their discovery burden "so long as they are transparent with the plaintiff about their computer-assisted review process," with Delaware counsel remaining closely involved in the review and sampling. If the plaintiff believed the resulting production was deficient, the court said, it could seek relief then.

Berger is a Delaware Court of Chancery decision resting on Court of Chancery Rule 26(b)(1), whose operative language — "any non-privileged matter that is relevant to any party's claim or defense and proportional to the needs of the case" — tracks the federal rule. Cite it for what it is.

The practical consequence: on this case law, a motion asking the court to dictate the other side's review technology faces long odds. A motion attacking the results — with metrics — does not.

Seed-set disclosure is not required, and it was never the real safeguard

Rio Tinto framed the transparency question squarely at 128: "One TAR issue that remains open is how transparent and cooperative the parties need to be with respect to the seed or training set(s)."

The opinion then catalogues the split it found. In Da Silva Moore, the defendant confirmed that "[a]ll of the documents that are reviewed as a function of the seed set, whether [they] are ultimately coded relevant or irrelevant, aside from privilege, will be turned over to" plaintiffs — which Rio Tinto describes as volunteered rather than ordered, though Peck had told the defendants at an earlier conference that if they used predictive coding they were "going to have to give your seed set, including the seed documents marked as nonresponsive," to opposing counsel. In In re Actos, the parties' protocol had experts from each side simultaneously reviewing and coding the seed set. In the second Biomet decision, Rio Tinto records, Judge Miller said he could find no authority that would allow him to require Biomet to share seed-set documents with plaintiffs' counsel, while suggesting Biomet rethink its opposition. Peck's conclusion: "where the parties do not agree to transparency, the decisions are split."

Two points from that passage matter more than the split itself.

First, the seed set stops being the interesting artifact under continuous active learning. Rio Tinto states that if the TAR methodology uses CAL, "as opposed to simple passive learning (SPL) or simple active learning (SAL), the contents of the seed set is much less significant." A protocol that spends its negotiating capital on seed-set production while the vendor is running CAL is fighting over the wrong document set.

Second, Rio Tinto identifies what substitutes for seed-set access: requesting parties "can insure that training and review was done appropriately by other means, such as statistical estimation of recall at the conclusion of the review as well as by whether there are gaps in the production, and quality control review of samples from the documents categorized as non-responsive." Id. at 128-29.

That is the negotiating list. Recall estimated at the end. Gap analysis against known documents. Sampling from the discard pile.

The court then declined to decide the transparency question at all: "The Court, however, need not rule on the need for seed set transparency in this case, because the parties agreed to a protocol that discloses all non-privileged documents in the control sets." Id. at 129. It added a warning that cuts against over-lawyering these protocols: "it is inappropriate to hold TAR to a higher standard than keywords or manual review. Doing so discourages parties from using TAR for fear of spending more in motion practice than the savings from using TAR for review."

TAR 1.0 versus continuous active learning changes the protocol architecture

The distinction is not vocabulary. It determines which paragraphs of a protocol are meaningful.

One point of status first. The protocol attached to Rio Tinto is a stipulation the parties negotiated and the court signed. Peck was explicit: "The Court is approving the parties' TAR protocol, but notes that it was the result of the parties' agreement, not Court order," and its approval "does not mean ... that the exact ESI protocol approved here will be appropriate in all [or any] future cases." Id. at 129. It is a well-documented model, not a judicial standard.

That protocol is a TAR 1.0 design. Paragraph 4(b) requires a Control Set drawn as a Statistically Valid Sample, with the number of documents and the coding disclosed and all non-privileged control-set documents produced. Paragraph 4(c) governs Seed Set identification; 4(d) governs Training Sets and requires disclosure of "the estimated rates of Recall and Precision with their associated error margins." The protocol defines "Statistically Valid Sample" as a random sample of sufficient size and composition to permit statistical extrapolation with a margin of error of plus or minus 2% at the 95% confidence level, and its footnote works the arithmetic: at 95% confidence, assumed richness of 0.5 and a population of 1,000,000, that is 2,395 documents. Id. at 131 & n.2. Da Silva Moore's ESI Protocol used a 2,399-document random sample on the same 95% / plus-or-minus-2% basis. 287 F.R.D. at 199 (ESI Protocol para. E.4).

Control-set architecture assumes you can afford to estimate prevalence up front. In a low-richness collection that assumption strains: a control set sized for a 3% prevalence population is dominated by non-responsive documents, and the recall estimates built on it carry intervals wide enough to be hard to argue about. CAL workflows are generally validated at the end against the unreviewed population rather than continuously against a fixed control set.

One question is worth putting in writing before a protocol is signed: does the workflow use continuous active learning, and if so, what does "control set" mean in your paragraph 4(b)? If the answer is that the tool has no control set, the paragraph is describing a process nobody will run.

Validation: the number is the interval, not the point estimate

No decision sets a recall floor. What courts have engaged with is whether recall was measured at all, and how.

Rio Tinto's paragraph 4(f) is the fullest published model. Before production, the responding party reviews a Statistically Valid Sample of the documents categorized as likely non-responsive, and discloses in writing the number of purported non-responsive documents, the size of the validation set, the number found responsive, and the implied rate of recall — producing any responsive, non-privileged documents found. Paragraph 4(e) closes the gap everyone forgets: the responding party discloses the number of documents the software cannot evaluate for any reason, "including the unavailability of machine-readable text or documents that could not be ranked," and reviews those in their entirety.

The sharpest illustration of why this matters is recent. In WorkCo, Inc. d/b/a Toku v. LiquiFi, Inc., C.A. No. 2024-1334-JTL, 2025 WL 1168234 (Del. Ch. Apr. 21, 2025), the Court of Chancery adopted a Special Discovery Magistrate's report and, on its basis, denied the requesting party's motion to compel. The report describes what untested search terms had produced: more than 12,000 documents reviewed, 332 produced, a responsiveness rate under 3% — and, once measured, a production representing "less than 4% of the responsive documents in LiquiFi's data set." The producing party had read its low hit rates as proof that responsive documents did not exist. Statistical sampling later put the collection's richness at roughly 2.5% to 3% at 95% confidence, meaning both sides' proposed terms "performed about as well as a random selection from the documents." After the search was redesigned around agreed topics — metadata searches, search terms tested and refined through sampling, and TAR — the production reached 9,426 documents, 8,267 of them responsive, with recall "estimated to range between 88% and 98%, to a 95% confidence level."

Read that last figure carefully. It is a range ten points wide, and it is the honest form of the answer. When a producing party reports "we achieved 82% recall," the follow-up questions are the sample size, the confidence level, and the interval — because a point estimate without them is not a measurement.

The report also frames the responding party's autonomy correctly. It quotes Sedona Principle 6 — "responding parties are best situated to evaluate the procedures, methodologies, and technologies appropriate for preserving and producing their own electronically stored information," The Sedona Principles, Third Edition, 19 Sedona Conf. J. 1, 118 (2018) — and then states the corollary: "Even where a responding party has demonstrated that the requesting party's proposed search terms are ineffective, the responding is not relieved of its obligation to develop search methods that are reasonably designed to identify responsive information."

Precision is largely the requesting party's problem, and at least one court has addressed it as a scheduling matter rather than by reopening methodology. In In re Domestic Airline Travel Antitrust Litigation, MDL No. 2656, Misc. No. 15-1404 (CKK) (D.D.C. Sept. 13, 2018), the plaintiffs represented that a defendant's TAR process had produced more than 3.5 million core documents of which only approximately 17%, or 600,000, were responsive, leaving them to sort the rest. Judge Kollar-Kotelly granted a six-month extension of fact discovery and related deadlines for good cause under Federal Rule of Civil Procedure 16(b)(4), holding that the volume of non-responsive material was unforeseen and that the plaintiffs had been diligent. She expressly set aside the producing party's argument that its precision level was reasonable as "irrelevant to the issue at hand." The methodology was not reopened; the calendar moved.

Negotiating a protocol that survives contact with a motion

Rio Tinto's attached stipulation is the best publicly available checklist, and the court's own criticism of it is instructive: Peck noted he had informed counsel that their stipulated protocol "was somewhat vague and generic, which is why they felt the need to accompany it with a cover letter" summarizing the TAR processes their respective vendors would actually run. Id. at 129. If a draft protocol could describe any tool, it describes none.

These are the operative terms of that approved protocol:

  • Initial disclosure before review begins. The software's name, publisher, version number and a description of the software and the process; the name and qualifications of the technical expert who will oversee implementation; a description of the document universe by custodian/source, data type and document count, in total and per custodian; and the responsiveness categories.
  • Culling criteria, disclosed and sampled. If search terms or other criteria reduce the universe before TAR runs, disclose the criteria in writing and the number of documents removed, review a statistically valid sample of the excluded documents, disclose that sample's size, and produce any responsive, non-privileged documents found. The reason this paragraph earns its place is that TAR applied only to keyword hits inherits the keywords' recall ceiling. Rio Tinto's footnotes record both sides of that problem: Biomet ran keyword culling followed by TAR, and in Progressive Casualty Insurance Co. v. Delaney, Magistrate Judge Leen held Progressive to a stipulated keyword-then-manual protocol and would not allow it to use TAR only on the positive keyword hits.
  • A defined "statistically valid sample," with a stated confidence level and margin of error, not the phrase alone.
  • Uncategorized documents reviewed in full, with their number disclosed.
  • An end-of-process validation set drawn from the purported non-responsive population, with the size, the responsive count and the implied recall disclosed.
  • The technical expert made reasonably available to answer questions about the technical operation of the software.

Delaware has since pushed the disclosure obligation further. In Wright v. SLWM, LLC, C.A. No. 2024-1339-JTL (Del. Ch. June 25, 2025), Vice Chancellor J. Travis Laster granted a motion to compel and wrote that "[a] cooperative discovery effort also requires that the producing party share details about the search and culling techniques that were applied and provide related metrics, including search set volumes and test results relating to the effectiveness of the searches," citing WorkCo as "a concrete example of how parties can use sampling and testing to collect documents in a cost-effective and efficient way." He added that "[e]mbarking on empirically and informationally blind negotiations over search terms will rarely be an effective way to proceed," and that a responding party should instead start from the categories of information sought and build a plan combining metadata searches, sampling and technology-assisted review. The court then said it would, by separate order, appoint an experienced discovery professional as a discovery facilitator to help the parties design and implement a reasonable protocol — naming Tara Emory, who had served as Special Discovery Magistrate in WorkCo. That is the same structural move discussed in working with a discovery special master.

Two related pieces are worth reading alongside this one: how generative AI is changing document review, because large-language-model review raises the validation questions above without the settled case law; and how to read a forensic report, because collections like the one in WorkCo run forensic examination and TAR on separate tracks — the magistrate's report expressly placed the forensic examinations outside its scope — and the metrics in one do not validate the other.

Learn more

Faculty whose practice touches these issues include Katherine Charonko, Partner and ESI Practice Group Leader at Bailey & Glasser; Gail Gottehrer, VP, Global Litigation, Labor & Employment, and Government Relations at Fresh Del Monte Produce, Inc.; and David Shonka, Partner and General Counsel at Redgrave LLP and three-time Acting General Counsel of the FTC. On how these disputes are resolved when they reach a neutral, see Hon. Gregory M. Sleet, a JAMS neutral who served 20 years on the U.S. District Court for the District of Delaware, seven as chief judge, and Hon. James Orenstein, a JAMS neutral and former U.S. Magistrate Judge in the Eastern District of New York.

This article is general information about publicly available court decisions and rules of procedure. It is not legal advice, and it does not create an attorney-client relationship.

Go deeper — courses on this

Keep reading

Get the next one by email

Plain-English analysis of the law-and-technology developments that change how you advise. No more than monthly, and you can leave whenever you like.