Technology-assisted review (TAR), often called predictive coding or computer-assisted review, has become a common method for identifying responsive documents in large productions of electronically stored information. Its defensibility turns on validation—the statistical measurement of how well the process performs. This article explains the two core metrics, recall and precision, and situates a validation protocol within the framework of the Federal Rules and court guidance on proportional, cooperative ESI discovery.
What Recall and Precision Measure
Reliability in a TAR process is determined by statistical methods that measure two distinct quantities. Recall is "the percentage of responsive documents in the entire data set that the computer has located," while precision is "the percentage of documents within the computer's output set that are actually responsive."
The distinction matters for validation because the two metrics test opposite kinds of error. As the commentary explains, "'recall' tests the extent to which the predictive coding system misses responsive documents, while 'precision' tests the extent to which the system is mixing irrelevant documents in with the production set." A validation protocol that reports only one metric leaves the other risk unmeasured. High recall paired with low precision means the output captures responsive material but drags in a large volume of irrelevant documents; high precision paired with low recall means the produced set is clean but incomplete.
How Validation Fits the TAR Workflow
Validation is not a bolt-on step; it is the mechanism that tells the parties when the iterative training is finished. In a common form of predictive coding, "representative samples of the electronic documents are identified," counsel review these "seed sets" and code each document, and the system generates a "training set" reflecting its determinations of responsiveness. Counsel then train the computer by evaluating where their decisions differ from the system's and making adjustments.
This cycle is repeated "until the system's output is deemed reliable," and reliability is precisely what recall and precision measure. In other words, the statistical metrics supply the stopping criterion for training. The resulting output "can be either produced as is or further refined by subsequent human review," so a validation protocol should also specify whether measured performance is acceptable on its own or triggers additional review.
Why Validation Is Worth the Effort
The impetus for TAR is cost. Discovery costs have "ballooned as people increasingly write, transmit, and store documents electronically," and predictive coding "holds the promise of dramatically reducing the costs of discovery in cases involving large volumes of data." In large productions the savings are significant: "If humans need not look at a significant percentage of the collected documents, the savings over millions of documents is tremendous."
There is also evidence that predictive coding, "when used properly and in the right circumstances—may be more reliable than the more expensive methods it replaces." But that qualifier—"used properly"—is where validation earns its keep. The technology is expensive to set up and "is not the right tool for every case," so a validation protocol should be scaled to the matter rather than applied reflexively.
Anchoring the Protocol in Proportionality and Cooperation
A TAR validation protocol operates within the discovery rules governing the production of ESI. Rule 34 permits a party to request electronically stored information "stored in any medium from which information can be obtained," and a request "may specify the form or forms in which electronically stored information is to be produced." The scope of any such request is bounded by Rule 26(b).
Court guidance emphasizes that the proportionality standard applies across the discovery lifecycle. The District of Maryland's ESI Principles direct that parties "apply the proportionality standard set forth in Fed. R. Civ. P. 26(b) to all phases of the discovery of ESI, including the identification, preservation, collection, search, review, and production of ESI." Search and review—the phases where TAR operates—are expressly named, so the recall and precision targets a party sets should be defensible as proportional to the needs of the case.
Those same principles call for cooperation. The court "expects cooperation on issues relating to the preservation, collection, search, review, production, integrity, and authentication of ESI" and "particularly emphasizes the importance of cooperative exchanges of information about ESI at the earliest stages of litigation." For a neutral, this is the practical center of gravity: agreeing early on the validation methodology, sample sizes, and metric thresholds is far more efficient than litigating adequacy after production. A protocol negotiated transparently under the cooperation principle reduces the likelihood of later disputes about whether the process was reliable.
This article is provided for general informational purposes only and does not constitute legal advice. Engagement of Daniel Garrie as a neutral is administered exclusively through JAMS.