Aug 13, 2026 |

When Legal AI Is Evaluated Like Legal Work, the Results Change

Tyler Alexander  VP, CoCounsel AI Reliability, ³ÉÈËVRÊÓÆµ

In my last post, I introduced CoCounsel Bench, or CoCoBench, and explained why legal AI needs to be evaluated against the work lawyers actually perform. Benchmarking means testing a system against a fixed set of tasks with a known standard for a correct answer, and it matters because the tasks you choose determine what ‘good performance’ even means: if the tasks don’t capture how legal work actually unfolds, a high score doesn’t tell you much.Ìý

Now we are beginning to see what happens when that standard is applied. In one recent evaluation, experienced attorneys reviewed CoCounsel Legal’s performance across 50 complex CoCoBench tasks. Each task was estimated to require a lawyer an average of six hours to complete.

CoCounsel Legal completed each task in less than eight minutes.

More significantly, attorneys determined that CoCounsel Legal produced a stronger response than the original expert-written reference answer on nearly 40% of the tasks. The speed is remarkable. But speed is not the most important finding.

The more consequential result is that an agentic legal system, evaluated by experienced lawyers against realistic legal work, can do more than produce a plausible response. It can produce work that attorneys judge to be complete, accurate, well-supported, and usable in practice. That is a materially different standard from performing well on a public benchmark or delivering an impressive product demonstration.

And it changes how legal AI performance should be understood.

From benchmark performance to professional performance

Traditional AI evaluation often begins with a predefined answer and asks whether the system reproduced the expected elements.

That can be useful for testing a discrete capability. But legal work is rarely a matter of locating one answer or checking one box.

A lawyer may need to review an unfamiliar complaint, identify the relevant claims, research the governing law, assess potential defenses, and translate the analysis into a memo appropriate for a client. The value of the final product depends on how well all of those steps work together.

A response can identify the right doctrine but apply it incorrectly. It can reach a defensible conclusion while omitting a material issue. It can cite a real authority that does not actually support the proposition attached to it.

Those failures can disappear inside a benchmark score reliant solely on binary criteria. They are much harder to hide from an experienced attorney reviewing the output as actual work product.

That is the shift CoCoBench is designed to make: from measuring whether data is present to determining whether the work is professionally usable.

What a realistic legal AI test looks like

Consider one of the tasks included in CoCoBench. The scenario begins when a client receives an antitrust complaint naming it as a defendant. The client asks counsel to assess the claims and identify potential defenses.

The agent receives the complaint as its sole source document. It must identify the salient facts, research the applicable law for each claim and defense, evaluate which defenses may be viable, and prepare a memo written for the client.

This is the type of assignment litigators routinely face at the beginning of a matter. The available information may be incomplete. The issues may be ambiguous. The relevant law is not packaged neatly inside the source materials.

The task was authored by Jon Faria, a Senior Specialist Legal Editor at Practical Law who previously practiced as an antitrust litigation and investigations partner at Kirkland & Ellis.

That background matters.

Jon is not constructing a test that merely resembles legal work. He is reconstructing a problem he encountered in practice and applying the expectations he would have brought to an associate’s work product.

That is fundamentally different from generating a synthetic scenario and using another model’s answer as the standard of correctness.

The difference attorney judgment makes

CoCoBench combines automated evaluation with direct review by experienced attorneys.

The automated layer allows every run to be evaluated consistently and at scale. It identifies regressions, isolates specific failures, examines citation support, and helps engineering and data science teams understand where performance is improving.

Attorney review asks a more demanding question: would this work hold up in practice? Attorneys assess outputs across four dimensions:Ìý

  • Correctness: Are the factual and legal claims accurate, and are the citations used properly?Ìý
  • Completeness: Does the response address the full assignment, including the analysis that materially affects the conclusion?Ìý
  • Readability: Is the output organized and clear enough for a practitioner to use without reconstructing it?Ìý
  • Overall judgment: Taken as a whole, is the work fit for professional use?Ìý

This review can reveal distinctions that a conventional benchmark may flatten. Two systems might both mention the relevant legal standard. One simply states it. The other explains its elements, applies them to the facts, addresses competing interpretations, and reaches a conclusion that a lawyer could defend.

A presence-based benchmark may reward both. A practicing attorney will not treat them as equivalent.

The hardest failures are often the ones that look right

One of the most important findings from our evaluation work is that legal AI failures are not always obvious.

A fabricated case is serious, but it is also relatively easy to recognize once someone attempts to verify it.

Misattribution can be more dangerous. A system may make a legally accurate statement and cite a real case, yet the cited passage does not actually support the claim. Everything appears credible: the authority exists, the proposition sounds plausible, and the citation is formatted correctly.

The failure lies in the connection between the claim and the authority. CoCoBench evaluates that relationship directly. It distinguishes among claims that are properly supported, claims that lack citations, claims based on incorrect reasoning, fabricated authorities, and propositions attributed to the wrong source.

It also differentiates between levels of misattribution. A passage that partially supports a reasonable inference is not the same as a claim whose real support appears only in an entirely different authority.

Those distinctions matter because they point to different technical problems and different levels of professional risk. A single accuracy score cannot show that.

Why the results matter

The initial results demonstrate what becomes possible when an advanced agentic system is paired with authoritative legal content, realistic evaluation, citation verification, and continuous attorney involvement. But they also expose a broader problem in the legal AI market.

Systems are often compared using measures that reward fluency, isolated task completion, or performance against synthetic reference answers. Those measures can create the appearance that several products perform at roughly the same level. When the evaluation moves closer to real legal work, the differences become clearer.

Can the system sustain accuracy across a six-hour assignment rather than a single turn prompt? Can it recognize which issues require deeper research? Can it connect each legal proposition to the authority that actually supports it? Can it produce a deliverable that an experienced lawyer would use as a credible starting point without redoing the core work? Those are much harder questions. They also require far more than access to a frontier model.

They require realistic legal tasks authored by practitioners. They require gold-standard responses grounded in substantive expertise. They require evaluation systems capable of testing both the final deliverable and the sources behind it. They require attorneys who can distinguish between an answer that sounds right and work that is right.

³ÉÈËVRÊÓÆµ can bring those elements together because legal expertise is not being added to the system at the end. It exists throughout the process.

It is embedded in Westlaw and Practical Law content authored, reviewed, and continuously maintained by attorney-editors. It guides the agents. It shapes the tasks used to evaluate the system. It informs the criteria against which outputs are judged. And it provides the continuous feedback used by our product, engineering, and data science teams to improve CoCounsel Legal.

That combination is difficult to build and even harder to sustain at scale.

A higher bar for legal AI

The question facing the legal industry is no longer whether AI can generate legal language.

It can. The question is whether an AI system can complete complex legal work accurately enough, thoroughly enough, and transparently enough for a professional to rely on it.

As agents take on longer workflows, involvement in evaluation becomes more important, not less. Errors introduced early can carry through research, analysis, drafting, and revision. A polished final answer can conceal weaknesses in the work that produced it. That is why the standard cannot be how intelligent the output sounds.

The standard must be whether the work can be examined, verified, explained, and defended.

The early CoCoBench results show that this standard is achievable. They also show why the way legal AI is evaluated will increasingly determine which systems are truly ready for professional work.

Because once legal AI is evaluated like legal work, the leaderboard changes.

Read the for a deeper look at CoCoBench’s attorney-authored tasks, automated evaluation framework, citation analysis, and attorney review methodology.Ìý

Share