Jul 31, 2026 | AI and product innovation
³ÉÈËVRÊÓÆµ Built Its Own AI Model That Now Ranks Among the World’s Best
The most capable AI models no longer come only from frontier AI labs. One now comes from ³ÉÈËVRÊÓÆµ.Ìý
Today, we are sharing early benchmarking results for Thomson, a first of its kind AI model. Across a range of benchmarks assessing legal and general capabilities, Thomson performed competitively with the strongest frontier models on the market, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro. Ìý
Why? Because it knows the work.Ìý
Launching later this summer, Thomson is the newest layer of the ³ÉÈËVRÊÓÆµ AI strategy, and a demonstration of what becomes possible when authoritative content, expert judgment, professional tools, and model development come together.Ìý
In 2024, ³ÉÈËVRÊÓÆµ acquired Safe Sign Technologies, an AI research company. At the time, the market was betting that access to increasingly powerful general-purpose models would be enough.Ìý
We made a different bet. We believed the future of professional AI would require more than general-purpose intelligence. It would require models built specifically for the domains, standards, and consequences of professional work. We believed that the distinct advantages of ³ÉÈËVRÊÓÆµ decades of world-class content and expertise could be best expressed in a model that we ourselves crafted.ÌýÌý
Thomson is the result of that bet. And it is why ³ÉÈËVRÊÓÆµ will continue to set the standard for Fiduciary-Grade AIâ„¢. Ìý
Meet ThomsonÌý
Thomson starts from a strong open-source foundation, so it performs general-purpose work just as effectively as the frontier models. It then goes further: trained using state of the art mid-training and post-training techniques on decades of authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, content professionals have staked their reputations on for generations.Ìý
That training was shaped by hundreds of subject matter experts who evaluated outputs, identified failure modes, and validated that the model reasons the way legal professionals actually work. The same professional standard governs how Thomson is deployed. Customer data is never used to train the model.Ìý
The result is a model that thinks and reasons like a lawyer while outperforming models multiple times larger on the work that matters.ÌýÌý
Thomson Matches the Best. And Beats the Rest.ÌýÌý
We evaluated Thomson against the leading general-purpose models on the market for general professional work and categories spanning:Ìý
| LegalÌý | CodingÌýÌý |
| TaxÌýÌý | MathÌýÌý |
| AccountingÌýÌý | MultilingualismÌýÌý |
| JournalismÌýÌý | Agentic tasksÌýÌý |
| SafetyÌý | Long contextÌýÌý |
| ReasoningÌýÌý | Following instructionÌýÌý |
Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate. ³ÉÈËVRÊÓÆµ has achieved that performanceÌý by combining exceptional AI talent with authoritative proprietary content and deep domain expertise. And with less than 10% of ³ÉÈËVRÊÓÆµ content used in its training so far, there remains significant opportunity to expand its capabilities.Ìý

Instruction Following is a composite average of the IFEval and FollowBench benchmarks. Ìý
Reasoning is a composite average of the GPQA Diamond, HLE, and MMLU-Pro benchmarks.Ìý Ìý
Coding is a composite average of the SWE-Bench Pro and Terminal-Bench 2.1. Ìý
Long Context is a composite average of the Infinity Bench as well as some internal benchmarks developed by ³ÉÈËVRÊÓÆµ.Ìý
These evaluations show Thomson’s competitiveness with industry recognized benchmarks. Our internal evaluation and training cover a wide range of scenarios, including carefully designed agentic use cases optimized for real-world professional work, tens of thousands of real-world queries written by experts, end-to-end deep research training with human-calibrated judges, and data-centric mid-training on a large amount of our content. To increase safety and robustness, we conduct training and evaluations consistent with ³ÉÈËVRÊÓÆµ values, and stress-test the models through human and automated red-teaming. Thomson is still early in its development. To date, less than 10% of ³ÉÈËVRÊÓÆµ content has been used in its training, leaving significant opportunity to expand its domain knowledge and capabilities through additional training, rigorous evaluation, and expert validation.Ìý
Thomson also showcases its strength when a native integration with ³ÉÈËVRÊÓÆµ content is added. When compared with leading frontier models given unrestricted access to the web, Thomson’s access to proprietary data sources such as Westlaw, Practical Law, and Reuters news ensures both superior completeness and factuality (i.e. the ability to back up claims through accurate citations to trusted sources).

This evaluation covers 53 legal research queries written by our internal Subject Matter Experts to represent real world questions. LLM’s are connected via an in-house agentic harness to Westlaw/Practical Law for TR content and Brave search engine for web search. Completeness and Factuality are scored using LLM’s as a judge. Completeness is scored based on SME-written rubrics that list every element that would be required for a good answer to the question. Factuality is based on extracting the claims made in each report and checking whether the cited sources provide evidence for each claim. These metrics were developed and calibrated against SME scoring.
The First Deployment. Not the Last.Ìý
Thomson’s first integration will launch in August inside Tabular Analysis in CoCounsel Legal, and that choice was deliberate. Tabular Analysis performs high-volume, structured document review against a clear, measurable accuracy standard. It is exactly the kind of work where a purpose-built model has a demonstrable advantage over a general-purpose alternative, and where that advantage is immediately visible to the professionals relying on the output. Thomson will become the default model powering Tabular Analysis, and over the next year we will continue to integrate it across the ³ÉÈËVRÊÓÆµ product portfolio in legal and tax.ÌýÌý
Content. Expertise. Tools. Now the Model.Ìý
³ÉÈËVRÊÓÆµ has always brought together authoritative content, deep domain expertise, and the tools professionals rely on every day. Thomson adds the fourth element: a model purpose-built to power it all, trained on content competitors cannot access and validated to the standard professional’s demand. It delivers frontier-level performance at a fraction of the size and operating cost of many general-purpose models. A model only ³ÉÈËVRÊÓÆµ could build.Ìý
This is our commitment to Fiduciary-Grade AIâ„¢ in action: AI designed for professionals with duties of care and accountability, where almost right is not good enough.ÌýÌý
General-purpose AI is built for everyone. Thomson is built for the professionals who cannot afford to be wrong.