Thomson Reuters builds its own legal AI language model
For about $40 million, Thomson Reuters built “Thomson,” its first proprietary large language model — based on Alibaba's open Qwen model and trained on decades of legal data from Westlaw.

Thomson Reuters has unveiled “Thomson,” its first proprietary large language model (LLM). It is built on Qwen, the open model from Chinese company Alibaba — the company says it currently uses the Qwen3.5-397B version. Thomson Reuters says it has invested about $40 million over two years in staff and compute for the project, according to a report by The Decoder.
The single final training run of the current version reportedly cost about $450,000 — a figure that gets cited more often but represents only a fraction of the real cost. Even the $40 million doesn't capture the real underlying value: decades of content from Westlaw, Practical Law, Checkpoint and Reuters, plus the work of hundreds of legal domain experts who contributed to the project over the years.
How an open model became a legal model
Working with Imperial College London, Thomson Reuters first realigned the open Chinese model for safety, ethics and political neutrality. This intermediate step carries the internal codename “Snowdon,” after a mountain in Wales. It was followed by pre-training on the company's proprietary content, post-training with domain experts, and agentic reinforcement learning inside the company's own tool environments, such as Westlaw. According to the company, less than 10% of the available content has been used for training so far.
CTO Joel Hron told The Decoder that the project's starting point was changed “about half a dozen times or so.” Head of research Jonathan Schwartz added that the real achievement is not a single model, but a reusable “model factory.”
Thomson needs to set the frontier of intelligence for legal. That's a different job than what I think a lot of the frontier labs are doing.
If you simply take the open-source model without any of these additional steps, it won't know as much. You won't necessarily be aligned with your values, and it won't be as good as using the tools that you've built later on.
Benchmarks that call for caution
On the company's own tests, Thomson scores 0.823 on the Stanford LegalBench benchmark, behind Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark, it lands just behind Claude Opus 4.8. Thomson leads on instruction following and on the difficult PrBench Legal, but falls clearly behind on reasoning and especially on coding. The comparison is methodologically biased: Thomson is tested with test-time scaling, while GPT-5.5 is tested without its reasoning mode.
On the company's internal Deep Research benchmark, Thomson scores 0.53 for factual fidelity with web access alone, against 0.65 for GPT-5.4. With access to the group's proprietary content, however, Thomson's score climbs to 0.83, edging past GPT-5.4's 0.82. Andrew Bean, who leads the evaluation, acknowledges that with web access alone the model is “certainly not yet” leading. Notably, GPT-5.4 also benefits heavily from access to proprietary content — data access counts almost as much as specialized training itself. SiliconANGLE notes these internal results have not yet received extensive independent validation; a technical report with more benchmark results is expected later. Thomson Reuters has started sharing the model with legal experts and academic institutions for testing.
Why build rather than fine-tune a frontier model
Why build a proprietary model instead of fine-tuning an existing frontier model from OpenAI or Anthropic? The company gives three reasons:
- Economics: standard fine-tuning often degrades a model's general capabilities and leaves the company dependent on inference costs and the vendor's roadmap; a smaller proprietary model is profitable for high-volume tasks like document review.
- Data: training inside proprietary tool environments like Westlaw produces exactly the performance jump observed — access the company doesn't want to grant to any third party.
- Compounding effect: every expert evaluation made during product updates becomes training data that accumulates like equity with a proprietary model, rather than benefiting an outside vendor.
The company acknowledges limits to the approach: it only generalizes to companies with three rare ingredients — exclusive databases, hundreds of in-house domain experts, and workflows where quality can be measured objectively. For a company with that profile, the open-source community only delays catching up with frontier labs by a few months, and $40 million can be enough for a competitive specialized model. For a company without proprietary data or evaluation infrastructure, a proprietary model would mostly cost in ongoing maintenance.
First deployment: inside CoCounsel
Thomson is first integrated into the “Tabular Analysis” feature of CoCounsel Legal, Thomson Reuters' AI assistant for legal work, handling high-volume document review — a use case where a smaller, cheaper model makes economic sense. CoCounsel remains multi-model: administrators can still choose other models. Thomson isn't meant to orchestrate the whole system, but to handle subtasks such as citation verification. According to the company, customer data isn't used for training.
A smaller version of the model will be released with open weights on Hugging Face under a non-commercial license, alongside a technical report and a developer portal. Preliminary, non-binding discussions are also under way with law firms about direct licensing.
Thomson Reuters points to the broader context: its flagship Westlaw platform comprises more than 40,000 individual databases and more than 150 years of legal publishing and editorial curation. The company says controlling its own model gives it more authority over deployment, governance and future development — an “AI sovereignty” argument. Discussions are already under way with major law firms and companies about direct access to the model, including the possibility that clients adapt Thomson to their own knowledge and workflows.
For law firms and corporate legal departments, the case shows under what conditions a homegrown language model is even economically worthwhile: only an organization with exclusive, decades-deep data, in-house domain experts and objectively measurable quality standards can build, on a comparatively modest budget, a specialized model that keeps pace with frontier labs. For most companies, integrating existing models through an agentic system is likely to remain the more pragmatic path, especially for automating recurring legal and administrative tasks.
Sources
- Warum Thomson Reuters 40 Millionen Dollar in ein eigenes Sprachmodell investiertThe Decoder · August 24, 2026
- Thomson Reuters launches proprietary AI model for legal workSiliconANGLE · August 24, 2026



