Thomson Reuters has started transferring some of its legal AI tasks from Anthropic’s Claude to Thomson-1, which is an in-house model that uses Chinese open-source technology. This change is expected to reduce its costs, but at the same time raises another question concerning the security of using Chinese AI solutions in the uncertain environment of the industry.
According to Business Insider, Thomson-1 will reportedly carry out the document-review procedures that Claude undertook before this. Legal professionals, accountants, and various companies using CoCounsel and similar products probably won’t notice any significant changes in their operations. However, the most telling signal is the fact that the Western information company trusts the capabilities of the Chinese model that has undergone serious adjustments in-house.
Expenses played an essential role in this case. CTO Joel Hron explained to Business Insider that the high prices charged by such labs as Anthropic and OpenAI inspired Thomson Reuters to consider the development of its own model. By developing its own model, the firm can leverage its IP rather than pay for service subs from third parties.
Hron explained that Thomson-1 won’t completely take the place of Anthropic. Thomson Reuters enhanced its partnership with Anthropic in May, and CoCounsel largely uses Claude. Instead, the company plans to take it step-by-step and only assign Thomson-1 tasks when there is a benefit to the company’s own expertise. Initially, the plan is to use Thomson-1’s system for “high-volume, structured document review.”
Thomson-1 is built on Snowdon, created by “realigning” an open-source Qwen model from Alibaba. Unlike Claude, which is accessed as a closed system through Anthropic, Qwen’s open weights can be downloaded and modified.
A Thomson Reuters–Imperial College London team spent months adapting Qwen into Snowdon, which Hron described as “ethically and politically de-biased and safe to use.” Cryptopolitan has previously reported that Western companies are increasingly adapting foreign open-weight models to gain more control over their AI stacks. Hron also played down dependence on Alibaba’s model, saying “there’s nothing that necessarily ties us to Qwen.”
Thomson Reuters says Thomson-1 performs competitively with Claude Opus 4.8 and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro across its broader evaluation suite. The results remain company-reported, with independent academic benchmarking still underway. Its published scores show a mixed picture rather than a clean sweep.
The savings come with political baggage. Anthropic has accused Chinese labs of illicitly “distilling” its model outputs and has pushed Washington for restrictions. Senator Tom Cotton has also raised security concerns about US companies using Chinese open-source models.
Stanford researchers see a different problem. HAI’s James Landay argued that downloadable model weights do not make a system fully auditable because users still cannot see its training data or understand every behavior. He calls that “open distribution” rather than open source. That makes claims about successfully removing bias difficult for customers to independently verify.
| Benchmark | Thomson-1-Large | Claude Opus 4.8 | GPT-5.5 | Claude Sonnet 5 | Gemini 3.1 Pro |
| Stanford LegalBench | 0.823 | 0.818 | 0.832 | Not published | 0.843 |
| PRBench Legal Hard | 0.352 | 0.315 | 0.333 | Not published | 0.293 |
| Harvey Legal Agent Benchmark | 0.857 | 0.869 | 0.781 | Not published | 0.555 |
| Instruction Following | 0.914 | 0.861 | 0.885 | Not published | 0.848 |
| Reasoning | 0.684 | 0.737 | 0.589 | Not published | 0.748 |
| Coding | 0.399 | 0.598 | 0.414 | Not published | 0.500 |
| Long Context | 0.753 | 0.741 | 0.703 | Not published | 0.750 |
| Thomson Reuters proprietary content used in training | <10% | N/A | N/A | N/A | N/A |
Companies are willing to consider the trade-off because Chinese models have rapidly caught up. Stanford’s 2026 AI Index said the US-China performance gap had “effectively closed.” On Arena’s Text leaderboard as of August 21, 2026, the leading US model scored 1,508 against 1,489 for the top Chinese model—a gap of just 1.3%. Alibaba’s Qwen3.8-Max scored 1,481. When capability differs by only a few points while costs remain far apart, document review becomes an obvious testing ground for cheaper alternatives.
If you're reading this, you’re already ahead. Stay there with our newsletter.