{"id":159457,"date":"2026-08-04T10:28:26","date_gmt":"2026-08-04T18:28:26","guid":{"rendered":"https:\/\/xira.com\/p\/2026\/08\/04\/thomson-reuters-says-its-homegrown-ai-model-now-rivals-the-frontier-labs-i-take-a-closer-look-at-the-benchmarks\/"},"modified":"2026-08-04T10:28:26","modified_gmt":"2026-08-04T18:28:26","slug":"thomson-reuters-says-its-homegrown-ai-model-now-rivals-the-frontier-labs-i-take-a-closer-look-at-the-benchmarks","status":"publish","type":"post","link":"https:\/\/xira.com\/p\/2026\/08\/04\/thomson-reuters-says-its-homegrown-ai-model-now-rivals-the-frontier-labs-i-take-a-closer-look-at-the-benchmarks\/","title":{"rendered":"Thomson Reuters Says Its Homegrown AI Model Now Rivals the Frontier Labs \u2013 I Take A Closer Look At the Benchmarks"},"content":{"rendered":"<p>Thomson Reuters says the large language model it has been quietly building in-house now performs competitively with the best general-purpose AI models in the world, and it has released its first benchmarking results to back up the claim. In a July 31 post on the company\u2019s innovation blog, Chief Technology Officer Joel Hron and Head [\u2026]<\/p>\n<p>Thomson Reuters says the large language model it has been quietly building in-house now performs competitively with the best general-purpose AI models in the world, and it has released its first benchmarking results to back up the claim.<\/p>\n<p>In a <a href=\"https:\/\/blogs.thomsonreuters.com\/en-us\/innovation\/thomson-reuters-built-its-own-ai-model-that-now-ranks-among-the-worlds-best\/\" rel=\"nofollow noopener\" target=\"_blank\">July 31 post on the company\u2019s innovation blog<\/a>, Chief Technology Officer <a href=\"https:\/\/www.linkedin.com\/in\/joel-hron-90a3421a\/\" rel=\"nofollow noopener\" target=\"_blank\">Joel Hron<\/a> and Head of AI Research <a href=\"https:\/\/www.linkedin.com\/in\/schwarzjonathan\/\" rel=\"nofollow noopener\" target=\"_blank\">Jonathan Schwarz<\/a> shared early benchmark data for Thomson, the eponymously named proprietary model the company has been developing since its 2024 acquisition of the AI research startup Safe Sign Technologies.<\/p>\n<p>Across a range of legal and general-purpose benchmarks, they wrote, Thomson performed competitively with Anthropic\u2019s Claude Opus 4.8 and ahead of OpenAI\u2019s GPT-5.5, Anthropic\u2019s Claude Sonnet 5, and Google\u2019s Gemini 3.1 Pro.<\/p>\n<p>\u201cThe most capable AI models no longer come only from frontier AI labs,\u201d Hron and Schwarz wrote. \u201cOne now comes from Thomson Reuters.\u201d<\/p>\n<p>I\u2019ll go into what the numbers actually show, how the tests were run, and what remains unverified, but first, some background on Thomson.<\/p>\n<h3><strong>What Thomson Is<\/strong><\/h3>\n<p>\u201cLaunching later this summer, Thomson is the newest layer of the Thomson Reuters AI strategy, and a demonstration of what becomes possible when authoritative content, expert judgment, professional tools, and model development come together,\u201d Hron and Schwarz say in their post.<\/p>\n<p>According to the post, Thomson starts from an open-source foundation model and is then trained, using what they describe as state-of-the-art mid-training and post-training techniques, on decades of proprietary content from Westlaw, Practical Law, Checkpoint and Reuters.<\/p>\n<p>Hundreds of subject-matter experts evaluated outputs, identified failure modes, and validated the model\u2019s reasoning, they say, adding that customer data is never used in training.<\/p>\n<p>Notably, they say that less than 10% of Thomson Reuters\u2019 proprietary content has been used in training so far \u2013 suggesting there is plenty of room for Thomson\u2019s capabilities to further improve.<\/p>\n<p>The model\u2019s first deployment will come in August, when Thomson becomes the default model powering Tabular Analysis in CoCounsel Legal.<\/p>\n<p>Hron and Schwarz say they chose Tabular Analysis because that type of high-volume, structured document review offers a clear, measurable accuracy standard where a purpose-built model\u2019s advantage should be immediately visible. Over the next year, they write, Thomson Reuters will integrate the model across its legal and tax product portfolio.<\/p>\n<h3><strong>The Numbers<\/strong><\/h3>\n<p><a href=\"https:\/\/i0.wp.com\/www.lawnext.com\/wp-content\/uploads\/2026\/08\/TR8347338_02A_1080x1080-1.jpg?ssl=1\" rel=\"nofollow noopener\" target=\"_blank\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-53933\" src=\"https:\/\/i0.wp.com\/www.lawnext.com\/wp-content\/uploads\/2026\/08\/TR8347338_02A_1080x1080-1-1024x1024.jpg?resize=1024%2C1024&#038;ssl=1\" alt=\"\" width=\"1024\" height=\"1024\" title=\"\"><\/a><\/p>\n<p>The benchmark that drove the headline pits Thomson (technically called Thomson-1-Large) against Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5 across three legal benchmarks and four general-domain composites. Thomson took the top score in three of the seven rows:<\/p>\n<p><strong>Legal domain:<\/strong><\/p>\n<ul>\n<li><strong>Stanford LegalBench:<\/strong> Thomson 0.823 \u2013 behind Gemini 3.1 Pro (0.843) and GPT-5.5 (0.832), ahead of Opus 4.8 (0.818).<\/li>\n<li><strong>PrBench Legal Hard:<\/strong> Thomson 0.352 \u2013 best in class, ahead of GPT-5.5 (0.333), Opus 4.8 (0.315), and Gemini (0.293).<\/li>\n<li><strong>Harvey Legal Agent Benchmark:<\/strong> Thomson 0.857 \u2013 a close second to Opus 4.8 (0.869), and well ahead of Gemini (0.555).<\/li>\n<\/ul>\n<p><strong>General domain:<\/strong><\/p>\n<ul>\n<li><strong>Instruction following<\/strong> (IFEval + FollowBench): Thomson 0.914 \u2013 best in class.<\/li>\n<li><strong>Reasoning<\/strong> (GPQA Diamond, HLE, MMLU-Pro): Thomson 0.684 \u2013 behind Gemini (0.748) and Opus (0.737).<\/li>\n<li><strong>Coding<\/strong> (SWE-Bench Pro, Terminal-Bench 2.1): Thomson 0.399 \u2013 last place, well behind Opus (0.598).<\/li>\n<li><strong>Long context<\/strong> (Infinity Bench plus internal TR benchmarks): Thomson 0.753 \u2013 best in class, by a hair over Gemini (0.750).<\/li>\n<\/ul>\n<p>In their blog post, Hron and Schwarz write, \u201cThomson matches the best. And beats the rest.\u201d Taking the benchmark scores at face value, that is a fair characterization, particularly given that Thomson is a fraction of the size and operating cost of the frontier models it was up against. For a domain-specific model built by a legal information company, this is a notable achievement.<\/p>\n<p>But the results are more mixed than the headline suggests. Thomson did not win a majority of the rows. Gemini and GPT-5.5 both beat Thomson in the best-known legal benchmark, LegalBench. Anthropic\u2019s Opus edged Thomson on the Harvey agentic benchmark and beat everyone on coding.<\/p>\n<p>In addition, the \u201cfine print\u201d in the above graphic reveals that the test conditions were not entirely equal. Thomson used test-time scaling \u2014 which gives it more time and computing power. Gemini and Opus ran in the more-powerful reasoning mode, but GPT-5.5 was tested in non-reasoning mode.<\/p>\n<p>While non-reasoning mode is faster and more efficient for routine tasks, it is typically less capable when handling complex, advanced or multi-step challenges. That suggests that the benchmark may have understated OpenAI\u2019s performance.<\/p>\n<h3><strong>Some Benefit to Thomson<\/strong><\/h3>\n<p>Through a Thomson Reuters spokesperson, I asked Hron about this. He said that Thomson did, in fact, benefit from test-time scaling on certain tasks, but that the overall effect was minor. \u201cIts grand average increased from 0.787 without test-time scaling to 0.789 with it,\u201d he said, \u201cand the conclusions from the benchmarks would not have materially changed without it.\u201d<\/p>\n<p>With regard to the use of non-reasoning mode for GPT-5.5, Hron said that the reasoning-mode evaluation took substantially longer to complete, and that those results were not ready in time for publication.<\/p>\n<p>\u201cWe expect reasoning mode to improve its performance but based on the results we have seen so far, not dramatically,\u201d Hron said. \u201cWe are continuing to run those evaluations and plan to include the results in the technical report.\u201d<\/p>\n<p>Hron said it is also important to note that reasoning mode and test-time scaling are not necessarily direct analogs.<\/p>\n<p>\u201cA model operating in a mode described as \u2018non-reasoning\u2019 may still benefit from test-time scaling or other inference-time techniques implemented behind the provider\u2019s API,\u201d he said. \u201cBecause external providers generally do not disclose the full details of those systems, we cannot know precisely which techniques were applied.<\/p>\n<p>\u201cWe made a best effort to ensure a fair comparison based on the information available and disclosed Thomson\u2019s use of test-time scaling in the interest of transparency.\u201d<\/p>\n<h3><strong>The Retrieval Test<\/strong><\/h3>\n<p>In a second evaluation, Hron and Schwarz compared Thomson\u2019s performance when integrated with Thomson Reuters\u2019 proprietary content with leading frontier models given unrestricted access to the web.<\/p>\n<p><a href=\"https:\/\/i0.wp.com\/www.lawnext.com\/wp-content\/uploads\/2026\/08\/TR8347338_02B_1920x1080-scaled-1.jpg?ssl=1\" rel=\"nofollow noopener\" target=\"_blank\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-53934\" src=\"https:\/\/i0.wp.com\/www.lawnext.com\/wp-content\/uploads\/2026\/08\/TR8347338_02B_1920x1080-scaled-1-1024x576.jpg?resize=1024%2C576&#038;ssl=1\" alt=\"\" width=\"1024\" height=\"576\" title=\"\"><\/a><\/p>\n<p>Specifically, they tested completeness and factuality on 53 legal research queries written by Thomson Reuters\u2019 internal subject-matter experts. Thomson was connected via an in-house agentic harness to Westlaw and Practical Law, while the competing frontier models were given unrestricted web access through the Brave search engine.<\/p>\n<p>Completeness was scored against SME-written rubrics listing every element a good answer would require. Factuality was scored by extracting each claim in a report and checking whether the cited sources actually supported it. Scoring was done by LLMs as judges, calibrated against expert scoring.<\/p>\n<p>Unsurprisingly, Thomson-plus-Westlaw beat frontier-models-plus-web-search on both dimensions.<\/p>\n<p>I say \u201cunsurprisingly\u201d because this is less a test of the models than of the retrieval sources behind them. It seems fair to assume that, for legal research questions, a model grounded in Westlaw and Practical Law should beat a model fishing the open web. It also seems that would be true regardless of whether the model was Thomson, Opus or GPT.<\/p>\n<p>In that sense, the evaluation was not so much of the Thomson model itself, but of the value of Thomson Reuters\u2019 content it drew on. What it does not tell us is what would happen if the frontier models were given the same Westlaw access.<\/p>\n<p>As Anthropic, OpenAI and others push deeper into legal workflows, and as tools such as MCP connectors make it increasingly feasible to pipe authoritative legal content into general-purpose models, that is a comparison I would like to see.<\/p>\n<p>Here again, I asked Hron about this. Here is his response:<\/p>\n<blockquote>\n<p><em>\u201cWe did test other frontier models with access to Thomson Reuters content across dimensions including factuality, completeness, conciseness, understandability, relevance and coherence. Importantly, these tests did not use the CoCounsel Legal harness, which includes a more sophisticated combination of models, tools and capabilities. Instead, we tested each model using a simpler agentic harness with native search access to our content.<\/em><\/p>\n<p><em>\u201cAcross the four models compared, overall scores ranged from 0.81 to 0.91, with Thomson scoring 0.89. This reinforces our view of the value of authoritative Thomson Reuters content, while also demonstrating that Thomson is broadly competitive with frontier models under equivalent conditions.\u201d<\/em><\/p>\n<\/blockquote>\n<h3><strong>Caveats and Context<\/strong><\/h3>\n<p>Needless to say, these benchmark results are entirely self-reported on the part of Thomson Reuters, with no mention of any independent third-party verification. Undoubtedly, as Thomson is put into wider use, there will be independent evaluations.<\/p>\n<p>Hron acknowledges this in a statement the company provided: \u201cInternal evaluations are necessary because they allow us to test against the real workflows, failure modes and quality standards our customers encounter. We also recognize the importance of credible third-party validation and expect that to be an important part of how Thomson is evaluated over time.\u201d<\/p>\n<p>When I <a href=\"https:\/\/www.lawnext.com\/2026\/06\/thomson-reuters-ceo-steve-hasker-on-the-next-generation-of-cocounsel-the-future-of-professionals-report-and-why-tr-is-building-its-own-llm.html\" rel=\"nofollow noopener\" target=\"_blank\">interviewed Thomson Reuters CEO Steve Hasker<\/a> in June, he talked about the development of Thomson and said one of the reasons for building it was that it provided the company with a degree of flexibility beyond reliance on the flagship LLMs.<\/p>\n<p>He also talked a lot about the concept the company has been emphasizing of delivering \u201cFiduciary-Grade AI,\u201d and Hohn and Schwarz use that phrase twice in their post, positioning Thomson as a development that is in furtherance of that commitment.<\/p>\n<p>It might also be a signal of where the legal AI market is headed.<\/p>\n<p>\u201cThese results show that the most capable AI for professional work does not have to come onlyfrom the largest frontier labs,\u201d Hron said. \u201cThomson is a fraction of the size and cost of many leading models, yet it performs competitively with the strongest models available and outperforms leading models in several evaluated categories. That is what becomes possible when you build the model around the work.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Thomson Reuters says the large language model it has been quietly building in-house now performs competitively with the best general-purpose AI models in the world, and it has released its first benchmarking results to back up the claim. In a July 31 post on the company\u2019s innovation blog, Chief Technology Officer Joel Hron and Head [&hellip;]<\/p>\n","protected":false},"author":3,"featured_media":159458,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"_et_pb_use_builder":"","_et_pb_old_content":"","_et_gb_content_width":"","_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[24],"tags":[],"class_list":["post-159457","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-lawsite"],"jetpack_featured_media_url":"https:\/\/i0.wp.com\/xira.com\/p\/wp-content\/uploads\/2026\/08\/Thomson_Reuters_Building169-UMwoic.jpg?fit=611%2C344&ssl=1","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/posts\/159457","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/comments?post=159457"}],"version-history":[{"count":0,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/posts\/159457\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/media\/159458"}],"wp:attachment":[{"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/media?parent=159457"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/categories?post=159457"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/xira.com\/p\/wp-json\/wp\/v2\/tags?post=159457"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}