AI for legislative content: new Large Language Models ranking

This article was updated on 03/08/2026.

The Regulatory Institute undertook another comparative Large Language Models (LLMs) test. Its purpose was to assess the ability of LLMs to identify and draft candidate policy and regulatory elements that make a law more effective. The results are surprising. Previous test results have mostly become obsolete, whilst our generic recommendations on how to use LLMs remain valid. 

OpenAI’s GPT 5.6‑Sol was this time the frontrunner, whilst its sisters, OpenAI GPT 5.6 Terra Thinking and OpenAI GPT 5.6 Terra were only mediocre. GPT 5.6-Sol took the place of our previous favorite, Anthropic’s Claude Sonnet (currently best: 5 thinking) which surprisingly dropped quite to the bottom. The same for Anthropic Opus 4.8. We haven’t tested Anthropic’s top model Mythos as it was not accessible to us. But Anthropic’s other new model Fable was accessible and performed fairly well, see just below.

GPT 5.6-Sol was followed by a group of six, comprising two other U.S. LLM, one French LLM and three Chinese LLMs:

  1. Nvidia Nemotron 3 Ultra (based on Llama),
  2. Anthropic Fable,
  3. Mistral Vibe (previously: Le Chat),
  4. Alibaba Qwen 3.7 Plus Thinking,
  5. Moonshot Kimi K2.6 Thinking (not to mix up with the less performing K2.5 or K2.4, whilst the brand-new Kimi K3 displayed impressive “thinking” capabilities, but did not produce any results, presumably for being overloaded),
  6. Z.ai GLM 5.2 (slightly less performing than the others).

Mistral was already in the top tier in previous tests and is now slightly outstanding in the group of five. Moreover, its servers are located in the EU where strict data protection rules apply. Therefore, we recommend Mistral Vibe for online use, together with GPT 5.6-Sol if accessible and not deemed too expensive.

We have good news for administrations and parliaments strictly forbidding online use: Mistral Vibe, Alibaba Qwen 3.7 and Z.ai GLM 5.2 can be installed with full confidentiality control on domestic computers. However, we cannot assess whether the downloadable versions are equally performing.

Microsoft Copilot was, as usual, in the lower tier, and Google’s Gemini 3.1 Pro Thinking was even at the very bottom. Their poor performance is even more deplorable as so many administrations and parliaments are bound to use these two LLMs. Using only these two LLMs for legislative content generation and legislative completeness checks causes bills to be unduly incomplete, thus not containing the elements necessary to reach the legislative goals, namely with regard to implementation. 

Other tested, but not to be recommended, LLMs were: Perplexity’s Sonar (based on Llama), OpenAI’s GPT Terra or Terra Thinking, Deepseek and Kimi K2.5.

As the so-far-impressive Chinese LLM Manus has been bought by Meta, whilst Meta is notoriously poor in terms of data protection, we have not even tested Manus again.

Our prompt and the various responses are available on request.

See also our previous test results and our general recommendations on how to use LLMs.

Leave a Reply

Your email address will not be published. Required fields are marked *

thirteen − ten =