The Regulatory Institute undertook another comparative Large Language Models (LLMs) test. Its purpose was to assess the ability of LLMs to identify and draft candidate policy and regulatory elements that make a law more effective. The results are surprising. Previous test results have mostly become obsolete, whilst our generic recommendations on how to use LLMs remain valid.
OpenAI’s GPT 5.6‑Sol was this time the frontrunner, whilst its sisters, OpenAI GPT 5.6 Terra Thinking and OpenAI GPT 5.6 Terra, were only mediocre. GPT 5.6-Sol took the place of our previous favourite, Anthropic’s Claude Sonnet (currently best: 5 thinking), which surprisingly dropped quite to the bottom. The same for Anthropic Opus 4.8. We haven’t tested Anthropic’s top models Fable and Mythos as they operate beyond our price threshold. However, we would not be surprised if these models were to produce very good results.
GPT 5.6-Sol was followed by a group of five, comprising one other U.S. LLM, one French LLM and three Chinese LLMs:
- Nvidia Nemotron 3 Ultra (based on Llama),
- Mistral Vibe (previously: Le Chat),
- Alibaba Qwen 3.7 Plus Thinking,
- Moonshot Kimi K2.6 Thinking (not to mix up with the less performing K2.5 or K2.4),
- Z.ai GLM 5.2 (slightly less performing).
Mistral was already in the top tier in previous tests and is now slightly outstanding in the group of five. Moreover, its servers are located in the EU where strict data protection rules apply. Therefore, we recommend Mistral Vibe for online use, together with GPT 5.6-Sol if accessible and not deemed too expensive.
We have good news for administrations and parliaments strictly forbidding online use: three LLMs from the group of five can be installed with full confidentiality control on domestic computers – only Nemotron 3 Ultra and Kimi K2.6 cannot. However, we cannot assess whether the downloadable versions perform equally well.
Microsoft Copilot was, as usual, in the lower tier, and Google’s Gemini 3.1 Pro Thinking was even at the very bottom. Their poor performance is even more deplorable as so many administrations and parliaments are bound to use these two LLMs. Using only these two LLMs for legislative content generation and legislative completeness checks causes bills to be unduly incomplete, thus not containing the elements necessary to reach the legislative goals, namely with regard to implementation.
Other tested, but not to be recommended, LLMs were: Perplexity’s Sonar (based on Llama), OpenAI’s GPT Terra or Terra Thinking, Deepseek and Kimi K2.5.
As the so-far-impressive Chinese LLM Manus has been bought by Meta, whilst Meta is notoriously poor in terms of data protection, we have not even tested Manus again.
Our prompt and the various responses are available on request.
See also our previous test results and our general recommendations on how to use LLMs.