Digest

Weekly Hallucinations: Claude Opus 5.5, GPT-6 Sol, and a Survey on How You Use AI at Work

Weekly Hallucinations: Claude Opus 5.5, GPT-6 Sol, and a Survey on How You Use AI at Work

Write a prompt, check the answer, rework the result. We wondered how much work time AI saves once all of that is taken into account, and what happens to quality. Let’s do the math together. First, though, here’s what happened over the past week.

Anthropic released Claude Opus 5.5 for Claude Code and the API. The company promises Fable 5.1-level performance on most tasks and says the main improvement is the cost of everyday work, rather than another few percentage points on a benchmark: at standard settings, a task should cost about 40% less than with Opus 5. API list prices fell 20%, to $4/$20 per million tokens. Cache reads cost $0.20 instead of $0.50.

image.png

Artificial Analysis tested the model at five reasoning effort levels. Max scored 58 on a composite index covering reasoning, coding, and document tasks, but spent $5.98 per task versus $5.86 for Opus 5 at its maximum setting. The new model generated roughly 119,000 output tokens per task instead of 73,000. In other words, the price per token fell, but the bill for the most demanding run did not.

Artificial Analysis Intelligence Index (28 Sep '26).png

Anthropic’s materials show a particularly clear shift toward agentic development. On Terminal-Bench 4.0, where an agent completes multistep tasks in a terminal, the company reports 66.4% for Opus 5.5 at the second-highest reasoning effort level, versus 52.3% for Opus 5 in its own run. Those figures come from different effort settings and the company’s own test configuration. Opus can now be used at Medium, with a more expensive setting reserved for tasks that get stuck.

OpenAI released GPT-6 Sol and Luna. Both build on Astra, but are designed for a higher volume of everyday requests. For short contexts, standard API pricing is $2/$10 for Sol and $0.10/$0.50 for Luna per million tokens; long contexts cost more. The models are available in Codex, ChatGPT Work, and the API, while Luna is also available on the free tier.

AutomationBench.png

Artificial Analysis notes an improvement in value for money for Sol and Luna. For a long research task or a complex migration, I’d first look at task success rates with the specific setup being used, along with total token consumption.

Xiaomi answered the price race: the MiMo-V2.6-Pro-RL weights and MiMo-V2.6-Flash-RL were released under MIT. The larger model has 1.02 trillion parameters in total, but activates 42 billion for each token. It is an MoE model, which uses only the relevant part of a large network. The model accepts text, images, audio, and video, and has a context window of 1 million tokens.

Xiaomi describes a unified RL loop: the model was trained using feedback on its actions in programming, tool use, visual tasks, and cybersecurity tasks. The team released some of the training recipes and environments alongside the weights. In Artificial Analysis’s measurements, MiMo scored 46 on the overall index at roughly $0.13 per task in its test suite.

SpaceXAI introduced Grok 4.7 as a model for long coding sessions and working with information, while keeping the stated price and speed at 4.6 levels. Artificial Analysis gave it 46 points on its overall index, just two more than the previous version. In the coding agent index, though, Grok paired with its own Grok Build setup rose from 47 to 56. Elsewhere, the results are mixed: the analysts saw gains in terminal tasks and declines in long-context evaluation and app automation. The model used noticeably more tokens per task.

HSwLcNqXIAAwNtz.png

Jev now has an interesting counterpart: the open CLM. We wrote about Jev last week and examined the solutions’ API and fifteen projects built on it.

CLM builds representations of the current state and possible actions separately, compares them, and returns scores through a TypeSafe-compatible interface. The published version is based on Qwen3-8B; its code is available under Apache-2.0, so you can deploy it yourself and test it on your own tasks. The authors report speeds up to nine times higher than Jev in certain scenarios and strong results when evaluating coding agents’ decisions. Their tests use small held-out sets, and there is no common independent test for the whole category yet.

At Connect, Meta described upcoming features for Muse. We wrote two weeks ago about the agent’s initial launch. Muse will get its own email address, so people will be able to write to the agent directly. The company announced voice conversations and new connections to stores and workplace services. Meta also promises to bring Muse to its glasses in the coming months, allowing the agent to take into account what the wearer is looking at.

We previously wrote about Muse’s separate virtual machine and its monitoring agent, Sentinel. The implications of that setup are clearer now. An agent with its own email address, purchasing access, and a camera on someone’s glasses will have far more opportunities to make mistakes than a chatbot in a separate window. Meta describes confirmation for sensitive actions and controls over access to passwords. These are useful engineering measures, but a presentation cannot show how well they work; real-world reliability will become clear from tasks and failures once access expands.

Gemini 3.8 Flash came out back on September 2. Enough measurements have now accumulated to place it between expensive flagship models and the cheapest assistants. Artificial Analysis gives it 41 points on its overall index and around 300 output tokens per second when measured through the Google API. Fast output does not mean a fast first response: on the same page, TTFT (time to first token) is noticeably above the median for comparable models.

On ARC-AGI, a set of unfamiliar puzzles that require finding underlying rules, Flash solved 89.2% of the second version’s tasks at high reasoning effort. On the third version, that falls to 10.4% with the standard setup and 35% with the provider’s setup, which preserves reasoning state between calls.

Qwen-Image 2.1, introduced on September 20, combines generation and editing in one model. Its generation component has 7 billion parameters; the model can produce an image with a transparent background directly, accept up to ten reference images, and edit selected areas.

example-15.jpg

There is a significant condition attached to the weights. The license text on Hugging Face permits use of the model materials for research and evaluation, but requires a separate license for commercial use. A Reddit discussion also debated rights to the resulting images.

I wonder, can it do this?

Коллаж_3x2.jpg

Hugging Face introduced a release candidate for tokenizers v1. A tokenizer turns text into IDs for pieces of text that a model understands. On an Apple M4 Max, the team measured a 3–30× speedup for this operation over version 0.23 across ten model families using a single thread. Token IDs stay the same, so the optimization should not silently change the model’s input. Reusing words and buffers that have already been processed cuts unnecessary work; the benefit is greater for identical prefixes and smaller for dissimilar texts.

photo_2026-09-28 17.59.14.jpeg

For now, this is a preliminary Rust version; the full Python API and stable 1.0 release are still in development. It is too early to rewrite a working service just to chase the best benchmark result. But if tokenization has already become a bottleneck in your pipeline, the authors have published their methodology and a command for reproducing the measurements on your own hardware.

We have a small announcement! We’ve launched a study called “AI at Work: Time Saved or More Work?”. We want to know how much time AI saves on real work tasks once preparation, checking, and rework are included. What happens to task quality and employees’ overall workload?

Принять участие.png

If you’re currently employed in Russia or elsewhere, tell us how your AI transformation is going. Share your experience, even if you don’t use AI at work at all. The survey takes about 10 minutes, is entirely anonymous, and we’ll report the findings in aggregate. We’ll publish the results at the end of October.

And yes, this isn’t an ad. We, Controlled Hallucinations, are conducting the study ourselves. We really want to measure AI’s actual impact, rather than put a flattering number in a report for stakeholders.

Stay curious.

I write about artificial intelligence, language models, and developer tools. I test models and services on real tasks and share my findings in my Telegram channel.

A take on AI and development: a deep dive into language models, gadgets, and self-hosting through hands-on experience.
© 2026 Gotacat Team