Digest

Weekly Hallucinations: Gemini 4 Argon, OpenAI DevDay and Sonnet 5.5's Big Appetite for Tokens

Weekly Hallucinations: Gemini 4 Argon, OpenAI DevDay and Sonnet 5.5's Big Appetite for Tokens

OpenAI did a lot at DevDay to make ChatGPT a place where you can do serious, sustained work. There are more useful capabilities now, but the high-capacity subscription became less generous at the same time. I like where the product is heading; I like the lower limits less.

Google announced Gemini 4 Argon, but for now only selected participants in the Fairwind cybersecurity program can try it. Access for developers, companies, and regular users is promised later, starting with the paid API and Google AI Ultra. Introductory API pricing is $2/$10 per million tokens and will rise to $4/$20 after the preview period. There is no date for when the discount ends. The output limit has increased from 64 thousand to one million tokens; this is separate from the context window, which is also one million. Long reasoning can be paused and resumed with the next request.

HTfV2KTa0AAVcsh.png

In Artificial Analysis measurements, Argon scored 53 at high reasoning effort, the same as GPT-6 Astra at maximum effort. The intelligence index combines results from several benchmarks covering knowledge, reasoning, and agentic work. Google is back among the leaders, but for now the economics are propped up by the discount. The average task in this suite costs $1.99, compared with $3.26 for Astra. Once the price doubles, Argon’s estimated task cost will rise to $3.98. It generates around 62 thousand output tokens per task on average, versus 27 thousand for Astra. So cheaper tokens help pay for longer reasoning; the reasoning itself has not become shorter.

Google is also showing internal applications of Argon. According to the company, agents found optimizations that freed up more than 300 TiB of memory in data centers. In the libgav1 video decoder, they rewrote 32 thousand lines of SIMD code—that is, operations on multiple values at once—in safe Rust with compiler auto-vectorization. The result was 2.7 times faster than the previous Rust port while producing identical video output. The comparison is specifically against the Rust port: the new version is still catching up with the optimized C++ implementation.

Pi has reached version 1.0. If you are not familiar with it, it is one of the most popular harnesses, known for being highly customizable and lightweight. The stable release adds Codemode, where the agent controls tools through code, native MCP support, and deferred loading for those tools. Descriptions of every available tool no longer have to be carried into every request. There is also Anthropic cache warming to reuse already processed context, plus the ability to add system messages during a conversation. Pi keeps a small core set of capabilities and leaves the rest to extensions.

image.png

Pi Durable was released as a separate experimental package for agent applications. It saves execution state at every step so a process can be stopped and resumed after a restart. But recovery distinguishes between an unfinished model request and an unfinished tool call. A model request can be sent again, while a tool is repeated only when it can be safely replayed. Saving history is not enough if the agent already sent an email but the confirmation got lost.

OpenAI DevDay on September 29 brought more than twenty announcements. At the center of the batch was an effort to turn ChatGPT into a place for ongoing work: models, background agents, shared documents, and developer tools. GPT-6.1 Sol replaces GPT-6 Sol, which I covered in the previous issue, just seven days after its release. Pricing remains $2/$10 per million tokens, while cache reads dropped from $0.20 to $0.10. The model is available through the API, Codex, and ChatGPT Work; it is not yet in regular Chat at launch.

HTavp4Wa4AANFTD.jpeg

Artificial Analysis tested the promise of “almost Astra at one-fifth the token price.” At maximum effort, Sol trails Astra by one point on the intelligence index, while the average task costs $0.72 versus $3.26. Compared with the previous Sol, task cost fell by 31%, even though the new model uses more output tokens. Efficiency comes down to answer quality and the cost of the entire run, so identical per-million-token pricing does not necessarily mean the same bill.

Artificial Analysis Intelligence Index (4 Oct '26).png

Dots now have their own cloud computer and persistent tasks. An agent like this can monitor connected sources, process feedback, prepare changes, and bring the result back for review. Autonomous background research is read-only: it cannot send messages or modify data in that mode. Tasks you explicitly assign run under the applicable access and approval rules. Dots are gradually rolling out to Pro and Business Premium users in supported countries, while enterprise customers have access to a beta that administrators can enable. Conversations with a dot do not count against the ChatGPT limit, but work it launches in Codex or ChatGPT Work uses the regular quota.

The results of this work can go into ChatGPT Space, a shared team workspace, and Pages, documents that people and agents can edit together. Collaborative presentations are promised in the coming weeks. Codex gained cloud environments, voice input in the console client, and multi-agent management; the Agents API can now work with a computer. Developers can extend the ChatGPT interface through plugins. With Sign in with ChatGPT, the usage included in a subscription can also be spent in participating third-party tools. There is also a new enterprise Marketplace, where part of a company’s OpenAI spending commitment can be directed toward approved partner software.

Интерфейс ChatGPT с погодной картой.png

The new ChatGPT Pro 500 subscription tier costs $500 a month. Pro 200 remains $200 with a smaller included allowance. According to Thibault, who leads Codex at OpenAI, Plus is the 1x baseline, Pro 100 provides 5x, the new Pro 200 provides 10x instead of the previous 20x, and Pro 500 is listed at 25x. These are relative quotas.

The previous quota is preserved through October 29, 2026, inclusive, for anyone who had Pro 200 active at any point between September 22 and September 29 at 10:00 a.m. Pacific Time. The subscription must be active to use the old allowance; an eligible resubscription on the same account also counts. After October 29, the allowance will decrease while the price remains $200.

image.png

Thibault explained the change in terms of the efficiency of the new models. The equivalent amount of included API spend is being cut in half, but the company promises that over time users will get more work done for the same money. The replies were rough. There are complaints about the worsened deal, declarations of moving to Claude, sarcasm about the promise of doing more work with half the quota, and personal attacks. People have already built the generous subscription quota into their workflows, and a promise of future efficiency does little to make up for the concrete reduction in today’s allowance. Tokens included in a subscription may have felt free once the monthly fee had been paid. After the Claude Code quota cuts I covered two weeks ago, OpenAI is now reducing the included allowance of its own high-capacity subscription. This trade does not look like a gift to me. The old 20x cost $200; the new 25x costs $500: the price is 2.5 times higher than the old plan, while nominal capacity is up by a quarter. The expensive tier has extra features, but for someone who mainly bought it for the generous quota, the math is unpleasant. I read this as tighter monetization of intensive use. OpenAI’s explanation about more efficient models can be true at the same time as the old deal becoming worse for heavy users.

One extra Pro 500 feature is Astra Ultrafast, a faster generation mode. At DevDay, OpenAI promised up to eight times the speed in Codex and up to six times in the API; Sol Ultrafast has only been announced for the future so far. Under the subscription, Astra Ultrafast consumes the included allowance eight times faster than Standard for the same model. Buying additional credits on Pro 100 or Pro 200 does not unlock the mode.

image.png

Anthropic released Claude Sonnet 5.5 at the same $2/$10 per million tokens and $0.20 for cache reads. The company says generation is more than 30% faster than Sonnet 5 and that the model uses fewer tokens per task. In Artificial Analysis testing at maximum effort, the picture looks different. Sonnet scored 56 on the intelligence index versus 58 for Opus 5.5, but used around 193 thousand output tokens per task. That is the highest usage in their measurements, roughly seven times Astra at maximum effort. The average task cost $7.60, about 50% more than the previous Sonnet.

These results are from a prerelease version in which a structured-output bug was later fixed; AA says it will rerun the evaluation. They still show clearly why “Sonnet is cheaper than Opus” cannot be checked from the price list alone. The effort setting determines how much the model reasons and double-checks itself. Claude Code and the apps use Medium by default, the platform uses High, while the record-breaking figures are discussed for Max. In AA’s testing, High was the most competitive setting in terms of quality versus cost.

The decision models I wrote about in the issue on Jev and its first alternatives now have new providers. You define the allowed answers, and the model returns a choice or probabilities instead of free-form text that then has to be parsed in code. At DevDay, OpenAI showed a Decisions API based on Luna for text and images; for now it is in limited preview. For routing a request between agents or classifying an inquiry, a dedicated step like this can replace a long conversation with a general-purpose model.

Cloudflare released Clef and Clef-flash with open weights under Apache 2.0 and hosting on Workers AI. They are compatible with the Jev API, understand images, and accept up to 64 thousand context tokens. In Cloudflare’s own comparisons, the models beat Jev on several datasets, though not on every benchmark. Having shared comparison tables is useful already. Previously, the alternatives mostly competed on their own examples.

Perplexity added a Decisions API with the pplx-decider-v1-27b model. It costs $0.04 per million input tokens, with output free. This is a fine-tuned Qwen3.8-27B with open weights under Apache 2.0, image support, and a claimed 250-thousand-token context window. The model card reports average accuracy of 85.71% across eleven benchmarks, versus 84.51% for Jev. On the difficult public section of JevBench, a set of structured decision-choice tasks, Perplexity trails: 70.30% versus 73.27%. So the overall average does not guarantee a win on your data. But you can now choose between a cloud API and running the same weights yourself instead of tying the entire decision step to a single service.

DeepSeek introduced Harness, an MIT-licensed shell built on the Cordis architecture. It works with code, files, spreadsheets, documents, and background tasks; there is a macOS app for Apple Silicon and Windows, while the web interface runs through Node.js. Tools, agent behavior, and the interface can all be extended with plugins. In Creator mode, the agent can write a new plugin from a description. This is still a public preview, while agent teams, schedules, and automatic approval checks are marked experimental. With this approach, an extension can add both an action and a way to display its result in the same application: a timer, a data panel, or a viewer for a generated document.

tg_image_830901153.jpg

Researchers from Meta Superintelligence Labs, together with colleagues from the University of Washington, MIT, and Trillium Labs, proposed Context Language Models. This is a way to manage the context of existing models, rather than a new standalone chatbot. The history the model sees on the next step becomes an editable file. The agent can remove unnecessary fragments, update a state table, or save important facts in a compact note; the changes are synchronized with its live context. Multiple such files can be maintained for multiple agents.

On EdgeBench, ten long-running program optimization tasks, this approach produced a result 5% better than the baseline method over twelve hours while using 59% fewer compute operations. The authors also improved cache reuse: after editing the history, computations for surviving fragments are preserved, including content after the modified section. A conventional cache reuses only a matching prefix.

Anthropic added Mods to Claude Code. Following support for AGENTS.md and cloud Projects, there is now a way to change the behavior of the client itself. A mod is a small TypeScript function inside a plugin and works in the CLI or the app. It can intercept an event before execution, after it, or in place of it. For example, it can rewrite a request to the model, block a tool call, strip a secret from a result, or replace a UI element. The built-in /diff has already become a mod that can be disabled or replaced with your own implementation.

For a team, this creates a place for custom rules and panels: build status alongside the conversation, a call log, or an extra approval step before changing a production config. Multiple mods handle an event in load order, so the order matters too. In managed installations, the built-in sec-default loads first and restricts dangerous overrides by user mods. The mods themselves, however, are not sandboxed and get the same access to the machine as Claude Code. An installed mod becomes part of the tool you have trusted with your project and working environment.

Last week, we launched the “AI at Work: Time Saver or New Workload?” study. We want to understand how much time AI saves on real work tasks once prompt preparation, checking, and rework are taken into account, and what happens to quality and employee workload. We have already collected good data, and your responses will help make the results more accurate.

Принять участие (1).png

If you currently work in Russia or abroad and have not taken part yet, fill out the survey. Even if you do not use AI at work, your experience matters to us too. The survey takes about 10 minutes and is completely anonymous. We will publish aggregated results at the end of October.

Stay curious.

I write about artificial intelligence, language models, and developer tools. I test models and services on real-world tasks and share what I learn in my Telegram channel.

A take on AI and development: a deep dive into language models, gadgets, and self-hosting through hands-on experience.
© 2026 Gotacat Team