Skip to content

· 8 min read

AI news: July 2026

Anthropic ships Claude Opus 5, OpenAI releases GPT-5.6 and cuts prices three weeks later, and Moonshot opens Kimi K3. July 2026.

Vintage egg illustrations arranged into the number five for Claude Opus 5

Image: Anthropic

July 2026 was a month about price. Anthropic shipped Claude Opus 5 on 24 July with an explicit pitch of near-Fable-5 performance at half the cost, and OpenAI, which had released GPT-5.6 on the 9th, cut two of its three models barely three weeks later. Underneath that, the open-weights releases kept coming: Kimi K3, Thinking Machines’ Inkling, Leanstral 1.5 and MiniMax H3.

Claude Opus 5, frontier-class performance at half the cost

Anthropic released Claude Opus 5 on 24 July, available the same day across all its platforms and set as the default model on Claude Max. Pricing is unchanged from Opus 4.8: $5 per million input tokens and $25 per million output, with a fast mode that runs 2.5x quicker for twice the base price. The numbers Anthropic quotes are its own: state of the art on Frontier-Bench and on the coding portion of GDPval-AA, three times the next-best model on ARC-AGI 3, and ahead of every model on OSWorld 2.0 at a third of Fable 5’s cost, coming within 0.5% of Fable 5 on CursorBench 3.2 at maximum effort. The release also brings effort settings, visual generation and, in beta, swapping tools mid-conversation plus automatic retries when a request gets blocked on safety grounds.

For teams already on Claude, the actionable part is not capability but the ratio against Fable 5: if the work holds up on Opus 5, those workloads land in a different billing bracket. Worth measuring on your own tasks before taking the vendor comparison at face value.

Source: Anthropic · also covered by 9to5Google and The Deep View

GPT-5.6: Sol, Terra and Luna, then a price cut three weeks in

OpenAI made GPT-5.6 generally available on 9 July as three models: Sol (the flagship), Terra and Luna. Launch pricing was $5 and $30 per million tokens for Sol, $2.50 and $15 for Terra, $1 and $6 for Luna. On 30 July, OpenAI cut Luna by 80% (to $0.20 and $1.20) and Terra by 20% (to $2 and $12), leaving Sol untouched, and put it down to efficiency gains made during development: 20% off the end-to-end cost of serving the model and more than 15% better token generation. Sol adds an “ultra” setting that coordinates four agents in parallel by default. On Artificial Analysis’ Coding Agent Index, Sol scores 80.0, 2.8 points above Claude Fable 5 while using less than half the output tokens.

Luna now sits at a different price point than it did on 9 July, which changes the arithmetic for teams running high volumes of routine work. The cut applies to the API rates; Sol’s pricing is unchanged.

Source: OpenAI · also covered by TestingCatalog and CNBC

Kimi K3, 2.8 trillion parameters with open weights

Moonshot published Kimi K3 on 17 July: an open-source model with 2.8 trillion parameters, native vision, a 1M-token context window and its own attention architecture, Kimi Delta Attention. Overall it lands behind Claude Fable 5 and GPT-5.6 Sol, scoring 67.3 on DeepSWE. The interesting part is the API pricing: $0.30 per million input tokens on a cache hit, $3.00 on a miss and $15.00 for output, with a cache hit rate above 90% on coding workloads according to Moonshot. It runs on Kimi.com, Kimi Work, Kimi Code, the API and the mobile apps; full weights went up on 27 July.

Open weights, a long context window and aggressive cache pricing together put it on the shortlist for sustained coding work, where the model replays most of the same context on every call.

Source: Kimi · also covered by TestingCatalog

Grok 4.5, xAI’s first model built for code and agents

xAI released Grok 4.5 on 9 July, its first model trained specifically for coding and agentic work. It shipped in Grok Build, across every Cursor plan and in the API console, with EU access expected mid-month. It offers a 500,000-token context window, function calling, structured outputs, web and X search, code execution and configurable reasoning (high by default), running at roughly 80 tokens per second. Pricing is $2 per million input tokens, $0.50 cached and $6 for output, with a higher rate above 200,000 tokens. The benchmarks cited are 62.0% on DeepSWE 1.0, 83.3% on Terminal Bench 2.1 and 64.7% on SWE-Bench Pro.

Landing in Cursor’s plans on launch day is what makes it easy to try in a real workflow: no tool change needed to compare it against whatever is already in use.

Source: TestingCatalog · also covered by Cursor

Inkling: Thinking Machines ships its first open-weights model

Thinking Machines Lab introduced Inkling on 15 July: an open-weights MoE with 975 billion total parameters and 41 billion active, a context window of up to 1M tokens, pretrained on 45 trillion tokens, with native text, image and audio. The lab’s own figures are 77.6% on SWE-bench Verified, 97.1% on AIME 2026, 87.2% on GPQA Diamond and 73.5% on MMMU Pro, and it matches Nemotron 3 Ultra on Terminal Bench 2.1 using around a third fewer tokens. It is available through the Tinker platform at a temporary 50% discount, in a free playground and via API on Together, Fireworks, Modal, Databricks and Baseten; the weights download in both the original format and NVFP4. Reasoning effort is adjustable between 0.2 and 0.99. An Inkling-Small preview also exists, at 276 billion total and 12 billion active parameters.

This is the first model out of Mira Murati’s lab, and it arrives with downloadable weights and a dial for reasoning effort, which is the part that lets you trade cost against latency without switching models.

Source: TestingCatalog

ChatGPT Work: OpenAI’s agent that hands back finished output

Alongside GPT-5.6 on 9 July, OpenAI launched ChatGPT Work: an agent built on Codex and GPT-5.6 that breaks a goal into tasks, works for hours and returns spreadsheets, slide decks, documents, dashboards or web apps. It asks for approval before sensitive actions and operates through a built-in browser that can click, type and move files. It connects to Slack, Microsoft Teams, Gmail, Google Drive, SharePoint, Salesforce, calendars and CRMs, invoked by typing “@” and the plugin name. It launched on Pro, Enterprise and Edu on web and mobile, with Plus and Business to follow; the desktop app for Windows and Mac is on every plan, including Free. Usage is metered the way Codex is, not flat-rate.

It is the pattern already familiar from coding tools, moved into office work. The thing to check before rolling it out to a team is the metering: cost tracks usage, not seat count.

Source: OpenAI · also covered by TestingCatalog and Bloomberg

Muse Image and Muse Video, Meta’s visual generation models

Meta introduced Muse Image and Muse Video on 7 July. Muse Image is an agentic image model: it calls search and coding tools, refines its own output and scales test-time compute rather than generating in a single pass. It is live in the Meta AI app, on meta.ai, in Instagram Stories in the US and in WhatsApp across a limited set of countries, with Facebook still to come. Meta places it second on Arena for text-to-image and editing as of 5 July, and it carries Content Seal watermarking. Muse Video is an early preview, third on Arena for text-to-video, with no availability date yet.

The model calling tools mid-generation is the real shift here: it brings image generation closer to the agentic pattern that already dominates text and code.

Source: Meta · also covered by TestingCatalog

GPT-Live: voice that listens and speaks at the same time

OpenAI introduced GPT-Live on 8 July, in two sizes, GPT-Live-1 and GPT-Live-1 mini. These are full-duplex voice models: they speak and listen simultaneously, handle turn-taking without explicit markers, slip in acknowledgements and can call tools several times per second mid-conversation. They replace Advanced Voice Mode in ChatGPT, with the mini model as the default and the larger one on paid tiers. At launch they run on GPT-5.5 underneath, not GPT-5.6. Since 31 July, the audio they produce carries SynthID watermarking.

For anyone building voice agents, full duplex removes the most visible flaw of the previous generation, which was turn latency. Worth testing before taking the natural-interruption claim as settled.

Source: OpenAI · also covered by TestingCatalog

Leanstral 1.5, Mistral’s model for formal proofs

Mistral published Leanstral 1.5 on 4 July under Apache 2.0: an open model for formal proof work in Lean 4, with 119 billion total parameters, 6.5 billion active and a 256,000-token context window. It saturates miniF2F, solves 587 of PutnamBench’s 672 problems, reaches 87% on FATE-H and 34% on FATE-X, and lifts FLTEval pass@8 from 31.9 to 43.2. Pointed at 57 Rust repositories, it flagged 47 potential violations of which 11 were genuine bugs, including 5 not previously reported on GitHub. It is available through the Labs API, in Mistral Vibe and on Hugging Face, free via Labs, with retirement announced for 30 September 2026.

Formal verification remains a niche, but the Rust result points somewhere broader: finding real bugs in existing code, not just proving theorems.

Source: TestingCatalog

MiniMax H3: 2K video with native audio, weights still pending

MiniMax closed the month with H3 on 31 July: a multimodal generation model that takes text, image, video and audio in, and produces video with native stereo audio of up to 15 seconds at 2K resolution. The company puts its per-second cost at 2K below a third of comparable models, and at 768p below half what 720p models charge. Inside there is a custom tokenizer, H3-VAE, that cuts sequence length by a factor of four, an H3-Omni transformer and an in-context regeneration technique. The weights were not published on announcement day: MiniMax said they would follow “in the coming days”.

Generating the audio in the same pass as the video removes a separate sound step. Until the weights are up, the only access is through the API.

Source: MiniMax · also covered by SCMP

Need help getting AI working in your company?

30 minutes, free, no commitment.

Book a call