Readplace

How Anthropic Nerfed Opus 4.6 Before the 4.7 Launch

fagnerbrack.com 3 min read
View original
  • current
Summary (TL;DR)
Anthropic shipped Opus 4.6 on Feb 5, 2026. Adaptive thinking became default Feb 9, letting the model choose its reasoning budget. On Mar 3, default effort dropped from high to medium. A beta header rolled out Mar 5-12 that hid reasoning from the UI. AMD's Stella Laurenzo analyzed her Claude Code sessions and found median visible thinking per turn dropped 73%, read-to-edit ratio fell from 6.6 to 2.0, and user interruption rate rose twelvefold. Edits to unread files climbed from 6.2% to 33.7%. Anthropic's Boris Cherny confirmed the defaults and said adaptive thinking sometimes allocates zero reasoning, causing the model to fabricate. On Apr 16, Opus 4.7 launched with a new 'xhigh' default effort. The version string and price never changed, but the product quality degraded for 70 days then felt restored. A separate viral claim about Opus 4.6 benchmark decline was debunked: it compared different test sets.

Want to come back later? Save this to readplace.com.

Straight to the meat:

Anthropic shipped Opus 4.6 on February 5, 2026 [4]. Four days later, adaptive thinking became the default [1] [7]. The model now chose its own reasoning budget per turn.

On March 3, the default effort dropped from high to medium. Boris Cherny, the Claude Code lead [6], said the change balanced intelligence, latency, and cost [1] [6].

A dialog showed up when you opened Claude Code [1]. Most people clicked through [9].

Between March 5 and March 12, a beta header called redact-thinking-2026-02-12 rolled out [1]. By March 8, 58% of responses had their reasoning hidden from the UI [1]. By March 12, effectively all of them [1].

None of these were weight changes [1]. All three were product decisions [1].

Stella Laurenzo, Senior Director of AI at AMD [5], noticed her Claude Code sessions getting worse [1] [5]. She exported 6,852 of them and ran the numbers [1] [5].

Median visible thinking per turn dropped 73% [1] [5]. In late January it was about 2,200 characters [1]. By mid-March it was 600 [1].

The read-to-edit ratio fell from 6.6 to 2.0 [1]. The model stopped looking at files before changing them [1].

The user interruption rate per 1,000 tool calls rose twelvefold, from 0.9 to 11.4 [1].

Early stopping hooks fired close to zero times per day before March 8, and about 10 per day after [1].

Edits to files the model had not recently read climbed from 6.2% to 33.7% [1]. That is more than five times worse.

Laurenzo posted the analysis as GitHub issue #42796 on April 2 [1] [5].

On April 6, Cherny responded [1] [6]. He confirmed the adaptive thinking default on February 9 [1].

He confirmed the effort=85 default on March 3 [1]. He said the redact-thinking header was a UI-only change that did not affect thinking budgets [1].

Then he said this:

The specific turns where it fabricated had zero reasoning emitted, while the turns with deep reasoning were correct. We’re investigating with the model team. [1]

Adaptive thinking sometimes allocates zero reasoning to a turn [1] [7]. When that happens, the model fabricates [1]. Cherny’s interim workaround was CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1 [1] [7].

Ten days later, on April 16, Anthropic released Opus 4.7 [3] [8]. The default effort on 4.7 was a new level called xhigh [3] [8].

Higher than the 4.6 default of medium [1] [3]. Higher than the 4.6 launch default of high [1] [4].

Users who upgraded felt the improvement [8]. Some of it was new weights [3]. Some of it was the effort default moving back up the curve [3] [8].

You cannot separate the two from outside the company [11].

In the same weeks, a second community story went viral [2]. BridgeMind posted a chart claiming Opus 4.6 had fallen from rank #2 to rank #10 on its benchmark, 83.3% down to 68.3% [2].

An independent AI researcher checked [2]. Yesterday’s score was from 30 tasks [2]. Last week’s was from 6 [2].

On the six common tasks the score was 87.6% then and 85.4% now [2]. The chart was comparing different test sets [2].

That piece did not hold up [2]. The Laurenzo analysis did [1] [5].

The version string did not change between February and April [3] [4]. The price did not change [3] [4]. What you got changed [1] [11].

I am not calling this deliberate planned obsolescence. The evidence does not establish intent. The structure produces the effect whether anyone plans it or not [10].

Three defaults changes ship quietly [1]. A UI redaction removes the main diagnostic signal [1] [11]. The new version launches with the defaults moved back up [3] [8].

Each decision defends on its own [1] [6]. The aggregate is a 70-day downgrade at a constant price [3] [4], followed by a release that feels like restoration [3] [8].

The version string is not the product. In an opaque platform, the defaults are.