# LiveNerf Will Wait 30 Days Before Calling Opus 5.5 Weaker

**Summary:** LiveNerf turns complaints about post-launch model decline into a pre-registered 30-day test. I trust its restraint more than an early dip, but its narrow Claude Code path cannot settle every user's experience.

- Canonical: https://markhuang.ai/news/livenerf-waits-30-days
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-30
- Section: News
- Tags: LiveNerf, Claude Opus 5.5, AI Evaluation, Claude Code, Model Monitoring
- Source: [LiveNerf on GitHub](https://github.com/ninjahawk/livenerf)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A sealed luminous model is measured repeatedly across a row of laboratory samples, with one amber fluctuation among blue traces](https://cdn.markhuang.ai/news/livenerf-waits-30-days/hero.webp)

*One odd reading is easy to notice. LiveNerf is designed to wait for a sustained change before calling it evidence.*

[LiveNerf](https://github.com/ninjahawk/livenerf) is running a 30-day test of whether Claude Opus 5.5 changes after launch. Its first 10 days form the baseline, followed by two 10-day comparison windows. Under the pre-registered rule, a change must clear a 99% confidence interval in both later windows, move by at least 3 percentage points, and not appear in the control arm. The earliest verdict comes at day 30.

My read is that LiveNerf is useful because it refuses to answer early. People notice a bad coding session immediately, while a vendor can change reasoning effort, prompts, routing, or infrastructure without changing the model name in the interface. LiveNerf holds one setup steady long enough to see whether the complaints persist. If it works, users get evidence they can inspect instead of another screenshot of one bad exchange.

## A rough session is evidence, not a measurement

The complaint behind the project is familiar: a newly released model feels sharp, then seems slower or less capable days later. The possible causes are tangled. Sampling varies. Workloads differ. A product update can alter the harness around the model. Heavy traffic may coincide with a bad run without causing it.

There is good reason to take user reports seriously without accepting the first explanation. In April 2026, [Anthropic traced Claude Code quality complaints to three product-layer changes](https://www.anthropic.com/engineering/april-23-postmortem): a lower default reasoning effort, a caching bug that repeatedly discarded older reasoning, and a system prompt that reduced verbosity enough to hurt coding quality. Anthropic said its API and inference layer were unaffected. To a user, though, the product had still become worse.

That distinction matters. LiveNerf measures Opus 5.5 as served through headless Claude Code on a Max subscription. It does not measure the raw API model, Claude chat, tool use, or long agent sessions. A detected shift would say that this served path moved. It would not identify whether the cause was weights, routing, a classifier, capacity, or a product setting.

## The benchmark freezes nearly everything it can

The repository says it screened 2,336 questions from GPQA Diamond, MMLU-Pro, competition math, and AIME 2025 and 2026. It kept 78 questions that Opus 5.5 answered correctly on some calibration attempts and missed on others. Those middling questions carry more information about movement than items the model always passes or always fails.

Each daily run sends those 78 questions to Opus 5.5 and a smaller GPQA control set to Opus 5. The prompts, effort level, Claude Code version, empty working directory, and scoring code are fixed. Tools, memory, hooks, repository instructions, and LLM judges are excluded. The project uses [Inspect](https://inspect.aisi.org.uk/), the open-source evaluation framework developed by the UK AI Security Institute and Meridian Labs, to run and record the samples.

The statistical design follows Evan Miller's [2024 paper on adding error bars to language-model evaluations](https://arxiv.org/abs/2411.00640). LiveNerf compares each question with its own baseline performance and clusters uncertainty by item. The project estimates that one daily pass can detect a change of about 7.5 points per 10-day window while consuming about 3.6% of the weekly Max-plan meter.

> **Info:**
>
> LiveNerf can test whether a narrow, pinned Claude Code path changes enough for this instrument to detect. It cannot prove why the change happened or settle whether every user's workflow improved or regressed.

## What the instrument misses

I like that the repository publishes the uncomfortable parts. Its validation found that low effort cut output tokens by 62% and reduced accuracy by 8.3 points relative to high effort. Yet swapping Opus 5 for Opus 5.5 was not distinguishable at the 99% threshold in that validation. The benchmark can catch a large change in thinking volume more readily than a modest same-family model swap.

The question set is also small and awkward by design. An audit of 80 selected or later-excluded items labeled 30 ambiguous and 8 as having suspect answer keys. LiveNerf keeps them and plans a sensitivity analysis without them, which is cleaner than quietly removing troublesome questions after seeing results. Still, the primary panel mostly tests knowledge and math under one-turn conditions. It says little about the long coding sessions that motivate many public complaints.

The [early Reddit discussion](https://www.reddit.com/r/ClaudeAI/comments/1wryrwx/is_opus_55_nerfed_new_benchmark_called_livenerf/) raises two fair objections. One commenter worries that a public benchmark could be recognized and treated differently. Others note that one run time may miss peak-load effects. The pre-registration acknowledges the second issue: results describe the model as served at the scheduled hour and do not generalize to peak hours.

The repository also says no license has been chosen. Anyone can inspect the method, but teams should not assume they have permission to reuse the code until that changes.

## The protocol is the useful lesson

LiveNerf will not tell me whether Opus 5.5 is still good at my work. The method is what I want to carry forward: capture a baseline before the system changes, pin the harness, repeat the same cases, keep the raw record, and decide in advance what would count as movement. I made a related argument when I wrote that [model comparisons need failure rates rather than a prettiest-run trophy](https://markhuang.ai/news/model-build-offs-need-failure-rates). Time adds another source of variance worth measuring.

For a production agent, I would run a smaller version against accepted patches, schema-valid outputs, recovery turns, and human corrections. I would keep provider status incidents beside the scores and separate model behavior from harness releases. A public benchmark cannot replace that local record, but it can show what honest monitoring looks like.

I am not reading the first dip in LiveNerf's chart as a verdict. The project had collected 6 of 30 days by its September 29 update and was still building the baseline. If the final result is no detected change, the public record will show that. If it finds a sustained shift, the fixed protocol will make the claim worth investigating. I would rather have that dated record than treat my worst session as a benchmark.
