# Claude Opus 5.5 Can End the Turn Before the Job Is Done

**Summary:** Anthropic says Opus 5.5 is over 30% faster, yet an unattended run can end after a progress report with work still open. I would make completion an explicit harness state, not infer it from end_turn.

- Canonical: https://markhuang.ai/news/claude-opus-5-5-end-turn-is-not-done
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-28
- Section: News
- Tags: Claude Opus 5.5, AI Agents, Agent Harnesses, Claude API, Developer Tools
- Source: [Claude Platform Docs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A robotic assembly arm continues building a luminous circuit bridge while an amber gate blocks the unfinished path](https://cdn.markhuang.ai/news/claude-opus-5-5-end-turn-is-not-done/hero.webp)

*The machine can keep working. The control loop has decided that the run is over.*

Anthropic says Claude Opus 5.5 produces output tokens more than 30% faster than Opus 5. Yet its new [prompting guide](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5) warns that an unattended agent can finish a turn with a progress report while parts of the job remain open. The API still returns `stop_reason: "end_turn"`.

That is the change I would put at the top of an upgrade review. Opus 5.5 may be faster and more efficient, but a direct API integration can mistake a conversational pause for completion. If the harness closes the job, releases resources, or tells a user the work is done, the model's extra capability cannot recover the missing steps.

My read is that this is less a prompt-engineering problem than a completion-contract problem. A production agent needs its own account of what remains, what counts as done, and when another turn is safe. The model's stop reason is only one input to that decision.

## The cheaper model still changes the contract

The launch case is attractive. [Anthropic lists Opus 5.5 at $4 per million input tokens and $20 per million output tokens](https://www.anthropic.com/claude-opus-5-5), down from $5 and $25 for Opus 5. It estimates a 40% lower cost on typical default workloads because the new model also uses fewer tokens per task. Anthropic's separate [cost walkthrough](https://claude.dev/blog/what-a-task-costs-on-opus-5-5/) makes the caveat explicit: the price change alone cuts its illustrative session by about 31%, while the rest depends on the work.

The settings are not a straight carryover either. Opus 5.5 uses `medium` effort by default instead of Opus 5's `high`, and adaptive thinking is always on. The [migration guide](https://platform.claude.com/docs/en/models/opus-5-5/migration-guide) says a direct Messages API client must read response blocks by type, preserve thinking blocks in tool loops, remove unsupported forced tool choices, and revisit token limits. Anthropic says Managed Agents need only a model-name change, so this concern lands mainly on teams that own their loop.

## `end_turn` is not a completion certificate

Anthropic describes a specific failure mode. During a long task, Opus 5.5 may report progress and end the turn before it has finished every part. A loop that equates any text-only `end_turn` with success stops there. The guide recommends keeping the task parts in a checklist, continuing when items remain and no blocker is stated, then stopping after two or three automatic continuations if the same task still does not finish.

```mermaid

flowchart TD
    A[Model returns end_turn] --> B{Required items complete?}
    B -->|Yes| C[Mark the job done]
    B -->|No| D{A blocker is stated?}
    D -->|Yes| E[Surface the blocker]
    D -->|No| F[Continue with the open items]
    F --> G{Continuation cap reached?}
    G -->|No| A
    G -->|Yes| H[Stop for review]
```

> **Warning:**
>
> I would treat `stop_reason` as a transport outcome. The job is complete only when its declared outputs and checks are complete.

The distinction matters because `end_turn` is not wrong. The model did end its turn. The product error appears when the harness gives that event a stronger meaning than the API promises.

## A checklist gives the loop something durable

I would make completion visible outside the model. The harness should record the requested deliverables and required checks, then compare artifacts and tool results with that record after each turn. If work remains, the next message should name the open items rather than ask the model vaguely to continue.

Anthropic also suggests using a smaller model to inspect the conversation and return a reason when the work is incomplete. That could help, but I would not let a second model become the only judge. Tests, expected files, structured task state, and explicit approvals are firmer evidence. The checker can catch an inconsistency; it cannot turn an ambiguous brief into an objective finish line.

The continuation cap is just as important. An automatic nudge can rescue an early stop, but unlimited nudges can burn tokens while a stuck agent repeats itself. Two or three retries, followed by a reviewable blocker, is a sensible default. Risky or irreversible actions still need their own confirmation gate.

## The interface can hide the work too

Opus 5.5 changes where progress updates appear. Text between tool calls now arrives in `thinking` blocks, and the default display mode leaves their text empty. A client that renders only ordinary text can look silent during a long run even while the agent is working. Anthropic says developers can request summarized or update-only thinking displays and render the nonempty blocks before each tool call.

This is separate from completion, but the failures reinforce each other. A user sees silence, then receives a polished progress report. The harness sees `end_turn`. It looks like a natural ending even when the checklist says otherwise.

I made a related argument about [Fable 5.1 making thinking history append-only](https://markhuang.ai/news/claude-fable-51-thinking-history-append-only): model upgrades increasingly change the conversation contract around the model. Reading the final text is no longer enough to operate the loop correctly.

## Public results do not test this boundary

One detailed [Reddit evaluation](https://www.reddit.com/r/ClaudeAI/comments/1woicc1/opus_55_is_the_real_deal_same_accuracy_as_fable/) ran 23 tasks and reported that Opus 5.5, Opus 5, Sonnet 5, and Fable 5.1 all passed them at their default settings. On 10 separate work-style tasks, Opus 5.5 passed every run at low and high effort, with low costing less. The author notes the small samples, single task author, and Claude Code harness.

A [Hacker News complaint](https://news.ycombinator.com/item?id=49821657) goes the other way, describing an Opus 5.5 code-review request that ended in a reasoning-related refusal while older models accepted the same prompt. That is one unverified report, not evidence of a general defect. It does show why a harness must distinguish success, refusal, an explicit blocker, and an unfinished progress report. They are different states even when every one of them ends a turn.

## I would upgrade the state machine first

Before moving production traffic, I would replay long tasks that need several tool rounds and seed some of them with an intentional blocker. I would check whether the loop resumes unfinished work, surfaces real blockers, stops at the continuation cap, and preserves the approval boundary. I would also measure cost per accepted task, because a cheaper model paired with avoidable continuation turns can give back part of the savings.

Then I would watch the failure rate that benchmark charts usually omit: jobs reported as complete while a required item remains open. That number should be zero before an unattended workflow earns more authority.

Opus 5.5 may finish more work in less time. I still would not let `end_turn` close the ticket. The harness has to prove that the job, not merely the turn, is done.
