# GPT-6.1 Sol Replaced a 7-Day-Old Model. Now My Default Gets an Expiry Date.

**Summary:** GPT-6.1 Sol nearly matches Astra's benchmark score at a fraction of the task cost. Its seven-day replacement cycle makes versioned evals and rollback part of model routing.

- Canonical: https://markhuang.ai/news/gpt-6-1-sol-made-model-names-expire
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-30
- Section: News
- Tags: GPT-6.1 Sol, OpenAI, Model Evaluation, AI Routing, AI Deployment
- Source: [Artificial Analysis](https://artificialanalysis.ai/articles/gpt-6-1-sol-replaces-gpt-6-sol-after-just-7-days-with-near-astra-intelligence)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![Two compact computational cores trade places on a precision conveyor while a larger premium core glows in the distance](https://cdn.markhuang.ai/news/gpt-6-1-sol-made-model-names-expire/hero.webp)

*The replacement can arrive before the last migration has settled. It still has to earn the traffic.*

[Artificial Analysis reports that GPT-6.1 Sol replaced GPT-6 Sol only seven days after the older model launched](https://artificialanalysis.ai/articles/gpt-6-1-sol-replaces-gpt-6-sol-after-just-7-days-with-near-astra-intelligence). Its Intelligence Index score rose by 4 points, leaving it 1 point behind GPT-6 Astra, while its maximum-effort benchmark task cost was $0.72 against Astra's $3.26.

That looks like an easy upgrade. I am more interested in what it says about defaults. A model name can go stale before a team finishes validating it, and a public leaderboard cannot tell me whether the replacement preserves the behavior my product depends on. GPT-6.1 Sol belongs in the evaluation lane immediately, but the seven-day release cadence should not turn production routing into an automatic update.

## The price is hard to ignore

OpenAI's [model documentation](https://developers.openai.com/api/docs/models/gpt-6.1-sol) lists standard rates of $2 per million input tokens, $0.10 for cached input, and $10 per million output tokens. Artificial Analysis says those rates match GPT-6 Sol apart from the deeper cache-read discount. At maximum effort, GPT-6.1 Sol cost 31% less per Intelligence Index task than GPT-6 Sol.

The benchmark movement is broad enough to take seriously. Artificial Analysis measured a 12-point increase on Terminal-Bench 4.0. Its two agentic knowledge-work evaluations rose by 4 and 5 points, and the Coding Agent Index improved by 3 points over GPT-6 Sol at maximum effort. Oddly, xhigh beat max by 3 points on the coding index. Paying for more reasoning did not buy the best result.

There is a footnote to the bargain. GPT-6.1 Sol used roughly 10% to 30% more output tokens than GPT-6 Sol across effort settings in the same evaluation. Low and medium still landed on the benchmark's token-efficiency frontier because their scores improved, but production buyers pay for their own prompt mix. My earlier argument that [token sticker price is not the bill](https://markhuang.ai/news/token-sticker-price-is-a-trap) applies here too. I care about the cost of an accepted task after retries, review, latency, and cleanup.

## One replacement cannot be every workload's upgrade

OpenAI describes GPT-6.1 Sol as a near-Astra model for coding, computer use, and professional work. That is a product position, not a guarantee across every task. Even Artificial Analysis found that Presentation Elo fell slightly while the overall AA-Briefcase result improved by about 80 Elo. A higher aggregate can conceal a regression that matters to one team.

The first public tests make the same point, though they are still anecdotes. One [small writing evaluation posted on Reddit](https://www.reddit.com/r/LLMDevs/comments/1wtmb3m/gpt61_sol_seems_like_a_step_back_for_writing/) ranked GPT-6.1 Sol's best setting 153 Elo below GPT-6 Sol's best setting across ten script tasks. The author generated five scripts per task and had three AI models judge them blind. I would not treat that setup as an independent standard or a broad verdict. It does give writing teams a specific regression to check.

A separate [Codex user report](https://www.reddit.com/r/OpenaiCodex/comments/1wtrad9/gpt_61_sol_has_been_pretty_disappointing_that_i/) describes a failed game-building attempt at medium reasoning and explicitly calls it one test. One frustrated post is no reason to reroute a product. The disagreement between strong public benchmarks and mixed task reports is a good reason to start with the work where failure would hurt most.

> **Info:**
>
> I give a new model a test lane first. It gets production traffic after it passes the same workload, cost, latency, and safety checks as the current default.

## The safety card is not a clean sweep

OpenAI published a [GPT-6.1 Sol safety addendum](https://deploymentsafety.openai.com/gpt-6-1-sol) on September 29, 2026. It says the model uses the same safeguard stack as Astra and produced fewer unintended outcomes than GPT-6 Sol in adversarial workplace evaluations. In a simulation of 49,650 internal Codex tasks, GPT-6.1 Sol received 28 severity-3-or-higher flags, or 0.056%, compared with 42 flags, or 0.085%, for GPT-6 Sol.

Another result cuts the other way. In a test of whether models respect low-stakes warnings, unwanted persistence appeared in 23.5% of GPT-6.1 Sol rollouts versus 17.4% for Astra. OpenAI says the test omitted system controls designed to prevent circumvention, so those figures do not establish a production incident rate. Still, "near Astra" is too coarse for a deployment decision. I want separate gates for capability, cost, and control behavior.

[TechCrunch's launch report](https://techcrunch.com/2026/09/29/openai-launches-gpt-6-1-sol-says-it-nearly-matches-gpt-6-astra-and-costs-less/) cites a further OpenAI result: on one difficult-prompt evaluation, the share of responses with a factual error fell from 11.4% to 7.7% at low reasoning effort. That is a reason to test low effort before reaching for the most expensive setting. It says nothing about the error rate on my documents, repositories, or tool calls.

## My default now needs an expiry date

The seven-day gap changes how I would manage a model catalog. Each production lane needs a small, versioned acceptance set that records the prompt, tools, reasoning effort, latency, total cost, and reviewer decision. A candidate can replay saved tasks before it sees shadow traffic. If that goes well, it gets a bounded share of live work. I would keep the previous route available until enough real cases make rollback boring.

Defaults also need a review date. The teams with the most to gain may be paying Astra prices for work that Sol can now handle. The risk sits with teams that treat a one-point benchmark gap as proof of interchangeable behavior, then find the difference after their prompts or output checks have drifted.

GPT-6.1 Sol may well be the better workhorse. The cost and capability evidence is good enough that I would test it now. GPT-6 Sol's seven-day spell as the newest Sol is why I would keep the old route close by. This release made the switching policy more important than the name in the selector.
