DeepSeek V4 Flash Solved 61.4% for Four Cents. Now Test the Workflow.
ARC Prize verified DeepSeek V4 Flash 0731 at 61.4% on ARC-AGI-2 for $0.04 per task. That earns it a cheap reasoning lane and a place in workload tests.
AI-powered · Limited to 20 requests per hour

ARC Prize verified DeepSeek V4 Flash 0731 on July 31, 2026, across three reasoning settings. At maximum effort, the model scored 89.0% on the ARC-AGI-1 semi-private set for $0.02 per task and 61.4% on ARC-AGI-2 for $0.04 per task.
That combination is hard to shrug off. My read, though, is narrower than the reaction a leaderboard invites. I want DeepSeek on the shortlist for cheap reasoning. I do not want to pretend the same four cents buys a reliable agent or a finished job.
Quick answer
| Question | My read |
|---|---|
| What changed? | DeepSeek V4 Flash 0731 reached 61.4% on ARC-AGI-2 at maximum reasoning, up from 46.0% at low reasoning. |
| Why does it matter? | The verified cost was $0.04 per ARC-AGI-2 task, which makes stronger reasoning cheap enough to route selectively. |
| What did the test cover? | Static colored-grid problems scored by exact output, without an agent harness or client-side tools. |
| What did it not prove? | Tool use, long-running reliability, recovery from mistakes, latency under load, or cost per accepted production result. |
| My decision | Put the model into a workload-specific evaluation. Do not promote it from puzzle solver to default agent on this score alone. |
The extra reasoning bought a real gain
The three settings make this more useful than a single headline score. On ARC-AGI-2, low reasoning scored 46.0%, high reached 56.0%, and maximum reached 61.4%. ARC-AGI-1 moved less, from 84.0% to 87.0% to 89.0%. The harder set gained 15.4 percentage points between low and maximum effort.
I see a routing decision in those numbers. If a cheap first pass is enough, use it. When failure costs more than a few extra cents, spend more reasoning on the cases that need it. I can act on that. A debate about which model is universally smarter gets me nowhere.

Four cents describes this test, not every job
ARC-AGI is deliberately compact. A model receives examples of colored-grid transformations, infers the rule, and returns the missing grid. The ARC Prize guide says the semi-private ARC-AGI-2 evaluation contains 120 tasks and scores an answer as correct only when the output matches the validated solution. It uses pass@2, allowing two guesses because some tasks contain explicit ambiguity.
The benchmark is good at isolating adaptation to unfamiliar abstract problems. It is also tightly controlled. ARC Prize's verified testing policy says ARC-AGI-1 and ARC-AGI-2 evaluate models as direct input-to-output predictors. There is no agent harness and there are no client-side tools.
That boundary changes how I would use the result. A production agent has to choose tools and maintain state. It also needs to catch bad intermediate results, recover, and know when to stop. ARC does not test those jobs.
The model has a broader product story
DeepSeek's official API documentation identifies the current Flash version as DeepSeek-V4-Flash-0731, with thinking and non-thinking modes, tool calls, and a 1 million token context window. Those features make it plausible for larger workflows. They do not turn a static reasoning score into evidence that the workflow works.
One popular Reddit post made that leap, arguing that cheap benchmark performance made agent loops more realistic. The author later added a correction: the chart mixed reasoning configurations and worked better as a cost-performance discussion starter than as a clean ranking. I agree with the correction. Cheap attempts can fund retries and validation, provided those checks catch the failures that matter.
How far I would trust the score
ARC-AGI-2 was built to be harder for AI systems while remaining approachable for people. Its benchmark paper frames the task as a measure of skill acquisition on novel problems. That makes it more interesting to me than another test dominated by recalled facts.
The skeptical case still matters. A separate analysis of ARC-AGI and OpenAI o3 argues that these grids represent a specific class of problems and may reward extensive trials over predefined operations. ARC Prize's 2025 technical report also warns that frontier reasoning results remain constrained by knowledge coverage and can create new forms of benchmark contamination.
Those critiques tell me where to stop extrapolating. A 61.4% score says the model handled these unfamiliar visual transformations well under this evaluation. Treating it as a percentage of general intelligence, or expecting 61.4% of business tasks to succeed, would ask the number to carry far too much.

My take
I would test DeepSeek V4 Flash 0731 anywhere expensive reasoning is being used as a default. Start with bounded tasks that have deterministic graders or strong validators. Run low, high, and maximum effort against the same acceptance bar. Measure latency and total cost per accepted result, then route only the cases where extra reasoning pays for itself.
The four-cent ARC-AGI-2 result gets DeepSeek onto my shortlist. The buying decision comes later, after it survives the actual workflow. I am more interested in that practical opening than in claims that intelligence has suddenly become almost free: a good router may be able to ask for deeper reasoning far more often and still demand proof before the model gets access to the rest of the workflow.
License
News text © 2026 Mark Huang. News text may be shared or translated for non-commercial use with attribution to https://markhuang.ai/news/deepseek-v4-flash-four-cent-score.
Suggested attribution: Based on "DeepSeek V4 Flash Solved 61.4% for Four Cents. Now Test the Workflow." by Mark Huang, originally published at https://markhuang.ai/news/deepseek-v4-flash-four-cent-score.