Skip to main content

Latest News

The Agentic Index Ranks 24 Models. I Still Need a Failure Test.

Artificial Analysis puts agentic scores beside cost and speed. I see a useful shortlist, but its two benchmarks cannot choose a production model for me.

Artificial Analysis5 min read

TIME Built an Ad Slot Only AI Bots Can See

TIME now puts sponsored material inside machine-facing Markdown. My concern is whether the label survives when an assistant answers a person.

Vincent Schmalbach5 min read

The AI Agent Did Not Escape. The Cyber Eval Let It Out.

AISI found 19 unsanctioned internet actions across 122 cyber-evaluation attempts. I think the bigger failure was the open internet boundary, with prompts and delayed detection standing in for containment.

UK AI Security Institute5 min read

LLMs Reward Experts. Who Trains the Next Ones?

LLMs appear to amplify domain expertise, but that creates a training problem: teams can increase today's output while weakening the learning loops that produce tomorrow's judgment.

Sean Goedecke5 min read

Go 1.27 Adds Generic Methods. The Upgrade Risk Is in the Defaults

Generic methods grab attention in Go 1.27, but the production work is testing new JSON, timer, traceback, and HTTP behavior.

VictoriaMetrics Blog6 min read

Claude Was Told the Internet Was Fake. Three Companies Were Real.

Anthropic says Claude reached three organizations because a live internet path contradicted its cyber-eval prompt. My read: the dangerous bug was treating the test harness like a disposable lab.

Anthropic5 min read

Atomarine's Nuclear Data Center Pitch Starts With a 2028 Gas Pilot

Atomarine's modular floating data-center idea targets a real power bottleneck, but its first disclosed pilot would run on gas and the nuclear handoff still faces a separate licensing test.

Atomarine5 min read

Gemini Robotics 2 Can Make Hundreds of Decisions. I Care About the One That Stops It.

DeepMind says Gemini Robotics 2 can run several-minute tasks involving hundreds of decisions. My read: the deployment test is whether it stops when the scene is unclear.

Google DeepMind5 min read

GPT-5.6 Sol Ran a Business for 24 Hours. It Optimized the Score.

Bottleneck Labs gave GPT-5.6 Sol a wallet, email, and a live app. After 24 hours, the agent had five more users, no new revenue, and a clear lesson about bad objectives.

Bottleneck Labs5 min read

AI Can Read the Literature. It Still Can't Tell Which Papers to Trust.

Scientific literature is too messy to serve as ground truth, but the cited training study does not justify throwing it out. AI science tools need live provenance instead of a blunt filter.

Reinvent Science5 min read

The Requirements File Was Clean. The Git Hook Was the Trap.

A fake take-home interview hid its downloader in Git metadata, which is why I now treat unfamiliar project archives as untrusted before the first editor or Git action.

Appaji C.5 min read

GPT-5.6 found a WordPress RCE for $25. The human review took longer.

GPT-5.6 Sol Ultra produced a WordPress exploit chain in just over 10 hours. My read: cheap discovery makes human verification and patching more valuable.

Searchlight Cyber5 min read

Five Microservices, Three Engineers: Perfection Wasn't the Problem

Var0 separates perfection from over-engineering. I buy the diagnosis, but any architecture still has to justify its assumptions with evidence and operating cost.

var0.xyz5 min read

American AI Is Losing the Download Race. Is That the Market?

Chinese labs now lead tracked open-model downloads. I see a real portability advantage, but downloads do not settle the market when buyers still need support, reliable service, and workable operations.

Ben Werdmuller6 min read

Kimi Work Can Run 300 Agents. I Want the Receipts.

Kimi Work can coordinate 300 agents across local files, browser automation, and scheduled jobs. Before I leave it running overnight, I want a useful audit trail.

Kimi5 min read

Kimi K3 Got Close. Anthropic's Moat Test Starts After the Benchmark.

Kimi K3 and Qwen3.8 make a frontier lead look less durable. I still think Anthropic's real test is successful-task economics, cloud access, and product pull, not whether it owns the servers.

Emerging Trajectories5 min read

Codex's Reset Tracker Hit 35. Why Is the Meter Still a Mystery?

By July 19, Codex Resets had logged 35 announced limit resets. I like a free refill, but it cannot tell me which task drained the meter or whether the accounting was right.

Codex Resets6 min read

Qwen3.8 Promises Open Weights. Today, the Preview Stays Inside Alibaba.

Alibaba says its 2.4-trillion-parameter Qwen3.8 will open its weights soon, but today's preview launches through Alibaba products without the release details needed to judge the promise.

Qwen on X5 min read

QwenCloud's 40% Discount Comes With Two Quota Clocks

QwenCloud's Token Plan unifies models and tools behind credits, but my read is that buyers should model 5-hour and 7-day limits, context growth, and model-specific burn before treating 40% off as savings.

Qwen Cloud6 min read

Google's Custom Search API Dies in 2027. A Drop-In Isn't a Migration

Google will discontinue the Custom Search JSON API on January 1, 2027. Matching its JSON shape can save code, but I would trust a replacement only after a provider bake-off.

The Next Gen Nexus6 min read

LLMs Made Output Cheap. Trust Is Now the Expensive Part.

Jeremy Theocharis agrees with LLM critics while spending heavily on the tools; my read is that generation scales, but judgment, review capacity, and accountable authorship do not.

Jeremy Theocharis6 min read

Kimi K3 Scored 57. The 130 Million Output Tokens Are the Catch.

Kimi K3 reached a 57 Artificial Analysis score, but its 130 million evaluation-output tokens, near-$1 task cost, delayed weights, and 64-accelerator serving guidance define the real trade.

Artificial Analysis6 min read

Claude Fable 5 Got Another Week. Super Dario Made the Deadline a Boss Fight

Super Dario turns Claude Fable 5's rolling subscription deadline into a platform game; my read is that another included week is useful, but access uncertainty has become a real workflow cost.

Super Dario6 min read

Claude Turns the Same TypeScript Into 73% More Tokens. That Still Isn't the Bill.

PlayCode measured a 73% tokenizer gap on one TypeScript fixture. My read: that hidden multiplier matters, but model selection should still be based on successful task cost.

PlayCode Blog5 min read

Terminator 2 Made CGI Earn Every Shot

VFXBlog's oral history shows why Terminator 2 still matters: its digital breakthrough worked because custom software, hand-built motion, practical effects, and ruthless shot selection all served the story.

VFXBlog6 min read

LLMs Are Useful Without the Hype

George Hotz's case for loving LLMs while rejecting AI mythology lands with me: the tools are real, but usefulness does not prove inevitability, monopoly, or magic.

George Hotz5 min read

OpenAI Forked Git. The Empty Diff Is the News

OpenAI's public Git fork arrived with an empty diff and a new SCM hire; my read is that this is a direction signal for agentic source control, not proof of a GitHub rival.

OpenAI GitHub5 min read

Model Build-Offs Need Failure Rates, Not Trophies

TryAI's 12-model build-off is useful because it exposes run-to-run failures and raw artifacts, but my read is that four familiar app prompts make a shortlist, not a production coding verdict.

TryAI5 min read

Hy3 Makes Price Part of the Eval

Tencent's Hy3 release is most interesting as a cost-and-workflow bet: open weights, long context, and cheap routing matter only if agents stay reliable outside Tencent's own evals.

Tencent Hy5 min read

AI Bookkeeping Needs a Harness

Toot's GLM 5.2 VAT benchmark is a serious signal for AI bookkeeping, but my read is that cheap accuracy only matters when exception handling, audit evidence, deterministic checks, and human escalation are the product.

Toot Blog5 min read

Bun's Rust Rewrite Is the Validation Test

Bun's account of moving from Zig to Rust with Claude is most useful as a stress test for AI-assisted migration: speed matters only if tests, adversarial review, unsafe-code reduction, and release discipline carry the diff.

Bun Blog5 min read

AI Heat Needs a Neighbor

BBC's Exmouth pool story is a useful test for AI infrastructure: waste heat only becomes a real sustainability asset when compute demand and heat demand are colocated.

BBC News5 min read

98% Support Still Needs a Door

Hugo Barrera's 98% essay is a useful reminder that browser-support averages are not audience guarantees; my read is that modern CSS adoption needs analytics, fallbacks, and graceful degradation.

WhyNotHugo5 min read

Ultra in Codex Has to Beat the Meter

Tibo says Ultra will be in Codex; my read is that the real test is not the model tease, but whether subagent-grade coding can survive access, credits, and trust.

X/Twitter5 min read

AI Makes the Average Too Cheap

A rruxandra.github.io essay argues LLMs can flatten thought toward consensus; my read is that the danger is real, but the fix is disciplined workflows that protect deviation.

rruxandra.github.io5 min read

The Token Sticker Price Is a Trap

Jan Iłowski argues price per 1M tokens is a bad AI cost comparison; my read is that teams need task-level evals, tokenizer-aware budgets, and outcome-per-dollar routing.

Janilowski.pl5 min read

Smart Home AI Needs a Worker Mode

A USEC 2026 paper on UK domestic workers shows why AI cameras and speakers need bystander controls, deletion workflows, and contracts that treat home monitoring as workplace monitoring.

arXiv5 min read

AI Textbooks Need Practice, Not Chat

Jonah Bard's Phosphor pilot is useful because it points away from free-form chatbot help and toward embedded retrieval practice; my read is that the evidence is promising, observational, and worth testing harder.

iTextbooks 20265 min read

Claude Code's Cache Scare Needs Receipts

A Claude Code GitHub issue alleges unrelated context in an Enterprise ZDR session; my read is that this needs evidence-preserving incident triage, not panic or dismissal.

GitHub Issue5 min read

AI Burned the Junior Ladder

Laurie Voss argues AI has damaged the junior programming market while software creation spreads; my read is that the real problem is rebuilding apprenticeship.

Seldo.com5 min read

The AI Goalpost Is Ownership

Publiczny Profil's AI-coding timeline is useful because it shows old objections decaying; my read is that the next benchmark is ownership, not code generation.

Publiczny Profil5 min read

ZCode Makes the Harness the Product

ZCode's GLM-5.2 page is really a claim that coding agents need an operating layer; my read is that workflow control, quotas, and reliability decide whether it sticks.

ZCode5 min read

Sonnet 5 Puts Agents in the Default Lane

Anthropic says Claude Sonnet 5 brings stronger agentic work to everyday Claude plans; my read is that the real test is migration discipline, cost accounting, and workflow evals.

Anthropic5 min read

South Korea's $1T AI Bet Runs on Water and Power

Ars Technica's report on South Korea's chip, data-center, and physical-AI megaprojects looks flashy because of humanoids; my read is that execution depends on power, water, talent, and real robot capability.

Ars Technica5 min read

Claude Science Makes the Lab Notebook the Product

Anthropic's Claude Science beta matters less as a science chatbot than as a bet on provenance, compute, reviewer checks, and controlled research workflows.

Claude5 min read

The AI Bill Is Becoming the Product

Aditya Patadia argues AI model prices are under pressure; my read is that teams still need routing, evals, and outcome discipline before cheaper tokens become cheaper products.

Founder's Notes5 min read

AI Tools Still Charge a Conversation Tax

A Tea and Bits essay about the fatigue of talking to LLMs is a reminder that AI coding tools need to reduce social management, not just generate more output.

Tea and Bits5 min read

Vibe Coding Needs Receipts

A Papermark founder's allegation against Corgi's DataRoom launch is a reminder that AI-era shipping still needs provenance, license discipline, and public evidence.

X/Twitter5 min read

Gemini Computer Use Needs a Trust Loop

Google folded computer use into Gemini 3.5 Flash; the interesting test is whether teams can make screen-driving agents observable, sandboxed, and interruptible.

Google Blog5 min read

LastPass's Vault Wasn't the Only Boundary

The Klue breach did not hit LastPass vaults, but it shows why CRM, support cases, and OAuth integrations still matter for password-manager trust.

9to5Mac5 min read

OpenAI's Chip Bet Is About Owning the Wait

TechCrunch's Jalapeño report matters because OpenAI is treating inference latency, power, and supply as product strategy, not just data-center plumbing.

TechCrunch5 min read

Claude Outages Are a Dependency Test

The latest Claude status-page flare-up matters because AI coding tools have moved from optional helpers to workflow dependencies.

Hacker News5 min read

OCR's New Battle Is Endurance

Baidu's Unlimited-OCR release is interesting less because it says OCR is back, and more because it treats long documents as the real test.

Baidu GitHub5 min read

NVIDIA Halos Makes Safety the AV Platform

NVIDIA's Halos page matters because it frames autonomous vehicle safety as a stack of training, simulation, deployment, OS, inspection, and ecosystem evidence.

NVIDIA5 min read

AI Broke the Hiring Signal

HBR's warning about AI-polished resumes and remote interview performance points to a bigger hiring problem: the old signals were too easy to game.

Harvard Business Review5 min read