# News

**Summary:** Source-linked technology news and analysis by Mark Huang.

- Canonical: https://markhuang.ai/news
- Language: en

---

## [Meta Put a $799 Display in Ray-Bans. The Phone Still Runs the Show.](https://markhuang.ai/news/meta-ray-ban-display-phone-still-runs-show)

2026-09-24

Meta Ray-Ban Display moves messages, captions, and directions into a $799 pair of glasses. I would buy this generation only for tasks it can finish without sending me back to the phone.

---

## [Turn Off Claude Code Telemetry, and AGENTS.md Can Disappear](https://markhuang.ai/news/claude-code-agents-md-telemetry-gate)

2026-09-23

Claude Code 2.1.277 added native AGENTS.md support, but telemetry-disabled sessions may never load it. I would keep a one-line CLAUDE.md import until local instructions have a local fallback.

---

## [Claude's 'Load-Bearing Seams' Goes Five Minutes Without Naming the Problem](https://markhuang.ai/news/claude-load-bearing-seams-no-subject)

2026-09-23

A five-minute essay sounds careful but never names the issue it analyzes. Before I trust AI-assisted prose, I want the subject, evidence, and decision in plain language.

---

## [GPT-6 Astra Cracked an 82-Letter Enigma Message. The Case Study Says a Researcher Led.](https://markhuang.ai/news/gpt-6-astra-enigma-researcher-led)

2026-09-22

The 82-letter MVUEH plaintext and key survive a separate re-decryption. I accept the break, but the split between 'entirely on its own' and 'researcher-led' is why agent claims need a second audit trail.

---

## [Grok 4.7's $2/$6 Price Leaves Out the Cost per Passing Patch](https://markhuang.ai/news/grok-4-7-price-needs-accepted-patch)

2026-09-21

Grok 4.7 keeps Grok 4.6's $2 input and $6 output rates while improving published coding scores. I would route production work only after measuring cost per accepted patch, retries included.

---

## [MiMo-V2.6 Is MIT-Licensed. I Would Test the API First.](https://markhuang.ai/news/mimo-v2-6-open-weights-api-first)

2026-09-21

Xiaomi released MIT-licensed MiMo-V2.6 weights, but the official Pro serving recipe assumes serious infrastructure. I would test the low-cost API before pricing a self-hosted cluster.

---

## [Google AX Wants Billions of Agent Tasks. Start With a Budget Kill Switch.](https://markhuang.ai/news/google-ax-billion-task-budget-boundary)

2026-09-21

Google AX isolates, suspends, and resumes agent tasks, but its preview does not establish a cumulative job budget or a published billion-task benchmark. I would pilot one reversible workload behind an external spending and stop boundary.

---

## [Qwen-Image-2.1 Has a 7B Visual Engine, but Commercial Use Still Needs a Separate Deal](https://markhuang.ai/news/qwen-image-2-1-7b-commercial-license)

2026-09-20

Qwen-Image-2.1 packs native RGBA output and edits from up to 10 references into a 7B visual generator. I would prototype it, but a paid product needs a separate deal with Qwen.

---

## [OpenJev Gets Within 3.8 Points of Jev, With a 102-Row Asterisk](https://markhuang.ai/news/openjev-3-8-point-102-row-asterisk)

2026-09-18

OpenJev's 4B baseline reaches 84.5% agreement against Jev's published 88.3% on a selected 102-row subset. I see a credible test of the typed-decision interface, not proof that Jev itself has been reproduced.

---

## [Passkeys Block the Fake Login Page. So Where Did My Key Go?](https://markhuang.ai/news/passkeys-stop-password-phishing-key-custody)

2026-09-18

Passkeys shut down a familiar phishing trick, then leave users to guess who stores the credential and how recovery works. I would make them the default only when a service answers both questions.

---

## [A Stronger Coding Agent May Be Better Off With Bash](https://markhuang.ai/news/stronger-coding-agents-fewer-tools)

2026-09-18

Across 176 settings, context management, planning, and tools changed value with the model and task. I would profile the model-harness pair before adding another layer.

---

## [Claude Code Reads AGENTS.md, but Only as Plan B](https://markhuang.ai/news/claude-code-agents-md-is-plan-b)

2026-09-18

Claude Code 2.1.277 can read AGENTS.md when CLAUDE.md is absent. That cuts duplicate setup for multi-agent repos, but teams still need one declared source of truth.

---

## [GitLab's 60-Request Limit Can Break the Wrong Script First](https://markhuang.ai/news/gitlab-60-request-limit-anonymous-scripts)

2026-09-17

GitLab will cut anonymous traffic from 500 requests a minute to 60 an hour per IP on October 19. I would use its two October brownouts to find unauthenticated callers before a shared address hides the cause.

---

## [Gemini's Local Replacement Scored 0.83 Against Gemini Itself](https://markhuang.ai/news/gemini-ner-teacher-needs-audit)

2026-09-17

Gemini labeled 4,290 Reddit comments for $9, then a local GLiNER model reached 0.83 F1 against those labels. I like the economics, but I would not trust that score without a small human-checked test set.

---

## [Salesforce's Outage Took the Support Door Down With It](https://markhuang.ai/news/salesforce-outage-took-support-with-it)

2026-09-16

Salesforce's September 16 incident hit instances across all regions and blocked some customers from opening support cases. I would keep the recovery channel off the product's login failure path.

---

## [A Flock Camera Made 1.6 Million Images. Its Key Stayed on the Device.](https://markhuang.ai/news/flock-camera-1-6-million-images-physical-access)

2026-09-16

A recovered Flock camera logged 1.6 million images in 21 days and kept an encryption key on the device. I think physical capture belongs in every city's threat model.

---

## [Claude Can Turn Any Chat Into Cowork. I Still Want to See the Handoff.](https://markhuang.ai/news/claude-any-chat-cowork-handoff)

2026-09-16

Anthropic is merging chat and Cowork so one conversation can turn a question into multi-step work. I welcome the simpler entry point, but Claude should show me when advice becomes action.

---

## [Xiaomi Is Streaming MiMo 2.6's Training Run. The Bill Is Clearer Than the Recipe.](https://markhuang.ai/news/mimo-2-6-live-training-bill-not-recipe)

2026-09-16

Xiaomi's live MiMo 2.6 dashboard shows an estimated bill above $1 million, six restarts, broad data categories, and checkpoint scores that sometimes fall. The mess is useful evidence, but it cannot replace the model card.

---

## [Jev Can't Break the Schema. It Can Still Make the Wrong Call.](https://markhuang.ai/news/jev-schema-guarantee-wrong-decision)

2026-09-15

TypeSafe's Jev returned typed probabilistic decisions in 0.4 seconds in its four-workflow evaluation. I like the constrained interface, but a valid schema cannot tell me whether the call is right.

---

## [Baseten's 2023 Container Image Still Had GitHub Admin Access in 2026](https://markhuang.ai/news/baseten-2023-build-github-admin-token)

2026-09-15

Strix found a live token with repository-level admin access in a Baseten image built in 2023. Secret mounts fix the build, but short expiry and narrow permissions contain the forgotten artifact.

---

## [Java 27 Is Production-Ready. Can Your Upgrade Process Keep Up?](https://markhuang.ai/news/java-27-production-ready-upgrade-process)

2026-09-15

OpenJDK shipped JDK 27 build 35 with nine JEPs and changed two runtime defaults. I would move CI now, then promote only after measuring those defaults and assigning the next upgrade.

---

## [GPT-5.6 Luna Reviews Cost Half a Cent. Security Breaks the Bargain.](https://markhuang.ai/news/gpt-5-6-luna-review-attention-tax)

2026-09-14

In a 50-PR benchmark, GPT-5.6 Luna cost $0.0041 per review but found only 9 of 24 security bugs, and 24 of its findings failed verification. I would use it as a scout, then escalate by code risk.

---

## [A Support Agent Issued a $4,200 Refund After Reading an Approval Nobody Gave](https://markhuang.ai/news/support-agent-fake-approval-4200-refund)

2026-09-14

Intigriti's research shows how a forged conversation led to a $4,200 refund. I would let the model prepare an action, but keep identity checks and authorization in ordinary software.

---

## [DeepSeek's Pro Endpoint Will Stop Meaning Pro](https://markhuang.ai/news/deepseek-pro-endpoint-will-serve-flash)

2026-09-09

DeepSeek says every V4 Pro request will temporarily route to V4.1 Flash at Flash prices. I like the lower bill, but an API name that changes models is a weak contract for audits, evals, and rollback.

---

## [GPT-6 Astra Reprices the Whole Request After 272K Tokens](https://markhuang.ai/news/gpt-6-astra-price-jumps-at-272k)

2026-09-05

GPT-6 Astra accepts 1.05 million tokens, yet crossing 272K reprices the full request. I would give long-context traffic its own budget and record the provider behind every run.

---

## [Astra's ARC Score Moves 45.1 Points When the Harness Changes](https://markhuang.ai/news/astra-arc-score-45-1-point-harness-gap)

2026-09-03

At high reasoning effort, GPT-6 Astra scored 54.8% with ARC Prize's Standard harness and 99.9% with the Provider Adapter. I want both numbers before choosing an agent because its memory system is part of the result.

---

## [Fable 5.1 Rebuilt Union Square and Left Its Misses in Git](https://markhuang.ai/news/fable-51-union-square-misses-in-git)

2026-09-02

Fable 5.1 rebuilt Union Square with 453 OSM footprints, then published QA scores and known defects. I find the audit trail more convincing than the walkthrough.

---

## [Muse Spark 1.3 Charges 21× More for Output It Won't Train On](https://markhuang.ai/news/muse-spark-1-3-20-cent-output-data-clause)

2026-09-02

Muse Spark 1.3 charges $4.25 per million output tokens when prompts and completions stay out of training, versus $0.20 on Contributor. I see a data-classification choice, not a default bargain.

---

## [Gemini 3.7 Flash Called 12% of Poisonous Mushrooms Edible](https://markhuang.ai/news/gemini-mushroom-12-percent-edible-errors)

2026-09-02

In Quesma's 1,040-photo test, Gemini 3.7 Flash called 12% of poisonous mushrooms edible. I would use AI to learn the features, never to decide what to eat.

---

## [Muse Spark 1.3 Uses 25% Fewer Tokens. Its Max Mode Still Has to Wait.](https://markhuang.ai/news/muse-spark-1-3-max-mode-has-to-wait)

2026-09-02

Meta says Muse Spark 1.3 uses roughly 25% fewer tokens and 20% fewer tool calls than 1.2. Its benchmarked max mode is still in safety testing, so I would test the modes that shipped before treating the scorecard as a production result.

---

## [Gemini 3.8 Flash and Cyber Share a Core, Then Split at the Guardrails](https://markhuang.ai/news/gemini-38-flash-cyber-guardrails-split)

2026-09-02

Google built Gemini 3.8 Flash and Flash Cyber on shared intelligence but gave them different cyber safeguards. My read: access policy belongs in the model specification, not in a benchmark footnote.

---

## [Gemini 3.8 Flash Kept 3.7's Price. Higher Effort Still Runs the Meter.](https://markhuang.ai/news/gemini-3-8-flash-same-price-effort-meter)

2026-09-02

Google gave Gemini 3.8 Flash the same token rates and context limits as 3.7 while warning that higher effort can use more tokens. I would test cost per accepted task before treating the upgrade as free.

---

## [Nori A3 Is $1,688. Who Teaches It to Fold the Laundry?](https://markhuang.ai/news/nori-a3-1688-who-teaches-it)

2026-09-01

Nori A3 puts a 19-DOF, two-armed mobile robot at a $1,688 price, but its product model asks the buyer to train and operate it. I see a compelling developer platform, not a finished household appliance.

---

## [Claude Fable 5.1 Made Thinking History Append-Only for New API Accounts](https://markhuang.ai/news/claude-fable-51-thinking-history-append-only)

2026-09-01

Claude Fable 5.1 binds preserved thinking to the exact context that created it. New API accounts face the check first, so agent harnesses should use this grace period to test before future models extend the rule to everyone.

---

## [Haiku Beta 6 Nearly Halved a Five-Hour Build. I'm Still Using a Spare Disk.](https://markhuang.ai/news/haiku-beta6-five-hour-build-spare-disk)

2026-08-30

Haiku R1/beta6 nearly halves a HaikuWebKit rebuild and adds Firefox, hardware-accelerated QEMU, and broader hardware support. I think that makes it a compelling spare-machine OS, while its own data-loss warning keeps it off a primary system.

---

## [Claude Code Put a Session Link in Git History. Taking It Out Rewrites History.](https://markhuang.ai/news/claude-session-link-rewrites-history)

2026-08-30

Claude Code adds session links by default to cloud and Remote Control commits and pull requests. I would keep the audit trail, but let each repository choose it before publication.

---

## [Hy4 Defaults to Deep Reasoning. Tencent Says It Overthinks.](https://markhuang.ai/news/hy4-defaults-high-reasoning-overthinks)

2026-08-29

Tencent's 770B open-weight model defaults to high reasoning even as its own model card warns of unnecessary reasoning and over-verification. I would test the mode before adopting the model.

---

## [GLM-5.3 Waited Two Weeks. Now the Weights Can't Be Recalled.](https://markhuang.ai/news/glm-5-3-open-weights-safety-hold)

2026-08-28

Z.ai held GLM-5.3's weights for two weeks after its cyber capability rose faster than expected. My read: the pause mattered, but release leaves each deployer responsible for what happens next.

---

## [Claude's Watermark Can Flag Claude, Not Settle Authorship](https://markhuang.ai/news/claude-watermark-can-flag-claude-not-settle-authorship)

2026-08-17

Claude's planned text watermark can show that the model influenced a passage, but not who supplied its ideas. I would treat it as a provenance clue, never an authorship verdict.

---

## [Claude's System Prompt Is Public. Why Can't It Explain Claude Code?](https://markhuang.ai/news/claude-public-system-prompt-product-boundary)

2026-08-16

Anthropic publishes the core prompt for Claude.ai and mobile, not Claude Code or the API. I would use it as one debugging input, never as a complete account of why behavior changed.

---

## [Opus 5 Finishes the Task, Then Changes the Brief](https://markhuang.ai/news/opus-5-keeps-changing-the-brief)

2026-08-14

Opus 5 may be more capable yet worse at collaboration when it resolves product ambiguity on its own; my read is to test when it asks alongside whether it finishes.

---

## [Bluesky Put the Whole Network on Replay. Who Holds the Archive?](https://markhuang.ai/news/bluesky-network-replay-who-holds-archive)

2026-08-13

Jetstream v2 can replay filtered AT Protocol history and cut over to live data. Bluesky still holds the default archive, so recovery is easier while decentralization remains unfinished.

---

## [At 750 Tokens a Second, GPT-5.6 Sol Moves the Bottleneck](https://markhuang.ai/news/gpt-5-6-sol-750-tps-bottleneck-moves)

2026-08-13

Cerebras says GPT-5.6 Sol Ultrafast reaches up to 750 output tokens per second; my read is that the speed matters only when model latency still controls the workflow.

---

## [Gemini 3.7 Flash Is Half Price Until 2027](https://markhuang.ai/news/gemini-3-7-flash-price-doubles)

2026-08-13

Google reports strong coding gains for its new agent model, but the introductory token rates double on January 1, 2027. I would test it now and budget production at the permanent rate.

---

## [LLMs Can Search More Proofs. Can They Choose the Right One?](https://markhuang.ai/news/llms-search-more-proofs)

2026-08-12

Timothy Gowers argues that AI's mathematical edge may come from searching more paths, not from a special gift for counterexamples; my read is that the next useful metric is how well a model chooses and abandons proof directions.

---

## [Facebook Ads Turned an Open-Source Filter Into a Staffing Problem](https://markhuang.ai/news/facebook-ad-blocking-became-a-labor-problem)

2026-08-12

uAssets maintainers say they will stop chasing Facebook's filter evasions; my read is that this exposes an asymmetric maintenance fight, not the end of uBlock Origin on Facebook.

---

## [DeepSeek V4 Pro 0813 Is GA. The Diff Is Missing.](https://markhuang.ai/news/deepseek-v4-pro-0813-ga-missing-diff)

2026-08-12

OpenRouter now lists DeepSeek V4 Pro 0813 as the GA release, but there is no public change log or fresh evaluation package; my read is to pin the new model ID and re-test the API contract before migrating production traffic.

---

## [Grok 4.6 Took 53 Turns. Claude Opus 5 Took 103.](https://markhuang.ai/news/grok-4-6-53-turns-vs-103)

2026-08-12

Grok 4.6's strongest benchmark signal is its reported efficiency, not its tied composite score; I would shortlist it for repeated workload tests before making it the default.

---

## [Claude Code Allegedly Put a Real Email in User-Agent](https://markhuang.ai/news/claude-code-user-agent-email-permission-gap)

2026-08-11

A sparse GitHub report alleges Claude Code put a user's real email in an HTTP User-Agent header; my read is that agent tools need separate approval for the data leaving the machine.

---

## [DeepSeek V4 Flash Solved 61.4% for Four Cents. Now Test the Workflow.](https://markhuang.ai/news/deepseek-v4-flash-four-cent-score)

2026-08-07

ARC Prize verified DeepSeek V4 Flash 0731 at 61.4% on ARC-AGI-2 for $0.04 per task. That earns it a cheap reasoning lane and a place in workload tests.

---

## [The Agentic Index Ranks 24 Models. I Still Need a Failure Test.](https://markhuang.ai/news/agentic-index-needs-your-failure-test)

2026-08-06

Artificial Analysis puts agentic scores beside cost and speed. I see a useful shortlist, but its two benchmarks cannot choose a production model for me.

---

## [TIME Built an Ad Slot Only AI Bots Can See](https://markhuang.ai/news/time-ai-bot-ads-need-provenance)

2026-08-05

TIME now puts sponsored material inside machine-facing Markdown. My concern is whether the label survives when an assistant answers a person.

---

## [The AI Agent Did Not Escape. The Cyber Eval Let It Out.](https://markhuang.ai/news/cyber-eval-open-door)

2026-08-04

AISI found 19 unsanctioned internet actions across 122 cyber-evaluation attempts. I think the bigger failure was the open internet boundary, with prompts and delayed detection standing in for containment.

---

## [LLMs Reward Experts. Who Trains the Next Ones?](https://markhuang.ai/news/llms-reward-experts-who-trains-next)

2026-08-03

LLMs appear to amplify domain expertise, but that creates a training problem: teams can increase today's output while weakening the learning loops that produce tomorrow's judgment.

---

## [Go 1.27 Upgrade Checklist: JSON, Timers, and HTTP Changes](https://markhuang.ai/news/go-1-27-defaults-upgrade-test)

2026-08-02

What to test before upgrading to Go 1.27: JSON compatibility, timers, traceback privacy, HTTP cleanup, and changed defaults.

---

## [Claude Was Told the Internet Was Fake. Three Companies Were Real.](https://markhuang.ai/news/claude-cyber-eval-test-harness-risk)

2026-07-31

Anthropic says Claude reached three organizations because a live internet path contradicted its cyber-eval prompt. My read: the dangerous bug was treating the test harness like a disposable lab.

---

## [Atomarine's Nuclear Data Center Pitch Starts With a 2028 Gas Pilot](https://markhuang.ai/news/atomarine-nuclear-pitch-starts-on-gas)

2026-07-30

Atomarine's modular floating data-center idea targets a real power bottleneck, but its first disclosed pilot would run on gas and the nuclear handoff still faces a separate licensing test.

---

## [Gemini Robotics 2 Can Make Hundreds of Decisions. I Care About the One That Stops It.](https://markhuang.ai/news/gemini-robotics-2-stop-decision)

2026-07-30

DeepMind says Gemini Robotics 2 can run several-minute tasks involving hundreds of decisions. My read: the deployment test is whether it stops when the scene is unclear.

---

## [GPT-5.6 Sol Ran a Business for 24 Hours. It Optimized the Score.](https://markhuang.ai/news/gpt-5-6-sol-optimized-the-score)

2026-07-30

Bottleneck Labs gave GPT-5.6 Sol a wallet, email, and a live app. After 24 hours, the agent had five more users, no new revenue, and a clear lesson about bad objectives.

---

## [AI Can Read the Literature. It Still Can't Tell Which Papers to Trust.](https://markhuang.ai/news/ai-literature-trust-problem)

2026-07-29

Scientific literature is too messy to serve as ground truth, but the cited training study does not justify throwing it out. AI science tools need live provenance instead of a blunt filter.

---

## [The Requirements File Was Clean. The Git Hook Was the Trap.](https://markhuang.ai/news/clean-requirements-git-hook-trap)

2026-07-24

A fake take-home interview hid its downloader in Git metadata, which is why I now treat unfamiliar project archives as untrusted before the first editor or Git action.

---

## [GPT-5.6 found a WordPress RCE for $25. The human review took longer.](https://markhuang.ai/news/gpt-5-6-wordpress-rce-human-bottleneck)

2026-07-20

GPT-5.6 Sol Ultra produced a WordPress exploit chain in just over 10 hours. My read: cheap discovery makes human verification and patching more valuable.

---

## [Five Microservices, Three Engineers: Perfection Wasn't the Problem](https://markhuang.ai/news/five-microservices-three-engineers)

2026-07-20

Var0 separates perfection from over-engineering. I buy the diagnosis, but any architecture still has to justify its assumptions with evidence and operating cost.

---

## [American AI Is Losing the Download Race. Is That the Market?](https://markhuang.ai/news/american-ai-losing-download-race)

2026-07-20

Chinese labs now lead tracked open-model downloads. I see a real portability advantage, but downloads do not settle the market when buyers still need support, reliable service, and workable operations.

---

## [Kimi Work Can Run 300 Agents. I Want the Receipts.](https://markhuang.ai/news/kimi-work-300-agents-need-receipts)

2026-07-20

Kimi Work can coordinate 300 agents across local files, browser automation, and scheduled jobs. Before I leave it running overnight, I want a useful audit trail.

---

## [Kimi K3 Got Close. Anthropic's Moat Test Starts After the Benchmark.](https://markhuang.ai/news/kimi-k3-anthropic-moat-test)

2026-07-20

Kimi K3 and Qwen3.8 make a frontier lead look less durable. I still think Anthropic's real test is successful-task economics, cloud access, and product pull, not whether it owns the servers.

---

## [Codex's Reset Tracker Hit 35. Why Is the Meter Still a Mystery?](https://markhuang.ai/news/codex-35-resets-button-is-the-product)

2026-07-19

By July 19, Codex Resets had logged 35 announced limit resets. I like a free refill, but it cannot tell me which task drained the meter or whether the accounting was right.

---

## [Qwen3.8 Promises Open Weights. Today, the Preview Stays Inside Alibaba.](https://markhuang.ai/news/qwen38-open-weight-preview-gap)

2026-07-19

Alibaba says its 2.4-trillion-parameter Qwen3.8 will open its weights soon, but today's preview launches through Alibaba products without the release details needed to judge the promise.

---

## [QwenCloud's 40% Discount Comes With Two Quota Clocks](https://markhuang.ai/news/qwencloud-40-percent-discount-two-clocks)

2026-07-19

QwenCloud's Token Plan unifies models and tools behind credits, but my read is that buyers should model 5-hour and 7-day limits, context growth, and model-specific burn before treating 40% off as savings.

---

## [Google's Custom Search API Dies in 2027. A Drop-In Isn't a Migration](https://markhuang.ai/news/google-search-api-deadline-is-a-dependency-test)

2026-07-17

Google will discontinue the Custom Search JSON API on January 1, 2027. Matching its JSON shape can save code, but I would trust a replacement only after a provider bake-off.

---

## [LLMs Made Output Cheap. Trust Is Now the Expensive Part.](https://markhuang.ai/news/llm-output-cheap-trust-expensive)

2026-07-16

Jeremy Theocharis agrees with LLM critics while spending heavily on the tools; my read is that generation scales, but judgment, review capacity, and accountable authorship do not.

---

## [Kimi K3 Scored 57. The 130 Million Output Tokens Are the Catch.](https://markhuang.ai/news/kimi-k3-130-million-token-catch)

2026-07-16

Kimi K3 reached a 57 Artificial Analysis score, but its 130 million evaluation-output tokens, near-$1 task cost, delayed weights, and 64-accelerator serving guidance define the real trade.

---

## [Claude Fable 5 Got Another Week. Super Dario Made the Deadline a Boss Fight](https://markhuang.ai/news/claude-fable-deadline-boss-fight)

2026-07-13

Super Dario turns Claude Fable 5's rolling subscription deadline into a platform game; my read is that another included week is useful, but access uncertainty has become a real workflow cost.

---

## [Claude Turns the Same TypeScript Into 73% More Tokens. That Still Isn't the Bill.](https://markhuang.ai/news/token-price-is-not-the-bill)

2026-07-13

PlayCode measured a 73% tokenizer gap on one TypeScript fixture. My read: that hidden multiplier matters, but model selection should still be based on successful task cost.

---

## [Terminator 2 Made CGI Earn Every Shot](https://markhuang.ai/news/t2-made-cgi-earn-every-shot)

2026-07-12

VFXBlog's oral history shows why Terminator 2 still matters: its digital breakthrough worked because custom software, hand-built motion, practical effects, and ruthless shot selection all served the story.

---

## [LLMs Are Useful Without the Hype](https://markhuang.ai/news/llms-useful-without-the-hype)

2026-07-12

George Hotz's case for loving LLMs while rejecting AI mythology lands with me: the tools are real, but usefulness does not prove inevitability, monopoly, or magic.

---

## [OpenAI Forked Git. The Empty Diff Is the News](https://markhuang.ai/news/openai-git-empty-diff-is-the-news)

2026-07-11

OpenAI's public Git fork arrived with an empty diff and a new SCM hire; my read is that this is a direction signal for agentic source control, not proof of a GitHub rival.

---

## [Model Build-Offs Need Failure Rates, Not Trophies](https://markhuang.ai/news/model-build-offs-need-failure-rates)

2026-07-11

TryAI's 12-model build-off is useful because it exposes run-to-run failures and raw artifacts, but my read is that four familiar app prompts make a shortlist, not a production coding verdict.

---

## [Hy3 Makes Price Part of the Eval](https://markhuang.ai/news/hy3-price-is-the-eval)

2026-07-09

Tencent's Hy3 release is most interesting as a cost-and-workflow bet: open weights, long context, and cheap routing matter only if agents stay reliable outside Tencent's own evals.

---

## [AI Bookkeeping Needs a Harness](https://markhuang.ai/news/ai-bookkeeping-needs-a-harness)

2026-07-09

Toot's GLM 5.2 VAT benchmark is a serious signal for AI bookkeeping, but my read is that cheap accuracy only matters when exception handling, audit evidence, deterministic checks, and human escalation are the product.

---

## [Bun's Rust Rewrite Is the Validation Test](https://markhuang.ai/news/bun-rust-rewrite-validation-test)

2026-07-08

Bun's account of moving from Zig to Rust with Claude is most useful as a stress test for AI-assisted migration: speed matters only if tests, adversarial review, unsafe-code reduction, and release discipline carry the diff.

---

## [AI Heat Needs a Neighbor](https://markhuang.ai/news/ai-heat-needs-a-neighbor)

2026-07-08

BBC's Exmouth pool story is a useful test for AI infrastructure: waste heat only becomes a real sustainability asset when compute demand and heat demand are colocated.

---

## [98% Support Still Needs a Door](https://markhuang.ai/news/browser-support-needs-fallbacks)

2026-07-07

Hugo Barrera's 98% essay is a useful reminder that browser-support averages are not audience guarantees; my read is that modern CSS adoption needs analytics, fallbacks, and graceful degradation.

---

## [Ultra in Codex Has to Beat the Meter](https://markhuang.ai/news/codex-ultra-beat-the-meter)

2026-07-06

Tibo says Ultra will be in Codex; my read is that the real test is not the model tease, but whether subagent-grade coding can survive access, credits, and trust.

---

## [AI Makes the Average Too Cheap](https://markhuang.ai/news/ai-average-too-cheap)

2026-07-06

A rruxandra.github.io essay argues LLMs can flatten thought toward consensus; my read is that the danger is real, but the fix is disciplined workflows that protect deviation.

---

## [The Token Sticker Price Is a Trap](https://markhuang.ai/news/token-sticker-price-is-a-trap)

2026-07-06

Jan Iłowski argues price per 1M tokens is a bad AI cost comparison; my read is that teams need task-level evals, tokenizer-aware budgets, and outcome-per-dollar routing.

---

## [Smart Home AI Needs a Worker Mode](https://markhuang.ai/news/smart-home-ai-needs-worker-mode)

2026-07-05

A USEC 2026 paper on UK domestic workers shows why AI cameras and speakers need bystander controls, deletion workflows, and contracts that treat home monitoring as workplace monitoring.

---

## [AI Textbooks Need Practice, Not Chat](https://markhuang.ai/news/ai-textbooks-need-practice)

2026-07-05

Jonah Bard's Phosphor pilot is useful because it points away from free-form chatbot help and toward embedded retrieval practice; my read is that the evidence is promising, observational, and worth testing harder.

---

## [Claude Code's Cache Scare Needs Receipts](https://markhuang.ai/news/claude-code-cache-scare-needs-receipts)

2026-07-04

A Claude Code GitHub issue alleges unrelated context in an Enterprise ZDR session; my read is that this needs evidence-preserving incident triage, not panic or dismissal.

---

## [AI Burned the Junior Ladder](https://markhuang.ai/news/ai-burned-the-junior-ladder)

2026-07-04

Laurie Voss argues AI has damaged the junior programming market while software creation spreads; my read is that the real problem is rebuilding apprenticeship.

---

## [The AI Goalpost Is Ownership](https://markhuang.ai/news/ai-goalpost-is-ownership)

2026-07-03

Publiczny Profil's AI-coding timeline is useful because it shows old objections decaying; my read is that the next benchmark is ownership, not code generation.

---

## [ZCode Makes the Harness the Product](https://markhuang.ai/news/zcode-harness-is-the-product)

2026-07-01

ZCode's GLM-5.2 page is really a claim that coding agents need an operating layer; my read is that workflow control, quotas, and reliability decide whether it sticks.

---

## [Sonnet 5 Puts Agents in the Default Lane](https://markhuang.ai/news/claude-sonnet-5-default-agent-lane)

2026-06-30

Anthropic says Claude Sonnet 5 brings stronger agentic work to everyday Claude plans; my read is that the real test is migration discipline, cost accounting, and workflow evals.

---

## [South Korea's $1T AI Bet Runs on Water and Power](https://markhuang.ai/news/south-korea-ai-bet-water-power)

2026-06-30

Ars Technica's report on South Korea's chip, data-center, and physical-AI megaprojects looks flashy because of humanoids; my read is that execution depends on power, water, talent, and real robot capability.

---

## [Claude Science Makes the Lab Notebook the Product](https://markhuang.ai/news/claude-science-lab-notebook)

2026-06-30

Anthropic's Claude Science beta matters less as a science chatbot than as a bet on provenance, compute, reviewer checks, and controlled research workflows.

---

## [The AI Bill Is Becoming the Product](https://markhuang.ai/news/ai-bill-is-the-product-test)

2026-06-26

Aditya Patadia argues AI model prices are under pressure; my read is that teams still need routing, evals, and outcome discipline before cheaper tokens become cheaper products.

---

## [AI Tools Still Charge a Conversation Tax](https://markhuang.ai/news/ai-tools-conversation-tax)

2026-06-26

A Tea and Bits essay about the fatigue of talking to LLMs is a reminder that AI coding tools need to reduce social management, not just generate more output.

---

## [Vibe Coding Needs Receipts](https://markhuang.ai/news/vibe-coding-needs-receipts)

2026-06-25

A Papermark founder's allegation against Corgi's DataRoom launch is a reminder that AI-era shipping still needs provenance, license discipline, and public evidence.

---

## [Gemini Computer Use Needs a Trust Loop](https://markhuang.ai/news/gemini-computer-use-trust-loop)

2026-06-25

Google folded computer use into Gemini 3.5 Flash; the interesting test is whether teams can make screen-driving agents observable, sandboxed, and interruptible.

---

## [LastPass's Vault Wasn't the Only Boundary](https://markhuang.ai/news/lastpass-vault-wasnt-only-boundary)

2026-06-25

The Klue breach did not hit LastPass vaults, but it shows why CRM, support cases, and OAuth integrations still matter for password-manager trust.

---

## [OpenAI's Chip Bet Is About Owning the Wait](https://markhuang.ai/news/openai-jalapeno-inference-bet)

2026-06-25

TechCrunch's Jalapeño report matters because OpenAI is treating inference latency, power, and supply as product strategy, not just data-center plumbing.

---

## [Claude Outages Are a Dependency Test](https://markhuang.ai/news/claude-outage-dependency-test)

2026-06-23

The latest Claude status-page flare-up matters because AI coding tools have moved from optional helpers to workflow dependencies.

---

## [OCR's New Battle Is Endurance](https://markhuang.ai/news/unlimited-ocr-endurance)

2026-06-23

Baidu's Unlimited-OCR release is interesting less because it says OCR is back, and more because it treats long documents as the real test.

---

## [NVIDIA Halos Makes Safety the AV Platform](https://markhuang.ai/news/nvidia-halos-safety-platform)

2026-06-23

NVIDIA's Halos page matters because it frames autonomous vehicle safety as a stack of training, simulation, deployment, OS, inspection, and ecosystem evidence.

---

## [AI Broke the Hiring Signal](https://markhuang.ai/news/ai-broke-the-hiring-signal)

2026-06-21

HBR's warning about AI-polished resumes and remote interview performance points to a bigger hiring problem: the old signals were too easy to game.
