Skip to main content

LLMs Made Output Cheap. Trust Is Now the Expensive Part.

Jeremy Theocharis agrees with LLM critics while spending heavily on the tools; my read is that generation scales, but judgment, review capacity, and accountable authorship do not.

Jeremy Theocharis6 min read
Share:
AI-Powered

AI-powered · Limited to 20 requests per hour

A cartoon engineer holds the key to a trust gate between useful AI-made parts and a flood of tangled machine output
When output becomes abundant, the scarce resource is the human judgment that decides what deserves to pass.

Jeremy Theocharis's essay about agreeing with LLM critics while using LLMs heavily is compelling because it refuses the easy team jerseys. He accepts the complaints about slop, open-source burden, weakened learning, dependence, and distorted thinking. He also says he spent almost $10,000 on tokens in June 2026 and still believes the tools help him produce fewer, better things.

I do not read that as hypocrisy. I read it as a useful description of where the bottleneck moved. Models made plausible output cheap. They did not make attention, taste, responsibility, or trust cheap. The more material they produce, the more valuable those human constraints become.

Answer Snapshot

QuestionMy read
What is the essay arguing?LLMs can amplify a person's thinking and craft, but they also amplify empty or careless work; the human cannot outsource judgment.
What problem does it expose?Readers and maintainers can no longer infer effort or understanding from polished output, so review capacity and contributor trust become the limiting resources.
Who benefits if the approach works?Experienced builders who can define the problem, inspect the result, and use models for iteration without surrendering ownership.
What is the strongest practical idea?Make the human state the problem and boundaries first, then force the model to question, criticize, and test those decisions.
What is overstated?Calling LLMs pure amplification is too clean: models also introduce defaults, errors, and persuasive nonsense that the user's existing judgment may not catch.
My thesisLLM fluency is not evidence of human care. Trust has to be earned through small scopes, visible checks, and accountable authorship.

The Contradiction Is the Useful Part

Theocharis opens with a scene from Local-First Conf in Berlin. He says attendees had coding agents open while applauding criticism of LLMs. He also recounts asking Armin Ronacher how the team behind Pi handles a flood of contributions; according to Theocharis, Ronacher said they auto-close almost all pull requests and issues while still encouraging people not to be discouraged.

The tension is sharper because Earendil, Ronacher's company, builds AI tools while its purpose statement says humans are the best agents and that people should wield the tool, not be wielded by it. I like that formulation because it makes human agency an operating constraint rather than a vague compliment.

The public policy response from some open-source communities shows the other side. Zig's official community rules prohibit LLM-generated code and prose, including editing, translation, and brainstorming shared back into its governed spaces. The Gentoo wiki contribution requirements likewise prohibit content created with NLP-based AI assistance.

I understand why maintainers reach for bright lines. A polished submission used to signal at least some investment by its author. Now the cost of creating a batch of plausible submissions can be lower than the cost of carefully reviewing one. Even if a blanket ban is difficult to verify, the policy is evidence that review overload and provenance are not imaginary objections.

A cartoon open-source maintainer inspects one machine part while a delivery chute overflows with polished-looking components
Generation scales faster than maintainer attention. That mismatch turns review into the real capacity limit.

Amplification Is Valuable, but It Is Not Neutral

Theocharis's positive case is not that a model contains a hidden expert. It is that a model can sharpen material already supplied by a human: brainstorming alternatives, checking grammar, testing a sentence, acting as a rubber duck, or taking the opposing side. His distinction between making more things and using more computation to make fewer things better is the most persuasive part of the essay.

That is also where I would tighten the claim. “Amplification” can sound as though the model merely turns up the volume on a clean signal. In practice it adds its own defaults. It can fill an underspecified decision with the familiar answer, smooth uncertainty into confidence, or make a mistaken premise feel finished. Theocharis recognizes this when he describes model agreeableness and warns that people can use the tools well only when they know what good looks like.

For me, that changes the correct unit of evaluation. I do not want to ask whether the output is eloquent or whether the code compiles once. I want to ask which decisions belonged to the human, what evidence checked the result, and how expensive a hidden mistake would be.

A cartoon creator feeds rough idea sketches into a small AI machine and personally selects one refined result
The model can multiply and refine options. The human still has to know which option deserves to become the work.

Productivity Evidence Explains Why Both Sides Feel Right

The public evidence does not support a single answer for every developer and task. METR's early-2025 randomized study found that 16 experienced open-source developers took 19% longer on 246 tasks when AI tools were allowed, even though participants believed the tools had made them faster. That is a strong warning about trusting perceived speed.

It is not a timeless verdict. METR now labels that result out of date. In a February 2026 update, the researchers said newer tools likely produced more speedup, but their next experiment could not measure the size reliably. Developers who valued AI most were less willing to join or submit tasks that might be assigned to the no-AI condition, and parallel agents made time accounting harder.

That measurement problem resembles Theocharis's dissonance. A person can be faster on some work, slower on familiar code, more willing to attempt neglected tasks, and happier while doing it. A single “productivity” number compresses those effects. It also misses the external cost when faster authors create more review work for maintainers.

Theocharis's Best Patterns Make Thinking Smaller

The essay's workflow examples work because they resist unlimited output. Theocharis describes a “grill me” prompt that asks one question at a time until the problem is understood. For coding, he borrows Basecamp's pitch structure and writes a short problem, what will ship, and what will not ship. He also uses fresh agents to attack a plan until their criticisms stop finding real defects.

I would not copy the machinery blindly. Multiple agents can repeat the same assumption, and an adversarial prompt can manufacture objections as fluently as an agreeable prompt manufactures praise. But the underlying shape is sound:

  • Make the human commit to a small problem statement before generation begins.
  • Separate decisions from facts the system can inspect.
  • Keep the review surface small enough that a person actually reads it.
  • Require an observable check: tests, screenshots, decoded output, or another task-specific acceptance condition.
  • Stop when the evidence supports the result, not when the prose sounds confident.

This is less glamorous than autonomous software creation. It is also closer to a workflow I can trust. The model expands the search; the human narrows the claim.

A cartoon engineer stress-tests an attractive AI-built bridge model and finds hidden weak joints while a robot assistant waits
A polished result is a candidate, not a conclusion. Verification is where confidence becomes deserved.

Trust Cannot Be Generated on Demand

Theocharis proposes a personal test for writing: would he read the words aloud in front of an audience without needing to explain what he really meant? I like that because it makes authorship costly again. The standard is not whether a detector sees a stylistic tell. It is whether a named person is willing to own every sentence.

For open source, ownership needs more visible evidence: a bounded change, a clear problem statement, tests, screenshots where useful, responsiveness to review, and a contributor who can explain the tradeoffs. None of those prove that no model was involved. They do make the contribution easier to evaluate on the dimensions maintainers actually need.

Local models may reduce dependence on a vendor, as Theocharis argues, but they do not solve this social layer. Open weights cannot tell a maintainer whether a submitter understood a patch. Cheap tokens cannot make a reader care. More adversarial agents cannot give an author taste they have not developed.

My Bottom Line

I agree with Theocharis that criticism and heavy use can coexist. In fact, the criticism is what makes serious use possible. The useful posture is neither surrender nor abstinence; it is to treat generation as abundant and judgment as scarce.

The mistake is to count output as progress before someone has paid the verification bill. A model can help me explore more options, state an idea more clearly, or implement a bounded task. It cannot lend me credibility. That still accumulates slowly, through work I understand, checks other people can inspect, and corrections I am willing to make.

LLMs made output cheap. The winning workflow will be the one that refuses to make trust cheap with it.

License

News text © 2026 Mark Huang. News text may be shared or translated for non-commercial use with attribution to https://markhuang.ai/news/llm-output-cheap-trust-expensive.

Suggested attribution: Based on "LLMs Made Output Cheap. Trust Is Now the Expensive Part." by Mark Huang, originally published at https://markhuang.ai/news/llm-output-cheap-trust-expensive.