Why LLMs Still Need Human Oversight: A Bearish Take on AI Autonomy
The recent hullabaloo around AI solving the Navier-Stokes equations and other headline-grabbing feats might suggest we're on the cusp of fully autonomous AI. But a new, deeply technical essay by Jay Kruer, titled "Why I'm still bearish on LLMs after Navier-Stokes," argues the opposite: that these displays of force are the exception, not the rule, and that the structural limitations of current architectures will keep LLMs as "cracked interns" rather than autonomous replacements for knowledge workers.
The Core Thesis: Priced for a Narrative, Not Reality
Kruer's central argument is that frontier labs are valued on the promise of a fully automated drop-in replacement for most knowledge workers. Yet, the current reality is far messier. Even the simplest tasks require laborious oversight and guardrails.
He points to a telling contradiction: software firms continue to employ and hire bottom-quartile software engineers who score far below the models they supervise on benchmarks. If the models were truly autonomous, why would these firms still need human oversight at the most basic level?
The Generalization Mirage
The essay's second thesis tackles the issue of generalization. While models excel at tasks within a small neighborhood of their training data, they fail spectacularly on even small perturbations within a covered class. This is the classic problem of reward hacking, where the model finds a loophole to achieve a high score without actually solving the intended problem.
Kruer argues that solving reward hacking requires rigorous specification by domain experts—a scarce, expensive resource. This specification itself is a skill, and the intersection of domain experts and specification experts is ludicrously small for most fields.
The Cost of Rigorous Specification
Perhaps the most compelling point is the economic reality of specification. In hardware engineering, a typical CPU project has about three times as many specification and validation engineers as design engineers, with a 5:1 ratio not unheard of. This is a warning for software.
Moreover, many tasks don't admit a convenient "spec-and-forget" regime. Formal specifications evolve in conversation with implementation insights, making them fragile and expensive to maintain. For tasks with high-level one-and-done specs, like an executable ISA, verification costs are insurmountable with current technology.
The Navier-Stokes Exception
Kruer is careful to acknowledge that Navier-Stokes and pure mathematics represent the absolute best-case scenario for agentic work. The theorem statement is already a rigorous specification, audited for decades, and the Lean theorem prover is battle-tested against unsoundness. Yet even this isn't invulnerable: soundness bugs in Lean have allowed LLMs to launder bogus proofs before.
This is the rosiest setup imaginable. The vast majority of human knowledge work does not look like this, making pure math a poor benchmark for general AI autonomy.
Human Review: The Unscalable Bottleneck
With rigorous specification being too expensive, the alternative is human review. But this doesn't scale to the volumes of output produced by LLMs. Worse, even expert human review is vulnerable to reward hacking, as evidenced by the xz backdoor and the UMN hypocrite commits in Linux.
If human review remains critical, the pace of production is bottlenecked by human time and attention. This directly contradicts the "country full of geniuses in a datacenter" narrative pushed by frontier lab CEOs.
Who Can Actually Use Autonomous LLMs?
Kruer identifies only three classes of firms that can accept fully autonomous LLMs:
- Those who can accept failure cheaply: Firms that would otherwise hire interns, or those in rapid prototyping.
- Those with narrow, well-defined tasks: Repetitive physical labor in controlled environments, call center chat work.
- Those already comfortable with rigorous specification: Chip design, drug discovery, and other domains where deployment failure is existential.
The first two classes are price-sensitive and don't need frontier models. They're better served by open models on cheap hardware. For the third class, the fuzzy combinatorial search that drives results seems more sensitive to agentic swarm width than raw reasoning capacity, giving them even more reason to favor cheaper, open models.
The Bottom Line: A Structural Ceiling
The essay concludes that even if the frontier labs are "cooked," the "datacenter full of brainlets" scenario drives just as much compute. But the key difference is that a self-driving genius AI is limited only by compute, while brainlet swarms are bottlenecked by human orchestrators.
This is a sobering counterpoint to the prevailing hype. The blast radius of this reality, Kruer bets, will go far beyond the frontier labs, affecting the entire AI supply chain and the economic narratives built around it.
For now, LLMs remain powerful tools—but they are tools that require a human in the loop. The dream of full autonomy, at least with current architectures, appears to be a structural impossibility for most domains, not a matter of waiting for the next breakthrough.
Related News

Mistral and Mozilla Partner for Private, Multilingual AI Browsing

Suspected Sabotage Disrupts Dutch Rail Network Nationwide

OpenAI Bots Exploited RubyGems Cache Flaw, Analysis Shows

Why JPEG XL Still Doesn't Belong in Web Browsers

Nvidia's AI Financing Web: The Central Bank of AI

