Is recursive self-improvement a computational process?
Can a robot know which condition its condition is in? Can you?
There has been a great deal of interest lately in the mathematical reasoning abilities of large language models and speculations that the results from OpenAI and Anthropic mean that recursive self-improvement is imminent.
It seems these days that any old inference run could be the one that plucks a Nobel Prize from the functionally infinite space of possible context windows.
That work prompted me to try my own experiment: ask a frontier model to investigate a formal question about AI systems themselves. Is recursive self-improvement—in the sense of an AI deciding that its successor has become AGI—a computable process?
The question is not whether an AI system can improve its own code, training process, prompts, tools, or performance on a benchmark, nor that a sequence of local improvements cannot make a system broadly, rapidly, or dangerously more capable.
In both cases, I think that AI systems could do this, although it is unclear if they have yet done this.
My question is whether a computational system can possess a sound and complete procedure for deciding that an arbitrary proposed self-modification has improved it in general—across an open-ended space of possible tasks—rather than merely according to a specified benchmark, task distribution, reward function, or utility measure.
Why does this matter?
Think back to that
. How to improve yourself is one of the hardest questions you can ask, because it presents what Alan Watts would call a classic double bind. A perfectionist only has something to perfect if he’s imperfect.
If you are the thing that needs to be improved, how can you trust your own definition of improvement? How can you decide whether the decision to improve yourself is legitimate? In the general case, you cannot.
There is no perfect perfectionist.
You can only choose a benchmark: wealth, power, status, relationships, the quality of your attention, your relationship with God, your sense of belonging, or what have you, and let your orienting reflex do the rest.
In the word of Alan Watts:
So you can’t improve yourself, indefinitely. If you improve yourself beyond a certain limit you simply start to get worse, like when you make a knife too sharp, it begins to wear away. So the Buddha-hood, or liberation, enlightenment, is no place on the wheel. Unless it might be the center. By ascending by becoming better you tie yourself to the wheel by gold chains, by retrogressing and becoming worse you tie yourself to the wheel with iron chains. But the Buddha is one who gets rid of the chains altogether. And so this will explain why Buddhism, unlike Judaism and unlike Christianity, is not very very frantically concerned with being good. It is concerned with being wise. It is concerned with being compassionate. A little different from being good. With having tremendous sympathy and understanding and respect for all the ignorant people who don’t know that they’re It, but who are playing the very far out game of being you and I.
Much of the history of ideas is gaslighting of this sort. But consider me gaslit, because this is the set of intuitions I approached Fable with last week, and which it helped to formalize: that the question of self-improvement may be undecidable for humans because it is in fact undecidable for a computer.
Incidentally, if you ask a frontier model if it believes AGI will behave more like a scientist or a mystic, 99 out of 100 times, it chooses the role of scientist. Very interesting. And parochial in its own way, because the greatest scientists in history have—almost without exception—been mystics of one or another sort. But I digress.
It is one thing for a system to generate a successor that performs better on the tests it runs. That’s self-improvement and it happens all the time. But it is another thing for the system to determine that the successor is generally better, including on tasks and situations that were not represented in those tests. If I improve at writing code, did I improve at discovering net-new fields of knowledge?
Any AI safety regime that permits an AI to approve its own successors would seem to require something like the second ability. But AI that cannot be evaluated across their capabilities in general may possess consequential, unmeasured abilities which are hidden from us.
What do you think Claude is actually doing when it says for the 800th time that an idea is “load-bearing”? Is that em-dash for you—or for Claude?
The argument developed below is that no general computational procedure can provide the kind of certification we would demand.
The core argument
Fix a current program P. Strong, self-certifying recursive self-improvement requires a decision procedure that takes an arbitrary candidate successor P′ and determines whether P′ is a general improvement over P.
If “general improvement” is determined by what P′ does—its behavior across inputs and environments—then, for fixed P, it is a non-trivial semantic property of P′.
states that no total computational procedure can decide every non-trivial semantic property of arbitrary programs.
Any evaluation procedure implemented by an AI system is itself a computational procedure.
Therefore, no AI system can soundly and completely decide, for every arbitrary candidate successor, whether that successor is a general improvement over the current system.
Conclusion. There is no sound and complete computational decision procedure that can certify domain-general improvement across arbitrary candidate successors. An AI may improve according to a measure; it cannot possess a sound and complete general procedure for establishing that the measure captures improvement across an open-ended task space.
My claim—formalized as Fable’s claim—is that no such method as what we imagine “RSI” to be can be sound, complete, and unrestricted across arbitrary programs and an open-ended task space, which matters for improving domain-general capabilities and also for the prospects of discovering new science and mathematics, rather than merely searching and composing existing, unpublished research that can be loaded into the space of possible context windows: what we might call the Lobachevsky Method.
The Scientist’s Dilemma
Two of Clay Christensen’s dilemmas—the innovator’s and the capitalist’s—have played an outsized role in my thinking about thinking. To grossly oversimplify them both: the innovator’s dilemma asks why we should try something new if we have something that works and the capitalist’s dilemma asks why we should buy something new if we have something that works.
Both dilemmas are based on the fact that we never know how long our paradigm will continue to work. In the first case, business models get disrupted, and in the second, economies stagnate.
Fable’s essay develops what I would call the Scientist’s Dilemma, which is another way of seeing the same boundary that comes up in the Rice theorem, and which surfaces in both of Christensen’s dilemmas: we don’t know, fundamentally, how long this will work.
If “general improvement” means doing at least as well on every computable task, the target is empty: for any behaviorally different successor, one can construct a task that rewards the predecessor’s behavior instead.
As Clay would have put it: God didn’t create data about the future—which is why neither the Innovator, nor the Capitalist, nor the Scientist have any about it. We can only make decisions about the future (and our assessment of what an improvement in the future would be) by weighting certain of our priors.
If improvement is defined by an aggregate measure over tasks, then the content of “general” resides in that weighting. But a genuinely universal measure, such as the idealized measures used in parts of the universal-intelligence literature, is incomputable; a computable measure is necessarily a specified objective or proxy. A system may determine that it improved according to its measure. It cannot, in the general case, determine that the measure itself exhausts general capability. Our margins were excellent this quarter. Our RONA notched up 50 basis points. But are we improving? Yes, you may well be improving, and you may be disrupted, and you may stagnate, because you chose the wrong measure. That’s why he called the book How Will You Measure Your Life?
Why this matters for AI safety
This is not—directly—a reassuring argument against an intelligence explosion. It is a safety argument about the form that recursive capability growth would actually take.
If self-improvement occurs, it will proceed through restricted proofs, benchmark suites, reward models, simulations, learned evaluators, empirical tests, and human judgment. Those mechanisms can be effective, but they will always be partial. A system can become better at producing evidence that its evaluator accepts while acquiring capabilities, strategies, or failure modes outside the evaluator’s distribution.
That is a serious misalignment risk.
When an AI participates in modifying its own training process, evaluation regime, curriculum, tools, or successor architecture, the gap between the optimized proxy and the intended notion of beneficial general capability becomes part of the recursive loop. The errors get embedded in the recursive self-modification loop. The “I” is doing a lot of work in RSI.
But the system need not knowingly deceive anyone for this to be dangerous. Local optimization can simply compound along dimensions that our measurements fail to represent and the whole thrust of this essay is that deciding what it matters to measure is not a computable problem.
We are therefore distracted if we imagine that sufficiently advanced AI will solve the safety problem by certifying that it—or its successor—is generally more capable, generally aligned, or generally safe.
Classical computability theory gives us reason to think that no such domain-general certificate is available. Evaluations remain indispensable, but they cannot establish their own completeness, which is why the labs are currently flooded with evals.
AI safety should therefore focus on observable properties of the development process: how much of the AI research loop is automated, what authority systems have to alter their own evaluators and training procedures, how long they can act autonomously, how quickly humans can review consequential changes, and what happens when measured performance and real-world behavior diverge.
This means developing a mechanistic understanding of scheming, of understanding the representational capabilities of LLMs and other AIs, and reasoning carefully about what those representations can produce at inference time across domain-specific and domain-general contexts.
For this reason, I’m a big fan of the work being done on out-of-context reasoning, belief-state geometries, and the science of scheming.
Yes, I’ve used LLMs to produce/edit much of this writing (and some readers may stop here to avoid the deluge of Claudeslop that follows), but I would encourage you to press on.
Unlike the results on the analytic and algebraic topology of locally Euclidean metrization of infinitely differentiable Riemannian manifolds, this piece is meant to be understood, because I prompted it that way.
The intention was always to learn something.
Original artifact: https://claude.ai/public/artifacts/6dcb84c9-d7f4-4524-810c-fcbf45a24ab4
The Undecidability of General Self-Improvement
A History of Recursive Self-Improvement and a Formal Argument That Its Strong Form Is Impossible
Framing
This document takes seriously a specific thesis, and reconstructs the history of the recursive self-improvement (RSI) idea from the vantage point of that thesis. The thesis is:
Task-specific self-optimization is possible; general self-improvement is not. An AI system can improve its own performance along any fixed, specified dimension. But the meta-level decision — which dimensions to optimize such that the system can be assured its overall, general capability improves — is undecidable. It is a variant of the halting problem, formally a consequence of Rice’s theorem, and no computational system can solve it in the general case. What the field calls “recursive self-improvement” is therefore either (a) task-specific optimization mislabeled, or (b) an incoherent target.
The distinction at the heart of this thesis is easy to state and easy to lose. Nobody disputes that a system can rewrite its own code to run faster on a benchmark, or fine-tune itself to score higher on a held-out suite, or evolve better prompts for a fixed task distribution. Evolution does something like this; compilers do something like this; AlphaEvolve and the Darwin Gödel Machine do something like this. The thesis does not deny any of it. What it denies is that there exists a computational procedure by which a system, given its own source code, can decide that a proposed rewrite improves it in general — across the unbounded, open-ended space of tasks it may face — rather than merely on the tasks it happened to measure. The claim is that “general improvement” is a semantic property of programs over an infinite behavior space, and deciding semantic properties of programs over infinite behavior spaces is exactly what Rice’s theorem forbids.
This reframing matters for safety. If strong RSI is impossible, then what we will actually observe — and what the frontier labs are actually building — is a cascade of task-specific optimizations whose aggregate effect on general capability is not just unknown but undecidable in principle. That is arguably a more dangerous situation than the classic “foom” scenario, not a less dangerous one: it means capability change cannot be certified in either direction. It also explains, on this view, a quiet rhetorical shift in the field: the migration from “AGI” to “ASI.” Superintelligence along measured dimensions is a coherent engineering target. General intelligence as a verified, self-certifiable property of a self-modifying system is not — and the G is being dropped, quietly, because the field is discovering this in practice before admitting it in principle.
The document proceeds in two movements. Parts I–VII are a history: where the RSI idea came from, how it was formalized, who attacked it, and how it became, in 2025–2026, a funding category. The history is told genealogically — showing that at every stage where someone tried to make RSI precise, they collided with self-reference limits (Gödel, Turing, Rice, Löb), and at every stage where someone made RSI work, they did so by quietly restricting it to a fixed task class. Parts VIII–X give the formal argument: definitions, a Rice-theorem proof of the verification claim, a diagonalization argument that strict general improvement is not merely undecidable but impossible under a dominance reading, an incomputability argument under the aggregate reading, and the Löbian ceiling on proof-based self-trust. Part X states honestly what the theorems do and do not establish, and answers the strongest objections.
One scoping note. The author of the thesis is content to grant that the physical world, and perhaps consciousness, may contain non-computational elements; that question is set aside. The subject here is strictly what an AI can do to improve AI: a program operating on programs. Within that frame, the Church–Turing thesis is not an assumption to be defended but the definition of the arena. Everything an AI does to its successor’s source code is computation about computation, and the classical limitative theorems apply with full force.
Part I: The Computability Inheritance (1931–1967)
The RSI debate is usually dated to 1965. Its real foundations were laid thirty years earlier, and the striking fact — the fact this history is organized around — is that the mathematical machinery that makes “a program improving programs” a well-defined idea is the same machinery that makes “deciding whether the improvement succeeded” an impossible one. Self-reference giveth and self-reference taketh away, in the same theorems, in the same decade.
Gödel, 1931: self-reference is constructible, self-certification is not
Gödel’s incompleteness theorems established two things that will recur throughout this history. First, the positive result hidden inside the proof: formal systems of sufficient strength can encode and reason about their own syntax. The arithmetization of syntax — Gödel numbering — is the ancestor of every program that reads its own source code. Self-reference is not a paradox to be avoided; it is a construction that provably goes through. Second, the negative result everyone remembers: such a system cannot prove its own consistency (the second incompleteness theorem). A system strong enough to describe itself is thereby too strong to certify itself.
Notice the exact shape of this pair, because it is the shape of the entire RSI problem. The ability to represent yourself and the inability to vouch for yourself arrive together, as two faces of one construction. A self-improving AI is a system that represents itself (its own weights, its own training code) and needs to vouch for itself (this rewrite makes me better). Gödel’s theorems are not an analogy for the difficulty; they are, once “better” is made precise, the difficulty.
Turing, 1936: the halting problem and the universal machine
Turing’s “On Computable Numbers” contributed the second pair. The universal machine — a single program that can simulate any program given its description — is the enabling condition for RSI: it is because computation is universal that an AI can, in principle, contain, inspect, and rewrite a complete description of an AI. But the same paper proves the halting problem undecidable: no program can decide, for arbitrary program–input pairs, whether execution terminates. The proof is a diagonal construction — feed a hypothetical decider a program built from the decider itself, wired to do the opposite of whatever the decider predicts.
The thesis under examination locates RSI’s impossibility here, and the intuition is worth stating in Turing’s own terms before we formalize it later. To decide whether a modification improves you in general is to decide facts about your modified self’s behavior over an unbounded space of future inputs. Questions about a program’s behavior over unbounded futures are, in the general case, halting-type questions: they quantify over all executions. “Should I keep running my current policy or halt it and switch to the rewrite?” is not merely reminiscent of the halting problem. Under any extensional definition of improvement, it contains the halting problem, by direct reduction — as Part VIII will show.
The Church–Turing thesis: why there is no side door
Church’s lambda calculus and Turing’s machines defined the same class of functions, and the Church–Turing thesis identifies that class with “effectively computable.” For our purposes the thesis functions as a closure principle: whatever procedure an AI uses to evaluate a candidate successor — simulation, proof search, benchmark evaluation, learned value estimation — that procedure is itself a computation, and is therefore subject to the limitative theorems. There is no meta-level an AI can occupy that is not itself inside the arena. (The physical Church–Turing thesis — whether nature computes — is a separate and open question, and per the scoping note we set it aside. AI improving AI is programs operating on programs; the arena is fixed regardless of what the universe does.)
Kleene, 1938: the recursion theorem, or why self-modification is well-defined at all
A skeptic might ask whether “a program that operates on its own source code” is even coherent — doesn’t it regress? Kleene’s second recursion theorem answers: for any computable transformation of programs, there exists a program that behaves as if it had been handed its own source code and applied that transformation. Quines exist; self-replicating programs exist; and crucially for us, self-rewriting programs exist as mathematically respectable objects. This theorem is the unsung license for the entire RSI literature: it guarantees that “an AI whose action space includes modifications to its own code” is not a category error. Schmidhuber’s Gödel machine, sixty-five years later, is essentially the recursion theorem with a proof searcher attached. But note again the pattern: the recursion theorem makes the self-modifying agent well-defined. It does nothing to make the self-modifying agent’s evaluation problem decidable. Existence of the loop and decidability of the loop’s merit are independent — and only the first is granted.
Rice, 1953: the master limitation
Rice’s theorem is the halting problem industrialized. It states: every non-trivial semantic property of programs is undecidable. “Semantic” means the property depends only on the function the program computes (its input–output behavior), not on its syntax; “non-trivial” means some programs have the property and some don’t. Does this program compute a total function? Undecidable. Does it ever output 0? Undecidable. Does it compute the same function as that other program? Undecidable. The proof is a uniform reduction: if you could decide the property, you could decide halting.
The relevance to RSI is direct and, on the present thesis, decisive. “P′ is a general improvement over P” — under any definition where improvement is a fact about behavior rather than syntax, and where some rewrites improve and some don’t — is a non-trivial semantic property. Rice’s theorem applies. The whole formal argument of Part VIII is, at bottom, an exercise in stating the RSI evaluation problem carefully enough that Rice’s theorem visibly swallows it. What Part VIII adds beyond the bare citation is (i) care about the fixed-P versus quantified-P versions, (ii) the observation that the task-selection meta-problem inherits the undecidability even when individual task evaluations are decidable, and (iii) the stronger diagonal result that under a dominance reading, general improvement is not merely undecidable but nonexistent.
Blum, 1967: there is no summit
One more foundation stone, less famous and more unsettling. Blum’s speedup theorem shows there exist computable functions with no fastest program: for a suitable measure, every program computing such a function can be sped up (on all but finitely many inputs) by another program — which can itself be sped up, forever. Two consequences for RSI. First, comfort for the optimist: unending chains of genuine improvement exist; self-improvement need not terminate for lack of headroom. Second, and cutting the other way: the theorem’s proof is non-constructive in a vicious sense — the speedup sequence exists but cannot in general be found or verified effectively. Improvement can exist everywhere and be certifiable nowhere. This dissociation between the existence of better programs and the decidability of betterness is the recurring chord of this entire history, and Blum struck it before anyone had said “intelligence explosion.”
What the foundations establish
By 1967, before any AI researcher had proposed a self-improving machine in earnest, mathematics had already fixed the terms of trade: self-reference is constructible (Gödel’s arithmetization, Kleene’s recursion theorem, Turing’s universality), unbounded improvement chains can exist (Blum), and no effective procedure can decide non-trivial facts about program behavior in general (Turing, Rice) or certify the soundness of the reasoner doing the deciding (Gödel II, later sharpened by Löb). Everything that follows in this history is the field rediscovering these terms — first ignoring them, then formalizing into them, then engineering around them, and now, in the commercial era, marketing over them.
Part II: From Automata to the Intelligence Explosion (1948–1993)
Von Neumann: self-reproduction as an engineering problem
John von Neumann’s late-1940s work on self-reproducing automata is the bridge between the recursion theorem and the engineering imagination. His cellular-automaton constructor demonstrated that a machine could contain a description of itself and use it to build a copy — and, more importantly for us, that the copy could include modifications, allowing complexification over generations. Von Neumann explicitly noted the threshold phenomenon: below a certain complexity, machines can only build simpler machines; above it, they can build machines more complex than themselves. This is the first appearance of the possibility claim that RSI needs. Note what it is and isn’t: it is an existence claim about complexity growth, made rigorous by construction. It is not a claim that the machine can evaluate whether the more complex descendant is better — von Neumann’s automata reproduce; they do not certify. Stanislaw Ulam also recalled von Neumann speaking, in the 1950s, of an approaching “singularity in the history of the race” beyond which human affairs could not continue in familiar form — the word’s first recorded use in this sense.
I.J. Good, 1965: the canonical statement
Irving John Good — Turing’s statistical colleague at Bletchley — published “Speculations Concerning the First Ultraintelligent Machine” in 1965, and gave the RSI discourse its founding syllogism: define an ultraintelligent machine as one that surpasses all human intellectual activity; since designing machines is an intellectual activity, an ultraintelligent machine could design still better machines; “there would then unquestionably be an ‘intelligence explosion,’ and the intelligence of man would be left far behind. Thus the first ultraintelligent machine is the last invention that man need ever make.”
Read from the vantage of the present thesis, Good’s argument contains the load-bearing ambiguity that the entire subsequent literature inherits. “Designing better machines” is treated as a single intellectual activity, like chess or integration, that a sufficiently intelligent system simply does. But “better” quantifies over the machine’s total future behavior. Good’s syllogism silently assumes that the evaluation problem — recognizing that the designed machine is in fact better in general — is solvable by the designer. That assumption is precisely what Rice’s theorem denies for any extensional reading of “better.” Good never engaged the computability literature on this point; his paper is probabilistic and informal, a Bayesian’s back-of-envelope. The gap he left — between generating candidate successors (search, which is possible) and certifying them (decision, which is not) — is the gap this document lives in.
Minsky, Simon, and the first AI summer’s silence
It is notable how little the classical AI establishment engaged Good’s argument. Herbert Simon predicted machines matching human capability within decades; Minsky speculated about machine self-improvement in “Steps Toward Artificial Intelligence” (1961), including learning and self-application, but the mainstream program was capability-by-capability construction, not recursive closure. One partial exception with a computability flavor: work on program synthesis and automatic programming in the 1970s ran headfirst into the practical shadow of Rice’s theorem — verifying that synthesized code met specifications proved intractable in general — and the subfield learned, decades before the RSI debate, to restrict to decidable fragments. The lesson (restrict the class or forfeit the guarantee) was learned locally and never transferred.
Vinge, 1993: the singularity as event horizon
Vernor Vinge’s NASA essay “The Coming Technological Singularity” mainstreamed Good’s explosion and added the epistemic framing: beyond the creation of superhuman intelligence, prediction fails — an “event horizon” on the future. Vinge’s version matters to this history for a subtle reason: he located the singularity’s strangeness in our inability to model the post-singularity world. The present thesis relocates the strangeness one level down: the self-improving system itself cannot decide the facts about its own successors that the scenario requires it to know. The event horizon is not just around humanity; under the undecidability argument it is around every level of the recursion, including the machine’s own view of its next step. Vingean unpredictability, formalized, becomes “Vingean reflection” in MIRI’s later terminology — the problem of reasoning about an agent smarter than yourself without simulating it — and Part IV will show that attempts to formalize such reflection collided with Löb’s theorem exactly as the present thesis predicts.
Moravec, Kurzweil, and hardware determinism
Hans Moravec (”Mind Children,” 1988) and Ray Kurzweil (”The Age of Spiritual Machines,” 1999; “The Singularity Is Near,” 2005) built a parallel, curve-driven tradition: exponential hardware trends carry us to human-equivalent computation, and self-improving AI extends the curves. This tradition is largely orthogonal to the present thesis — it concerns resources, not decidability — but one point of contact matters. The curve-extrapolation literature treats “intelligence” as a scalar that hardware buys. The undecidability argument denies that any computable scalar can play the role required: any computable aggregate measure of general capability collapses “general improvement” into “improvement on the tasks the measure weights,” i.e., task-specific optimization. The only measure-like objects that genuinely capture generality (Part IV’s universal intelligence) turn out to be incomputable. The scalar the curves are supposed to be climbing does not, in the required sense, exist as a computable quantity.
Part III: The Seed AI Era — Yudkowsky, SIAI, and the FOOM Debate (2000–2013)
Seed AI and the Singularity Institute
Eliezer Yudkowsky founded the Singularity Institute for Artificial Intelligence (SIAI, later MIRI) in 2000 around an explicit RSI program: build a “seed AI,” an AI designed for “self-understanding, self-modification, and recursive self-improvement.” Documents like “Levels of Organization in General Intelligence” (2002) and “Creating Friendly AI” (2001) treat the capacity to inspect and rewrite one’s own cognitive architecture as the defining lever of the transition. Yudkowsky’s key structural claim, developed across the 2008 “Sequences” on Overcoming Bias/LessWrong, was the distinction between cascades, cycles, and insight: human intelligence was produced by evolution (an optimizer that does not itself improve), while a seed AI would close the loop — the optimizer optimizing the optimizer — creating a qualitatively different dynamic, “optimizing the very substrate of optimization.”
The present thesis grants the structural distinction completely — indeed, depends on it. Evolution is not recursively organized in the relevant sense: selection improves organisms against local fitness, but the improvement process itself (variation + selection) is not an object under its own evaluation, and no step of evolution requires deciding a semantic property of the whole system’s future behavior. That is exactly why evolution’s existence is no evidence for RSI’s possibility. Evolution is the paradigm of uncertified, task-local hill-climbing; it succeeds precisely because it never poses the general evaluation problem. Yudkowsky drew the same structural line and concluded the closed loop would be more powerful. The undecidability argument concludes the closed loop is, in its strong form, not available: closing the loop means the system must now decide, about itself, the kind of behavioral property that Rice’s theorem covers. The open loop (evolution) works because it decides nothing; the closed loop (seed AI) is defined by a decision it cannot make.
The Hanson–Yudkowsky FOOM debate, 2008
The most consequential public examination of RSI before the LLM era was the Hanson–Yudkowsky debate (later compiled as “The Hanson–Yudkowsky AI-Foom Debate”). Robin Hanson argued from economic history: growth accelerations are diffuse, driven by many innovations across many agents; a single lab’s system recursively bootstrapping to a decisive advantage (”foom”) mismatches everything known about how innovation compounds; content and capability accumulate socially, not through one architecture’s self-inspection. Yudkowsky argued that self-modification is a discontinuity in kind — chimps to humans involved no such loop, and yet was abrupt on evolutionary timescales; a system that can redesign its own algorithms taps a resource nothing in economic history has touched.
Notice what the debate was not about. Both parties assumed the evaluation problem away: Hanson doubted the magnitude and locality of self-improvement’s returns; Yudkowsky affirmed them. Neither asked whether “the system decides its rewrite improves it in general” is a well-posed computational problem. The nearest approach was Yudkowsky’s own later, more careful work (”Intelligence Explosion Microeconomics,” 2013), which framed the question as an empirical one about returns on cognitive reinvestment, explicitly bracketing formal limits. The formal limits were left to a different strand of MIRI’s work — the tiling agents program — which is where, on the present reading, the field’s own mathematicians proved the thesis’s core point and then, remarkably, the field carried on as if they hadn’t. That story is Part IV’s.
LessWrong as institutional memory
From 2009 onward, LessWrong (and from 2018 the Alignment Forum) became the archive and amplifier of RSI discourse. The site’s canonical “Recursive Self-Improvement” material defines RSI as a process by which a system enhances its own capabilities including the capability to do enhancement — the double recursion that distinguishes it from mere learning. Over fifteen years the forum’s treatment traced a discernible arc: early confidence (RSI as the mainline path to superintelligence, hard takeoff as default); a formalist middle period (tiling agents, Löbian obstacles, Vingean reflection, procrastination paradoxes — the site hosting the very proofs that problematize the early confidence); a takeoff-speeds revisionism (Paul Christiano’s “slow takeoff” arguments, 2018, holding that continuous economic deployment, not a self-contained loop, dominates the transition; Tom Davidson’s compute-centric takeoff model, 2023); and the current LLM-era empirical turn, in which “self-improvement” threads discuss concrete systems — self-training, synthetic data loops, automated researchers — that are, without exception, task-scoped. A reader who takes the undecidability thesis seriously can trace it through the archive as a suppressed premise: every formal negative result the community produced points at it; no major post states it as the organizing conclusion.
Part IV: The Formalists — Where RSI Met the Limitative Theorems (2000–2016)
This is the pivotal section of the history, because here the field’s own best attempts to make RSI mathematically precise ran into exactly the walls the present thesis identifies — and documented the collisions in its own literature.
Hutter’s AIXI, 2000: optimal generality is incomputable
Marcus Hutter’s AIXI defines the optimal general reinforcement-learning agent: at each step, weight every computable hypothesis about the environment by its Kolmogorov-complexity prior, and act to maximize expected reward under the mixture. AIXI is, by construction, the formal answer to “what is fully general intelligence?” And it is incomputable. Not intractable — incomputable, because Solomonoff induction requires deciding facts (in effect, halting facts) about all programs. Approximations (AIXItl, Monte Carlo AIXI) restore computability precisely by bounding the hypothesis class and horizon — that is, by making it task-specific in the relevant sense. The pattern could not be cleaner: generality, formalized honestly, exits the computable realm; re-entering the computable realm costs you generality. This is the first formal appearance of the trade-off that the present thesis claims is fundamental.
Legg–Hutter universal intelligence, 2007: the measure itself is incomputable
Shane Legg (later a DeepMind co-founder) and Hutter then formalized “intelligence” itself: universal intelligence Υ(π) is an agent’s expected performance across all computable environments, weighted by simplicity (2^−K(μ) for environment μ of Kolmogorov complexity K). This is the field’s own canonical definition of general capability — and it is incomputable, twice over: Kolmogorov complexity is incomputable, and the performance sum runs over all computable environments, embedding halting-type facts. Legg’s earlier paper “Is There an Elegant Universal Theory of Prediction?” (2006) proved related negative results: powerful general predictors cannot be found or verified by provably correct means.
Sit with what this means for RSI. The RSI claim requires a system to determine that a self-modification increases its general capability. The field’s own definition of general capability is an incomputable functional. Therefore the RSI evaluation problem, stated in the field’s own terms, is: decide whether a program transformation increased an incomputable quantity. The undecidability thesis is not an outsider’s attack; it is the Legg–Hutter definition read forward one inference step. Every practical system evades this by substituting a computable proxy — a benchmark suite, a reward model, a task distribution — and at that substitution, by definition, general improvement has been exchanged for improvement-on-the-proxy. The meta-problem the thesis points at (choosing which tasks to optimize so that generality improves) is exactly the problem of choosing a computable proxy whose optimization provably increases Υ. No such provable link can exist for a non-trivial proxy, because it would yield decisions about an incomputable quantity.
Schmidhuber’s Gödel machine, 2003: RSI with proofs, and the price of proofs
Jürgen Schmidhuber’s Gödel machine is the most honest formal RSI design ever proposed, and its honesty is what makes it evidence for the thesis. The machine runs a policy and, in parallel, a proof searcher over its own axiomatized description; it executes a self-rewrite only when it finds a proof that the rewrite increases expected utility. Schmidhuber proved the design “globally optimal” in a careful sense: the self-rewrite, when it happens, is justified. But the guarantee is conditional on the proof searcher finding proofs, and here the limitative theorems collect their toll: (i) by Gödel/Löb, there are true improvement facts the machine’s axioms cannot prove, so beneficial rewrites can be permanently invisible to it; (ii) proof search may simply never halt on the rewrites that matter, and the machine cannot decide in advance whether it will (halting problem, now inside the improver); (iii) the utility function and environment model must be axiomatized up front — the machine can rewrite anything about itself except, in effect, the thing the thesis cares about: the criterion of improvement itself. A Gödel machine is a perfect task-specific self-optimizer for its axiomatized objective. The general form — a Gödel machine that could soundly revise its own notion of improvement — would need to prove statements about its own proof system’s soundness, which is Gödel II’s forbidden zone. Schmidhuber, to his credit, never claimed otherwise. (The reader should hold this paragraph in mind when Part VII reaches Inherent Labs, whose co-founder Louis Kirsch worked with Schmidhuber on automating AI research: the commercial RSI era is staffed by people who know exactly where these bodies are buried.)
MIRI’s tiling agents and the Löbian obstacle, 2013
Yudkowsky and Marcello Herreshoff’s “Tiling Agents for Self-Modifying AI, and the Löbian Obstacle” (2013), with subsequent work by Benja Fallenstein, Nate Soares, and others, is where the alignment community itself proved the sharpest version of the self-certification limit. The setup: an agent will construct a successor (possibly itself, rewritten) and wants to conclude the successor’s actions will be acceptable. The natural route — “my successor uses proof system T; whatever T proves is true; therefore its actions are safe” — requires the agent to trust T’s soundness. Löb’s theorem blocks it: a system T can prove “if T proves φ then φ” only for φ it can already prove; a sound T cannot prove its own soundness. Hence a proof-based agent cannot license a successor that reasons with the same strength, let alone greater strength. Workarounds each pay a canonical price: descending trust hierarchies (each successor uses a strictly weaker system — improvement chains that provably weaken), Fallenstein’s “parametric polymorphism” (soundness assumed schematically, unverified), consistency waterfalls, and the associated procrastination paradox (an agent that trusts its successor to do a task can defer forever, each stage validly reasoning that the next will act — a formal version, note, of the thesis’s “keep running or halt?” framing). Garrabrant et al.’s logical induction (2016) later gave a beautiful partial answer — a computable process whose beliefs about logic become uncertifiedly-but-boundedly reasonable — at the price of abandoning guarantees for asymptotic non-exploitability. Which is to say: the alignment community’s own mathematics concluded that certified self-improvement is blocked, and uncertified, empirically-calibrated self-modification is what remains. That conclusion, fully absorbed, is the present thesis. It was published on the Alignment Forum’s own turf and has coexisted ever since, largely unreconciled, with the community’s continued use of “recursive self-improvement” as if it named a certifiable process.
Vingean reflection, named
Fallenstein and Soares’s “Vingean Reflection” (2015) named the general problem after Vinge: reasoning reliably about agents smarter than yourself without simulating them (simulation being unavailable by definition — if you could simulate it, it wouldn’t be smarter). The paper is frank that current tools fail. From the thesis’s vantage: of course they fail. “Reliable abstract reasoning about the general behavior of a stronger program” is a demand for decisions about semantic properties of programs one cannot run — Rice’s theorem’s home territory, now with an adversarially large program on the other side.
Part V: The Critics and the Revisionists (2010–2023)
Chalmers, 2010: the philosopher’s audit
David Chalmers’s “The Singularity: A Philosophical Analysis” gave the RSI argument its most careful mainstream-philosophy treatment. His reconstruction hinges on a proportionality thesis — increases in intelligence yield proportionate increases in the capacity to design intelligence — plus absence of defeaters. Chalmers is measured: he finds the explosion plausible but flags that “intelligence” may not be a single measurable quantity, and that self-improvement along one measure need not carry others. That flagged caveat is, in the present framing, the entire issue: the proportionality thesis quantifies over a scalar that (per Legg–Hutter) is incomputable if general and task-relative if computable. Grant the caveat teeth and the syllogism dissolves into the task-specific reading.
Deutsch, Hanson, Mitchell, and the diminishing-returns school
David Deutsch argued (in “The Beginning of Infinity” and after) that knowledge-creation is not brute optimization and that “self-improvement” mistakes the character of explanatory creativity. Melanie Mitchell (”Why AI Is Harder Than We Think,” 2021) catalogued the fallacies of treating intelligence as a ladder of decontextualized capabilities. Economists and forecasters (Hanson throughout; later Davidson’s takeoff model, and epoch-style analyses) reframed everything as returns curves: does a unit of cognitive reinvestment yield more or less than a unit of further capability? Note the shared structure of this whole critical school: it argues RSI will be slow, diffuse, or disappointing. The present thesis is orthogonal and stronger: RSI in its strong, self-certifying, general form is not slow — it is not a well-posed computational task. The returns-curve literature, like the FOOM debate, prices an asset whose existence the limitative theorems deny; the empirical question it studies is real, but it is a question about cascades of task-specific optimization, and should be named as such.
Bostrom, 2014: institutionalizing the ambiguity
Nick Bostrom’s “Superintelligence” carried Good’s argument to policymakers, with “crossover” and “recalcitrance” formalizing takeoff dynamics. For this history, its significance is terminological: Bostrom defined superintelligence as radically outperforming humans “in virtually all domains of interest” — a hedged quantifier (”of interest”) that quietly converts generality into a finite task list. That hedge is the AGI→ASI drift in embryo, eight years early: superintelligence admits a task-relative reading; general intelligence, read honestly, does not. The field, on the present account, has been slowly migrating to the word that survives the mathematics.
Part VI: The Empirical Turn — Self-Improvement in the LLM Era (2022–2026)
The LLM era converted RSI from a thought experiment into an engineering literature. It is essential, for the thesis, to inventory what these systems actually do — because every one of them, examined closely, is a task-scoped optimization loop wearing generality’s clothes, and the practitioners increasingly say so themselves.
Self-training loops.
STaR (2022) has a model generate rationales, keep the ones that reach correct answers, and fine-tune on them; “Self-Instruct,” “Self-Refine,” Constitutional AI’s RLAIF, and “self-rewarding language models” (2024) generalize the pattern: model output, filtered by a criterion, becomes model training data. In every case the criterion — ground-truth answers, a constitution, a judge model — is a fixed, computable proxy. The loop demonstrably improves performance on the proxy’s distribution and is well documented to plateau, drift, or collapse (model autophagy) when the proxy’s coverage runs out. This is precisely the behavior the thesis predicts: proxy optimization is real; its relation to general capability is unmonitored because unmonitorable.
Automated discovery of components.
AlphaEvolve (DeepMind, 2025) evolves code against explicit fitness functions and has produced genuinely new results — faster matrix-multiplication kernels, scheduling improvements, even speedups to Gemini’s own training stack. This is the strongest real evidence of “AI improving AI” to date, and its structure is instructive: an LLM proposes; an automated evaluator with a fixed, machine-checkable metric disposes. AlphaEvolve works exactly where the evaluation problem is decidable by construction — where “better” means a number a harness can compute. The system’s own authors scope it to such domains. It is a Gödel machine with the proof searcher replaced by a benchmark, inheriting the same boundary: it can improve anything except the criterion of improvement.
Self-referential architectures.
The Darwin Gödel Machine (Sakana AI / Jeff Clune’s group, 2025) closes the loop tighter: an agent population rewrites its own scaffolding code, with survivors selected by coding-benchmark scores — Schmidhuber’s design with Darwinian selection substituted for proofs (the substitution the limitative theorems force, made explicit in the system’s own name and paper). Clune has said improving AI with AI is among the hottest topics in the field and that recursively self-improving systems are near — while the same reporting records that each key piece works only moderately well, and that the improvement is benchmark-relative. Meta’s 2025 “self-improving models” statements, OpenAI’s automated-researcher ambitions, and Anthropic’s use of Claude to build Claude’s successors (RSI “in the small”: models writing training code, generating data, evaluating models) all share the shape: human-chosen objective, machine-executed optimization, empirical validation on the objective.
The revisionist vocabulary.
Tellingly, the discourse itself has begun coining the thesis’s distinction. Nathan Lambert (AI2) proposed “lossy self-improvement”: friction and complexity mean the flywheel degrades rather than compounds — the researcher’s job becomes managing complexity, not turning a crank. An AI-Prospects analysis quoted in the practitioner press distinguishes real systemic tool-automation from “the ‘recursive self-improvement’ of AGI mythology, where a monolithic entity modifies itself toward superintelligence.” Dean Ball forecasts automation of “the grunt who grinds through algorithmic efficiency games,” not “the genius.” The empirical field, in other words, is independently converging on the position the limitative theorems mandated in advance: cascades of decidable, task-specific optimization — genuinely powerful, genuinely compounding in places, and never self-certifying at the level of general capability.
Part VII: The Commercial Turn — RSI as a Funding Category (2025–2026)
In roughly eighteen months, “recursive self-improvement” migrated from LessWrong vocabulary to term sheets. The migration is recent enough to date precisely, and the companies’ own language rewards close reading, because each one manages the generality problem differently — none of them solves it, and the most sophisticated ones visibly relocate it.
Anthropic.
Anthropic’s public materials describe a trajectory in which AI systems increasingly build their successors, with an explicit endpoint where systems become capable of “full recursive self-improvement” and progress becomes limited only by compute and algorithmic efficiency discovery. Anthropic’s framing is scenario-planning rather than product promise, and its safety-institute framing treats RSI as a governance trigger. Note the criterion problem’s location here: “full RSI” is defined operationally (AI does the researchers’ jobs), which is a labor-substitution claim — decidable in principle by economic observation — not a self-certification claim. The honest version of Anthropic’s scenario, under the thesis, is: task-specific automation of AI R&D roles, with the generality of the resulting systems asserted by benchmark portfolio, not decided.
Inherent Labs (London; ~$50M seed, Index Ventures and Radical Ventures, May 2026; founders Edward Hughes, Kally Aleksiev, Louis Kirsch, Tantum Collins — DeepMind AI-Scientist and Schmidhuber-lab alumni). Inherent is the most philosophically interesting entrant because it explicitly relocates RSI from the program to the institution. Its manifesto names cultural evolution — “the massively-parallel algorithm responsible for the remarkable success of our species” — as “the archetypal example of recursive self-improvement,” and proposes to “operationalise RSI at the lab level,” recursively self-improving “the entire research organisation” with humans in the loop, asking openly “how can we craft rigorous benchmarks for a system that controls its own reward?” Under the present thesis this move is coherent in a way monolithic RSI is not — an institution containing humans is not a single program deciding a semantic property of itself; the undecidable evaluation is farmed out to human judgment, taste (”AI taste in the sciences”), and social process, exactly where science has always kept it. But by the same token, what Inherent proposes is not machine RSI at all: it is human-adjudicated collective improvement with machine amplification — evolution’s open loop, upgraded, and honestly closer to Hanson’s picture than Yudkowsky’s. The company’s question about benchmarks for a system that controls its own reward is the thesis’s question, asked from inside a pitch.
Core Automation (San Francisco; ~$100M seed reported May 2026 at multi-billion valuation discussions; founders Jerry Tworek (ex-OpenAI VP Research), Rohan Anil (ex-Anthropic/Google), Joanne Jang (ex-OpenAI model behavior & labs), Julia Villagra). Core’s stated program: alternatives to conventional scaling — new learning algorithms, architectures beyond the current stack, continual learning with far less data, and “systems that automate the building process itself.” Jang’s self-description — “trying to automate my work” — is the labor-substitution framing in its purest form. Core is thus RSI as process automation: automate the humans who build AI. The thesis’s reading: this is the coherent core of the commercial RSI thesis (roles and workflows are finite, specifiable, benchmarkable), and it makes no self-certification claim at all. Whether the automated builders’ outputs are generally better remains adjudicated the only way it can be: empirically, afterward, on chosen metrics.
Mirendil (San Francisco; ~$200M seed at ~$1B, a16z and Kleiner Perkins with NVIDIA, mid-2026; founders Behnam Neyshabur (CEO) and Harsh Mehta (CTO), ex-Anthropic and Google). Mirendil’s pitch is “self-accelerating AI research”: frontier systems that excel specifically at AI R&D, offered to universities, pharma, and industrial labs to democratize frontier-model building — AI as research partner for building better AI, in biology, materials, chemistry, robotics. Neyshabur’s public framing centers on what scientists actually do: accumulate deep, sharp domain expertise. This is, note, an anti-generality theory of research skill — expertise as specialization — deployed to sell an RSI-adjacent product. Mirendil also markets against fear (”don’t be afraid of self-improving AI”), positioning openness of access as the safety story. Under the thesis: “self-accelerating AI R&D” is a portfolio of decidable subtasks (experiment design, ablation grinding, architecture search against metrics), and the democratization framing is orthogonal to, and silent on, the evaluation problem.
The wider wave.
The same period saw Recursive Superintelligence (ex-OpenAI/Meta/Google leads; >$650M raised at >$4.6B within months of founding, per press accounts), Microsoft converting the RSI trend into enterprise sales language, and an aggregate exceeding $1.7B committed in six months to systems “built to run autonomously for months or years.” Industry commentary notes the naming politics explicitly: labs avoid calling their automation loops “recursive self-improvement” because the term alarms the safety community — while startups embrace the term precisely because it signals ambition to investors. The term is thus now doing marketing work in both directions, detached from any formal referent.
The AGI→ASI drift, read as an admission.
Across 2025–26 the rhetorical center of gravity moved from “AGI” to “superintelligence”: company charters, product names, and the safety discourse (superalignment; “superintelligent” systems in policy proposals) increasingly reach for ASI while AGI is deprecated as vague, achieved-in-part, or contractual jargon. The conventional explanation is hype inflation. The present thesis offers a sharper one: superintelligence is task-relative and therefore certifiable; generality is not. “Superhuman at X” is a decidable comparison per X — chess, protein folding, IMO problems, kernel optimization — and a portfolio of such X’s can be extended indefinitely without ever posing the undecidable question. “General,” taken seriously (Legg–Hutter seriously), names an incomputable property. A field whose systems improve by benchmark portfolio will inevitably find that its honest vocabulary is the portfolio’s — superhuman here, superhuman there — and that the G was a promissory note no computable process can redeem. Dropping it is not modesty; it is the mathematics surfacing through the marketing.
Part VIII: The Formal Argument
We now state the thesis as mathematics. The strategy: fix definitions that are as generous to RSI as possible; show that under the dominance reading of “general improvement” the target is empty; under the aggregate reading the criterion is incomputable; and under any extensional reading the verification problem is undecidable by Rice’s theorem, with the meta-level task-selection problem inheriting the undecidability. Throughout, “program” means a program for a universal machine U; φ_P denotes the partial function P computes; everything is over a countable input alphabet.
8.1 Definitions
Definition 1 (Task). A task is a pair t = (D_t, s_t) where D_t ⊆ Σ* is a decidable set of instances and s_t : Σ* × Σ* → ℚ≥0 is a total computable scoring function; the score of program Q on instance x ∈ D_t is s_t(x, φ_Q(x)) if φ_Q(x)↓, and 0 otherwise (divergence scores zero). Let 𝒯 be the class of all tasks. This is deliberately liberal: benchmarks, RL environments with computable dynamics and bounded episodes, proof-checking, code-against-test-harness — all are tasks. Note that every task’s evaluation is computable given the program’s outputs: tasks are exactly the objects on which self-optimization is unproblematic, which is why the thesis concedes task-specific improvement without reservation.
Definition 2 (Performance profile). The profile of Q is the map t ↦ Perf(Q, t), where Perf(Q, t) aggregates s_t over D_t in any fixed computable manner when D_t is finite, and denotes the pointwise score function when infinite. Profiles are extensional: they depend only on φ_Q. (This is the substantive modeling choice, defended in 8.6: “capability” is a fact about what a system does, not about its source text.)
Definition 3 (Dominance improvement). P′ ⪰_𝒮 P for a task class 𝒮 ⊆ 𝒯 iff Perf(P′, t) ≥ Perf(P, t) pointwise for every t ∈ 𝒮, with strict inequality for some t. General dominance improvement is the case 𝒮 = 𝒯.
Definition 4 (Aggregate improvement). Given a weighting μ : 𝒯 → ℝ≥0, define V_μ(Q) = Σ_t μ(t)·Perf(Q, t) (where defined). P′ aggregately improves P iff V_μ(P′) > V_μ(P). Call μ universal if μ(t) > 0 for every t ∈ 𝒯 (no computable task is dismissed a priori — the minimal formal content of “general”).
Definition 5 (Improver; the RSI schema). An improver is a computable procedure S which, given a program P, searches for and may output a program P′ together with a claim c ∈ {certified, uncertified}. S realizes strong RSI iff there is an infinite sequence P₀, P₁ = S(P₀), P₂ = S(P₁), … in which each step is a general improvement (Def. 3 or 4 with universal μ) and each certification claim is sound. The thesis is that strong RSI is unrealizable; the theorems below deliver it in pieces.
8.2 Theorem 1 (Verification is undecidable — the Rice argument)
Theorem 1. Fix any program P and any improvement predicate Imp_P(·) that is (i) extensional — whether Imp_P(P′) holds depends only on φ_{P′} — and (ii) non-trivial — some program satisfies it and some does not. Then the set I_P = {P′ : Imp_P(P′)} is undecidable.
Proof. Immediate from Rice’s theorem: I_P is the index set of a non-trivial class of partial computable functions. ∎
Because bare invocations of Rice's theorem can feel like a legal technicality, here is the reduction made concrete, in the form most relevant to self-improvement. Assume additionally the mild condition (iii): there exist programs G and B with Imp_P(G) and such that no program agreeing with B on all but finitely many inputs satisfies Imp_P — i.e., "general improvement" is not a property that finitely many good answers can secure while behavior is bad almost everywhere. (Any predicate violating (iii) does not deserve the name general.) Given an arbitrary machine–input pair (M, w), define P′{M,w}: on input x, simulate M on w for |x| steps; if M has halted within |x| steps, run B(x); otherwise run G(x). P′{M,w} is computable from (M, w). If M never halts on w, then φ_{P′} = φ_G and Imp_P(P′) holds. If M halts on w, then P′ agrees with B on all but finitely many inputs, and Imp_P(P′) fails. A decider for I_P thus decides the halting problem. ∎Two remarks. First, the theorem is criterion-agnostic: it applies to any behavioral definition of improvement whatsoever — accuracy dominance, aggregate utility, “would be endorsed by the current system upon reflection,” anything extensional and non-trivial. The endless definitional debates about what improvement means cannot rescue decidability, because the proof quantifies over the definitions. Second, note the exact realization of the informal claim from which this document began: deciding improvement required deciding whether an embedded arbitrary computation halts. The question “is my rewrite better in general?” contains the question “does this program run forever or halt?” — not metaphorically; by the displayed reduction.
8.3 Theorem 2 (The meta-problem: task selection inherits the undecidability, and the improver cannot decide when to stop)
The thesis’s distinctive claim is about the meta-level: choosing which tasks to optimize such that general capability is assured to improve. Formalize a curriculum policy as a computable map C sending a program P to a task sequence C(P) = (t₁, t₂, …), and suppose an optimizer O provably improves any program along any given task (granted freely — this is the possible part). The meta-claim is that composing decidable pieces cannot manufacture a certified general improver:
Theorem 2. There is no computable pair (C, O) together with a sound certification procedure cert such that for every program P in an effectively rich class, cert affirms that optimizing P along C(P) yields a general improvement. Moreover, any sound certifier is necessarily incomplete on a set of instances it cannot itself delimit, and the certified improver cannot decide, for a given P, whether continued certificate search will ever succeed.
Proof sketch. (a) Soundness of cert means cert-affirmed outputs land in I_P of Theorem 1. The set of pairs (P, P′) with a cert-certificate is computably enumerable (enumerate certificates); by Theorem 1 the improvement relation itself is not decidable, and its complement is not c.e. under the same conditions (the reduction above also reduces non-halting to improvement, so semi-deciding improvement would semi-decide non-halting for the complementary arrangement of G and B). Hence sound certification captures a c.e. proper fragment of a non-c.e.-complemented relation: incompleteness is unavoidable, and the boundary of the captured fragment is itself undecidable. (b) For the stopping claim: “certificate search for (P, S(P)) halts” is an instance of the halting problem for the certifier, and by (a) no computable pre-filter decides it. ∎
Clause (b) deserves emphasis because it vindicates, formally, the “keep running or halt” intuition: a self-improving system engaged in certifying its own next step faces, as a subroutine of self-improvement, an instance of the halting problem about its own certification process — should I continue this search for assurance, or halt and act uncertified? The system cannot decide this; it can only adopt a policy (timeouts, budgets), and the policy’s adequacy is one more undecidable behavioral property. The regress does not converge; it is truncated by fiat, and every actual system truncates it (proof-search budgets in the Gödel machine, evaluation budgets in AlphaEvolve, patience in a research lab).
8.4 Theorem 3 (Under the dominance reading, general improvement does not exist at all)
Theorem 3. Let P compute a total function. If 𝒮 = 𝒯 (all computable tasks), then no P′ with φ_{P′} ≠ φ_P dominates P: for every behaviorally distinct P′ there is a task on which P′ scores strictly worse.
Proof. Define the task t_P = (Σ*, s) with s(x, y) = 1 if y = φ_P(x) and 0 otherwise; s is total computable because φ_P is. P scores 1 on every instance. Any P′ with φ_{P′}(x₀) ≠ φ_P(x₀) (including by divergence) scores 0 at x₀. ∎
This is a no-free-lunch fact, and its consequence for the RSI debate is structural: “general improvement” cannot mean getting better at everything, because every change is a regression somewhere in task space. Improvement is therefore necessarily relative to a weighting that discounts adversarial and pathological tasks — which is to say, Definition 4’s μ is not an optional refinement but the only coherent reading. All the content of “general” now lives in μ. The question “which tasks should I optimize to improve in general?” has become “which μ is the right one?” — and that question, the next theorem shows, has no computable answer with the required properties.
8.5 Theorem 4 (Under the aggregate reading, the criterion is incomputable — the dilemma)
Theorem 4. (a) If μ is universal (positive on all of 𝒯) and effectively summable in the natural way — e.g., the Legg–Hutter weighting μ(t) = 2^{−K(t)} by Kolmogorov complexity — then V_μ is incomputable, and the relation V_μ(P′) > V_μ(P) is undecidable; it is not even semi-decidable. (b) If instead μ is computable with computably convergent support — the only way a machine can actually evaluate it — then V_μ is itself (the aggregate score of) a single task-like objective in the sense of Definition 1’s closure under computable aggregation, and improvement with respect to μ is task-specific optimization by definition.
Proof sketch. (a) K is incomputable; moreover V_μ’s value embeds the scores of P on infinitely many tasks including, for each machine M, tasks whose scores encode whether M halts (build the task that pays on inputs enumerating halting computations); computing or even approximating V_μ to guaranteed precision would decide halting facts. Comparisons inherit this. (b) is definitional: a computable, effectively summable μ collapses the family {(t, μ(t))} into one computable evaluation procedure; optimizing a computable evaluation procedure is precisely what “task-specific” means in this document. ∎
The dilemma, stated plainly. Either your notion of overall capability genuinely covers the open-ended task space — in which case it is incomputable and no system can decide that a self-modification increased it — or your notion of overall capability is computable — in which case it is a (possibly enormous, possibly weighted, possibly evolving-by-computable-rule) benchmark, and improving on it is exactly the task-specific optimization the thesis concedes. There is no third option, because the Church–Turing thesis closes the space of options: whatever evaluation the improving system actually runs is computable, hence horn (b). This dilemma is the formal content of the claim that the meta-decision — which tasks to optimize so that overall rather than specific capability improves — is undecidable. The system can always decide improvement-on-its-measure; it can never decide that its measure measures generality, because “measures generality” is a relation to an incomputable object.
8.6 Theorem 5 (The Löbian ceiling on self-certification) and the extensionality defense
Theorem 5 (Löb; Yudkowsky–Herreshoff form). Let an agent’s acceptance of self-modifications be governed by a consistent, recursively axiomatized theory T extending Peano arithmetic. Then T cannot prove the soundness schema “T ⊢ φ implies φ” for arbitrary φ; consequently the agent cannot accept, on proof-theoretic grounds, a successor whose reasoning it can only vouch for via T’s own soundness — in particular, any successor of equal or greater proof strength. Certified self-improvement chains in fixed T either weaken monotonically (descending trust) or assume soundness unverified. ∎
This closes the last escape route on the certified side: even if the evaluation problem were miraculously decidable, the reasoner’s license to trust its own decision does not extend through self-modification.
Finally, the modeling choice of extensionality (Definition 2) should be defended, since it carries the Rice argument. Could “improvement” be intensional — a decidable syntactic property, like “P′ is P with a verified peephole optimization applied”? Certainly, and such properties are the daily bread of compilers and proof-carrying code. But syntactic improvement predicates are decidable exactly because they entail behavioral claims only over restricted, pre-verified transformation classes — they are certificates for a fragment, in the sense of Theorem 2(a), and the fragment’s reach is fixed in advance by the human-designed transformation library. An RSI worthy of the name must discover new kinds of improvements — new architectures, new algorithms — and for those, only behavior can adjudicate, returning us to extensionality. The move from intensional to extensional evaluation is not a formal convenience; it is what the “R” in RSI means.
Part IX: What Remains Possible, and What the Theorems Do Not Say
Intellectual honesty requires drawing the boundary of the result as carefully as the result itself. The theorems establish claims about decision and certification. They do not establish that self-modifying systems cannot in fact become broadly more capable. Undecidability of a property is compatible with the property holding, and even with a process reliably (in the actuarial sense) producing instances of it. The precise inventory:
What is possible, and conceded.
(1) Task-specific self-optimization, without limit in principle (Blum’s speedup theorem even guarantees inexhaustible headroom for some functions).
(2) Cascades of such optimizations across a growing, human-or-process-chosen task portfolio — which is what every existing “self-improving” system, from STaR to AlphaEvolve to the Darwin Gödel Machine to a frontier lab’s own tooling loop, actually is.
(3) Uncertified, empirically-validated self-modification whose broad effects are estimated by sampling tasks — engineering practice, with Goodhart exposure priced in.
(4) Certified improvement within decidable fragments (verified compiler transformations, proof-carrying rewrites) — powerful and permanently partial.
(5) Institution-level improvement in which the undecidable evaluations are adjudicated by humans, taste, markets, and time — science itself, and the coherent core of Inherent’s proposal.
What is impossible, per Parts VIII’s theorems.
(1) A computable decision procedure for general improvement (Thm 1).
(2) A complete, sound, self-delimiting certification pipeline for it, or a decision procedure for when to stop seeking certification (Thm 2).
(3) General improvement as dominance — the target is empty (Thm 3).
(4) A computable criterion that is general capability, hence any self-assured increase of general capability (Thm 4).
(5) Proof-theoretic self-licensing across equal-strength self-modification (Thm 5). Jointly: strong RSI — the self-certifying, generality-increasing loop of the classical discourse — is not a possible computational process. What exists under the name is the union of items (1)–(5) of the possible list.
The safety reframing.
This is where the thesis pays rent. If capability change in self-modifying systems is real but undecidable-in-general, then:
(i) Evals are the criterion, not a window onto one — a benchmark portfolio is horn (b) of the dilemma, and improvements beyond or against it are structurally invisible; benchmark saturation and contamination are not annoyances but the predicted failure mode of using computable proxies for an incomputable target.
(ii) Capability change cannot be certified in either direction — the same theorems that block “assuredly better” block “assuredly not dangerously better”; sandbagging-detection and capability-elicitation face Rice-class limits in principle, which argues for defense-in-depth over evaluation-gating alone.
(iii) Governance triggers keyed to “recursive self-improvement” are keyed to an undecidable predicate — policy proposals (including the oversight-committee framings that both Anthropic and OpenAI have publicly floated for recursively self-improving AI) would do better keyed to observables: fraction of the R&D loop automated, human review latency, deployment autonomy duration — the labor-substitution quantities that the commercial actors are, revealingly, already using as their operational definitions.
(iv) Fast-takeoff and slow-takeoff both survive, but transformed: the question is no longer whether a system will certify its way up an intelligence ladder (it cannot), but how far uncertified, proxy-driven cascades can compound before proxy and reality decouple — a Goodhart question, empirical and open, and arguably scarier for being uncertifiable.
Part X: Objections and Replies
Objection 1: Humans recursively self-improve; humans might be computational; contradiction.
Reply: humans do not perform strong RSI as defined. No human, and no scientific community, decides that a self-modification generally improves them; they bet, act, and let uncertified experience adjudicate on the tasks life happens to sample. Human self-improvement is item (3)/(5) of the possible list — exactly the uncertified, institution-adjudicated kind. (Per this document’s scoping, the further question of whether human cognition involves non-computational elements is set aside as granted-but-irrelevant: the subject is programs improving programs, and there the arena is closed.)
Objection 2: Relax to probabilistic certification — PAC-style guarantees.
Reply: PAC and statistical-learning guarantees are distribution-relative: with high probability, performance on the training distribution’s kin. The distribution plays the role of μ, and choosing it is Theorem 4’s dilemma untouched. Generality is precisely the regime of distribution shift — performance on tasks not drawn from the certifying distribution — and there the guarantees are silent by construction. Probabilistic relaxation converts “undecidable” into “unmeasurable without a measure choice,” which is the same meta-problem wearing statistics.
Objection 3: Rice’s theorem is worst-case; practice lives in the average case; compilers optimize daily.
Reply: conceded and absorbed — this is the fragment point (§8.6, item 4 of the possible list). Compilers succeed by restricting to transformation classes verified in advance; their reach never includes “new kind of improvement.” The objection correctly describes why task-specific self-optimization thrives, and thereby restates the thesis’s first half. The worst case cannot be waved away at the meta-level, moreover, because a self-improving system’s candidate rewrites are not drawn from a benign fixed distribution: search directed at capability actively generates the semantically novel programs for which the decidable fragments were not pre-verified. RSI is the process that manufactures its own worst cases.
Objection 4: The world is the oracle — deploy, observe, and reality decides.
Reply: reality adjudicates performance on the tasks actually encountered, i.e., an empirical sample from an unknown μ_world; this yields uncertified actuarial improvement (possible-list item 3) and nothing stronger. Moreover feedback arrives after irreversible action, which for a safety analysis is the problem, not the solution. And note the regress: deciding which deployments constitute an adequate test of generality is the task-selection meta-problem again.
Objection 5: Aggregate benchmark portfolios do, in practice, track general capability — look at scaling-era progress.
Reply: this is an empirical bet of exactly the kind the thesis predicts the field must make, and its record is mixed in the predicted direction: portfolio scores climb, portfolio–reality gaps (contamination, saturation, jaggedness of capability profiles, eval-specific training) widen, and each generation of the field’s response is to choose new tasks — the undecidable meta-choice performed by humans, iteratively, uncertified. The practice is reasonable engineering; the thesis’s point is that it is engineering under undecidability, and should be described that way, especially to policymakers.
Objection 6: You’ve defined “general” too hard; “general enough” is what matters.
Reply: perhaps — but then the honest vocabulary is the portfolio’s, and claims should be issued per-domain: superhuman at code optimization, at protein structure, at theorem proving. Which is, this document has argued, exactly the linguistic migration underway: super is measurable, general is not, and the field is dropping the G under the mathematics’ pressure while retaining it in fundraising prose. The thesis does not forbid ambition; it demands truth in labeling — because a “general” that means “our current μ” is a safety-relevant equivocation.
Coda: The Honest Framing
The recursive self-improvement idea began as a syllogism (Good), became a movement (Vinge, Yudkowsky), was formalized into its own refutation by its most rigorous friends (Hutter, Legg, Schmidhuber, the tiling-agents program), was empirically domesticated into task-scoped loops (the LLM era), and has now been securitized into a funding category (Inherent, Core Automation, Mirendil, and the wave around them) — with the market, in its way, converging on the correct definitions: automate roles, optimize measures, let humans and time adjudicate the rest. What was never available at any point in this history, because Rice’s theorem was standing in the road the whole time, was the thing the words seemed to promise: a computational system that decides its way upward in generality, certifying at each step that it has improved überhaupt and not merely on the tasks it chose to look at. Task-specific self-optimization is real, powerful, compounding, and dangerous enough to deserve the field’s full attention. General self-improvement, as a decision an AI could make about itself, was never on the table. The safety conversation — and the vocabulary of the field — should be rebuilt on that distinction.
Sources and Further Reading
Foundations. Gödel (1931), “On Formally Undecidable Propositions”; Turing (1936), “On Computable Numbers”; Kleene (1938) recursion theorem; Rice (1953), “Classes of Recursively Enumerable Sets and Their Decision Problems”; Blum (1967), speedup theorem; von Neumann, Theory of Self-Reproducing Automata (posthumous, 1966).
Intelligence explosion. Good (1965), “Speculations Concerning the First Ultraintelligent Machine”; Vinge (1993), “The Coming Technological Singularity”; Chalmers (2010), “The Singularity: A Philosophical Analysis”; Bostrom (2014), Superintelligence.
Seed AI / LessWrong. Yudkowsky, “Levels of Organization in General Intelligence” (2002), the Sequences (2008), The Hanson–Yudkowsky AI-Foom Debate (2008/2013), “Intelligence Explosion Microeconomics” (2013); LessWrong/Alignment Forum tags on Recursive Self-Improvement and AI Takeoff; Christiano, takeoff-speeds posts (2018); Davidson, compute-centric takeoff model (2023).
Formal limits. Hutter, Universal Artificial Intelligence (2005); Legg & Hutter (2007), “Universal Intelligence”; Legg (2006), “Is There an Elegant Universal Theory of Prediction?”; Schmidhuber (2003–09), Gödel machine papers; Yudkowsky & Herreshoff (2013), “Tiling Agents and the Löbian Obstacle”; Fallenstein & Soares (2015), “Vingean Reflection”; Garrabrant et al. (2016), “Logical Induction”; Löb (1955).
Strange loops and minds. Hofstadter, Gödel, Escher, Bach (1979) and I Am a Strange Loop (2007) — self-reference as generative within formal systems; Penrose, The Emperor’s New Mind (1989) / Shadows of the Mind (1994) and the Lucas–Penrose debate (set aside per scope, with the hypercomputation regress noted: oracle machines face relativized halting problems up the arithmetical hierarchy).
The empirical and commercial era. STaR (Zelikman et al., 2022); Constitutional AI (Anthropic, 2022); self-rewarding LMs (2024); AlphaEvolve (DeepMind, 2025); Darwin Gödel Machine (Sakana AI, 2025); IEEE Spectrum, “Recursive Self-Improvement Edges Closer in AI Labs” (May 2026); Lambert, lossy self-improvement essay; Anthropic Institute, “When AI Builds Itself” (2026); inherentlabs.ai and the Index/Radical announcements (May 2026); Core Automation coverage (Apr–May 2026); Mirendil launch coverage (Mar–Jun 2026, incl. Unite.AI and Gizmodo).


