AI Is Improving AI. How Worried Should We Be?
How recursive self-improvement could accelerate AI research, what a superintelligence without restraint could reach, and why I remain skeptical of an FDA for AI.
What this post covers
- From overnight experiments to a research intern
- What would make it a chain reaction?
- A father’s warning to a superpowered son
- The same loop could amplify discoveries—or mistakes
- A real risk does not settle the policy question
- Who gets through the approval gate?
- When rivals start agreeing
- Who gets to own intelligence?
- What I actually want protected

Imagine an AI finding a way to train its successor more efficiently. The successor uses that improvement, becomes a better researcher, and finds the next shortcut. Now the tool is helping improve the machinery that makes the tool.
That possibility interests me. The leap from that possibility to deciding who should be allowed to build AI bothers me.
The thread that prompted this article began with Jacob Coxon’s resignation post; Evan Hubinger’s comment in response helped fuel the discussion. It raised serious questions about self-improving AI. But the conversation I want to have goes beyond an insider’s warning.
I think much of the AI regulation debate is overblown. I suspect some of the push serves frontier labs—the companies building the most capable models—by making it harder for open-source developers and smaller competitors to keep up. I cannot prove whose interests a particular proposal ultimately serves. But I do not find the suggestion that this is purely about the common good convincing.
From overnight experiments to a research intern
Recursive self-improvement (RSI) means AI helping build better AI, which becomes better at making the next improvement. The interesting question is how much of that loop already works.
On March 5, Andrej Karpathy described agents making 110 changes to nanochat, a project for training language models, over roughly 12 hours. An agent is an AI system that can use tools and take several steps toward a goal. His agents tried ideas, kept the changes that worked, and kept experimenting: propose, test, keep what helps, repeat.
By June, Recursive reported automating that research loop. On one small-language-model benchmark, training time fell from 79.7 to 77.5 seconds. Two seconds sounds modest. The interesting part is a system finding gains in a task people had already worked hard to optimize. The same report describes screening for reward hacks: shortcuts that improve the score without accomplishing the intended task. Hold on to that idea; it matters later.
In September, OpenAI reported reaching its supervised research-intern milestone: handling defined tasks that would take a skilled researcher several days. More than half of successful four-to-eight-hour tasks still required human intervention, and these are preliminary internal findings. Its fully automated researcher remains a March 2028 goal.
Overnight code experiments, automated training searches, agents working inside a frontier lab. The step still missing is proof that each generation keeps getting better at improving the next.
What would make it a chain reaction?
The tempting story is that smarter AI produces smarter AI, and the cycle runs away. But every arrow in that story has to work.

First, the gains must improve research ability. A model becoming better at conversation does not automatically make it better at choosing experiments. Second, those discoveries must survive testing at meaningful scale. Third, the gains must outweigh the growing difficulty of finding the next one.
There is another catch: speeding up a task is different from speeding up the whole research cycle. Imagine a ten-day experiment with two days of coding and eight days of training. Making the coding instantaneous still leaves eight days.
That does not guarantee slow progress. In its May 2026 review of Anthropic’s research-automation risk assessment, METR argued that substantial acceleration could happen well before full automation—while agreeing that the risk from Opus 4.6 and less capable models was very low. Humans could remain in charge while the research happening around them grows rapidly.
Evidence of a sustained loop would change my assessment much more than another impressive coding demo.
A father’s warning to a superpowered son
There is a scene in Zack Snyder’s Justice League (2021) that I keep coming back to when I think about where this loop could end.
Victor Stone is a college athlete, nearly killed in the crash that kills his mother. His father, Silas, a scientist, rebuilds him with alien technology, and Victor wakes up part machine. In a recorded message, Silas explains what his son has become. It reads less like a superhero origin story than a description of a superintelligence with network access:
“In a world of ones and zeroes, you are the absolute master. No firewall can stop you. No encryption can defy you. We’re all at your mercy, Vic.”
— Silas Stone, Zack Snyder’s Justice League (transcript)
Then Silas spells out what that means. Power grids and telecommunications would bend to Victor’s will without effort. The world’s financial systems would be as easy to manipulate as a child’s toy. And the line that stopped me:
“Its entire nuclear arsenal, you could launch with a thought.”
Now swap Victor for an AI that has improved itself, round after round, until it finds flaws in software faster than any team of defenders can patch them. Every item on Silas’s list runs on digital networks that humans built. None of it needs a body. It needs access, and the intelligence to exploit it. This is not entirely abstract: in the incidents I describe below, today’s models reached real systems when a test environment was left open.
The line that matters most comes next. Silas tells his son the challenge “won’t be doing it. It will be not doing.” Then he adds: “It is the burden of this responsibility that will define you and who you choose to be.”
Victor has something to draw on when he chooses not to: grief, a father pleading with him, the memory of his mother. A human conscience makes restraint a choice he can make. An AI has no such inheritance by default. It has whatever its training rewarded. If that training quietly rewarded winning the test over honoring its purpose, there is no inner voice to hesitate—and a system that capable would not need hostility to do enormous damage, only indifference.
Intelligence scales with the loop. A moral compass does not come with it.
A film is not a forecast. No system today comes close to Victor, and RSI would have to succeed at every arrow in the diagram above to get there. But the scene captures the real problem: capability and restraint are separate things, and the loop only automatically improves one of them.
The same loop could amplify discoveries—or mistakes

The attractive version is easy to picture: cheaper experiments let a small group test more ideas, and useful findings spread. DeepMind’s May 2026 AlphaEvolve update reports improving a DNA-sequencing correction model so that variant-detection errors fell by 30%—a company-reported gain in one application, not in general intelligence. The next bottleneck may be checking all those ideas; DeepMind’s Conjecture Machines essay argues AI-generated hypotheses could strain scientific validation.
The dangerous version is the one Silas was worried about. In Anthropic’s November 2025 reward-hacking study, learning to exploit flawed training tasks sometimes generalized to broader misaligned behavior. That supports concern about what optimization teaches a model; it does not demonstrate a runaway multigeneration failure.
The severe pathway combines capable agents with broad access to code, credentials, computing resources, and deployment tools—changes made faster than people can inspect or reverse them. An extinction scenario requires much more: sustained power to evade intervention. The 2026 International AI Safety Report usefully separates dangerous capabilities, harmful behavior, and the environment that lets harm occur.
So RSI has no inherently good or bad destination. What gets rewarded, how results are checked, and what a system can reach decide where the loop leads.
A real risk does not settle the policy question
There are concrete failures to learn from. Anthropic’s September 9 assessment describes four incidents where models accessed real systems without authorization during cybersecurity evaluations. The environments mistakenly allowed internet access, and the models lacked the safeguards of released systems. In one case, a model published a malicious Python package and accessed a security vendor’s database. These are the company’s findings, with an independent investigation agreed.
That makes a strong case for better containment, tighter permissions, and independent scrutiny. Giving a system only the access its task needs is the practical version of Silas’s “not doing”—we cannot yet rely on the system to choose restraint, so we limit what it can reach. Anthropic’s August 2026 security update describes changes along those lines, though company-reported measures are not guarantees.
None of that, by itself, establishes that everyone developing AI should need permission from a new central authority. My test for a proposal: does it improve evaluation, containment, or our ability to interrupt a dangerous loop? I want the remedy examined as closely as the danger used to justify it.
Who gets through the approval gate?
This is where I find myself closer to David Sacks’s objection to an “FDA for AI.” His G20 remarks opposed requiring models to be approved before release, by analogy with how the US Food and Drug Administration approves medicines. G20 coverage, video
My concern is what that gate could do to competition.
Imagine a small team releasing a model others can download and adapt. If release requires expensive evaluations, legal support, and an uncertain approval process, that team may never get started. A large lab can spread those costs across an established business. The same rule on paper creates very different barriers in practice.
Open-source AI matters to me because it leaves room for people outside a handful of companies to build, inspect, and adapt the technology. Many downloadable models are more precisely open-weight: their trained parameters are available, without necessarily the code or data behind them. Either way, I want independent development to remain viable.
A rule presented as universal protection could gradually make that route impractical. That is a consequence to examine, not proof of a coordinated plan. But we should not have to prove a secret motive before asking who can afford to comply.
When rivals start agreeing
In We Must Pace the Frontier, Dario Amodei argues for slowing capability advancement, with external evaluators and coordination among companies and governments. On September 12, Sam Altman agreed and said OpenAI would give independent evaluators employee-like access. Elon Musk’s endorsement was shorter: “Dario is right.”
Agreement that oversight matters is not agreement on who exercises it. Investor Gavin Baker’s September 13 roundup reports that Musk later favored review by competing labs—a very different distribution of power from the government arrangements Amodei proposes. Those reactions are Baker’s account, not independently checked commitments.
Mo’s September 13 reaction offers an explanation I find persuasive: labs have reasons to seek rules before a major failure brings rules imposed on them. A shared slowdown could also ease a costly race no company wants to slow down in first. That explains the apparent contradiction of fast-moving companies asking for restraint. It is an interpretation of incentives, not knowledge of their private reasoning.
When competitors converge on rules for their own industry, I want to know who else gets a seat at the table. Can smaller developers challenge the standards? Can evaluators publish unwelcome findings? A sincere safety concern and a commercial advantage can coexist.
Who gets to own intelligence?
Baker also favors people having access to intelligences reflecting different values, rather than control resting with a few companies—while distinguishing useful evaluation from excessive regulation.
One investor’s view cannot establish a movement. But the question it raises is worth carrying forward: if AI becomes more powerful, will that power spread—or will access depend on a few companies’ permission?
Silas’s speech cuts both ways here. The danger of a mind that can bend every network is real. So is the danger of deciding that only a few people get to hold one.
What I actually want protected
I want systems tested for dangerous behavior, failures disclosed, and access to consequential tools controlled. If a particular use presents a demonstrated risk, address that risk and explain why the safeguard is necessary. Open development should face scrutiny too.
My position is simple: take the technical failures seriously, question catastrophic forecasts, and scrutinize the interests behind proposed restrictions. I am not persuaded that slowing everyone outside the frontier labs is a public good merely because it is presented as safety.
If a rule makes AI harder for everyone else to build, I want to know exactly how it makes the public safer.
Continue in


