AI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.
The acceleration is real: Anthropic says more than 80% of its merged code is now written by Claude, with internal research speedups around 52x. But the jump from that to autonomous research just met its first controlled test. A Princeton and UK AI Security Institute shadow evaluation ran frontier agents on two unpublished NeurIPS submissions. It is early evidence, not a verdict, but the split was clean: they handled every piece of the engineering and could not answer the research question. Anthropic's own post says the same thing: we are not there yet.
"AI is starting to build AI: frontier systems are automating meaningful parts of AI research, with Anthropic delegating a growing share of its own development to AI and OpenAI reporting a model that saved researchers several weeks, feeding the view that autonomous AI research is within reach." [SOURCE ↗]

THE CLAIM. AI is beginning to build AI. Anthropic reports it is delegating a growing share of its own development to AI systems, OpenAI says a model helped post-train a smaller one and saved researchers several weeks, and the wider read is that autonomous AI research is within reach.
THE CHECK. the acceleration is documented and real. What was missing was any controlled test of the leap from acceleration to autonomy. A Princeton and UK AI Security Institute team built one, a shadow evaluation, where a frontier agent takes on the central open question of a high-quality unpublished paper and the paper's original authors grade the output. They ran it on unpublished NeurIPS 2026 submissions with frontier agents given six days and thousands of dollars of compute each. The agents completed all of the engineering without human help, ran the literature searches, debugged the GPU code, compiled full papers, and could not make substantial progress on the actual research questions. Both papers were unambiguously rejected. The authors catalogued five recurring failure modes, starting with poor judgment about what clears the bar for publishable research. Even Anthropic's post agrees on the ceiling: we are not there yet.
What the headlines say
The story of the summer is that AI has started building AI. Anthropic's post 'When AI builds itself' put numbers on it, OpenAI said one of its models helped post-train a smaller one and saved researchers several weeks, and the compressed version traveling the timeline is that autonomous AI research is nearly here. As the new study's own abstract puts it, 'Forecasts of explosive AI progress hinge on AI agents automating AI research.' So the question is not rhetorical. It is the load-bearing assumption under a lot of the hype, and until now it had barely been tested.
The acceleration is real
Start by granting the strong part, because it is true. Anthropic reports that 'As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude,' and that its internal research loop reached roughly 52x speedups on some tasks. That is not a rounding error and it is not marketing fluff. AI genuinely does a large and growing share of the engineering labor inside a frontier lab. If the claim were only 'AI massively accelerates the engineering of research,' this desk would have nothing to autopsy.
The first controlled test
The claim is bigger than that, though, and a team from Princeton and the UK AI Security Institute built the first clean test of the bigger version. They call it a shadow evaluation: an agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors, the people who actually solved it, grade the agent's output. They ran it on two unpublished NeurIPS 2026 submissions, 'giving frontier agents six days and thousands of dollars of compute.' This is not a toy benchmark scraped from public repositories. It is a real open question with a known good answer that the model could not have memorized.
Engineering yes, research no
The result is clean and it splits exactly where the hype blurs. In the authors' words, 'The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions.' Coverage of the study fills in the picture: the agents 'ran literature searches, debugged GPU code, completed hundreds of experiments and robustness tests, and compiled full papers in LaTeX.' Every piece of the machinery worked. And then 'both papers were unambiguously rejected by the authors.' One got a Strong Reject. The team catalogued five recurring failure modes, and the first one is the whole story: poor judgment about the bar for publishable research. The agents could not tell a real contribution from a narrowed restatement of a falsified hypothesis, could not back out of dead ends, and drifted off instructions. They did the work. They could not do the science.
Anthropic already said the quiet part
Here is the part the headlines skip. Anthropic's own post does not actually claim autonomous research. It hedges, plainly: 'We are not there yet, and recursive self-improvement is not inevitable.' The lab shared real acceleration data and an explicit ceiling in the same document. The overclaim was manufactured downstream, in the retelling that kept the 80% and dropped the caveat. The study did not catch Anthropic lying. It caught the ecosystem rounding a careful claim up into an inevitability.
The honest steelman
Two papers is a small sample, the models will improve, and automating the engineering of research is itself a large deal that frees scarce human attention for the judgment calls. All fair. But the burden here runs the other way. The confident version of the claim, that autonomous research is within reach, is the one making forecasts and raising money, and it is the one that just failed its first controlled test on real problems with generous compute. Early evidence is still evidence, and right now the evidence says the hard part is untouched.
What would settle it
Run more shadow evaluations, across more papers and labs, and show frontier agents producing work the original authors would accept rather than reject. Until that exists, the accurate sentence is the boring one: AI does the engineering of AI research well and, on the problems tested so far, the research itself not yet, and the distance between those two facts is the entire debate.
There are two honest sentences here and the headline keeps only the first. One: AI now does the engineering of research at a level that genuinely compresses weeks of grunt work, which is a real and useful capability. Two: on the part that makes research research, forming a judgment about what is worth pursuing and recognizing when a result clears the bar, the same agents produced papers their own authors rejected. Automating the labor is not the same as automating the science, and on the two papers tested the gap was total, not marginal.
The 80% and the 52x will keep getting read as a countdown to AI that does science by itself, and this study is the first hard data point that the countdown is measuring the wrong thing. Expect the engineering-automation half to keep improving fast, because that is what the productivity numbers actually track. Expect the research-judgment half, forming good questions and knowing when an answer is real, to be the slow, sticky part, and expect the next round of self-improvement claims to keep quietly resting on the engineering half while the headline implies the whole.
Flips toward BS if the acceleration claims are shown to be cherry-picked engineering wins with no path to research judgment. Flips toward HOLDS UP if a follow-on shadow evaluation, on a larger set of papers, shows frontier agents producing work the original authors would accept, not just engineering that runs.
RECEIPTS (9) · CONFIDENCE HIGH
every URL below answered a live HTTP check before publish · sweep 2026-08-28
- ▼ arxiv.org ⧉ · "The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions."
- ▼ arxiv.org ⧉ · "As a result, both papers were unambiguously rejected by the authors."
- ● arxiv.org ⧉ · "We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute."
- ● arxiv.org ⧉ · "Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle."
- ● the-decoder.com ⧉ · "The agents managed all engineering work without human help. They ran literature searches, debugged GPU code, completed hundreds of experiments and robustness tests, and compiled full papers in LaTeX."
- ▲ the-decoder.com ⧉ · "OpenAI claimed that GPT-5.6 Sol helped with post-training a smaller model and saved researchers several weeks."
- ▲ anthropic.com ⧉ · "at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work."
- ▲ anthropic.com ⧉ · "As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude."
- ● anthropic.com ⧉ · "We are not there yet, and recursive self-improvement is not inevitable."






