The pitch is that frontier models can generate genuine research ideas. A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.
Reconstruction, posted to arXiv on August 17, is built to defeat contamination: it withholds the seed paper and all later literature, hands a model only a frozen, anonymized bibliography, and asks it to propose the hypothesis the paper actually made. Across 643 papers in six domains, seven frontier models scored 3 to 15 percent, near the floor. A multi-agent pipeline running hypotheses through cross-model review and a Swiss tournament lifted that to 23 to 42 percent, a real 2.4x gain that still leaves most ideas unrecovered.
"Frontier AI models can do genuine scientific reasoning and generate real research ideas, the capability labs increasingly market as AI accelerating science." [SOURCE ↗]

THE CLAIM. AI is becoming a scientific research partner, able to reason over the literature and generate genuine new ideas. Labs sell this constantly, from AI accelerating research to models proposing hypotheses.
THE CHECK. a benchmark called Reconstruction, posted to arXiv on August 17, tests the cleanest version of that claim. It gives a model only a paper's pre-publication bibliography, with the seed paper and all contemporaneous or future literature withheld and reference IDs anonymized, then asks it to propose the paper's actual core idea, which an independent model judge scores against the held-out truth. The anti-leakage design means a high score cannot come from having memorized the answer. Across 643 papers in six scientific domains, seven frontier models landed at approximately 3 to 15 percent. A reference-only multi-agent pipeline, cross-model review plus a Swiss tournament over competing hypotheses, raised that to roughly 23 to 42 percent, a genuine 2.4x lift that still leaves the majority of ideas unrecovered. Reproducing the paper's specific hypothesis from a reading list is the part the models mostly cannot, with the caveat that a model proposing a different but valid idea would also score zero here, so this measures recovery of a known answer rather than open idea generation.
The pitch, and the test built to check it
The story labs keep telling is that frontier models are turning into research partners: reasoning over the literature, proposing hypotheses, accelerating science. It is a hard claim to test honestly, because most benchmarks are built from public papers whose ideas already sit in the training data, so a high score can be memory dressed as insight. A benchmark called Reconstruction, posted to arXiv on August 17, is built specifically to close that loophole. In its own words, it is 'a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature.' The model gets one thing: the paper's frozen, anonymized reference list. From that alone, it has to propose the hypothesis the paper actually made, and an independent model judge scores the proposal against the held-out real idea.
The result
The number is stark. 'Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%).' Read that against the marketing. On the specific act of recovering the new idea that a real paper contributed, from the same references that paper was built on, the best models the industry ships land it between three and fifteen times in a hundred, with the seed idea absent from what they were given. Coverage of the result put it plainly: the models recover research ideas from bibliographies 'at just three to fifteen percent.' This is not a model failing at arithmetic or clocks. It is a model failing to land the specific idea a real paper contributed, the cleanest available proxy for the research-partner pitch even if it is not the whole of it.
Why it is so low
The reason is in the shape of the task. A bibliography gestures at a problem space, the neighborhood of prior work a paper is standing on. But the creative leap that connects that neighborhood to a specific, non-obvious hypothesis is exactly what the paper itself supplies, and exactly what a reference list withholds. Retrieval and synthesis, the parts models are genuinely good at, get you to the neighborhood. They do not get you to the idea. Reconstruction isolates that gap and finds it wide.
The part that works
Here is the honest other half, because the benchmark reports it too. A multi-agent pipeline that 'combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search' does much better: match rates climb 'to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline.' That is a real, large gain, and it is worth sitting with. It comes not from a smarter base model but from structure, many models proposing, critiquing, and filtering hypotheses in a process that mimics distributed peer review. It is the same lesson showing up across agent research this year: on the hard, open-ended tasks, how you organize the models matters as much as which model you use. And still, the best orchestrated system on this benchmark misses the majority of ideas.
What it means
The defensible read is a split, not a dunk. Frontier models are strong research assistants for the retrieval, synthesis, and drafting around science, and a well-built multi-agent process can more than double their idea-recovery. But recovering the specific hypothesis a paper chose is not solved, and single models sit near the floor when they cannot lean on memorized answers. Read carefully, the benchmark scores proposals against one held-out target, so a model that generates a different but valid hypothesis also scores zero, and the honest claim is about matching the known answer, not a proof that models cannot have ideas. The gap between the 42 percent headline and the 3 to 15 percent baseline is the gap between an orchestration win and a capability that is actually there in the model.
What would settle it
Run the research-generation claims contamination-blind, the way Reconstruction does, and report the single-model number next to the multi-agent one. Until a frontier model recovers ideas at a high rate with the seed literature genuinely withheld, the accurate sentence is the smaller one: AI can survey the field and cannot yet have the idea, and the progress that exists is mostly in how we wire the models together, not in the models themselves.
There are two honest readings and the marketing keeps the wrong one. The models are genuinely useful at the retrieval and synthesis around research, and the multi-agent 2.4x lift shows that structure, not raw scale, buys real gains here. But on recovering the specific hypothesis a given paper landed on, from its reference list alone, single frontier models perform near the floor, and even the best orchestrated pipeline misses most of the time. One caveat the benchmark cannot rule out: a model could propose a genuinely good, original hypothesis that simply is not the one this paper pursued, and still score zero, so the low numbers measure idea-recovery against a fixed target, not open-ended idea generation. Even read that narrower way it punctures the pitch: when a lab says its model can generate research ideas, ask whether it can land the actual idea when the answer is not already in its training data, and the first blind test says mostly not. And keep the evidence in scale: this is one unreviewed preprint, three days old, with no independent replication yet, so treat the exact percentages as first data, not settled fact.
Expect the 42 percent multi-agent number to travel as proof that AI is nearly a research scientist, and the 3 to 15 percent single-model floor to stay in the appendix. The durable finding is the split: retrieval and synthesis are largely solved, genuine idea generation is not, and the gains that do exist come from orchestration and review rather than a smarter base model. Watch for lab demos that quietly let the model see contemporaneous literature, which reintroduces exactly the leakage Reconstruction was built to remove, and treat any research-generation claim that has not been run contamination-blind as unmeasured.
Flips toward HOLDS UP if frontier models reach high single-model match rates on a contamination-blind idea-recovery test, showing genuine hypothesis recovery rather than retrieval. Flips toward BS if the benchmark's LLM judge is shown to be lenient or biased, inflating even the modest scores it now reports. Flips toward our own overclaim if the benchmark is shown to penalize valid non-target hypotheses, meaning it measures idea-matching to a fixed answer rather than idea generation.
RECEIPTS (6) · CONFIDENCE LOW
every URL below answered a live HTTP check before publish · sweep 2026-08-28
- ▼ arxiv.org ⧉ · "Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%)."
- ▲ arxiv.org ⧉ · "Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline."
- ● arxiv.org ⧉ · "We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature"
- ▼ techtimes.com ⧉ · "Solo scores of three to fifteen percent on a contamination-resistant task"
- ● techtimes.com ⧉ · "tests seven frontier models against 643 papers across six scientific domains"
- ● emergentmind.com ⧉ · "combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search"