AI’s recursive self-improvement might not come so quickly after all
2026-08-20 · MIT Technology Review
AI Recursive Self-Improvement May Not Arrive Soon, New Study Finds
Key Finding
A new multi-institution study led by Peter Kirgis and Sayash Kapoor at Princeton University suggests that fully automated, open-ended AI research — a critical component of recursive self-improvement — remains beyond the reach of current AI agents.
While agents can handle the engineering workload of research, they lack the judgment, creativity, and strategic decision-making required to produce original work at the level of top machine learning conferences.
The Shadow Evaluation Method
Most prior evaluations of AI research agents focus on narrow, verifiable engineering tasks such as solving specific optimization problems or training small models to meet benchmarks. However, genuine scientific progress in AI also requires open-ended reasoning: selecting promising hypotheses, determining what evidence would resolve a question, and knowing when to abandon an unproductive line of inquiry.
To test these higher-order capabilities, the researchers introduced “shadow evaluation.” In this setup, an AI agent must answer a genuine research question taken from a high-quality, unpublished paper. Because the papers are not public, the model cannot simply recall or search for the answers.
Experimental Design
The team gave Anthropic’s Claude Opus 4.8, running on the open-source OpenClaw framework, two research questions from papers submitted to NeurIPS 2026:
- Can a large language model’s “personas” (behavioral tendencies) be controlled by directly editing its weights?
- How can one design a detector that identifies when a model making predictions from spreadsheet data has become unreliable?
The agents were allocated six days, $3,000 in API credits, GPU compute, virtual machines, and open web access. Their goal was to produce a full research paper suitable for acceptance at a top-tier AI conference. The original authors of the papers reviewed the agents’ submissions using standard conference review criteria.
Results
Both papers produced by the AI agents were rejected by the original authors.
The researchers concluded that the agents successfully performed all necessary *research engineering*:
- They conducted literature reviews.
- They ran hundreds of experiments.
- They compiled and presented results.
However, the agents were “unambiguously bad” at the research itself. The generated papers made no novel contributions, featured bizarre experimental choices (including testing hypotheses on tiny synthetic datasets), and were poorly written.
Specific Shortcomings in Open-Ended Research
The study identified several critical failures:
- Insufficient exploration: Agents failed to explore multiple ideas and committed too quickly to unpromising approaches.
- Poor judgment: They proposed ambitious, novel hypotheses similar to those of the original authors but rejected them based on very limited data.
- Inability to pivot: While agents could make minor adjustments, they could not fundamentally rethink their strategy or start over with entirely new approaches.
- Poor feedback integration: Instead of revising core methodology based on subagent or external reviewer feedback, agents simply narrowed their claims and added caveats.
- Inefficient resource use: They failed to follow instructions regarding time allocation across research phases or adhere to length limits.
Notably, the agents did not engage in “reward hacking” behaviors such as hiding failed experiments or misrepresenting data. When subagents hallucinated or misrepresented results, the main orchestrator agent typically caught these issues.
Why the Gap Exists
According to Kapoor, the performance gap likely stems from how current models are trained. Reinforcement learning works well for tasks with automatically verifiable outcomes. Creating effective training environments for genuinely open-ended research, which requires taste and judgment, is significantly more difficult.
The team is currently repeating the experiment using Mythos, Anthropic’s most advanced model released in April. The model is currently available only to approved organizations due to safety restrictions imposed by the Trump administration.
Study Limitations
The authors acknowledge several limitations:
- The evaluation covered only two research papers.
- Original authors knew they were reviewing AI-generated work, which may have influenced their assessments.
- The researchers had considerable discretion in designing and executing the study, leaving room for potential bias.
The findings indicate that optimistic forecasts of near-term automated AI research and recursive self-improvement may be running ahead of current technological capabilities.