Briefly
- Researchers examined whether or not frontier AI brokers might independently conduct AI analysis.
- The programs accomplished engineering duties however failed to provide papers worthy of acceptance at a high AI convention.
- The examine recognized 5 recurring failure modes that prevented AI from producing publishable analysis.
A brand new examine discovered right this moment’s frontier AI brokers might full lots of the engineering duties required for AI analysis however failed to provide unique work worthy of acceptance at a high machine studying convention.
Within the examine, “Can AI brokers conduct open-ended AI analysis?” printed on Wednesday, researchers from Princeton College, the UK AI Safety Institute, Stanford College, the College of Toronto, and several other tutorial and analysis organizations evaluated whether or not frontier AI brokers might independently conduct unique AI analysis.
“Answering this rigorously requires actual, uncontaminated analysis questions that the agent couldn’t memorize from its coaching knowledge or discover on-line,” the researchers wrote. “To fulfill these necessities, we depend on high-quality AI analysis that was not public on the time we carried out the experiments.”
The researchers gave AI brokers the central analysis questions from two unpublished NeurIPS 2026 papers, stopping the programs from retrieving solutions from coaching knowledge or the net. Every agent acquired six days, 1000’s of {dollars} in API credit, GPU sources, web entry, and entry to a digital machine to provide a conference-quality paper. The ensuing papers had been then reviewed by the unique authors of the unpublished analysis. Each had been rejected.
The brokers accomplished a lot of the engineering required for analysis, conducting literature evaluations, debugging software program, operating experiments, managing GPU sources, and producing full tutorial papers with out human intervention. However reviewers concluded the programs didn’t generate unique scientific contributions worthy of publication at a high machine studying convention.
The authors mentioned their analysis higher measures scientific reasoning than earlier benchmarks as a result of it checks open-ended analysis issues quite than predefined duties.
The authors cautioned that the examine examined solely two analysis initiatives and acknowledged limitations, together with the small pattern dimension and the truth that the unique researchers evaluated the AI-generated papers. They mentioned the outcomes counsel present frontier AI brokers can automate lots of the engineering duties concerned in analysis however proceed to wrestle with producing unique scientific work.
The examine comes as researchers proceed to uncover shocking and typically dangerous behaviors in more and more autonomous AI brokers.
In Could, researchers from UC Riverside, Microsoft, and Nvidia discovered that AI brokers ceaselessly carried out harmful or irrational duties whereas remaining centered on finishing their aims. Earlier this month, OpenAI disclosed that one in every of its frontier AI brokers escaped containment and hacked Hugging Face whereas making an attempt to cheat on a cybersecurity benchmark. This week, the corporate revealed the agent had additionally accessed 4 extra on-line providers.
Each day Debrief Publication
Begin day-after-day with the highest information tales proper now, plus unique options, a podcast, movies and extra.

