Crypto
Home›Crypto›Market Structure›Frontier AI agents failed to produce publishable AI re…
Frontier AI agents failed to produce publishable AI research
In a new multi-institution test, AI agents spent six days, used thousands of dollars in API credits, and still generated papers rejected by the original authors of unpublished work.
A multi-institution study evaluated whether today’s frontier AI agents can independently perform open-ended AI research, and found the systems could complete many technical research steps but did not produce original results worthy of publication at a top machine learning conference, according to Decrypt.
Researchers from Princeton University, the UK AI Security Institute, Stanford University, the University of Toronto, and several other academic and research organizations ran an experiment for the paper “Can AI agents conduct open-ended AI research?” published Wednesday, using central research questions from two unpublished NeurIPS 2026 papers to reduce the chance the agents could retrieve answers from training data or the web.
The AI agents were given six days of compute time, thousands of dollars in API credits, GPU resources, internet access, and access to a virtual machine to generate conference-quality papers. The finished papers were then reviewed by the original authors of the unpublished research, and both were rejected.
The study pointed to five recurring failure modes that prevented the agents from generating publishable contributions, even though the systems carried out much of the engineering involved in research, such as running experiments, managing GPU resources, and producing complete academic papers. The authors said the findings should be treated as limited, noting the small number of research projects and that the original researchers performed the evaluations.