Princeton and UK AI Security Institute test finds AI still can't do autonomous research
Anthropic and OpenAI have promoted their models as able to speed up, and eventually run, AI research on their own. A new experiment from Princeton and the UK AI Security Institute suggests the research judgment behind that claim is not there yet. The team built a method they call Shadow Evaluation. An agent receives the central research question from an unpublished paper, then the paper's original authors, who spent months on it, review the result as conference reviewers would. Because the pape

















