A 150M model hits 29.5% on ARC-AGI-1 for $0.0007 a task
Pathway's 150M-parameter BDH-CQ scores 29.5% on ARC-AGI-1 at $0.0007 a task, about 57x cheaper than GPT-5.6 Luna (Low).
Articles
Pathway's 150M-parameter BDH-CQ scores 29.5% on ARC-AGI-1 at $0.0007 a task, about 57x cheaper than GPT-5.6 Luna (Low).
Anthropic and OpenAI have promoted their models as able to speed up, and eventually run, AI research on their own. A new experiment from Princeton and the UK AI Security Institute suggests the research judgment behind that claim is not there yet. The team built a method they call Shadow Evaluation. An agent receives the central research question from an unpublished paper, then the paper's original authors, who spent months on it, review the result as conference reviewers would. Because the pape
A neurosurgery resident with no formal advanced mathematics training has solved Crouzeix's conjecture, a 22-year-old open problem in numerical linear algebra, using a 16-hour autonomous run of GPT-5.6 Sol in ChatGPT Work mode. Dr. Shanmu Jin, a postdoctoral researcher and neurosurgery resident at Peking Union Medical College Hospital, prompted ChatGPT 5.6 on July 30, 2026 to work on the conjecture. The model ran autonomously for approximately 16 hours, exploring proof strategies through a branc
Every model on the leaderboard is within a few points of every other. Here's what that does — and doesn't — mean for which one you should ship.
Years of patient mech-interp work just turned a corner. The result is the most concrete reason for optimism in alignment research in years.