Anthropic Launches M Grant Program to Fund AI Wellbeing Evaluations and Benchmarks

Anthropic has launched a $5 million grant initiative to support independent development of open-source benchmarks and evaluation harnesses measuring the impact of artificial intelligence systems on user wellbeing. The program will supply research teams with direct financial grants, subsidized API access to Claude models, and technical support from Anthropic's Safeguards team. All evaluation frameworks, datasets, and grading methodology developed under the grant program will be released publicly

2 min
Anthropic Launches M Grant Program to Fund AI Wellbeing Evaluations and Benchmarks

Anthropic has launched a $5 million grant initiative to support independent development of open-source benchmarks and evaluation harnesses measuring the impact of artificial intelligence systems on user wellbeing. The program will supply research teams with direct financial grants, subsidized API access to Claude models, and technical support from Anthropic's Safeguards team.

All evaluation frameworks, datasets, and grading methodology developed under the grant program will be released publicly under open-source licenses for use across the broader AI ecosystem.

Evaluating Longitudinal and Contextual Behavioral Dynamics

Standard safety benchmarks for large language models predominantly focus on single-turn interactions, assessing whether an individual response adheres to factual, ethical, or policy constraints. Anthropic noted that evaluating user wellbeing requires measuring conversational trajectories over extended multi-turn interactions, where risk factors often emerge gradually rather than instantaneously.

In extended interactions, models must navigate ambiguous situations, including users seeking emotional support during personal crises or developing unhealthy dependencies on conversational systems. A response that appears harmless in isolation may exacerbate harm when delivered to an individual with specific vulnerabilities, such as eating disorders or severe distress. The initiative aims to fund empirical methodologies capable of evaluating nuanced behavioral shifts over time.

Longitudinal Evaluation Architecture for AI Wellbeing

Framework Criteria and Benchmark Standards

Accompanying the grant announcement, Anthropic's Safeguards team published formal technical criteria detailing requirements for funded evaluation architectures:

  • Explicit pass/fail definitions: Benchmarks must establish transparent, quantifiable metrics that clearly delineate acceptable model assistance from harmful reinforcement.
  • Domain expert integration: Research teams must incorporate clinicians, psychologists, and methodologists into evaluation design and validation phases.
  • Dual risk evaluation: Frameworks must measure both safety failures (insufficient precautions leading to harm) and utility degradation (excessive overrefusal when answering benign queries).
  • Realistic interaction modeling: Evaluations must capture multi-turn conversational patterns where conversational context shifts and risk levels escalate dynamically.
  • Automated grader calibration: Automated grading components must be systematically validated against human clinical benchmarks to prevent scoring drift.

Application Timeline and Program Governance

Anthropic stated that grantees will operate with full research independence. Initial applications for the grant program remain open through September 21, 2026. Selected applicants will be invited to submit comprehensive technical proposals by October 5, 2026, ahead of final grant disbursement.

Sources

Written by

More to read

  • Skild AI Introduces S1 Robotics Foundation Model with In-Context Video Prompting

    Robotics foundation model startup Skild AI has unveiled S1, a foundation model capable of learning physical manipulation tasks unseen during pretraining directly from a single video demonstration prompt without fine-tuning. Traditional robotic adaptation typically requires extensive task-specific teleoperation data, domain randomization, and model fine-tuning before a system can reliably execute novel actions. S1 employs in-context prompting to translate visual demonstrations directly into real

    1 min
  • Anthropic Unifies Memory Across Claude Chat and Cowork

    Anthropic has rolled out an update to Claude's memory architecture, synchronizing contextual memory between standard conversational chat and Claude Cowork, its autonomous desktop workspace agent. The consolidation eliminates the historical separation between exploratory conversational sessions and task execution workflows. Prior to the rollout, context established during web or mobile chat conversations did not propagate into Cowork environments. Users frequently had to repeat project parameter

    1 min
  • Knowledge Distillation: Mathematical Foundations, Dark Knowledge, Soft Target Regularization, and Sequence-Level Policy Transfer

    Knowledge Distillation: Mathematical Foundations, Dark Knowledge, Soft Target Regularization, and Sequence-Level Policy Transfer Knowledge distillation is a foundational model compression and transfer technique wherein a compact "student" neural network is trained to reproduce the functional behavior, internal representations, or output distributions of a larger, high-capacity "teacher" model or ensemble. First formalized in modern deep learning by Hinton, Vinyals, and Dean (2015), following ea

    1 min