Anthropic announced a $5 million grant program on August 25, 2026, to support independent research into how AI models affect user wellbeing. Applications are due by September 21, 2026, and the program is designed to produce open-source evaluations that developers can use to measure model behavior in real interactions.
Funding comes with access to Anthropic’s tools
Selected researchers will receive direct funding, access to Anthropic’s models and technical support. Anthropic says grant recipients will conduct their work independently before publishing the resulting evaluations as open-source projects. The intended outcome is therefore a shared measurement resource rather than a private assessment limited to Anthropic’s own systems. The program details are set out in Anthropic’s announcement.
That distinction matters for developers and researchers trying to assess wellbeing risks across AI systems. A public evaluation can be reused by other teams, but its usefulness depends on whether its measures are clear enough to interpret outside the project that created them.
The research must test more than obvious failures
Guidance from Anthropic’s Safeguards team asks applicants to define precisely what an evaluation measures, what qualifies as success or failure and why the result matters. The proposed work should also involve clinical and subject-matter experts in both the design and validation of the evaluations.
The program’s scope extends beyond checking whether a model avoids an obviously unsafe response. Anthropic wants researchers to examine protective behavior alongside potential harm, including over-obedience, where a model follows a user too readily, and excessive refusal, where safeguards prevent too much of a requested interaction.
Researchers are also asked to use multi-turn scenarios that reflect how people actually use AI. In these longer conversations, context can change and risk can rise over time. Testing only isolated prompts would miss that progression and could make a system appear safer or more useful than it is in a sustained exchange.
Expert validation is part of the requirement
Applicants must validate their evaluators against real experts in the relevant field. A scoring system may look rigorous while tracking the wrong signal; comparison with domain experts gives an open-source evaluation a stronger basis for judging whether a model’s behavior is safer or more harmful in context.
This requirement also sets a practical limit on what the grants can deliver. Publishing an evaluation makes it broadly available, but openness alone does not establish that its scoring reflects user wellbeing. The measures, success criteria and expert review still determine how much confidence developers can place in the results.
Applications are due September 21
The application deadline is September 21, 2026. Candidates invited to submit a full proposal will be notified on October 5, 2026. Those dates give researchers a defined window to shape projects around the program’s requirements, including independent work, open-source publication, multi-turn testing and expert validation.
For developers, the immediate change is a funded path toward reusable wellbeing evaluations, not an announced change to Anthropic’s models or a new consumer feature. The program could make it easier to compare model behavior using public tools, while leaving researchers to solve the harder question of how wellbeing should be measured across complex conversations.
Anthropic is putting independent, public measurement at the center of this research program. The next concrete milestones are the September 21 application deadline and the October 5 notification date; the broader value of the initiative will depend on whether its evaluations capture real user interactions and hold up under expert scrutiny.
