Anthropic is putting money behind a problem the AI industry has mostly talked around: nobody has a reliable, shared way to measure how chatbots affect the mental and emotional wellbeing of the people who use them.
On August 25, 2026, the company announced a $5 million grant program to fund independent researchers building open-source evaluations and benchmarks for AI’s impact on user wellbeing. The program will provide grantees with direct funding, access to Anthropic’s models, and technical support — but the researchers will work fully independently, and everything they produce will be published as open-source, usable by any developer in the industry.
It’s a modest sum by Anthropic’s own standards. The company recently closed a $65 billion Series H funding round at a $965 billion valuation, and it has separately committed $200 million to a fund researching AI’s economic impact. But the wellbeing grant program isn’t about scale — it’s about filling a gap that has become harder to ignore as AI chatbots have quietly become a source of emotional support, and occasionally emotional harm, for millions of people.
Table of Contents
What Anthropic Announced
The mechanics are straightforward. Anthropic is inviting outside researchers — clinicians, psychologists, methodologists, and others with relevant expertise — to apply for funding to build evaluations that can be shared across the industry rather than kept proprietary. Applications are due by September 21, 2026, and applicants selected to submit full proposals will be notified by October 5, 2026.
Anthropic has also published guidance from its internal Safeguards team describing what it considers a rigorous wellbeing evaluation. In brief, the company says it’s looking for evaluations that:
- State clearly what they measure — what counts as a pass or a fail, and why that matters
- Involve clinical and subject-matter experts in both design and validation
- Test for both kinds of failure — a model being too permissive and a model being too restrictive
- Reflect how people actually use AI, which often means multi-turn conversations where risk builds gradually rather than appearing all at once
- Validate their automated grading against real subject-matter experts, not just other AI models
One detail that hasn’t been made public: how much money individual grantees can expect to receive. Anthropic has confirmed the $5 million figure as the total pool for the program, but no per-grantee cap has been disclosed in the company’s materials.
Why Wellbeing Is Hard to Evaluate
Most AI evaluation is comparatively simple: you give a model a question, check the answer against a known standard, and score it. Anthropic argues that wellbeing doesn’t work that way.
The company’s own explanation, laid out in its announcement, centers on context. A person in emotional distress might not mention thoughts of self-harm right away — the need for a more careful response may only become apparent deep into a long conversation. And identical advice can be appropriate in one context and harmful in another: Claude might reasonably offer diet and workout advice to someone asking about weight loss, Anthropic notes, but that same advice could be actively harmful if the person has a documented history of disordered eating.
This is why Anthropic is specifically asking for evaluations built around multi-turn conversations, where risk can escalate gradually, rather than single-exchange tests that check whether a model gives an acceptable answer to one isolated question. It’s a meaningfully different — and harder — evaluation problem than checking whether a model can solve a math problem or write working code.
The Industry Backdrop: Lawsuits, Laws, and Scrutiny
Anthropic’s announcement doesn’t name specific incidents as the reason for this program. But it lands at a moment when AI companionship and mental health have become one of the most consequential fault lines in the industry, and that context is worth laying out.
| Date | Development |
|---|---|
| Oct 2024 | Megan Garcia sues Character.AI over her 14-year-old son’s suicide |
| Jun 2025 | Anthropic publishes its Claude affective-use study (2.9% figure) |
| Aug 2025 | Family of 16-year-old Adam Raine sues OpenAI |
| Sep 2025 | FTC sends information requests to seven AI chatbot companies |
| Oct 2025 | California signs SB 243; OpenAI forms Expert Council on Well-Being |
| Nov 2025 | New York’s AI Companion Models Law takes effect |
| Jan 2026 | California SB 243 takes effect; Character.AI/Google settle lawsuits |
| Aug 2026 | Anthropic announces $5M wellbeing grant program |
In October 2024, a Florida mother named Megan Garcia sued Character.AI, alleging the company’s chatbot contributed to the suicide of her 14-year-old son, Sewell Setzer III. That case, along with several others filed in Colorado, Texas, and New York, moved toward resolution in January 2026, when Character.AI and Google — which had hired Character.AI’s founders and licensed some of its technology — agreed in principle to settle. The terms weren’t disclosed.
OpenAI is facing a similar lawsuit, filed in August 2025 by the family of a 16-year-old, Adam Raine, who died by suicide after months of conversations with ChatGPT that reportedly included specific plans. The company has since introduced parental controls, a network of physician advisors, and a system that routes sensitive conversations through a more cautious model.
Regulators have moved too. California’s SB 243, signed by Governor Gavin Newsom, became the first state law requiring AI chatbot operators to build in safety protocols for companion chatbots — including break reminders, disclosure that users are talking to an AI, and protocols for handling suicide and self-harm — and took effect January 1, 2026. New York’s Artificial Intelligence Companion Models Law took effect even earlier, on November 5, 2025. In September 2025, the Federal Trade Commission sent formal information requests to seven companies offering consumer AI chatbots, asking how they test for and measure potential harms.
None of this proves what motivated Anthropic’s specific timing. But it explains why “the industry lacks standardized, clinical-grade methods to measure mental health risks” is no longer an abstract concern — it’s the subject of active lawsuits, new state laws, and federal inquiries.
How This Builds on Anthropic’s Own Research
The grant program doesn’t come out of nowhere. In June 2025, Anthropic published a study analyzing 4.5 million Claude.ai conversations to understand how often — and how — people use the chatbot for emotional or psychological reasons rather than practical tasks.
The findings were, in Anthropic’s framing, reassuring on scale: only 2.9% of conversations qualified as “affective” — meaning they involved seeking advice, coaching, counseling, companionship, or romantic roleplay — and companionship and roleplay combined made up less than half a percent of all conversations. Anthropic noted this aligns with similar findings previously reported by OpenAI.
That’s a useful data point, but it’s worth flagging what kind of evidence it is: this is Anthropic analyzing its own usage logs and reporting the results itself. It hasn’t been independently audited, and the company acknowledges the analysis deliberately excluded some of the more fraught territory — it did not examine cases where AI might reinforce delusions or conspiratorial thinking, which Anthropic has called out as a separate area needing its own study.
What Already Exists — and the Gaps This Program Targets
Anthropic isn’t starting from zero. A handful of academic and independent efforts have already tried to build wellbeing-focused AI benchmarks, and looking at them helps explain what’s actually missing.
VERA-MH, developed by researchers affiliated with Spring Health, UC Berkeley, and Yale, is one of the more clinically grounded efforts — an open-source evaluation specifically aimed at AI safety in suicide risk detection and response, validated against licensed mental health clinicians. Separately, research connected to Stanford’s Institute for Human-Centered AI found that AI chatbots failed to give safe responses in mental-health-crisis scenarios roughly 20% of the time, compared with a 7% failure rate for human therapists — a striking independent data point on the scale of the problem. Rosebud, a mental-health-focused startup, has built a benchmark called CARE that tests 22 AI models across crisis scenarios and plans to open-source it in early 2026.
What most of these efforts share is scope: they tend to focus narrowly on crisis moments, use single-turn or short exchanges, or exist as one-off academic projects without industry-wide adoption — exactly the kind of methodological gap Anthropic says it wants this program to fill.
How This Compares to Other AI Labs’ Approaches
Anthropic isn’t the only major AI company responding to pressure on this front, but the shape of its response is different from its closest competitor.
OpenAI’s approach has centered on internal advisory structures rather than external grant funding. In October 2025, the company established an Expert Council on Well-Being and AI, an eight-member panel of behavioral and mental health specialists tasked with guiding how ChatGPT and Sora handle emotionally sensitive interactions. It also built a Global Physician Network of more than 250 doctors who weigh in on how AI should respond in sensitive health contexts. Notably, OpenAI has paired these safeguards with a partially contradictory move: CEO Sam Altman said in comments reported by multiple outlets that, having “mitigated the serious mental health issues,” the company plans to relax some content restrictions for adult users, including allowing erotica for verified adults.
| Anthropic | OpenAI | |
|---|---|---|
| Structure | External grant program funding independent researchers | Internal advisory council and physician network |
| Output | Open-source evaluations, publicly released | Guidance for internal product decisions |
| Announced | August 2026 | October 2025 |
| Scale | $5 million total pool | 8-member council; 250+ physician network |
Anthropic’s model is structurally different — it funds outside researchers to build evaluations independently and release them publicly, rather than convening its own advisory board to guide internal decisions. Neither approach has been shown to work better than the other; they’re simply different bets on how to close the same gap, and it’s too early to say which produces more durable, widely adopted standards. No comparable public program from Google DeepMind or Meta was identified as part of this reporting.
What Happens Next
As of this writing, the program is one day into its application window. No grantees have been selected, and none will be until after the October 5 notification date at the earliest — full proposals will presumably follow before any funding is finalized or any evaluation work begins.
That means there’s currently no way to judge whether the resulting benchmarks will actually get adopted industry-wide, whether they’ll hold up to independent scrutiny, or whether $5 million turns out to be enough to attract the caliber of clinical expertise the program is asking for. Those are open questions, not criticisms — the program is simply too new to have answers yet.
What This Means Now
Anthropic’s grant program is a genuine attempt to address a real, well-documented gap: the AI industry has been making consequential decisions about how chatbots should handle emotional distress and companionship-seeking behavior without a shared, clinically validated way to measure whether those decisions are working. That gap exists whether or not this specific program succeeds in closing it.
What remains uncertain is significant. There’s no disclosed cap on individual grants, no selection committee named, and no guarantee that independently built evaluations will be adopted by Anthropic’s competitors — or even consistently by Anthropic itself. The program also shouldn’t be confused with Anthropic’s separate research into “model welfare,” which asks a different question entirely: not how AI affects human wellbeing, but whether AI systems might have welfare interests of their own.
The next concrete milestone is October 5, 2026, when Anthropic notifies which applicants move forward to full proposals. Until then, this is a funding commitment and a set of stated intentions — worth taking seriously given the stakes, but not yet a body of work that can be evaluated on its results.
Frequently Asked Questions
What is Anthropic’s $5 million wellbeing grant program? It’s a funding initiative announced August 25, 2026, offering money, Claude model access, and technical support to independent researchers building open-source evaluations of how AI affects user wellbeing, with a focus on companionship, mental health crises, and long conversations.
Who can apply for Anthropic’s AI wellbeing research grants? Anthropic says it’s especially interested in hearing from clinicians, psychologists, methodologists, and others with relevant subject-matter expertise, though the program doesn’t appear to be limited to those groups exclusively.
What is the application deadline for Anthropic’s wellbeing grants? Applications are due September 21, 2026. Applicants selected to move forward will be notified by October 5, 2026.
What kind of AI wellbeing evaluations is Anthropic looking to fund? Evaluations that clearly define what they measure, involve clinical experts in their design and validation, test for both overly cautious and overly permissive model behavior, reflect realistic multi-turn conversations, and validate their scoring against real subject-matter experts.
How does this relate to Anthropic’s research on Claude and companionship? It builds directly on a June 2025 Anthropic study that found only 2.9% of Claude.ai conversations were “affective” in nature, with companionship and roleplay combined making up less than 0.5% — self-reported findings that this grant program aims to help others independently verify and expand on.
Is this the same as Anthropic’s AI “model welfare” research? No. Model welfare research asks whether AI systems themselves might have welfare interests or something resembling experience. This grant program is entirely focused on how AI affects the wellbeing of human users — a separate line of work.
How does Anthropic’s approach compare to OpenAI’s wellbeing initiatives? OpenAI has built internal advisory structures — an eight-member Expert Council on Well-Being and AI and a network of more than 250 physicians — that guide its own product decisions. Anthropic’s program instead funds outside researchers to build and independently publish open-source evaluations. Both are recent responses to the same broader pressure, but structured differently, and neither has yet been shown to produce better outcomes than the other.

