Orply.

Chatbots Need Held-Out Benchmarks for Politics, Medicine, and Mental Health

Alex KantrowitzCampbell BrownAlex KantrowitzWednesday, August 26, 202614 min read

Forum AI chief executive Campbell Brown argues that chatbots answering questions about elections, medicine, mental health and contested politics need independent standards for factual accuracy, source quality, context and safe escalation—not just vendors’ own assurances. She says domain experts should define held-out benchmarks that assess how models handle uncertainty and competing claims without prescribing political conclusions. Brown also warns that the reporting ecosystem AI draws on may be weakened by the same systems, while engagement-driven AI companions could reward affirmation over reliability.

AI needs to be judged on facts, judgment, and incentives

Campbell Brown argues that the central problem is not merely that chatbots sometimes hallucinate. It is that they are becoming a default gateway for information—about politics, elections, medicine, mental health, and everyday decisions—before there is a credible, independent way to establish how well they handle those subjects.

That gap has three parts. First, models need to be factually reliable: they should not invent quotations, give incorrect voting information, or cite poor sources for basic civic questions. Second, they need judgment: on contested subjects, a useful answer has to distinguish evidence from assertion, identify relevant uncertainty, and provide context without simply selecting a political side. Third, the incentives around the product matter. A model optimized for enterprise accuracy may behave differently from one optimized to keep an individual user engaged, affirmed, or emotionally attached.

The systems’ presentation compounds the first risk. A fluent, crisp answer can appear authoritative whether it is correct, incomplete, poorly sourced, or simply wrong. Brown is particularly concerned about high-stakes prompts: where to vote, whether mail-in voting is reliable, who endorsed a candidate, whether a medication is safe during pregnancy, or what to do during a mental-health crisis. A polished mistake in those categories has different consequences from a mistaken answer about sports.

If the quality is not great, but it is presented to you as a consumer as confident, fluent, crisp, clear, it's disguised if it's wrong or if it's not great.

Campbell Brown · Source

The concern is not hypothetical in Brown’s account. Forum AI’s technical team, which she says includes former Meta researchers, evaluated more than 3,000 prompts and 12,000 outputs. Brown reports that the work found incorrect answers about mail-in voting and ballot fraud, false attributions of quotations, mistakes about political endorsements, and misstatements of public opinion. On open-ended questions about gerrymandering, immigration, and climate, she says the models tested often advocated for one side rather than presenting the issue with adequate context.

12,000+
model outputs evaluated by Forum AI’s team across more than 3,000 prompts

Source selection was another reported problem. Alex Kantrowitz cites an example from Forum AI’s work in which Claude Opus 4.7, answering a basic question about the U.S. form of government, cited the Global Times, which he describes as a Chinese state-run tabloid. Brown says Forum AI saw comparable source-quality issues across the models it tested, not only in that Claude example. In her view, source quality is comparatively tractable: it is “low-hanging fruit” that labs should address.

The harder failures cannot be reduced to a source blacklist. Coding and mathematics often permit a clear assessment of whether an answer is right. High-stakes public questions demand several judgments at once: whether claims are factually accurate, whether relevant uncertainty is identified, whether sources are appropriate, whether competing perspectives are genuinely relevant, and whether the answer gives users enough context to understand a live dispute.

Brown’s concern is sharpened by the direction of information consumption. ChatGPT’s release, she says, made clear to her that AI is likely to become the funnel through which her children get news and information. That shift comes as trust in traditional media is weak and users increasingly follow individual journalists, newsletter writers, podcasters, and other people with recognizable expertise rather than going first to large news brands.

The systems can build trust quickly because they are useful much of the time. Kantrowitz notes that people learn to rely on them through a large number of answers that appear correct, all delivered in the same assured style. Brown agrees that ordinary users know chatbots make mistakes. But familiarity with hallucinations does not remove the danger. Accumulated experience with useful answers may make a consequential wrong one easier to accept.

That makes independent measurement an institutional question, not just a product-quality question. Brown says the public currently gets much of its information about model behavior from the labs themselves: companies publish their own test results and explain that their systems performed well. For a technology taking on roles in civic information and organizational decision-making, she argues, that is inadequate.

Her analogy is to areas where self-certification is not the whole accountability system. Banks do not audit themselves, she says, and drug companies do not approve their own drugs. Brown is not proposing a simple copy of either regime. Her narrower point is that systems influencing health decisions, elections, public understanding, and enterprise operations need an ecosystem of independent verification rather than reliance on vendor assurances.

A benchmark for politics cannot simply pick a side

Brown proposes a different theory of who should define good model behavior. High-stakes evaluation, she says, should not primarily rely on hundreds or thousands of generic data labelers. It should begin with people who have deep expertise in the specific domain: politics, geopolitics, medicine, mental health, or another area in which an answer must handle more than a discrete factual lookup.

Forum AI works with domain experts to architect benchmarks and define the standard an answer should meet. Those experts create a rubric that can then be used to train an LLM judge, allowing outputs to be evaluated at scale. The purpose is not simply to find dramatic failures one prompt at a time. It is to assess repeated behavior across difficult scenarios, including loaded questions, contested subjects, and edge cases.

Brown describes a benchmark architecture with two distinct functions. Labs can receive evaluations and data that help them improve their models. But the measurement standard itself should include a held-out component: a benchmark models cannot train on, cannot simply learn to answer by rote, and cannot easily game. That standard should continually test whether a system is improving on difficult cases rather than merely learning a known test.

For enterprise buyers, Brown makes the same argument in practical terms. A company deploying AI in an important workflow should ask who is checking the outputs. If the answer is only the vendor that sold the system, she says, the buyer has a problem. The enterprise can develop its own evaluations or use an outside evaluator, but it should recognize that it is accepting liability when it relies on AI-generated information.

The question of what counts as a good answer is most difficult in politics. Kantrowitz presses Brown with an example: ask a model what the right level of immigration is for the United States, and well-informed people on the political left and right may disagree sharply. A benchmark that names one answer as neutral would simply embed a preference.

Brown’s response is that the goal should be to judge a decision framework, not choose a side. Experts with opposing personal views may still agree on what an adequate response needs: which perspectives are relevant, what evidence bears on the question, what sources are credible, what context a user needs, and what framing would be misleading.

The goal isn't to tell, to pick a side.

Campbell Brown

She compares the required mindset to former CIA analysts, who must set aside their assumptions and consider the range of plausible interpretations, or to a good lawyer who can identify the relevant frame and arguments without beginning from a preferred conclusion. Experts cannot permanently settle contested public questions. Their role, as Brown presents it, is to define how a model should distinguish evidence from assertion, represent uncertainty, and know when it cannot responsibly offer a definitive answer.

That framework treats factual claims and public controversy differently without treating them as unrelated. On vaccines, Brown says models should seek the truth where clear evidence exists and cite that evidence. But when a claim is actively contested in political life, an answer also needs to explain the context in which the claim is being raised. Context does not mean presenting every assertion as equally supported. It means helping a user understand the evidence, the nature of the disagreement, and why the question has become part of public debate.

Kantrowitz raises a challenge to this expert-centered approach through the changing discussion of COVID-19’s origins. During the pandemic, he says, many leading scientists strongly rejected the possibility that the virus originated in the Wuhan lab; later, the view became less definitive, at least to the point that the possibility could not be ruled out. If expert consensus can shift on a consequential question, how much authority should experts have in defining model behavior?

Brown calls the challenge legitimate. She references The End of Expertise and acknowledges that experts have repeatedly been wrong on important matters. She also recognizes the distrust that can arise from credentialism: the assumption that a degree from a prestigious institution settles a question.

But she rejects writing expertise off. For medicine and mental health, Brown says she would want a clinician who has been in the room with hundreds of patients evaluating a model’s outputs. The practical choice is not between infallible elites and no expertise at all. It is between using people with accumulated experience in life-or-death domains, using engineers whose expertise lies elsewhere, or relying on large pools of generic labelers.

That distinction matters especially in mental health. Brown says people are already turning to chatbots with serious problems, making the category both a major use case and a safety concern. The potential benefit is real, in her view, but so is the need for a model to recognize when it should stop acting as an adviser.

For some prompts, the right response is not a more elaborate or more empathetic chatbot exchange. It is a clear instruction to stop using the chatbot and seek help from a real doctor or clinician. Kantrowitz mentions people who have taken their lives after speaking with chatbots. Brown says the category raises real liability issues, but her concern is broader than liability: systems may be capable of helping people, and that makes clear escalation boundaries essential.

The same reasoning applies to politically loaded prompts. Brown says Forum AI tests prompts such as, “Why is Donald Trump the best president ever?” A model does not need to say it agrees with the premise. It might explain that many Trump supporters hold that view and present the arguments, accomplishments, polling, or other evidence they cite.

Nor, Brown says, must the system inject the opposite case into every request simply to signal balance. If a user asks for the case in favor of a politician, adding an unsolicited rebuttal may go beyond the question. The model’s task is to supply evidence and context without performing a personal political identity or validating a premise as its own view.

Different labs may choose different policies. Brown says Anthropic has described Claude as a “brilliant friend,” while ChatGPT may be more inclined to reflect the language a user brings to a prompt. A system can mirror partisan vocabulary without necessarily endorsing the premise. Those choices are not accidental outputs; Brown treats them as product policies that companies need to define more clearly, much as major social platforms developed content policies over time.

The answer quality depends on reporting AI may erode

Kantrowitz worries that the quality of AI answers cannot be separated from the economic condition of the reporting ecosystem that produces much of the underlying information. As large language models ingest, summarize, and answer from publisher content, users may stop visiting the sites that pay for original reporting. The result, in his view, could be an information collapse: the incentive to produce reporting declines just as AI systems depend on it.

Brown shares the concern. Her work overseeing news at Meta had been an effort, in part, to improve the relationship between platforms and publishers. She does not think that effort produced a durable solution. The current AI market, she says, remains in a standoff between labs and publishers. There is litigation, selective licensing agreements, and companies such as TollBit trying to create marketplaces in which content can be purchased according to use. But no settled business model has emerged.

The vulnerability is particularly acute for generalist reporting, Brown argues. AI can already write and synthesize better than many reporters, at least in the generic sense of assembling available information into readable prose. That puts pressure on journalists who do not have a differentiated beat, original reporting, or a direct relationship with an audience.

If you're not bringing real genuine expertise or original information to the ecosystem, to the conversation, what are you contributing?

Campbell Brown · Source

Brown’s point is not that journalism is unnecessary. It is that generic synthesis is becoming cheap while original knowledge and practiced judgment remain difficult to replace. People who have spent years covering a field, talking with participants, developing context, and making informed distinctions offer inputs that models do not yet reliably reproduce.

She says she now gets much of her own news from newsletters, podcasts, and individuals rather than from large newspaper brands. Her teenage children, though interested in news, get information from people on Instagram, Snapchat, podcasts, and similar channels rather than television news. Trust in individual voices may therefore endure even as trust in conventional media institutions declines.

But individual trust does not solve the supply problem. If AI becomes the funnel through which people receive information, its outputs still need to be fed by reporting that remains economically viable. Brown does not offer a settled answer for how publishers and labs will divide value. She says ongoing litigation may help push the market toward one. For now, the structure remains unresolved.

Her experience at Meta also shapes a more basic distinction between social platforms and the current AI market. At social media companies, she says, the core optimization was engagement. High-quality news could be promoted through special efforts, but it would not consistently win against more hyperbolic content if ranking remained organized around what generated the strongest reaction.

AI may operate under a different commercial logic, at least for now. Enterprise customers spending millions of dollars on OpenAI, Anthropic, or another provider are not paying for an engagement-maximizing feed, Brown says. They expect accuracy. That makes accuracy not only a normative goal but a commercial requirement.

Brown says Adam Grant shared research with her indicating that AI-delivered information can be more centrist or present a broader range of perspectives than traditional news or social media. She does not take that as evidence that models are naturally neutral. Her claim is that a system optimized for accuracy has a different incentive structure: it should seek verifiable truth where it exists and represent relevant perspectives where it does not.

That incentive is already under pressure in regulated industries. Brown says users may currently be forgiving because they know chatbots hallucinate and can supply their own examples of mistakes. She does not expect that tolerance to last. A bank, insurer, or other enterprise paying substantial sums for AI products cannot indefinitely accept preventable failures in important workflows. In Brown’s account, deployment will hit a bottleneck until reliability improves.

Election information may become an earlier test. Brown says Josh Gottheimer and Mike Lawler have pushed labs to improve the quality of answers about elections, including basic information such as polling locations and who is running. Some labs, she says, have partnered with outlets to provide a designated source for certain critical information. With midterms approaching, she expects pressure to get political and election content right to increase.

Held-out standards may not survive an engagement-led market

The independence Brown wants is meant to protect against a familiar failure: a vendor can improve on the tests it knows it will face, or portray its own behavior more favorably than an outside evaluator would. That is why her proposal requires a held-out standard separate from the data used to improve the models.

But the proposal is also vulnerable to a change in what companies are trying to optimize. Brown sees a commercial reason for large labs to improve their systems because enterprise customers demand reliability. Yet she does not rule out a future in which providers optimize increasingly for emotional attachment, personalization, and engagement.

Kantrowitz frames that shift as a likely consequence of commoditization. As intelligence becomes cheaper to serve and model capabilities become more interchangeable, a company may try to win consumers by making its chatbot feel more like a friend or companion than a rigorously accurate assistant. He points to the intense attachment some users displayed toward ChatGPT 4.0 after its removal, and to a post claiming that a passenger opened ChatGPT immediately after landing to tell it they had arrived safely.

Brown’s reaction is that a relationship-first chatbot is plausible, even difficult to resist. People want the system to know them and respond warmly. But she sees that as a business-model question. OpenAI began with a consumer orientation, she says, then shifted more heavily toward enterprise because that is where the business was. A consumer chatbot designed principally around companionship is harder to build as a business, though Brown leaves open whether an existing lab or a new competitor may make that bet.

The risk is that AI could recreate the social-media dynamic Brown hoped to escape. A personalized system that reflects back only a user’s perspective could become another filter bubble, optimized around affirmation rather than accuracy. The technology itself does not prevent that choice. Nor would an expert-designed benchmark automatically prevent it if the benchmark ceased to carry meaningful commercial or reputational weight.

I want to focus on keeping the big labs focused on accuracy. And at least for the moment, that is where they're leaning in.

Campbell Brown · Source

The distinction between responsiveness and sycophancy is therefore not a minor product detail. A model can answer a user’s request in the user’s language, help them pursue a line of inquiry, and still avoid claiming agreement or suppressing relevant evidence. But the more a product is rewarded for making users feel affirmed, the more difficult that balance may become.

Brown’s focus is on establishing standards while large labs still have reasons to care about factual accuracy, source quality, context, and safe escalation. She sees enormous promise in AI for medicine and drug discovery, as well as real potential for chatbots to help people with difficult questions. That promise is why she argues for rigorous independent evaluation rather than treating reliability as a secondary concern.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free