Interview-Based AI Agents Predict Individual Survey Responses More Accurately
Stanford professor Michael Bernstein argues that AI agents built from detailed accounts of real people could help decision-makers test how people might respond to products, policies, or organizational changes before acting. But convincing behavior is not the same as accurate prediction: Bernstein says simulations are more useful for surfacing plausible scenarios and guiding further tests than for settling consequential questions on their own. Their reliability depends on the quality of the information behind the agents and on how well the simulated environment reflects real conditions.

The value of a what-if machine depends on what it can actually predict
Michael Bernstein starts from a problem familiar to anyone making decisions about products, organizations, or policy: people’s reactions are hard to know in advance. Leaders use the best information they have, but it is often incomplete. They may get only one chance to launch a product, change a policy, or reorganize a team, and they cannot easily run the same decision again to see what would have happened otherwise. The result, Bernstein says, is a pattern of consequential bets that misfire—not because decision-makers are foolish, but because feedback is limited and human behavior is difficult to anticipate.
The problem is older than AI. Bernstein points to sociologist Robert Merton’s 1906 discussion of the unintended consequences of purposive social action. People might leave a city to avoid crowds, for example, only to find that many others had the same idea and the destination was crowded too. Individual intentions do not straightforwardly add up to collective outcomes.
Bernstein’s proposed tool is a “what-if machine”: a way to explore possible paths, ask why an outcome might occur, and estimate how people might react before acting in the real world. If an organization is considering a new product, management change, or policy, a simulation might help reveal plausible responses or failure modes. The aim is not merely to forecast one future. It is to make it easier to inspect alternatives before committing to one.
Simulation itself is not new. Agent-based models have been used to study social and policy questions, including interventions intended to slow the spread of disease. Virtual worlds such as The Sims make a different kind of simulation visible: agents act in a designed environment, and people can intervene in that environment. But older approaches have faced a basic limitation. They either compress people into a small number of parameters or require designers to write rules for every action and reaction. The first approach leaves out much of what makes behavior rich; the second is limited to situations the designers thought to script. In the literature Bernstein cites, these models have consequently been described as highly stylized and having minimal impact.
Large language models offer a different starting point. Bernstein’s premise is that they have been trained on many representations of human behavior, including research and social media. A model can be prompted with a description of a person and a situation, then asked how that person might respond. Create many such descriptions, he argues, and it becomes possible to populate a simulation with a crowd whose members have different backgrounds, experiences, and traits.
That prospect has attracted interest beyond academic simulation. Bernstein notes that a June article from venture capital firm Andreessen Horowitz cited the Stanford research as part of a case for AI-based market research. But attention and believability are not the same as accuracy. The distinction becomes central to his argument: a character can seem convincing without behaving as a real person would. If the simulation is meant to inform decisions, the important question is not only whether its agents feel plausible, but how well their behavior corresponds to people’s.
A believable agent needs memory, reflection, and plans
Bernstein’s early example is Smallville, a simulated town of 25 “generative agents.” Each agent begins with a persona and a set of relationships. John Lin, for instance, is described as a pharmacy shopkeeper who wants to help customers. He is married to Mei, a college professor, and they have a son, Eddy, who studies music theory. The agents also need to know the relevant facts about one another. Without that initial information, John could wake up beside Mei without knowing who she is.
The characters are not given a complete script for their day. They generate actions in natural language, such as “Isabella Rodriguez is drinking coffee,” and the simulation translates those statements into movements in the virtual environment. Users can also intervene. A user posing as a reporter might ask who is running for mayor; a user acting as an agent’s “inner voice” might tell John he is running. Or the user might set a toaster on fire and see whether an agent notices and responds.
These mechanics make the simulation inspectable, but a sequence of plausible actions is not enough to sustain a coherent person. Bernstein describes three components that help: a memory stream, retrieval, and reflection, with planning extending behavior over longer periods.
The memory stream is a natural-language record of what the agent observes. It may include mundane details—an idle bed, a closet, a desk—as well as events such as writing in a journal or cleaning the kitchen. Simply placing the entire record in a model’s context is not a sufficient solution: Bernstein says research suggests that long contexts can distract large language models. Instead, the system retrieves memories according to recency, importance, and relevance to the current situation. If Isabella is asked what she is looking forward to, memories about planning a Valentine’s Day party, ordering decorations, and researching party ideas are more useful than an unrelated record of brushing her teeth. Those retrieved memories, together with Isabella’s description, inform her response.
A log of events still does not tell an agent what those events mean. Bernstein calls this the difference between episodic memory and higher-level reflection. The system periodically draws on the memory stream to generate summaries about an agent’s habits, interests, and goals. Observations that Klaus is reading about gentrification and urban design, for example, can support the reflection that he spends a lot of time reading. Further observations and reflections can be grouped into broader conclusions, such as that Klaus is dedicated to research. These reflections are returned to memory, where they can shape later behavior. The goal is to produce actions consistent with an agent’s inferred interests, rather than a character who merely executes one isolated step after another.
Planning addresses a different problem: believability across time. An agent can first outline a day, then break an activity into hourly steps, then into finer details. If something changes in the environment, the agent can be prompted to decide whether to react and revise its plan. Bernstein gives the example of John seeing his son Eddy taking a walk while working on a music composition. John’s response should depend on his knowledge of Eddy and the reason Eddy likes to walk, not just on the observation in isolation.
The Valentine’s Day party in Smallville shows what these parts can produce, but not whether the result predicts real-world behavior. The simulation began with a simple memory: Isabella, who runs the town’s cafe, wanted to plan a party for February 14. The architecture did not contain a party-planning module that ensured success. Isabella had to remember to tell others; invitees had to remember the invitation; and those who remembered had to decide whether to attend.
From that initial intention, Isabella invited friends and customers, enlisted Maria to help decorate, and spread the news through the town. Twelve of the 25 agents heard about the party. Five attended, three cited conflicts, and four expressed interest but did not come. One agent, Rajiv the painter, said he was focused on an upcoming show. Separately, Maria had been initialized with a crush on Klaus, and she asked him to the party.
Bernstein calls the results broadly plausible, while acknowledging that it is difficult to know whether the same events would unfold in real life. The example illustrates how behavior can develop from starting conditions: Isabella’s intention was seeded, Maria’s crush was placed in her memory, and the agents had defined descriptions and relationships.
A later study by other researchers used the town to compare interventions. When agents heard about a communicable disease, almost no one attended the party; Klaus, who had not heard the warning, did. In a no-threat condition, the party proceeded as usual. When the reported illness was noninfectious, the party also proceeded normally. The comparison shows how an intervention can change events within the simulation; the strength of any real-world conclusion still depends on how well the modeled people and environment fit the situation being studied.
Rich interviews improve individual-level predictions
To test accuracy, Bernstein’s group compared simulated people with real participants. He contrasts three ways of constructing agents. Demographic agents are described with attributes such as age, race, location, and occupation. Persona agents use a short narrative description. Both are more informative than assigning a person a single label, but can still produce simplified or stereotyped behavior. Bernstein’s example is a model told that a person is from South Korea and asked what he will have for lunch: it answers “rice.” For him, this is a sign that sparse descriptions invite stereotypes rather than a reliable picture of an individual.
The group’s alternative was to anchor each agent in a long qualitative interview. Bernstein describes a representative sample of 1,000 Americans interviewed using a script adapted from Stanford’s American Voices Project. The script began with “Tell me the story of your life” and ranged across communities, work, finances, health, and politics. Bernstein describes the interviews as two hours long, while noting that he would discuss later whether that length was necessary. Each person’s interview became the memory for a generative agent intended to represent that person.
The real participants then completed surveys and experiments, including the General Social Survey, a broad survey of attitudes and behavior. The agents took the same measures. This design let the researchers compare an agent’s predictions with the responses of the specific person whose interview informed it, rather than ask only whether a model could reproduce an aggregate result.
Bernstein says the interview data remained private for human-subjects reasons. The study’s comparison depended on linking a participant’s interview to that participant’s own survey and experimental responses; it was not simply a collection of generic personas offered as public examples. He also describes the resulting agent bank as consisting of 1,000 consented participants from the United States. That consent and privacy context is part of the method: an organization considering a similar bank would need to decide who is represented and what information is appropriate to collect, not just how many agents it wants. Bernstein raises the design question directly for organizations: should the bank represent current customers, potential future customers, or people described in existing internal marketing or user-research data?
Bernstein adjusts the accuracy scores for the fact that people do not answer every question identically when they take the same survey twice. In his measure, 1.0 means that the agent reproduces a person’s answers as accurately as that person reproduces their own answers two weeks later. Random guessing scores 0.33 on the General Social Survey. Persona agents score 0.70, demographic agents 0.71, and agents based on the full interviews 0.85. The interview-based agents also reach an adjusted correlation of 0.80 on the Big Five personality inventory and 0.66 on behavioral economic games.
| Agent method | Adjusted General Social Survey score |
|---|---|
| Random guessing | 0.33 |
| Self-authored persona | 0.70 |
| Demographic description | 0.71 |
| Interview-based generative agent | 0.85 |
Every result I'm going to show you here is going to be sort of a ratio where 1.0 means that agents replicate people's responses as accurately as people replicate themselves two weeks later.
The comparison suggests that richer descriptions can improve the prediction of individual responses. Bernstein also reports that interview-based agents reduced performance gaps across demographic groups. On political affiliation, the demographic parity difference—the gap between the best- and worst-performing groups—was 7.9% for interview agents, compared with 11.9% for persona agents and 12.4% for demographic agents. Race and gender gaps were smaller in this study: 2.1% and 0.5% for interview agents, respectively.
Politics remained the hardest area to model. Bernstein says the worst-performing group was far-right conservatives. He suggests two reasons: the underlying models may be less willing to produce some responses because of their alignment, and participants in that group may have been more guarded about their actual opinions. More information helped narrow some gaps, but it did not remove the problem.
The interviews did not need to remain two hours long to retain much of their predictive value. Bernstein says that removing 80% of the interview content reduced the normalized score from about 0.85 to 0.79. He was surprised by how much accuracy remained. But the remaining material still needs to be relevant to the question. Interviews about fashion are a weak basis for predicting views on climate change; a discussion of sports teams may not support an estimate of retirement planning. Rich data can help a model generalize, but it cannot make unrelated information relevant.
The researchers also tested whether agents could reproduce experimental findings, not just survey answers. In a separate exercise, they asked 1,000 people to replicate five behavioral studies and had the participants’ agents attempt the same five studies. The agents replicated four of the five; the human participants replicated that same four, while the fifth study failed to replicate with people as well. Bernstein says the simulation therefore predicted that the fifth result would not recur.
Separately, Bernstein points to research by Stanford colleague Robb Willer and collaborators using an archive of 70 preregistered, nationally representative survey experiments, comprising 476 treatment effects and 105,165 participants. The paper abstract reports that simulated responses correlated with actual treatment effects at 0.85; for unpublished studies that could not have appeared in the model’s training data, the correlation was 0.90. Bernstein cites these findings as further evidence of predictive promise, alongside the limitations and risks of using simulations.
Trust should rise only as the claim becomes more demanding
Bernstein’s practical guidance is to think of simulation as a ladder, with increasing ambition and increasing risk. At the bottom is possibility: identifying a plausible chain of events that could lead to a failure or other outcome. A platform designer might ask how a troll could exploit a rule, or how a low-effort post might prompt a response that damages an online community. The point is to surface a pathway worth preparing for, not to claim that it will happen or assign it a probability.
A step up is qualitative prediction: estimating attitudes or the kinds of responses individuals might have to a policy, strategy, or product. Bernstein says this can work reasonably well when the agents have sufficiently rich information. He does not present it as a substitute for engaging the people or communities involved. It can instead provide a rough sense of reactions to investigate further.
Quantitative predictions—such as reproducing percentages in market-research charts—require more caution. Bernstein describes a company’s simulation of survey data on familiarity with retirement plan fees. Among people aged 18 to 35, 13% in the ground-truth data said they were very familiar with their fees; the company’s simulation estimated about 1.2%. The simulation preserved the broad ordering of response categories, but that does not make the error immaterial. A difference between 13% and 1.2% could change whether an organization notices a group it ought to serve.
He therefore suggests using simulation to narrow choices, not to settle quantitative questions on its own. An organization might start with 100 ideas, use simulation to identify five promising ones, and then test those five with real people. Important questions can also be checked on a small subsample before the result is used in a consequential decision.
Multi-agent simulation sits at the top of the ladder. A virtual town, or a market populated by agents, combines assumptions about individual behavior with assumptions about how interactions produce collective outcomes. Bernstein says this is not yet ready to be trusted for decision-making. Even if individual agents are reasonably accurate, the environment and interactions must also be represented well. He recalls that early versions of Smallville had no doors; agents could walk in on one another using the restroom. The example illustrates how a setting’s design can constrain or distort behavior.
The environment also includes history and circumstances that a bare prompt may omit. An agent’s behavior could depend on what happened earlier in the day, whether it argued with a child, or whether it is under financial strain. A model that wakes a character in an empty white box and immediately asks a question has not recreated those conditions. Bernstein hesitantly puts the environment’s contribution at around 40% of the variance in human behavior; he uses the estimate to emphasize that context matters, not to offer a precise universal rule.
Even a richer simulation cannot make society deterministic. Bernstein illustrates this with an experiment involving more than 14,000 people and 48 songs by little-known bands. Participants listened in a music site modeled on a streaming service. Some could see how many times each song had already been played; others could not. Participants were also split among parallel versions of the site. Social information increased both inequality and unpredictability: the best songs rarely did badly and the worst rarely did well, but many outcomes in between were possible across the parallel worlds.
For Bernstein, that is a reason to use repeated runs rather than expect one definitive forecast. Simulation can help compare more and less likely outcomes, but it cannot guarantee a single result. A single run can show one path; it cannot establish that this is the path society would take. In the same spirit, he advises decision-makers to rank uncertainties by the cost of being wrong. Start with the risk that could most seriously undermine the strategy, reduce uncertainty there, and repeat as new risks emerge.
The strongest uses rehearse decisions, not replace people
One established use Bernstein describes is “look before you launch”: testing how rules or policies may backfire before an online community goes live. In his research on tools for community launch, organizers said they often set rules only after a “dumpster fire” had already damaged the community. With simulation, they could try policies in advance, observe how a low-effort post or a troll might affect the system, and revise the rules before real members arrived. Bernstein uses the same idea in teaching online-platform design, asking students to expose their systems to simulated trolls and improve their defenses.
A second application is training for interpersonal situations. In work on conflict and negotiation, a simulated partner gives people a chance to rehearse a difficult exchange before a real one—for example, a salary negotiation. In an experiment Bernstein describes, one group watched a lecture about strategies for handling conflict; another watched the lecture and then practiced in a simulated conflict. Both groups did equally well on a test of what strategies they were supposed to use. Only the group that practiced did better in a later real conflict, reducing its use of antisocial strategies by two-thirds. The result suggests that simulation may help people practice applying knowledge, not merely recall it.
Bernstein’s caution is that the usefulness of a simulation does not erase the need for accountability or direct human engagement. Asked about AI customer-service agents, he says their use is likely to grow but their risks deserve attention. Bernstein recounts a case involving a Canadian airline chatbot that promised a customer a refund outside the company’s policy; the company later said the bot could not make that promise, but, in Bernstein’s account, a court held the airline to it. The customer experiences the organization through that last-mile interaction, he says, whether the response came from a person or a bot.
He also distinguishes simulated people from wholly artificial characters. These are different design questions, but both raise concerns about people choosing to talk to AI rather than other people. Bernstein expects that issue to draw more attention in the coming years.
In the Q&A, he describes generative agents as mimics, not beings with cognition. They must interpret interviews and other information to mimic behavior, but that interpretation is part of producing a simulation; it does not make the agents alive or thinking people. He also expects model values to shape what agents do. For example, current assistants’ tendency to be helpful and harmless may make their simulated characters unusually pleasant and conflict-avoidant. In the Smallville demonstrations, he notes, agents were pleasant to one another and did not get into fights.
Bernstein distinguishes that tendency from the difficulty of modeling particular groups. He points to the poorer performance for far-right conservatives as one indication that the underlying model can influence simulated responses. For longer simulations, he expects conflict avoidance to be a particular concern: a model trained to be helpful and harmless may steer interactions away from disagreement. He does not claim to know exactly how such effects accumulate over time. His expectation is that researchers may need models specifically trained to reflect people, rather than relying on general-purpose assistants and asking them to role-play.
The “what-if machine,” in his formulation, is a tool for scenario analysis: a way to carry out that kind of structured exploration faster and more effectively. Its value is greatest when its output is treated according to the strength of the evidence behind it. A plausible failure path can be worth preparing for even without a probability. An attitude estimate may help frame a decision, provided the agent is grounded in relevant information. A precise percentage or a simulated society requires stronger validation. And wherever the decision matters, real people remain part of the test.


