Why persona prompting is not a population
Asking a model to act like a 34-year-old shopper in Lyon gives you one confident voice. Research needs a distribution. Here is what has to happen in between.
There is a shortcut circulating in research teams right now. You open a chat model, describe a consumer, and ask it what she would think of your packaging. The answer arrives instantly, it is fluent, and it is frequently plausible. It is also a single sample from a model's idea of an average person, and the confidence of the prose hides how little is behind it.
The gap between that and research is not a prompt-engineering problem. It is a population problem.
One voice is not a distribution
A research answer is only useful if it carries variance. You need to know not just that people prefer option B, but by how much, in which segments, and how stable that preference is when the framing shifts. A single generated persona has none of those properties. Run it twice and it may answer differently, for reasons that have nothing to do with the population you care about.
The question is never what a plausible consumer would say. It is how the answers are distributed across the people who actually exist.
Where the structure has to come from
Population statistics give you margins: how many people fall in each age band, each region, each income bracket, each household type. What they rarely give you is the joint distribution, the combinations. Knowing the share of people over 60 and the share of single-person households does not tell you how many people are both.
Recovering those combinations is the part of this work that is statistics rather than language. Agents are generated so that the synthetic set reproduces the known margins and the dependencies between them. That is what makes a population weighted rather than merely populated.
Three properties worth insisting on
- Traceability. Every agent should be attributable to the data that produced it, not to a creative writing step.
- Stability. Rerun the same study and the topline should hold. If it moves, the population layer was doing less work than claimed.
- Inspectability. You should be able to open one respondent and read the reasoning that led to its answer.
Reasoning sits on top, not underneath
Language models do have a role here, and it is a real one. Once an agent exists with a defined profile and context, multi-step reasoning is what turns that profile into a considered response to a specific stimulus. The important distinction is ordering. The population is built first, from data. The reasoning runs on top of it. Reverse those two and you are back to one confident voice wearing a demographic costume.
See this run on one of your own studies
Bring a study you have already fielded. We simulate it blind, then you compare it against what the panel told you.