A synthetic user is a language model prompted to respond as a defined buyer: a segment description, a context, and a set of questions. Ask it to watch your ad and it will tell you what it understood, what it felt, and what it would do next. The output reads like an interview transcript, and that resemblance is exactly what makes the method both useful and dangerous.
Useful, because the questions a good interviewer asks are questions your team is too close to the work to ask. Dangerous, because fluent, confident, human-sounding text is easy to mistake for evidence about humans. It is not. It is a hypothesis generator that happens to speak in the first person.
This guide covers how to run synthetic users on ad creative specifically, which is a narrower and more tractable job than general synthetic research. It sits inside pre-spend creative intelligence: work you do while the ad is still editable and no budget has moved, so the team can decide what deserves spend. For the wider methodological picture, including where the approach comes from and where researchers disagree about it, start with synthetic users for marketing research.
What synthetic users can actually test on an ad
Be specific about the job, because the method is far better at some questions than others.
Comprehension. Does the ad communicate what you think it communicates? This is where synthetic users are strongest and where teams are most consistently wrong. You know the product, so you cannot un-know it while watching. A synthetic buyer who has never seen it will tell you what a first viewing actually conveys, and the gap is often uncomfortable.
Audience fit. Does the framing land differently for different segments? Running the same creative past a price-sensitive first-time buyer and a category-experienced repeat buyer surfaces where a single message is trying to serve two audiences and serving neither.
Objections. What is the reason not to act? Synthetic buyers are good at articulating hesitation, partly because language models have absorbed an enormous amount of text in which people explain their reluctance.
Clarity of the ask. Is the next step obvious, and does it feel proportionate to what the viewer has been given? A weak call to action often reads as weak in simulation for the same structural reason it reads weak in market.
What they cannot test: how many people will buy, which variant wins, what your CTR will be, or how a real person's attention behaves in a crowded feed. Those are population and behaviour questions, and simulation does not answer them.
The workflow
Six steps. The discipline is in steps two and three; skipping them is what produces the plausible-sounding mush that gives the method a bad name.
Step 1: define the segments before you write a single question.
Two to four segments, each defined by something that would plausibly change how the ad is read: relationship to the category (new versus experienced), the job they are hiring the product for, price sensitivity, or the objection you most suspect. Write each as two or three sentences of real specificity. "Women 25 to 45" is a media-buying target, not a segment definition; it gives the model nothing to reason with.
Include at least one segment you expect to dislike the ad. Teams naturally define personas who would love the work, and a panel of enthusiasts tells you nothing.
Step 2: ask every segment the same questions, in the same order.
This is the step that converts a chat into a study. Fixed questions make responses comparable across segments and across variants, and comparability is the entire analytical value. A rough default set:
- What is this ad for? (comprehension)
- Who is it for? (targeting fit)
- What is it asking you to do? (clarity of the ask)
- What would make you hesitate? (objections)
- What would you do next, if anything? (intent)
Ask them in that order every time. Do not improvise follow-ups for one segment and not another; you will read the extra detail as a difference between segments when it is a difference in your prompting.
Step 3: separate comprehension from emotion from action.
These three fail independently and require different fixes, so keep them apart in your analysis.
A viewer can understand the ad perfectly and feel nothing (comprehension fine, emotion flat: the problem is the idea, not the execution). They can be moved but not know what to do (emotion fine, action unclear: fix the ask, leave the story alone). They can misunderstand the offer entirely and still say something warm about the brand (a comprehension failure hiding behind a positive tone, and the most commonly misread output of all).
If you collapse these into an overall sentiment read, you lose the diagnosis and keep only the mood.
Step 4: compare variants, do not score them.
The method is more reliable at relative comparison than at absolute judgement. Run two or three cuts through the identical question set and read across, not down.
What you are looking for is divergence: the segment where variant B's comprehension is clean and variant A's is muddled, the objection that appears for one cut and vanishes for another. Consistent objections across every variant usually point at the proposition rather than the execution, which is a more expensive finding and a more valuable one.
Step 5: separate the objections worth acting on.
You will get more objections than you can address. Sort them with two questions: does this objection recur across segments, and is it about something you can change?
Recurring, changeable objections are your work list. Recurring but unchangeable objections (price, category skepticism, a real product limitation) belong in the brief as things the creative must pre-empt rather than problems the creative can fix. One-off objections from a single segment are noise until something else corroborates them.
Step 6: verify against real outcomes.
This is the step almost everyone skips, and it is what separates teams who get value from the method from teams who slowly drift into fiction.
Keep a running note of the objections synthetic users raised and what happened when the creative ran. Over a handful of campaigns you will learn which categories of synthetic finding hold up for your audience and which are systematically off. That calibration is the actual asset, and it is specific to you: nobody can hand it to you, and no vendor's validation substitutes for it.
A worked example
An agency has a new concept for a client's supplement brand: a fifteen-second cut opening on a lifestyle scene, product visible at second nine, offer in the final frame. The strategist likes it. The account lead is nervous. Nobody can articulate why.
They define three segments: a first-time buyer skeptical of supplement claims, an experienced buyer who already uses a competitor, and a lapsed customer who tried the category and stopped. Same five questions to all three.
The results diverge in a useful way. All three segments answer question one differently, which means the ad is not communicating a single clear idea. The skeptic's objection is about substantiation. The competitor user's objection is that nothing distinguishes this from what they already use. The lapsed buyer's response is the sharpest: they understood the ad fine and felt no reason to reconsider.
None of that is evidence. It is three hypotheses, and they point in one direction: the ad has no differentiating claim, and the product arrives too late to carry one. The agency re-cuts with the differentiator in the first three seconds, and the client is now arguing about the right thing.
Note what the exercise did not do. It did not predict performance, rank the concepts, or tell them the campaign would work. It turned "something feels off" into a specific, testable claim about the creative in under an hour.
What good output looks like
Judge synthetic output on three properties.
Specific to your creative. Responses should reference what is actually in the ad: the sequence, the claim, the offer. Output that would apply equally to any ad in the category means the model is generating category-plausible text rather than reacting to your asset.
Divergent across segments. If your segments produce near-identical answers, either the definitions are too similar or the model is not using them. Genuine divergence is the signal that the segmentation is doing work.
Falsifiable. A useful objection can be checked. "I would not trust this claim without proof" points at something. "The ad feels a bit generic" does not.
Grounding is what produces those properties, and it is why the strongest implementations do not run synthetic interviews on a blank slate. In SaliencyLab, BuyerLens opens only from a scored RoastIQ result, so its 36 buyer personas already have the creative's KPI scores, the weak dimension, and the platform context before the interview starts. Anchoring to a scored artifact is what keeps the conversation about your ad rather than about ads in general. Deciding when that second layer is worth running is its own question, covered in when to use BuyerLens after a score.
Limitations, stated plainly
Every honest account of this method includes this section. Here is ours.
It is simulation, not sampling. Synthetic responses do not represent a population. There is no sampling frame, no margin of error, and no basis for saying "62% of buyers would object". Any output presented as a percentage of a market is a fabrication with a number attached.
It inherits the model's biases. Responses reflect patterns in training data, which over-represents some populations, contexts, and languages and under-represents others. Segments furthest from that distribution get the least reliable simulation, which often means exactly the audiences you most need to understand.
It leans agreeable. Language models tend toward plausible, cooperative answers. Ask whether an ad is compelling and you will frequently be told it is. Prompting for objections directly, and defining at least one skeptical segment, counteracts some of this. It does not eliminate it.
It is weakest where novelty is highest. For a genuinely new product or an unfamiliar category, the model has little to reason from and will confidently interpolate. Novel propositions are where synthetic output should be trusted least and where teams are most tempted to lean on it.
It cannot observe behaviour. Stated intent and actual behaviour diverge in real consumer research, and simulation only ever produces the stated kind. Nothing here forecasts what anyone will do in a feed.
Fluency is not accuracy. The output reads like an interview. It is not one. This is the failure mode that matters most, because everything else on this list is easier to remember than the fact that confident prose is the method's default output regardless of whether it is right.
Where this fits with everything else
Synthetic users are one layer of a pre-launch process, not the process. The sequence that works: review structurally, score to catch structural problems, run synthetic buyers when the team cannot agree on why an audience would resist, then put real budget behind the survivors and let a live test decide. Each step is cheaper than the next and removes work the next would waste money on. The full pre-launch workflow covers the whole sequence.
It is worth being clear about the evidence boundary across that whole stack. Model-scored predictions in this category, ours included, correlate moderately with public engagement and click-intent outcomes: held-out, out-of-sample Spearman correlation of +0.30 to +0.32 as of May 2026, with sample sizes and method on the methodology page. Those figures describe scoring, not synthetic interviews, which have no equivalent validation and should not be presented as if they do. Neither predicts sales, ROAS, or brand recall.
Facts, in this method, are limited to what is in your creative. Predictions come from validated models and carry known, moderate error. Everything a synthetic buyer says is a hypothesis. Teams that keep those three apart get real value from synthetic users. Teams that blur them end up making budget decisions on the basis of confident fiction.
Frequently asked questions
What are synthetic users in ad testing? Synthetic users are AI-simulated buyer personas that respond to ad creative in a structured interview format, describing what they understood, how they reacted, and what would make them hesitate. They are used before launch to surface comprehension gaps and objections while the creative is still editable.
Are synthetic users accurate? The question is misframed, because accuracy implies a measurement. Synthetic users generate hypotheses about how a segment might interpret an ad. They can be useful and directionally right, and they have no sampling validity, so no percentage of a market can be inferred from them. Judge them on whether their objections hold up against real outcomes over time, not on a claimed accuracy figure.
Can synthetic users replace real consumer research? No. They are best used before real research, to sharpen what you would put in front of people, and alongside it, never instead of it. Anything where representativeness matters, such as claims substantiation, sizing, or a novel category, needs real respondents.
How many synthetic personas do I need? For ad testing, two to four well-defined segments produce more usable signal than dozens of thin ones. Specificity beats volume: a larger panel of vaguely defined personas mostly produces variations on the same answer.
How do I stop synthetic users from just agreeing with me? Define at least one segment predisposed to object, ask directly what would make them hesitate rather than whether they like the ad, and compare variants against each other instead of asking for a verdict on one. Agreeableness is a known tendency of language models, and prompt structure is the main defence.
What is the difference between synthetic users and predictive creative scoring? Scoring produces a structured, validated estimate of how a creative is likely to perform on defined dimensions. Synthetic users produce reasoning about why an audience might respond a certain way. Scoring tells you a dimension is weak; synthetic buyers suggest what a viewer might say about it. They answer different questions and the second is most useful after the first.
Should synthetic user findings go in a client deck? Only if labelled honestly. Presented as "simulated buyer reactions we plan to verify", they are a legitimate and useful input. Presented as research findings or with percentages attached, they misrepresent the method, and that misrepresentation tends to surface at the worst possible moment.
