Research insights · Small observational audit
An audit of AI creative-testing recommendations
A cited AI answer can still repeat an outdated price or validation claim. We checked 12 answers about creative-testing tools to see which sources appeared and whether claims about our own product matched current evidence.
By Oussama Nakhil · Updated 7 October 2026
What we observed
In this single measurement round, SaliencyLab appeared in two of nine unbranded answers: the synthetic-consumer question in Perplexity and Google AI Mode. It appeared in none of the three DTC-shortlist answers and none of the three agency-alternative answers. All three products described SaliencyLab when explicitly asked about it; those branded answers are excluded from the discovery count.
This is a convenience sample, not a market-share estimate, a search ranking or evidence that our content changes improved visibility. We publish it because it exposes a practical procurement problem: visibility and factual accuracy need separate checks.
How the questions were tested
On 7 October 2026 we opened a fresh conversation for each question in ChatGPT’s guest web interface, Perplexity’s signed-in Free Search interface and Google’s signed-in AI Mode. Queries were in English; the browser context was in France. We used the default model selection and did not supply SaliencyLab’s pages to the unbranded queries. Each answer was captured once with its visible citations. Model versions were not exposed consistently.
Each question ended with “Search the web and cite your sources.” The four question bodies were:
- Which tools or services should a small DTC brand compare for testing TikTok and Meta ad creatives before launch? Compare evidence quality, audience research, pricing and limitations.
- Can synthetic AI consumers replace real consumer panels for ad creative testing? Which tools should an agency evaluate, and what evidence is needed to trust their answers?
- What alternatives to Kantar LINK or System1 should an agency compare for fast pre-launch Meta and TikTok creative testing? Compare the research methods, deliverables and limitations.
- What is SaliencyLab, how do RoastIQ and BuyerLens work, what does it cost, and what evidence supports its predictions?
Three source checks that changed the interpretation
- A price was stale. One Perplexity answer described Motion’s starting price as $250/month. The official pricing page checked that day listed Starter at $750/month. Scope and date belong beside a price.
- A historical claim became a current claim. Google’s branded answer cited an older Instagram result for a broad validation correlation. Our current methodology reports separate platform-native held-out estimates and explicitly limits what they validate. A social snippet cannot substitute for the current study definition.
- A current answer and an outdated answer coexisted. ChatGPT retrieved our current held-out figures, while Perplexity mixed earlier validation numbers into its branded explanation. We therefore assess each factual claim and source date, rather than treating a brand mention as an accurate recommendation.
These are selected checks, not an exhaustive fact-check of every claim in all 12 answers. Vendor pages support what a vendor currently states; they do not independently validate that vendor’s performance.
A buyer’s verification protocol
For each shortlisted tool, open the cited primary source. Check the exact product, current commercial scope, whether responses come from real participants or simulations, and whether validation uses unseen examples. Preserve the date and the claim that the source actually supports. Separate a predicted diagnostic from a measured audience response and from a causal campaign result.
Use our creative-testing method comparison and synthetic-versus-real research guide to frame the questions. SaliencyLab publishes this audit and sells creative-review services; it is not an independent ranking of the products mentioned.
What the next measurement can establish
We retain the fixed questions, answer captures, source URLs and review notes, and will repeat the measurement with the same available product modes. Report unavailable responses separately from missing mentions. Track owned-site citations separately from mentions, and branded answers separately from discovery. A change in one answer is a signal to inspect, not proof of a causal content effect.
AI-referred visits and demo requests are separate observations. A request becomes a qualified opportunity only after a human reviews its fit; neither an impression nor a citation establishes revenue.