Most creative testing advice assumes you have already launched. Run an A/B test, wait for significance, pick the winner. That is real testing, but it happens after the budget starts moving, which makes it a way of comparing bets rather than a way of avoiding bad ones.
This is about the other window: the creative exists, nothing has been spent, and changes are still cheap.
Why before launch is a different problem
After launch you have behaviour. Real people, real impressions, real outcomes. The evidence is strong and the cost is that you paid for it.
Before launch you have no audience, so you cannot measure. What you can do is inspect. The question changes from "which variant won" to "is anything structurally wrong with this asset", and structural problems are visible without an audience.
That is a narrower question, but it is the one that determines whether spending money on this creative is a reasonable idea at all.
The four methods, honestly compared
| Method | Cost | Speed | What it gives you | Main weakness |
|---|---|---|---|---|
| Internal review | Free | Minutes | Team judgement and brand fit | Subjective, seniority beats evidence |
| AI creative scoring | Near zero | Under 90 seconds | Dimensional diagnosis, benchmark context | Prediction, not measurement |
| Synthetic audience | Low | Minutes | Articulated objections and buyer reasoning | Simulation, not real people |
| Small paid test | Media spend | Days | Real behavioural signal | Costs money, needs volume |
Most teams use only the first and the last, which leaves the widest gap in the middle: internal review is fast but unreliable, and a paid test is reliable but slow and expensive. The two methods between them exist to make that gap cheaper to cross.
Method 1: structured internal review
Internal review is not worthless, it is just usually run badly. The failure mode is a room reacting to a cut with no defined question, where the most senior opinion wins.
It improves immediately with two rules. Name the decision before anyone watches, and read the creative in a fixed order rather than reacting holistically. The opening carries the most decision risk, so it goes first.
This costs nothing and removes a surprising amount of drift. It still cannot tell you how the creative compares to anything outside the room.
Method 2: AI creative scoring
Scoring evaluates the asset against defined perception dimensions and returns a diagnosis rather than an opinion.
In RoastIQ that is five KPIs, weighted into a composite: Beat the Skip (25%), Get Noticed (20%), Brand Impact (20%), Sell Proposition (20%), and Build Brand (15%). The output is a verdict, Scale, Sharpen, or Rebuild, plus the specific dimension that is weakest.
The value is the specificity. "This ad feels slow" is not an edit instruction. "Sell Proposition is 53, driven by a weak CTA while Persuasion is strong" tells the editor to leave the argument alone and rebuild the ask.
The limitation is that it is a prediction. Held-out out-of-sample validation as of May 2026 shows Spearman correlation of +0.31 against TikTok engagement (n=700), +0.30 against TikTok CTR (n=691, 5-fold cross-validation), and +0.32 against YouTube view counts (n=403, 5-fold cross-validation). Those are moderate relationships against public engagement and click-intent outcomes, not forecasts of sales or ROAS.
Method 3: synthetic audience testing
Where scoring tells you a dimension is weak, a synthetic panel tells you what a buyer might say about it.
In SaliencyLab this is BuyerLens, and it never runs standalone. Every study opens from a RoastIQ result, so the personas already know the KPI scores, the weak dimension, the evidence timeline, and the platform. Without that grounding a synthetic interview produces plausible-sounding text about nothing in particular.
It is simulation, not a consumer panel, and it does not reproduce survey methodology. Treat convergence across personas as a strong prompt to look at the creative again, not as evidence of what the market will do.
Method 4: a small paid test
Sometimes you should just spend a little money. If two creatives both clear the structural bar and the decision is genuinely close, a small live test is the only thing that separates them honestly.
The mistake is reaching for this first. Running a paid test to discover that one variant had a broken opening is paying for information that inspection would have surfaced for free.
The sequence that actually works
- Structured internal review. Name the decision, read the opening first
- Score the creative. Cut anything that returns Rebuild
- Fix the weak dimension the score identified, then rescore to confirm the edit landed
- Run a synthetic panel only if the team still disagrees about why an audience would resist
- Paid test the survivors when two viable options need separating
Each step is cheaper than the one after it, and each removes work the next step would otherwise waste money on. The order is the point.
What to expect from it
Pre-launch testing will not tell you what your campaign will earn. Anything claiming otherwise is describing a product that does not exist.
What it reliably does is stop obviously broken creative from receiving budget, and turn vague dissatisfaction into a specific edit instruction while the file is still open. That is a smaller promise than the category usually makes, and it is the one worth building a workflow around.
If you want the decision-making side rather than the methods, Is Your Ad Ready to Launch? covers how to run the review meeting itself. For how the scoring layers fit together, methodology has the detail.
