Skip to main content
RoastIQBuyerLensHugoPricingBlogAbout
Book a demoSign inStart free →
Back to blog
JournalAugust 1, 2026 · 7 min read

How to Reduce Wasted Ad Spend with Pre-Spend Creative Testing

Oussama Nakhil

Written by

Oussama Nakhil

Founder & CEO

Founder at SaliencyLab · Previously L'Oreal & NielsenIQ

Most creative-driven waste sits in the window between spend starting and reliable data arriving. Screening before launch removes the losses that were visible in the asset all along.

How to Reduce Wasted Ad Spend with Pre-Spend Creative Testing

In short

Most wasted ad spend is not a targeting problem. It is a creative problem discovered after the money has moved, in the window between spend starting and the first reliable performance data. Screening creative before launch removes the structurally broken work while it is still free to change, which makes it the cheapest lever in ad spend optimization.

Every performance team has a version of the same story. A campaign goes live, spends for a week, and the numbers come back soft. The post-mortem finds the hook was weak or the product reason arrived too late. Everyone agrees, the creative gets recut, and nobody mentions that the diagnosis was available before the money moved.

That gap is where most creative-driven waste sits, and it is what this article is about: where the waste concentrates, which failures were actually predictable, what screening is worth in money terms, and how to make that case to whoever owns the budget. The category this belongs to has a name: pre-spend creative intelligence is the structured, model-based evaluation of an ad before budget is committed, built for marketers and agency teams deciding which creative deserves spend. The full definition and how it differs from A/B testing and panels is covered separately; this piece stays on the money.

Where the waste concentrates

Wasted spend gets blamed on targeting, bidding, or platform algorithms. Those matter. But they operate on whatever creative you give them, and a strong algorithm distributing weak creative distributes weak creative efficiently.

The expensive window is specific. It opens when budget starts flowing and closes when you have enough data to act on. Depending on volume, that is a few days to a few weeks, and during it three things are true at once:

  1. The money is being spent
  2. The creative is locked and in market
  3. You do not yet have reliable signal about which creative is the problem

Everything spent in that window on a structurally broken asset is avoidable waste, because the structural problems were visible in the file before launch.

The window also scales the wrong way. Small accounts sit inside it for weeks because volume is low and statistical confidence arrives slowly. Large accounts exit it faster but burn more per day while inside. Either way, the window is priced in media, and the diagnosis it eventually delivers was available for close to nothing.

Which failures are structural, and which are genuinely unpredictable

Not every failure is predictable, and pretending otherwise is how this category oversells itself. Some creative underperforms for reasons no pre-launch process could catch:

  • The audience did not want the offer at that price
  • A competitor undercut you mid-flight
  • A seasonal or cultural shift moved attention elsewhere
  • The landing page, not the ad, broke the funnel

Those are market discoveries. You pay for them because there is no other way to learn them, and that spend is not waste. It is the cost of information you could not have had earlier.

But a large share of creative failure is structural, and structural problems are properties of the asset itself:

  • The opening does not earn attention, so nothing after it gets seen
  • The brand arrives after the viewer has already decided to scroll
  • The message stack asks for too much at once and lands none of it
  • The product reason appears past the point where anyone is still watching

None of those require an audience to detect. They are inspectable before launch. Spending a week of budget to discover them is paying market prices for information that inspection would have surfaced for free.

The practical discipline is sorting your recent failures into these two buckets. In our experience most teams who run this exercise find the structural bucket is the larger one, which is worth verifying against your own post-mortems rather than taking on faith. Only the structural bucket is addressable before spend, but that is precisely why it is the one worth attacking first.

The economics: screening vs discovering in-market

Put the two ways of learning that a creative is broken side by side.

Discovering in-market costs the media burned during the detection window, plus the opportunity cost of the stronger creative that did not run, plus the production cost of the recut, plus the calendar time lost. On even a modest campaign, that is thousands of euros per broken asset, and the bill repeats every time an unscreened asset launches.

Screening before launch costs minutes per asset and close to nothing in cash. In RoastIQ that means uploading the creative and getting a 5-KPI diagnostic in about 90 seconds, with a verdict of Scale, Sharpen, or Rebuild. A Rebuild verdict, meaning a composite below 55 or two or more KPIs below 45, is the cheapest possible signal that an asset should not receive budget in its current state. The verdict logic also catches a case teams miss: Scale requires a composite of 70 or above and no individual KPI below 55, so a creative can average well and still be held back by one structural hole.

The asymmetry is the whole argument. Screening does not need to pick winners to pay for itself. It only needs to stop a fraction of the obvious losers from receiving budget, because each one it stops saves an entire detection window of spend. A filter with moderate accuracy applied to a cheap decision point beats a precise measurement applied after the money has moved.

There is a second-order saving too. Live A/B testing is expensive because half the spend funds the loser. When only screened, structurally sound variants reach the live test, the test resolves faster and the losing arm is a genuine near-miss rather than a broken asset that inspection would have caught. Screening does not replace live testing; it makes live testing worth its price. The scoring vs A/B testing comparison covers that division of labour in detail, and the full pre-launch workflow shows where screening sits in the end-to-end process.

What the evidence supports, and what it does not

Being direct about the boundary, because this is where the category tends to oversell.

The scores are model predictions validated against public engagement and click-intent outcomes. Held-out, out-of-sample validation as of May 2026: Spearman correlation of +0.31 against TikTok engagement (n=700), +0.30 against TikTok CTR (n=691, 5-fold cross-validation), and +0.32 against YouTube view counts (n=403, 5-fold cross-validation), across a benchmark pool of more than 1,200 ads with public outcome data. Top-vs-bottom quintile lift by predicted score is 6.5x. The scores do not predict sales, ROAS, attributed conversion, or brand recall, and no budget model should be built on the assumption that they do. The full validation approach is documented in the methodology.

What that means in practice:

  • Fact: the scores rank creative by predicted engagement and click intent with moderate, measured correlation. Moderate is the honest word.
  • Recommendation: use them as a screen, to stop funding structurally broken work, not as a forecast of campaign results.
  • Hypothesis: your category may show stronger or weaker fit than the pool average, which is why feeding live results back into your read of the scores is part of the job.

The saving does not come from predicting the winner. It comes from not funding the obvious losers.

Making the waste argument to a stakeholder

If you need budget or buy-in for screening, do not pitch it as performance prediction. That claim invites the correct rebuttal that nothing predicts in-market results reliably. Pitch it as waste removal, in four steps:

  1. Quantify the window. Pull the last three underperforming campaigns and count the days between launch and the decision to pause or recut. Multiply by daily spend. That number is the detection cost you are currently paying per broken asset.
  2. Sort the post-mortems. For each failure, ask one question: was the problem visible in the asset before launch? Weak hook, late branding, buried offer means yes. Offer rejection or competitor move means no. The "yes" share is your addressable waste.
  3. Price the alternative. Screening every asset costs minutes of someone's time. Compare that line to the detection cost from step one. The ratio usually ends the discussion.
  4. Propose a measured pilot. Screen every creative for one quarter, block nothing at first, just record verdicts. Then compare in-market outcomes of assets that would have been flagged against the rest. You are building the internal evidence that makes the policy permanent, using your own data instead of a vendor's.

This framing survives finance scrutiny because it claims nothing about winners. It claims that a specific, recurring category of loss, budget spent discovering inspectable problems, can be removed at near-zero cost.

The honest summary

Pre-spend screening does not eliminate risk, and any vendor promising that is selling a forecast they cannot deliver. Market failures will still happen, and paying to discover them is legitimate research spend.

What screening removes is the avoidable category: budget spent discovering structural problems that were visible in the asset before launch. That is a narrower promise than "reduce wasted ad spend" usually implies, and it is the one the evidence actually supports.

If you want to see what the diagnostic looks like before making the case internally, the example report shows the KPI pattern, the verdict, and the evidence behind it on a real creative.