Attention heatmaps have become standard in creative review, and they are widely misunderstood. Teams present them as though a room full of people watched the ad and a camera recorded their eyes. That is not what happened.
This is an explainer of what the technology actually does, what the colours mean, and where the output stops being trustworthy.
What an attention heatmap is
An attention heatmap is a colour-coded overlay on a frame. Warm areas (red, orange) mark zones predicted to draw high visual attention. Cool areas (blue, green, or no overlay) mark zones predicted to be largely ignored.
The overlay comes from a saliency model. The model was trained on datasets where researchers recorded real human eye movements across thousands of images, then learned the visual patterns that predict where gaze concentrates: contrast, faces, motion, text, edges, and isolation against a background.
When you feed it a new frame, it produces a probability surface. Nobody looked at your ad. The model is estimating, from learned visual regularities, where people would probably look.
Prediction versus measurement
This is the distinction that matters, and it is the one most often blurred in sales material.
Eye tracking is measurement. Real participants, a calibrated camera, recorded fixations. It produces ground truth about those specific people in those specific conditions, and it costs time and money to run.
Saliency prediction is inference. A model estimates a likely attention distribution in seconds at near-zero marginal cost. It generalises from population-level patterns, so it has no opinion about any individual viewer.
Both are useful. They are not interchangeable, and a heatmap presented as eye tracking overstates the evidence behind a decision.
How RoastIQ generates its heatmap
RoastIQ uses TranSalNet, a transformer-based saliency model, run per frame across the creative. Two things come out of it.
The first is the visual overlay in the evidence timeline, so you can scrub the video and see how predicted attention shifts as the cut progresses.
The second is a numeric attention score that feeds the KPI layer. Attention contributes 55% of Beat the Skip, which carries the heaviest composite weight at 25%. It is also one of three components inside Get Noticed. So a structurally weak opening does not just look bad on the overlay, it propagates into the verdict.
What the heatmap answers well
- Is the brand cue in a high-attention zone? A logo placed in a cold corner is technically present and functionally invisible
- Does the product reveal land where attention already is? If the reveal happens in a low-attention frame region, the sequencing is fighting itself
- Is attention concentrated or scattered? A frame with four competing warm zones usually means the viewer has no idea what to look at
- How does attention move over time? The timeline view matters more than any single frame, because ads are watched, not viewed
What the heatmap cannot answer
- What one person actually looked at. The output is a distribution, not an observation
- Why attention landed somewhere. The model predicts location, not motivation. A face draws attention whether or not it helps your message
- Attention in a real feed. The model sees your frame in isolation. It does not know what sat above or below it while someone scrolled
- Audio-driven gaze shifts. A visual model does not account for a sound cue pulling the eye
- Whether attention converts to anything. Attention is a precondition for the message landing, not a substitute for the message being good
That last point is where heatmaps get abused. A creative can score beautifully on predicted attention and still fail, because attention was never the problem. Heatmaps diagnose visibility. They say nothing about persuasion.
How to read one without over-reading it
Treat the heatmap as a strong prior, not a verdict. A useful sequence:
- Look at where the warm zones sit before you look at scores
- Ask whether the elements that carry the commercial job (brand, product, offer) sit inside them
- Check the timeline for whether attention holds or collapses after the opening
- Only then compare against the Beat the Skip and Get Noticed scores
If the overlay and the scores disagree, trust the overlay to describe visibility and the scores to describe consequence.
What validation actually exists
Saliency models are benchmarked on academic datasets such as MIT300 and SALICON, where performance means agreement with aggregated human fixation data on those datasets. That is a legitimate measure. It is not a claim about accuracy on your particular creative, and it transfers less cleanly to advertising content with rapid cuts, heavy text, or unusual composition.
For the downstream scores that attention feeds, held-out out-of-sample validation as of May 2026 shows Spearman correlation of +0.31 against TikTok engagement (n=700), +0.30 against TikTok CTR (n=691, 5-fold cross-validation), and +0.32 against YouTube view counts (n=403, 5-fold cross-validation).
Those are real but moderate relationships, validated against public engagement and click-intent outcomes. They are not forecasts of sales, ROAS, or brand recall, and holding them at the right strength is the difference between a useful tool and an overclaim.
The practical position
An attention heatmap is a cheap, fast, repeatable answer to a narrow question: where is this creative likely to be looked at, and is that where the important things are?
Asked that question, it is genuinely useful and worth running on every cut. Asked to predict business outcomes, it will happily produce a colourful picture that means nothing of the sort.
If you want to see one in context rather than in the abstract, the example report shows the overlay sitting alongside the KPI pattern it feeds.
