What TranSalNet does
RoastIQ is the scoring engine behind SaliencyLab, a pre-spend creative intelligence platform: it scores a creative before media budget is committed. Part of that analysis uses TranSalNet, a transformer-based saliency prediction model, to analyze each frame of your video. The model outputs a heatmap showing which areas of the frame are predicted to attract visual attention.
This is a prediction, not a measurement. No eyes were tracked. No participants watched your ad. The model was trained on large datasets of human fixation data and learned to predict where people are likely to look based on visual features like contrast, faces, text, and motion.
What it can tell you
- Which areas of the frame are predicted to attract attention. Useful for checking whether your brand logo, product, or CTA falls in a high-attention zone
- How attention distributes across frames over time. The evidence timeline shows when attention peaks and drops
- Whether key creative moments coincide with attention moments. If your product reveal happens in a low-attention frame, that is a structural problem
What it cannot tell you
- What any specific person actually looked at. Saliency prediction is probabilistic, not individual
- Why someone looked at a specific area. The model predicts where, not why
- Attention in context. The model sees the frame in isolation, not as part of a social media feed with surrounding content
- Audio-driven attention shifts. The visual model does not account for sound cues that redirect gaze
How accurate is it?
TranSalNet reaches state-of-the-art performance on standard saliency benchmarks such as MIT300 and SALICON.
The important thing is what that number means. Benchmark performance describes agreement with aggregated human fixation data on academic datasets. It is not an accuracy rate for your specific creative, and it does not transfer cleanly to advertising content that uses unusual visual techniques, rapid cuts, or text-heavy compositions. Treat the heatmap as a strong prior about where attention tends to go, not as a per-frame guarantee.
How RoastIQ uses attention data
Attention prediction feeds into two KPIs:
- Beat the Skip (25% of composite): the attention score contributes 55% of the Beat the Skip calculation. High visual attention in the first 2 to 3 seconds supports a stronger Beat the Skip score
- Get Noticed (20% of composite): Average Ad Viewed, which is attention-derived, is one of three sub-KPI components
The attention heatmap is also surfaced directly in the evidence timeline, so you can scrub through the video and see which frames are predicted to attract the most attention. For a practical guide to reading the overlay itself, see ad attention heatmaps explained.
The KPI scores that attention feeds are themselves predictions. In held-out, out-of-sample validation as of May 2026 they correlate +0.30 to +0.32 (Spearman ρ) with public engagement and click-intent outcomes, not with ROAS, sales, or brand recall. The validation record is on the methodology page.
Why we insist on the distinction
We call it visual attention prediction, not eye tracking, because that is what it is. The distinction matters for trust. If you present saliency predictions as measured eye tracking data, you overstate your evidence. If you present them as predictions trained on fixation data, you give the viewer the right calibration for how much to trust the signal.
This is also why attention is one input among five rather than the headline number. RoastIQ treats attention prediction as evidence that feeds the diagnostic. It is not the verdict.

