Start with one decision
A review of five ads can easily become five separate conversations. Write the shared question first: “Which cut should receive the next investment for this offer and audience?” Name the primary business objective and any constraints, such as production time or a required product condition.
If the candidates advertise different offers to different audiences, they are not a clean test of creative alone. You can still choose between them, but the recommendation must explain that broader trade-off.
Use the decision matrix
Download the editable CSV scorecard. Copy it for the actual asset versions under review.
| Candidate | Creative evidence | Buyer question | Comparable outcomes | Decision |
|---|---|---|---|---|
| A | Record an observable strength and concern | Record a hypothesis, if needed | Add source, period and uncertainty | Back / change / hold |
| B | Use the same review criteria | Ask the same neutral question | Check comparable conditions | Back / change / hold |
| C | Keep the exact asset version | Preserve disagreement | Leave unknown fields unknown | Back / change / hold |
Do not total incompatible evidence into a homemade certainty score. A model diagnostic, a synthetic quote and a campaign conversion are different observations. The matrix makes them visible without pretending they share a scale.
Read the evidence in layers
First inspect the ad itself. Is the offer readable? Does the export match the intended placement? Does a recommendation refer to a real scene? Next read the diagnostic explanation. Then consider any buyer hypotheses and real customer evidence. Finally, inspect comparable outcomes where they exist.
A weak model score can reveal a useful question, but it should not automatically overrule sound campaign evidence. Conversely, one strong observed result may reflect a different audience, budget or time period. Ask what else could explain the difference.
A worked decision without invented winners
Consider a fictional review in which A clearly demonstrates the product but omits an offer condition, while B includes the condition but uses a crowded final frame. Neither has campaign data. The immediate decision could be to revise A’s condition treatment while preserving its demonstration. It is not defensible to call A the measured winner.
After an appropriately designed live comparison, the team may have enough evidence to back one version. That decision should cite the metric, period, comparison conditions and uncertainty. Google’s video experiments are one way to structure a platform-specific comparison.
Include an insufficient-evidence option
A useful review sometimes ends with a bounded information request: verify the landing-page offer, recruit relevant customers for a comprehension question, or wait for the planned experiment window. Specify what information could change the decision. Avoid an open-ended instruction to keep testing forever.
If the decision cannot wait, state the uncertainty and choose a reversible next step with an owner. That is different from making a confident outcome claim the evidence cannot support.
Deliver the recommendation
Use one sentence: “Back [version] for [objective], because [strongest evidence]; preserve [strength], change [specific issue], and revisit if [defined contrary evidence].” This is a writing template, not a customer result.
Read a RoastIQ report to see how diagnostics support the matrix. Book a demo to discuss the creative decision your team needs to make.