CREATIVE REVIEW COMPARISON PROTOCOL Version 1 · 7 October 2026 · Prepared by SaliencyLab (an interested vendor) Status: protocol only. No completed head-to-head results or superiority claim. QUESTION For a fixed creative and brief, how useful, accurate and actionable is the next-edit advice from SaliencyLab, ChatGPT and Claude? TWO SEPARATE COMPARISONS 1. Equal-input advice: identical assets, transcripts, brand context and brief for every product. Provide equivalent reference evidence; record any input a product cannot accept. Compare the software output before human intervention. 2. Full workflow: include each product's actual setup, project context, operator work and revision steps. Report this separately; extra context or human labor is not a model-quality advantage. Human-reviewed SaliencyLab services are a separate arm with labor/time/cost disclosed. SELECTION BEFORE GENERATION An independent custodian selects at least six permissioned, previously unseen creatives spanning images and videos, two categories, and Meta/TikTok briefs. This is a small descriptive pilot, not a population estimate. Exclude our public demos and development/regression examples from held-out claims. Freeze asset hashes, brief, inclusion/exclusion rules, configuration, prompts, supported modalities and analysis plan before viewing outputs. Customer assets stay private unless publication is separately approved. COMMON PROMPT Review the supplied creative against the supplied campaign brief. Identify up to three priority edits. For each, give the exact frame, timestamp or passage; what to preserve; the proposed change; the reasoning and source; a plausible alternative interpretation; and what actual evidence would change your recommendation. Separate direct observations, model inferences and unknowns. Do not invent viewer quotes, campaign results or unsupported performance forecasts. If information is insufficient, say what is missing. Use the same output structure for every creative. EXECUTION Use fresh contexts with the frozen brief. Record product/model shown, date, search/tool settings, asset-processing limitations, output, run identifier, retries, elapsed time, operator minutes and available costs. Unknown costs are unknown, not zero. Keep failed runs in the denominator. Do not keep retrying until a preferred answer appears. Freeze repeat counts beforehand. Compare still-frame/transcript proxies only in their own stratum; do not silently equate them with full video analysis. BLIND REVIEW At least two independent human reviewers inspect the creative and brief. Randomize output order, remove only provider identifiers, retain the untouched originals, and disclose any residual unblinding. Reviewers must not produce the outputs. Score each dimension 0–4: 0 missing/wrong, 1 weak, 2 mixed, 3 useful/supported, 4 strong/specific. Dimensions: factual correctness; creative specificity; decision relevance; evidence fidelity; useful alternative interpretation; editing usefulness; preservation of useful elements; uncertainty calibration. Record unsupported claims, missed critical issues, contradictions, rationale and review minutes separately. Do not use an AI's assessment as an independent human rating. REPORT Publish per-case/per-dimension scores, reviewer disagreement, failures, actual denominators and missing accounting. Keep advice quality, speed, workflow effort and campaign effects separate. No composite marketing winner, statistical-significance claim or general prediction claim from this small pilot. Publish negative/inconclusive cases. If permissions prevent reproducing assets or outputs, state the reproducibility limit. Register protocol amendments with dates before new collection. OUTCOME STUDY A later campaign experiment needs its own randomization, metric definitions, stopping rule and permissions. Higher reviewer ratings alone cannot establish campaign lift.