How Many Ad Creatives Should You Test? A Framework
Test several distinct concepts, not one polished ad. Here's how to size your creative testing pipeline by budget and how to iterate on winners.
Test several genuinely different concepts at a time — enough that at least one has a real chance of hitting, but few enough that each gets meaningful spend. For most small accounts that means a handful of distinct angles per testing round, each with a couple of hook variations; larger accounts run continuous pipelines with new creatives entering every week. The exact number depends on your budget and margins, but the principle is fixed: most creatives lose, so the account that tests more distinct ideas finds winners faster.
Why does creative volume matter so much?
Because creative performance follows a hit-driven distribution — a small share of ads drives most of the results, and nobody can reliably predict which ad that will be. Media buying levers have largely been automated away by the platforms; the creative is the targeting now. Testing one ad at a time means betting the account on a single guess. Testing a steady stream of concepts turns finding winners from luck into a process.
What counts as a "different" creative?
This is where most testing goes wrong. Ten versions of the same video with different background music is one test, not ten. Structure your rounds in two layers:
- Concepts (angles): genuinely different arguments — problem/solution, us-vs-them, founder story, demo, objection-handling. Test a few of these per round.
- Variations: within a promising concept, vary the hook, the opening visual, the actor, or the length. A hook generator is useful for producing these quickly.
Find the winning angle first; then multiply it with variations. Iterating on a loser wastes budget no matter how many versions you cut.
How do budget and margin change the number?
Each creative needs enough spend to produce a readable signal before you judge it — so your testing count is capped by budget divided by that per-creative minimum, which itself depends on your price point and conversion path. Higher-margin products can afford wider tests; thin-margin products should test fewer creatives more decisively. There's no universal number: work backward from what your economics allow, and sanity-check the math with a ROAS calculator so you know what a "winner" even means at your margin.
How do you judge the test?
Kill fast on leading indicators, confirm on money metrics. Hook rate tells you within a small amount of spend whether the opening works; watch-through and click-through tell you if the body holds; ROAS over a full round decides who survives. Losers get cut without sentiment. Winners get a fresh set of variations in the next round.
What's the bottleneck — and how do you remove it?
Almost always production. Teams know they should test more; they just can't make ads fast enough. That's the case for generating instead of shooting: Polaris connects to Claude via MCP, and you generate ads directly inside Claude by attaching a product photo and describing the angles. Batches render in parallel, so a full testing round of UGC-style ads is a chat session, not a production sprint — and the first batch of 12 is free, which conveniently is a testing round on its own.