Falling behind on​​​​‌‍​‍​‍‌‍‌​‍‌‍‍‌‌‍‌‌‍‍‌‌‍‍​‍​‍​‍‍​‍​‍‌​‌‍​‌‌‍‍‌‍‍‌‌‌​‌‍‌​‍‍‌‍‍‌‌‍​‍​‍​‍​​‍​‍‌‍‍​‌​‍‌‍‌‌‌‍‌‍​‍​‍​‍‍​‍​‍​‍‌​‌‌​‌‌‌‌‍‌​‌‍‍‌‌‍​‍‌‍‍‌‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‍‌‌‌‍‌​‌‍‍‌‌‌​​‍‌‍‌‌‍‌‍‌​‌‍‌‌​‌‌​​‌​‍‌‍‌‌‌​‌‍‌‌‌‍‍‌‌​‌‍​‌‌‌​‌‍‍‌‌‍‌‍‍​‍‌‍‍‌‌‍‌​​‌​‍‌​​​‌‍​‌‍​‍​‌​​‍‌‌‍​‌‌‍‌‍​‍‌​‍‌​‌​‌‌‌‍‌​​‍‌​‌​​​‌​‍​​​​​‍‌‌‍​‍‌‍‌​​‌‌​‍​​‍‌‌‍​‍​‌​​‌​‌‍‌‍​​‍‌‍‌‍​​‌​‌‌​‌‌‍‌‌‌‍‌‌‌‍​‍​‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌​​‌‍​‌‌‍‌‌‍‌‌​‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌​‌‍‌‌‌‍​‌‌​‌‍‍‌‌‍‌‍‍‌​​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌​‌‍​​​‌‌‍​​‌‍‌‍‌‌​‌​​‌‌‌‍​​​‍‌‍‌‍​‍‌​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍‍​‌‍‌‌‌‍​‌‌‍‌​‌‍​‌‍‍‌‌‍‍‌‍‌‌​‌‍​‍‌‍​‌‌​‌‍‌‌‌‌‌‌‌​‍‌‍​​‌​‍‌‌​​‍‌​‌‍‌​‌‌​‌‌‌‌‍‌​‌‍‍‌‌‍​‍‌‍‌‍‍‌‌‍‌​​‌​‍‌​​​‌‍​‌‍​‍​‌​​‍‌‌‍​‌‌‍‌‍​‍‌​‍‌​‌​‌‌‌‍‌​​‍‌​‌​​​‌​‍​​​​​‍‌‌‍​‍‌‍‌​​‌‌​‍​​‍‌‌‍​‍​‌​​‌​‌‍‌‍​​‍‌‍‌‍​​‌​‌‌​‌‌‍‌‌‌‍‌‌‌‍​‍​‍‌‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌​​‌‍​‌‌‍‌‌‍‌‌​‍‌‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌​‌‍‌‌‌‍​‌‌​‌‍‍‌‌‍‌‍‍‌​​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌​‌‍​​​‌‌‍​​‌‍‌‍‌‌​‌​​‌‌‌‍​​​‍‌‍‌‍​‍‌​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍‍​‌‍‌‌‌‍​‌‌‍‌​‌‍​‌‍‍‌‌‍‍‌‍‌‌​‍‌‍‌​​‌‍‌‌‌​‍‌​‌​​‌‍‌‌‌‍​‌‌​‌‍‍‌‌‌‍‌‍‌‌​‌‌​​‌‌‌‌‍​‍‌‍​‌‍‍‌‌​‌‍‍​‌‍‌‌‌‍‌​​‍​‍‌‌ Amazon?

Benchmark your digital shelf against top sellers
Back to all posts

The Creative Supply Math: What a Weekly Testing Cadence Actually Costs in Creators and Briefs

Most teams don't have a creative testing problem, they have a supply problem. Here is the arithmetic behind a weekly testing cadence — briefs, assets per brief, and the yield number nobody measures — and how to run it on your own figures.

Black-and-white photo of a videographer carrying a cinema camera on his shoulder
Zach ChmaelSep 3, 2026

A premium activewear brand went from testing 10 to 15 things a week to testing 50 things a day.

That is Rhone, and it is published on their case study.

The description of what changed is worth reading closely, because it is not about testing methodology at all: they built an extensive library of content, scaled the amount they had to work with, and let the data drive performance from there. The reported outcome was a 3–4× increase in ROAS versus internal content, off 737 assets.

Here is the question every creative-testing framework skips.

What has to be true upstream for a team to test 50 things a day?

Not how you should structure the test. Where the creative comes from, and how much of it you need. That is an arithmetic problem, and most teams have never done the arithmetic.

Your replenishment rate is briefs per month × assets per brief × usable-asset yield. Most teams set a testing target, hit a wall, and diagnose creative fatigue, when the real constraint is upstream: not enough briefs, not enough creators, or too many assets failing review. Calculate the three inputs before buying another testing tool.

The number most teams have never calculated

Your sustainable testing cadence is set by three inputs multiplied together. If you have not measured all three, your cadence is being set by your bottleneck rather than your strategy.

Briefs per month. How many distinct creative requests you put into market. Not campaigns. Briefs, meaning a specific request with deliverables, direction and requirements attached.

Assets per brief. How many finished assets come back per brief. This depends on how many creators you activate per brief and how many deliverables each one owes.

Usable-asset yield. The percentage of delivered assets that clear brand, legal and platform review and actually run in market. This is the input nobody tracks, and it is the one that quietly destroys testing plans.

Multiply them and you get your replenishment rate: usable assets added per month.

Replenishment rate = briefs/month × assets/brief × usable-asset yield

Run it backwards to size a target:

Briefs/month needed = target usable assets/month ÷ (assets/brief × yield)

Worked example. Say you want 120 new usable assets a month. If a brief returns 8 assets and your yield is 60%, each brief nets you 4.8 usable assets, so you need 25 briefs a month. If your yield is 30%, each brief nets 2.4 and you need 50 briefs to hit the same target. The yield number doubled the brief requirement without anyone touching the testing plan.

Those percentages are illustrative inputs, not benchmarks — we are not going to hand you a yield figure we cannot defend.

That is the whole argument, though: two teams with identical testing ambitions and identical brief volume can end up with wildly different creative supply, and the difference is invisible unless yield is measured.

One distinction worth holding onto. Replenishment rate is a flow, not a stock. Rhone's 50 tests a day ran against a library of 737 assets, which means tests are combinations of asset, audience, placement and hook drawn from a library — not one new asset consumed per test. Library depth and replenishment rate are two different numbers, and both of them bind.

The three inputs, and how each one gets mismeasured

InputWhat it isWhere teams get it wrong
  • Briefs per monthDistinct creative requests in marketCounted as campaigns, which hides how few actual requests exist
  • Assets per briefFinished deliverables returnedConfused with creators activated; one creator can owe several assets
  • Usable-asset yieldShare clearing brand, legal and platform reviewNot measured at all, so shortfalls get blamed on creative quality
  • Replenishment rateUsable assets added per monthThe output nobody calculates before setting a testing target

Why "creative fatigue" is usually a supply diagnosis

We would argue that fatigue is a symptom of an unsolved replenishment rate, not a property of the creative.

To be fair to the agency frameworks that dominate this topic: their diagnostic work is good. They are correct that performance decays, correct that decay curves differ by placement, and correct that you need to distinguish creative decay from audience saturation. If you want to know whether your creative is fatiguing, those frameworks will tell you.

What they do not tell you is why it keeps happening.

Here is the pattern.

A team diagnoses fatigue, refreshes creative, performance recovers, and four to six weeks later they are back in the same conversation. Nobody solved anything. They metabolised a batch. If your replenishment rate is lower than the rate at which assets exhaust, fatigue is not an event you manage. It is your steady state.

This is our point of view rather than a measured finding, so treat it as a frame to test against your own numbers.

But the test is easy: the next time you diagnose fatigue, check whether you had a replenishment plan or a refresh. If it was a refresh, expect to be here again.

Usable-asset yield: the input nobody tracks

Usable-asset yield is the percentage of delivered creator assets that clear brand, legal and platform review and are actually usable in market. Everything downstream of your creative plan depends on it, and most teams cannot state theirs.

That definition is the term we would like to see the category adopt, because the alternative metric — assets delivered — is the one that flatters everybody and predicts nothing. More creator activity is not the same as more usable content.

The gap between delivered and usable is created by ordinary things:

  • The asset does not match a requirement the brief specified, or specified vaguely
  • Brand review rejects it on something the creator could not have known
  • Rights were never scoped for the channel you now want to run it on
  • Music or third-party IP makes it unusable in paid even though it is fine organically
  • It duplicates something you already have, so it adds no testing value

Two of those are brief problems, and that is the leverage point most teams miss. A vague requirement produces a resubmission cycle at best and an unusable asset at worst, and it does so at the same cost as a specific one. Brief Analysis gives feedback on briefs before they launch, including on compensation and creator expectations, which is the upstream half of the problem. Asset Analysis checks submitted assets against the requirements the brand set, which is the downstream half.

Neither of those makes yield a number you can skip measuring.

Measure your own over one quarter: assets that ran in market, divided by assets delivered.

What variant differentiation actually means

Volume of near-identical assets is not a testing programme. It is one asset with a hundred crops.

This section exists because the argument so far can be misread as "get more content," and that is not the claim.

A test only teaches you something if the variants differ in a way you chose deliberately: a different hook, a different creator demographic, a different setting, a different problem framing.

OrangeTheory is the cleanest published example we have. Over six months they ran more than 100 location-specific UGC videos and reported a 68% decrease in cost per lead on TikTok paid social. Single brand, TikTok paid social only, customer-reported. The number worth noticing is not 100. It is that the differentiating variable was location, which is a real axis of variance for a franchise business and not something you get from re-cutting one shoot.

Rack Room Shoes points at the same thing from a different angle: 1,000+ UGC photos and videos plus 130+ creator posts, with 59% more reach and 110% more engagement reported on Instagram against studio assets. The comparison there is against studio content, which is usually high-craft and low-variance.

So add a fourth input to your model, even though it does not multiply cleanly: how many genuinely distinct variants can your supply produce?

If the answer is "we have 400 assets and about six ideas," your constraint was never volume.

Cost per tested asset, honestly

Unit cost is the wrong metric if your yield is low, and it is the metric every vendor competes on.

The arithmetic is unforgiving.

An asset at $40 with a 30% yield costs you $133 per asset that actually runs. An asset at $80 with an 80% yield costs $100. The cheaper asset is 33% more expensive in the only currency that matters.

Cost per usable asset = cost per delivered asset ÷ usable-asset yield

That is the number to take into a vendor conversation, and almost nobody will be able to answer it about their own programme.

For a published unit-cost reference point, Samsonite generated 87 photos from regionally diverse creators at under $50 each, in under four weeks, refreshing Samsonite Canada's web, email and organic social. Photos only, single brand, customer-reported — video economics are different and this figure should not be read across to them.

On the production-model side, T3 Micro reported saving over $450,000 from two content campaigns run largely through product exchange rather than cash compensation. Single brand, customer-reported, and the saving is against their own prior production approach rather than a market benchmark.

One caveat we will state rather than bury: product exchange only works when the product carries enough retail value to function as compensation. T3 Micro's did. If yours is a $12 consumable, that model does not transfer, and the unit economics look completely different.

Where each production model breaks

Every model has a ceiling. Knowing which one you are about to hit is more useful than knowing which model is best.

Production models: ceilings, cost behaviour and failure modes

ModelThroughput ceilingCost behaviour at volumeRights handlingWhere it breaks
  • In-house studioLow. Bounded by crew, studio days and post capacityHigh fixed cost, low marginal cost per asset once runningClean. You own everythingVariance. High craft, low differentiation, and it cannot produce 40 distinct creator perspectives
  • Agency retainerMedium. Bounded by the retainer's scopeSteps up in blocks; overages are where budgets dieUsually clean but often scoped per campaignCost per asset at testing volume, and turnaround time against a weekly cadence
  • Self-serve marketplaceMedium to high on raw deliveryLow unit cost, and it stays lowThe common failure point. Often per-asset or per-channelYield and rights. Cheap delivered assets, expensive usable ones
  • Creator platformHigh. Bounded by brief throughput and review capacityLow unit cost with workflow overheadHandled in the workflowReview capacity, and the brief-writing bottleneck. Also a poor fit for one-off or low-volume needs

Where Cohley breaks, since we should say it. If you need one or two assets, occasionally, with no ongoing testing cadence, a platform is the wrong purchase and a marketplace or a freelancer is the right one. Our own buyer's guide says this. The model earns its keep when there is a repeating loop to run, and if there is no loop, the workflow is overhead you are paying for and not using. There is also a real fork between running it yourself and having it run: managed services exists because brief throughput is a staffing question as much as a software one.

Who this is for: if your testing cadence is weekly or faster and you are activating more than a handful of creators a month, the platform row is your comparison set. If you test monthly, the agency or studio rows probably still win on craft, and you should not let a volume argument talk you out of them.

Closing the loop: last week's winners become next week's briefs

The compounding part of this is a practice you build, not a feature you buy.

Once your replenishment rate is stable, the next gain comes from the loop: generate, approve, activate, learn, then write the next brief with what you learned. That last step is where most programmes leak, because the performance data lives in the ad platform and the brief lives somewhere else, and nobody owns the handoff.

Concretely, what you want written down after each cycle:

  1. Which hooks in the first three seconds outperformed, stated specifically enough to put in a brief
  2. Which creator profiles produced usable assets, not just applications
  3. Which requirements caused the most resubmissions, because those are yield losses with a named cause
  4. Which variants you have now exhausted, so the next brief does not reproduce them

A note on AI-generated variants

Synthetic variants change this math for some asset types and not others, and the tooling is not fully available yet.

Meta announced Muse Image in July 2026, with advertiser access through Advantage+ creative described as coming in the following weeks. As of writing, it is not shipped to advertisers.

The likely effect is uneven. Background variation, aspect-ratio adaptation and text treatments are cheap to synthesise. A real person using your product in their actual kitchen is not, and that is most of what creator content is for. Our expectation is that synthetic generation compresses the cost of variations on an asset while leaving the cost of new source material roughly where it is. Which, if right, makes the yield question more important rather than less.

That is a forecast, not a finding. Revisit it when the tooling actually reaches advertisers.

FAQ

How many creative variants should I test per week?

There is no honest universal answer, and anyone publishing one either has undisclosed data or is guessing. Your ceiling is briefs per month multiplied by assets per brief multiplied by usable-asset yield. Calculate those three and you will have a number that reflects your actual programme rather than someone else's.

What is usable-asset yield and how do I measure it?

It is the percentage of delivered creator assets that clear brand, legal and platform review and actually run in market. Measure it over a quarter: assets that ran, divided by assets delivered. Anything below your expectation is usually a brief-specificity or rights-scoping problem rather than a creative-quality one.

Is creative fatigue a creative problem or a supply problem?

Usually supply, in our view. The agency diagnostic frameworks are right that performance decays, but fatigue recurs when replenishment rate is lower than the rate at which assets exhaust. If your last fix was a refresh rather than a rate change, expect it to return in four to six weeks.

How many creators do I need for a weekly testing cadence?

Work backwards. Divide your target usable assets per month by assets per brief and by your yield to get the briefs you need, then multiply by creators per brief. The creator count is an output of the model, not an input you should guess at.

What does creator content cost per asset?

It varies by format, compensation model and product value, and unit cost is the wrong question anyway. Cost per delivered asset divided by your yield gives cost per usable asset, which is the figure that predicts your actual spend. A cheap asset with poor yield is expensive.

Can AI-generated variants replace creator content in testing?

Partly, for some asset types. Synthetic generation is well suited to background, format and text variation. It does not produce new source material of real people using a product, which is most of what creator content supplies. Meta announced Muse Image in July 2026 with advertiser access described as forthcoming; it is not yet available to advertisers.

When is an in-house studio still the right answer?

When craft matters more than variance and your cadence is monthly rather than weekly. Studios produce high-quality, low-variance work, which is exactly right for hero assets and exactly wrong for a testing programme that needs 40 genuinely different perspectives. Many brands should run both.

Next step

Run the formula on your own numbers. If the brief count it returns is higher than what your team can currently produce, that gap is your real constraint.

See what your testing cadence would require →