Asset performance — Gen 5, _260730 worldwide · three creatives in one CBO ad set

Lane: reference. Data through 2026-08-01 16:00 PT. Campaign bailingxia_meituan_ww_cvr_260730, one ad set ww_broad_purchase_adv_cbo150, CBO, Purchase event, worldwide, Android. Three creatives: ww_bodysuit_10s, ww_bikini_10s, ww_pool5_15s. Live 2026-07-30 15:00 PT → 2026-08-01. Roughly $1,072 of media across three calendar days.

This report runs on the ad account's own clock, UTC-7, and that is correct for what it measures. The daily campaign reports moved to a UTC day on 2026-08-05 to line up with Adjust, so a reader arriving from one of those will find these figures on a different basis. Everything here is built on budget regimes and hourly delivery cells keyed to the account's clock, and a budget change is an event at a wall-clock instant on that clock (daily_report_format.md § the report clock).

The question as designed: which of the three creatives wins?

Emphasis (which parts of creative_value_model.md apply): §7, the test-versus-bandit conflict inside a pooled ad set — and this report is where that section stops being a caveat and becomes the entire result. The ranking sections of the method are deliberately not applied, because the precondition they assume — that the arms were given comparable delivery — is false here and measurably so. What replaces them is a cohort measurement built on a new instrument (§4), which is the first thing in this account able to file revenue against the traffic that produced it.

1The result in one line

The test cannot answer its own question, and the reason is not sample size. Meta gave the three creatives disjoint delivery windows, so every per-asset number the design produces — CTR, cost per purchase, ROAS, and the estimator's Net per $ — is a statement about when Meta chose to run an ad rather than how good the ad is. On the one window where two of them ran the same hours with both cohorts drained, they are indistinguishable.

The test did, however, produce three things worth more than the answer it was built for: a measured allocation mechanism (§2), a working per-user attribution instrument (§4), and four reversed verdicts that each failed the same way (§6).

2Snapshot

ww_bodysuit_10sww_bikini_10sww_pool5_15s
lifetime spend$457.15$285.63$36.63
share of ad-set spend58.7%36.6%4.7%
attributed installs2,337822157
delivery state at 2026-08-01 17:00 PTOffOffActive, alone

pool5 took 4.7% of the money across the whole run and has never delivered since 07-30 21:00.

3What actually happened: the allocation, hour by hour

when (PT)what happened
07-30 15:00–16:00all three go live, budget $150/day
07-30 21:00last hour pool5 or bikini delivers anything
07-30 22:02budget $150 → $250
07-30 22:00 → 07-31 17:00bodysuit alone, 100% of delivery, ~19 hours
07-31 07:19budget $250 → $500
07-31 17:21budget $500 → $875
07-31 18:00bikini delivers again — 39 minutes after the $875 step — takes 27% of the day
08-01 00:00–04:10bikini takes ~73% of the block
08-01 04:10budget $875 → $500; five on/off toggles between 03:40 and 06:58
08-01 06:55bikini switched off (operator)
08-01 ~17:00bodysuit switched off (operator); pool5 alone

This timeline was corrected on 2026-08-01 from Ads Manager's activity history. The version this report and three other documents carried — "$150 → $875 at 07-31 07:19" — was wrong on the value, the count and the time. Budget history is not readable from any export: Ad set budget is current-state stamped onto every historical row, so an export taken today reports 875 Daily back to 07-01, before the campaign existed. It lives only at manage/campaigns/history (See history / Ctrl+I, campaign selected). Three days of analysis were built on an unverified budget timeline.

Selection happened on 07-30 and the budget step did not cause it. pool5 and bikini both went to zero at 21:00 on 07-30 — ten hours before the budget moved. Meta had already collapsed to one winner at $150/day and rode it straight through the step.

Winner-take-all inside an ad set is normal, and the peer account is more concentrated than we are. Inside each of SeedMotion's three ad sets the top ad takes 85.1%, 85.9% and 79.4% of lifetime spend against our 58.7%, and their 0726-9 fell to 0.9% and then took a hard zero for three straight days exactly as pool5 did. Their account looks spread out only because it is three winner-take-all races in parallel, each fed $50–148/day.

What the budget decided was how far down an already-frozen ranking Meta had to reach. bodysuit's frequency went 1.16 → 1.51 on 07-31 as Meta re-showed it to the same people, and even with bikini helping the campaign placed only $525 of the $875. The winner could not absorb the budget, so the overflow went to whichever ad the optimiser had left alive at 5–10%. In a one-ad-set account your number two is whatever is still breathing; in a three-ad-set account it is another ad set's proven winner.

4Why no per-asset verdict from this design is safe

The three creatives never shared a delivery window until 08-01. Every headline number computed across their own windows therefore carries the window, not the creative:

the claimthe number that made itwhat it was actually measuring
bikini has 2 points less CTR8.39% vs 10.41% on 07-31bodysuit ran all 24 hours; bikini only 18:00–23:59. Hour-matched over the hours both ran: 8.39% vs 8.92% — a 0.5-point gap
bikini returns 0.93× lifetimelifetime ROASdominated by 07-31 evening, the day's worst window, which bodysuit was not confined to
bikini absorbed cheap inventoryCPM collapse 18:00–23:59bikini paid a premium in nine of the ten shared hours — but that premium tracks how hard Meta was pushing it (0.98× at $1.6/h, 1.43× above $20/h, R² = 0.79) and is not a property of the creative. See the §6 note
pool5 was dropped for being narrowhighest CTR = narrowest audiencebodysuit's CTR on 07-30 was 14.04% against pool5's 13.59%. The highest-CTR ad is the one Meta kept
pool5 is gated out by its 15s lengthit never deliveredplacement × day export: its mix matches the 10s ads (FB Reels 34.3% vs 35.7%/37.8%). It served 13 of 18 placements; the 5 it missed total under 1.5% of any ad's impressions

The decay is real. It is neither placement mix nor auction depth — it is COUNTRY mix, and the budget is its dial.

Same creative, same hours of the clock, comparable install volume, only the budget differing:

regimeIndia share of installs
$250/day, overnight (R2)19.5%
$500/day, daytime (R3 and R5 — two independent days)30.8% / 31.0%
$875/day, overnight (R4)41.3%

India more than doubles from $250 to $875 while Spain and Chile go to zero and Turkey and Indonesia roughly halve. Drop the budget back to $500 and India returns to within 0.3pp of where the same budget put it a day earlier — the whole mix reproduces within ~2.5pp. It is a dial, and it reverts.

That is one mechanism for what were being treated as three separate mysteries. India is our weakest market (#42: $1.57 d7 against US $4.05), so a shift into it lowers CPM (cheap inventory), CTR (different audience) and revenue per install together. At $875 Meta cannot fill the budget from the mix it buys at $500, so it goes where volume is cheapest.

Which means the campaign never "decayed". It was moved into a worse country mix by the budget and moved back out again. No creative fatigue and no learning damage is needed to explain the three-day slide, and the budget ceiling is a mix control — a better reason to hold $500–600 than the DNU cap.

Placement mix moved too and is not the mechanism: holding the creative constant bodysuit went Facebook 40.6% → 56.2% → 63.5% and Instagram 59.3% → 42.2% → 33.6%, but CTR fell within each platform independently (Facebook 15.73% → 10.93% → 7.29%, Instagram 11.37% → 8.21% → 6.59%) and the drift ran toward Facebook, the higher-CTR surface.

⚠ And this confounds every per-asset comparison in this report. bikini_10s ran almost entirely inside R4, the 41%-India regime; bodysuit_10s's better lifetime figures draw substantially on R2 and R3 at ~20–31% India. A meaningful share of "bikini is worse than bodysuit" may be "R4 is worse than R2". The two were never given the same country mix to compete in, and the creative question is reopened, not settled — see §7.

5The instrument this test produced

Answering any of this needs revenue filed against the install that produced it, which no Meta export can do: Meta's Purchases column is its own 7-day-click attribution and carries no click timestamp. Two sources exist and one is dead.

ad_ops/lum_notifications/parse_export.py reconstructs both and emits payments_for_estimator.csv in the schema load_payments() wants. 695 of 696 payments join to their originating install. Four traps had to be handled, each of which silently corrupts the result rather than failing:

  1. Every payment is announced twice, in a long form carrying 账户ID and a short one-line form carrying 用户ID plus the tracker. Counting rows doubles revenue.
  2. The one-line form is comma-separated, so a field regex terminating only at \n makes the first field swallow every field after it. This was corrupting transaction_id and therefore breaking the dedupe in (1).
  3. Amounts arrive in the payer's local currency — 33 of them, INR alone being 60% of transactions — on regionally-priced SKUs. Summing raw amounts is meaningless; FX conversion is correct and validates against Meta's USD conversion value at 0.91–1.06 on the large samples.
  4. Telegram Desktop stamps a fixed standard-time offset and ignores DST. All 23,522 messages read UTC-08:00, including July and August, which are PDT (UTC-07:00). Converting with the stated offset puts every row an hour late. asset_estimator's payments_timebase_lag() caught it at exactly −1 h, r = 0.9947; it now reads +0 h, r = 0.9903. That hour was not cosmetic — it moved the between-arm payer-rate heterogeneity from p = 0.018 to p = 0.136, i.e. it was manufacturing a significant difference that does not exist.

The accrual curves, and why they must be measured on complete cohorts only

Built from installs ≥72 h old — young cohorts contribute their first-hour payments and none of their later ones, and including them front-loads the curve and understates every projection made from it:

by hourcumulative share of revenuecumulative share of payers
137.3%56.8%
964.3%73.6%
2475.2%79.1%
4884.1%86.8%
7288.7%90.9%

A quarter of revenue arrives after 24 hours. Any figure read off a same-day window is censored, and the censoring is largest for whichever arm is still buying.

6Per asset, on the corrected basis

Install cohorts, censoring-corrected for both revenue and payer count. Revenue is FX-converted gross; the store cut is not applied (see caveats).

assetinstall dayinstallspayer rate (corrected)projected cohort ROAS
pool5_15s07-301518.4%3.01×
bodysuit_10s07-30798.9%2.33×
bodysuit_10s07-311,2776.3%2.57×
bodysuit_10s08-019815.9%1.99×
bikini_10s07-313594.3%1.67×
bikini_10s08-014355.8%2.24×

The one clean comparison in the whole test — 08-01 installs 00:00–06:59 PT, the only window where two creatives ran the same hours and both cohorts are equally aged:

spendinstallspayer rateARPUARPPUCPIprojected ROAS
bodysuit_10s$64.472113.8% obs$0.431$11.36$0.3062.02×
bikini_10s$160.434164.1% obs$0.516$12.63$0.3861.91×

Indistinguishable — 6% apart on 8 and 17 payers.

bodysuit's decline is real but is one step, not a slide: payer rate 8.9% → 6.3% across the budget step, then flat at 5.9%. The apparent third leg down to 4.2% was the youngest cohort not having paid yet.

pool5 is the best asset on every measure available — highest corrected payer rate, cheapest installs at $0.243, highest projected return — on 151 installs and $36.63 of lifetime spend. That is 11 payers, so the interval is wide and the figure was earned in the $150/day regime. It is the most promising and least tested thing in the account.

7Four verdicts this test reversed, and the single reason

verdictreversed bywhy it failed
bikini should be cut (CTR two points low)hour-matchingmeasured bodysuit's full day against bikini's evening
bikini should be kept (Net per $ 1.15, interval clear of 1.00)its own revenue per installNet per $ scores each asset's funnel against the tier-pooled revenue per install. bikini's own is $0.354 against a $0.539 pool — 34% below. Corrected: 0.73×, a CUT
the budget step is the dominant cost driver (+54% cost per payer)ARPPU over the same stepARPPU rose +58% alongside it. Cohort ROAS came out flat, 1.91× → 1.96×
bodysuit is now underperforming badly (payer rate 7.6% → 4.2%)censoring correctioncorrected, 8.9% → 6.3% → 5.9%. One step at the budget change, flat since

Every one failed the same way: a ratio read on one side only, over a window the optimiser chose.

The estimator was changed on 2026-08-01 as a result, and re-run on this same data now returns not comparable for both delivering arms instead of KEEP. Four fixes: the ~16-payer gate counts lifetime payers rather than in-window ones (which is what wrongly gated bikini at 12 when it had 32); the gate now refuses the verdict instead of reporting a shortfall beside a KEEP; revenue per install is shrunk by a measured weight (payer_rate × ARPPU) rather than switched at a threshold, and both bases are printed; and delivery_overlap() marks arms sharing under 50% of their delivering hours as not comparable. Fixes 1–3 alone move bikini from KEEP to CUT; fix 4 is what turns that into the honest answer. Guards and the failure each encodes: asset_estimator/README.md. Which is also the answer to why bikini was cut at the moment it looked, on the data

then visible, better than the ad that was kept:

at 06:55 PT, installs 00:00–06:59spendrevenue visibleROAS visiblematured
bikini_10s$160.43$147.490.92×1.34×
bodysuit_10s$64.47$36.400.56×1.41×

68% of bikini's eventual payers had already paid — the cohort was mature enough. The cut was made against its lifetime 0.93×, which was the hour confound carried forward. And Meta's own headline pointed the same wrong way for a second, independent reason: Purchases is a transaction count. Purchases per dollar made bodysuit look 37% better while revenue per dollar made bikini 64% better, because a $2.99 credit pack and a $48.99 year-sub count identically and bikini skews to the larger SKUs (pack_b + pack_c are 22% of its transactions against bodysuit's 6%). Optimising on cost per purchase in this account optimises against its own revenue mix.

86b. What the corrected instruments say, and what is still not settled

Rebuilt on fixed-age hourly cohorts (§4), aligned on each ad's own first delivering hour, raw beside maturation-projected, coverage stated:

first 18 delivering hours, each from its own launchspendpayersROAS@6hROAS@24hproj@48h
bodysuit_10s$177.26391.96×2.14×2.38×
bikini_10s$309.40300.98×1.26× (97% projected)1.59×

Over the one continuous run where both delivered — 07-31 18:00 → 08-01 03:59, the whole of R4 — bikini returns 1.13× against bodysuit's 1.45× on identical hours at identical age. It was also turning down at the moment it was cut: rolling-3 ROAS peaked at 1.54 near 02:00 and fell to 0.91 by 03:00.

So the cut was reasonable and the number quoted for it was not. Bikini cleared break-even and ran roughly 20–25% behind bodysuit; it was never the 0.82× that justified switching it off.

Still unsettled, and the reason this report cannot close the creative question: both figures above pool across regimes, and the two ads did not run in the same ones. The outstanding test is each ad's country mix within R4 alone — if they match there, the creative comparison survives; if bikini skews further into India even inside the regime, then it does not, and every creative verdict in this account is a geography verdict wearing a creative's name.

That test was run on 2026-08-01 and settles this section (regime_matched_redo_2026-08-01.md, DECISION_LOG.md #48). Inside R4, bikini drew 31.8% India against bodysuit's 39.3% (−7.5pp, p = 0.007) — less of our weakest market. And that mix difference is worth −0.9%, 95% CI [−21.2%, +27.5%]: real in composition, indistinguishable from zero in money. The geography hypothesis is refuted, and there is no correction to apply in either direction.

It also retires the "20–25% behind" figure, in both directions. Revenue per install here has a coefficient of variation of 6.68; resolving −19% needs ~19,000 installs per arm and this window had 732 and 488, giving bootstrap ROAS 1.18× [0.64–1.82] against 1.45× [0.54–2.58]. The §5 "indistinguishable" verdict on 8 and 17 payers was right, and it applies to every per-asset ROAS in this report, including the ones computed on the corrected basis.

Nor does the delivery side separate them. On every stage a creative controls the two are indistinguishable — CTR −7.0% [−24.9%, +15.1%], click→install −14.9% [−33.7%, +9.3%], payer rate p = 0.31. The only separator is CPM, which the auction sets and which appears when Meta ramps an ad and vanishes when it stops: 0.98× at $1.6/h, 1.43× above $20/h, R² = 0.79 on the ad's own hourly spend. Meta's own quality rankings are blank on every export.

So the cut had no evidential basis in either direction. It may have been right; it was not justified. Region-held payer rate leans the other way (MH OR 3.17, p = 0.096) and bodysuit's edge is three payers in BE/SI/CH worth 55% of its R4 revenue.

§2's allocation timeline should also not be read as judgement. At the $875 step Meta held two parked arms — pool5 (13 purchases, ROAS 1.67) and bikini (1, 0.58) — reactivated bikini, gave pool5 nothing further, and took it to 76% of budget while its cumulative ROAS sat at 0.21. That was capacity: $875/day needs $36.5/h and bodysuit peaked at $28.5.

9Recommendations

  1. Never run a creative comparison inside a CBO ad set again. Selection converges in about six hours on CTR alone, before a single purchase has reported, and it is sticky. A read on new creative needs a test ad set with forced equal budget, or its own campaign.
  2. Give pool5 its own campaign and its own budget, at $150/day — the regime its 3.01× was measured in, so the result is comparable rather than a fresh unknown. It cannot be tested by being left switched on in the existing ad set; it has been on for two days at zero.

Amended 2026-08-01. The budget matters here far less than this assumed: priced at account-wide country values, the $150 regime's audience is worth 1.48 expected ROAS against the $500 regime's 1.45 — a 2% difference. pool5's number is not a regime artifact, and it need not be re-run at $150 to be comparable. It is also not yet a result: the 07-30 cohort has now realised 2.52× ($92.23 on $36.62, essentially complete — 100% of its revenue arrived within 48 h), on 11 payers of whom one account is 26% of the revenue.

  1. Set the budget to what the winning asset can absorb. Overshoot and Meta pulls back in ads it has already deprioritised, at a worse price — that is the whole mechanism of §2.
  2. Stop reading cost per purchase. Use revenue per dollar and the corrected payer rate.

Superseded 2026-08-01 by #48. Revenue per dollar cannot rank creatives at this account's volumes either — see the §6 note. Rank on CPM and CTR, which resolve a 0.5pp difference on one night's impressions; use revenue to decide whether the account is worth running, which is one number rather than a comparison between two. Run ad_ops/lum_notifications/test_power.py before designing a test.

  1. Repair adjust_sink — check the callback URLs in the Adjust dashboard before anything else. Until then this class of analysis needs a manual Telegram export and reaches back only to 2026-07-11.

10Caveats — what would change these answers

11How to reproduce

# 1. reconstruct the notification channels (Telegram Desktop HTML export of LumUser + LumPayment)
python ad_ops/lum_notifications/parse_export.py <LumPayment_dir> <LumUser_dir> \
    --out _runs/<run>/lum --tz America/Los_Angeles

# 2. run the estimator with payers attached
export ADJUST_API_TOKEN=$(bwx field suite.adjust.com api_token | tail -1)
python -m ad_ops.asset_estimator \
  --export ad_ops/_data/b1_meta_exports/B1-Ads-Jul-31-2026-Jul-31-2026.csv \
  --pt-day 2026-07-31 \
  --payments _runs/<run>/lum/payments_for_estimator.csv \
  --tiers '{"WW broad": "bailingxia_meituan_ww_cvr_260730"}' --out _runs/<run>/est_0731

Confirm time-base check vs Adjust: best lag +0 h before trusting anything downstream. Cohort, maturity and at-cut analyses: _runs/2026-08-01_account_recheck/lum_*.py. Meta exports for the window, including placement × day: _data/b1_meta_exports/.

Internal — noindex. Not for distribution outside the team.