The regime-matched redo — what survives when budget, hours and cohort age are all held

Lane: reference. Data through 2026-08-01 16:38 PT. Campaign bailingxia_meituan_ww_cvr_260730, ad set ww_broad_purchase_adv_cbo150, CBO, Purchase, worldwide, Android. Every figure here comes from python -m ad_ops.asset_estimator hourly — see Reproduce.

This carries out the redo that DECISION_LOG.md #47 ordered: recompute every per-asset verdict on regime-matched, fixed-age hourly cohorts rather than citing the old ones. The test #47 named as decisive was bikini against bodysuit inside the $875 regime alone — if their country mixes matched there, the creative comparison survived; if bikini skewed further into India even inside the regime, every creative verdict was a geography verdict wearing a creative's name.

1The result in one line

The geography hypothesis is refuted, and the comparison it was meant to rescue turns out not to exist either. Inside the $875 regime bikini drew less India than bodysuit — 31.8% against 39.3%, p = 0.007 — so it did not skew into our weakest market. And the value of that mix difference is −0.9%, 95% CI [−21.2%, +27.5%]: real in composition, indistinguishable from zero in money. There is no geography correction to make, in either direction.

The deficit it was supposed to explain cannot be read anyway. Revenue per install here has a coefficient of variation of 6.68 — separating two creatives at the observed gap needs about 19,000 installs per arm, and the best-powered window this account has ever produced had 732 and 488, with bootstrap ROAS 1.18× [0.64–1.82] against 1.45× [0.54–2.58].

What does resolve on one night's data is the delivery side. bikini was charged 38% more per impression than bodysuit, dearer in nine of the ten hours they shared, at equal CTR. That is measured on 57,743 impressions rather than 41 payers, and it is a sufficient reason to have cut the ad — unlike the 0.82× it was actually cut on.

2The test #47 called for

Regime R4 ($875/day, 07-31 17:21 → 08-01 04:10 PT) is the one window where both creatives delivered at volume, on the same ten hours. Install-level country comes from the Lum registration channel; each install is attributed to the delivery hour that produced it.

bikinibodysuit
installs in R4732488
India share31.8% [28.6–35.3]39.3% [35.1–43.7]−7.5pp, p = 0.007
US share1.9% [1.1–3.2]2.7% [1.6–4.5]−0.8pp, p = 0.38
what that mix is worth, at pooled country prices$0.483$0.504−0.9%, CI [−21.2%, +27.5%]

bikini skewed away from India, not into it — and the difference is not worth a measurable amount. The India gap itself is solid, on 732 and 488 installs. Its price is not. The country values it would have to be converted with are themselves estimates: India's revenue per install rests on a few hundred installs, the US on a few dozen, and a single large payer moves either. Resampling the pooled installs and rebuilding the country table from each resample puts the value gap at −0.9%, from −21% to +28% — spanning zero, and spanning it widely.

That is a better answer than either direction would have been: the creative comparison needs no geography correction, because the geography difference cannot be shown to be worth anything.

The exposure confound was nonetheless real — which is why the question was worth asking:

share of each ad's lifeR1 $150R2 $250R3 $500R4 $875R5 $500pooled Indiaregime-standardised India
bikini3%0%0%88%8%31.5%23.6%
bodysuit1%15%28%21%35%29.9%32.0%

bikini lived almost entirely inside the $875 regime while bodysuit spread across four. So every pooled lifetime comparison of the two was confounded exactly as #47 suspected. Standardising both onto the account's regime distribution moves bikini from 31.5% to 23.6% India and bodysuit the other way. The confound was there; its sign was the opposite of the one assumed.

3Pricing the mix — and what it is worth

A shift in country shares is only a confound if it moves money. Each ad's own mix is valued at pooled country prices built from the same cells: shares from the ad, prices from the pool. An ad's revenue inside one country is not measurable here at all, so it is never used.

R4, H = 12 hinstallspayersspendCPI obsCPI expcost effRPI obsRPI exprev effROAS95% CI
bikini73227$286.17$0.391$0.3231.21$0.459$0.4830.951.180.64 – 1.82
bodysuit48814$144.09$0.295$0.3180.93$0.429$0.5040.851.450.54 – 2.58

Two things fall out, and one non-thing:

  1. The entire ROAS gap is acquisition cost. bikini paid 21% more per install than its own audience implies while bodysuit paid 7% less — a 30% swing on the cost side, against a revenue side where both sit below their mix expectation by a similar margin.
  2. The revenue side does not separate them. The bootstrap intervals overlap across almost their whole length, which §4 shows is not a property of this window but of the measure.
  3. There is no mix correction to apply. Both arms' expected revenue per install land within 4% of each other, and that 4% carries an interval from −21% to +28%. A "mix-neutral ROAS" computed from a single country table reads −14% or −22% depending on which population the table is built from; neither number is real. The raw −19% stands as the gap, uncorrected, because geography has nothing measurable to say about it.

4Where the cost difference actually sits

CPI = CPM/1000 × 1/CTR × 1/(click→install). Over the same ten hours:

R4, 18:00 → 04:00spendimpressionsclicksinstallsCPMCTRclick→installCPI
bikini$286.1734,3452,616732$8.337.62%28.0%$0.391
bodysuit$144.0923,3981,840488$6.167.86%26.5%$0.295
bikini relative+35%−3%+6%+33%

The hook is not the problem. CTR differs by 3% and the two trade places hour by hour — bikini is ahead in five of the ten. Click-to-install is 6% better for bikini. The whole gap is that Meta charged 35% more per impression to deliver it.

That survives every control available:

adhour-cellsspend rangeslope95% CI
bodysuit39$3.7–37.9 (10×)0.9950.77 – 1.23
bikini11$4.5–39.2 (9×)0.9450.84 – 1.05

Both intervals span 1.00: impressions scale linearly with spend across a nine- to tenfold range, so CPM is flat within an ad. bikini's premium is a stable property of the creative, not a penalty for being favoured.

5The finding that governs the rest — revenue cannot referee this

The R4 window is the best-powered creative comparison this account has run: one regime, matched hours, both cohorts fully mature, the two largest install cohorts on record. Bootstrapping over installs:

ROAS95% CI
bikini1.180.64 – 1.82
bodysuit1.450.54 – 2.58

The intervals overlap across almost their whole length. Revenue per install is a near-zero vector with rare large entries — 4.2% of installs pay anything — so its dispersion is enormous relative to its mean:

n = 3,179 · payers 4.2% · mean $0.454 · sd $3.029 · coefficient of variation 6.68

to resolve a difference ofinstalls per armspend per arm at $0.35 CPIboth arms
50%2,798$979$1,958
30%7,772$2,720$5,440
20%17,487$6,120$12,241
10%69,945$24,481$48,962

Resolving the −19% actually observed needs ~19,000 installs per arm. The window had 732 and 488 — 4% and 3% of the requirement. No ROAS comparison between two creatives in this account has ever been close to resolvable, including the ones computed in this document. The delivery side, measured on tens of thousands of impressions rather than tens of payers, resolves a 0.5pp CTR difference on a single night.

Read the requirement as an order of magnitude, not a figure. CV falls as the horizon lengthens and more installs have paid, so the 20% requirement runs ~18,600 at H = 6, ~17,500 at H = 12 and ~11,600 at H = 24. Beyond that the account cannot measure it — only 40 vid_ww_* installs are 48 h old and one of them has paid. It does not drop below ~10,000 per arm anywhere in that range, and 732 is 4–6% of it throughout.

Horizon coverage is the other half of this. A ROAS quoted at a horizon says nothing about the spend sitting in cells too young to have reached it:

H = 1H = 6H = 12H = 24
bikini0.87 (100%)1.00 (100%)1.18 (95%)1.43 (1%)
bodysuit1.02 (97%)1.51 (75%)1.64 (63%)1.98 (40%)
pool50.89 (100%)1.02 (100%)1.02 (100%)1.22 (100%)

Parenthesised is the share of that ad's lifetime spend old enough to have reached the horizon. bodysuit's whole-life figures describe 63% of its money at H = 12 and 40% at H = 24, because most of its spend is recent; bikini's H = 24 number describes 1% and is not a number at all. Only the R4 window above has both arms at full coverage, which is why it is the only comparison quoted.

So creative selection runs on CPM and CTR. ROAS sizes the account, not the creative.

6The account, event by event

Budget changes are not the only events. An asset starting or stopping delivery changes the auction for every other asset in the ad set, so a window straddling one compares two different competitive fields. Cutting at the union of both — every budget step, every ad's first and last delivering hour, every gap where the optimiser parked one — gives the run's actual shape.

All three creatives were live from launch. The only operator switch in this window is bikini at 06:55 on 08-01; bodysuit followed at ~17:00, past the data cut. Everything else below is Meta reallocating inside the CBO ad set — an ad at zero delivery is still live and still competing, it simply got no budget that hour. That distinction is the Gen 5 finding: the disjoint delivery windows were the optimiser's doing, not anyone's hand. Operator on/off is in no export either; Ad delivery is stamped at export time, reading active for the same hour in a file pulled while the ad ran and inactive in one pulled after it was paused.

window PTbudgetwhat changedadspendinstCPMCTRIndiaaudience value
07-30 15:00–16:00$150 ←Meta starts pool5pool5$4.8814.8410.9%0%$0.277
07-30 16:00–20:00$150Meta starts bikini, bodysuitbikini$3.52811.548.2%11%$0.359
bodysuit$4.13813.0415.3%24%$0.657
pool5$23.914313.0713.4%14%$0.392
07-31 00:00–07:19$250Meta stops bikini and pool5bodysuit$107.434213.9612.0%18%$0.576
07-31 07:19–17:21$500 ←bodysuit$187.161510.9610.9%30%$0.445
07-31 18:00–08-01 04:10$875 ←Meta resumes bikini, 39 min after the stepbikini$286.27328.337.6%32%$0.483
bodysuit$144.14886.167.9%39%$0.504
08-01 04:10–07:00$500 ←bikini switched off (operator, 06:55)bikini$14.6517.558.4%35%$0.406
bodysuit$10.5436.628.4%28%$0.316
08-01 07:00–17:00$500bodysuit$254.47708.748.8%29%$0.468

marks the segment in which that budget took effect. Audience value is the segment's own country mix priced at pooled country rates — the only column not confounded by which ad earned it.

Meta parked two of three creatives on its own, at $150/day, within six hours of launch — and brought bikini back 39 minutes after the $875 step. No one touched the ads. That is why the three arms have disjoint delivery windows, and why every per-asset figure computed across them describes Meta's schedule rather than the creative.

What the regions did. India tracks the budget and reverts, exactly as #47 found: 18% at $250 → 30% at $500 → 39% at $875 → 29% back at $500, holding the creative (bodysuit) fixed throughout. The composition effect is unambiguous.

What it was worth is a different matter. Across the whole run the audience's value moves between $0.316 and $0.657 — but the extremes are the two thinnest segments (43 and 38 installs), and across the four large ones it sits in $0.445–$0.576, a spread of about 25%. Over the same span CTR falls from 12.0% to 7.6% and comes back to 8.8% — a 37% swing on a quantity measured over tens of thousands of impressions rather than hundreds of installs.

So the region story is the smaller half of the story. The budget does move the geography, the geography does move the audience's value, and that channel is worth roughly a quarter at the extremes. The larger and far better-measured effect is that at $875 Meta bought cheaper impressions from much lower-intent people inside the same countries — CPM fell 60% while CTR halved, so cost per install rose anyway.

The one real off-switch. After bikini was switched off at 06:55 on 08-01, bodysuit alone at $500 shows CPM $8.74 against $6.62 in the preceding both-live segment, on a near-identical India share (29% against 28%). It inherited the budget bikini had been absorbing and had to buy further up its own cost curve. ⚠ That segment is also daytime where the previous one is overnight, and §6 shows hour-of-day moves CPM on its own, so this is consistent with a real effect rather than evidence of one — the clean version needs an off-switch held at matched clock hours.

7The budget, re-examined

#47 established that the budget moves the country mix and read the account's decline off that. The mix moves; the question left open was what it is worth. Pricing every regime's pooled mix at account-wide country values:

regimebudgetinstallsIndiaexpected RPIexpected CPIexpected ROAS from mix alone
R1$15020114.9%$0.433$0.2941.48
R2$25036219.6%$0.561$0.3361.67
R3$50065229.1%$0.493$0.3261.51
R4$8751,21634.6%$0.497$0.3241.53
R5$50088529.9%$0.452$0.3121.45

The mix swings hugely and is worth very little. India runs 14.9% to 34.6% across the regimes — a 20pp swing — while the value of the resulting audience spans only 1.45 to 1.67, and the $875 regime prices out 6% above the $500 one. The countries that replace India are not systematically better: India's revenue per install ($0.410) is only 12% below account average, and its cost per install ($0.347) is 7% above, so it is a worse market by about a fifth, not by a multiple.

The budget's real damage is visible when the same creative is held on the same clock hours:

bodysuit, 00:00–03:59 PTCPMCTRIndiaexpected RPIobserved RPICPIROAS
R2, $250, 07-31$14.0113.09%19.1%$0.506$0.501$0.2921.72
R4, $875, 08-01$5.606.35%42.9%$0.443$0.330$0.3350.98
change−60%−51%+23.8pp−12%−34%+15%−43%

Tripling the budget did not buy the same people dearer — it bought far more, far cheaper impressions from a far lower-intent audience. CPM fell 60% while CTR halved, so cost per install still rose 15%. Country mix accounts for only −12 of the −34 percentage points of lost revenue per install; the rest is audience quality within the same countries, which is what the CTR collapse is measuring. Country targeting would recover about a third of what the budget step cost.

The revenue columns in that table are not reliable and the delivery columns are. An equal-budget control — same ad, same clock hours, $500 on both days — moves observed RPI by −75% and ROAS by −73%, as much as the budget contrast, while CTR moves only −7% and India only +2.7pp. At ~150-install cells the revenue side is noise. The CPM, CTR and mix figures are the findings here; the ROAS figures are illustration.

8pool5, before it is read again

pool5's 3.01× was a projected cohort ROAS on 151 installs. Two days on, it can be checked against realised money instead of re-argued. Its 07-30 cohort is now essentially complete — 100% of its revenue arrived within 48 h, 50% within the first hour.

pool5, 07-30 cohort
installs151
payers11 (7.3%)
revenue to date$92.23
Meta full-day spend$36.62
realised return2.52×
Meta's own day ROAS1.67×
largest single account26% of cohort revenue

Not a mix artifact. The $150 regime it ran in prices at 1.48 expected ROAS against the current $500 regime's 1.45 — a 2% difference. The regime was not unusually generous, so the number is not explained by where it was shown.

Also not a result. Eleven payers, one of whom — a Cambodian account making four purchases — is a quarter of the revenue. Against §4's table, 151 installs resolves nothing at all. pool5 has been running alone at $500/day since ~17:00 on 08-01 and none of that run is in this data; the Telegram reconstruction stops at the 08-01 16:38 export.

9What this changes

  1. The cut of ww_bikini_10s was right, for a reason nobody gave at the time. Not 0.82× (a calendar-day artifact), not 1.13× vs 1.45× (unresolvable), but a 38% CPM premium at equal CTR, dearer in nine of the ten shared hours, stable across a ninefold spend range. That is measurable in a night.
  2. Stop ranking creatives by ROAS. At $0.35 CPI, separating two creatives on revenue costs $12,000 of media for a 20% difference. Rank on CPM and CTR; use revenue to decide whether the account is worth running, which is a question about one number, not a comparison between two.
  3. #47's mechanism is right in direction and wrong in weight. The budget does move the country mix, near-deterministically and reversibly. But that channel is worth ≤15% of expected ROAS, while the budget step actually halved CTR. The dominant harm is within-country audience quality. The budget ceiling is still the right control — for a bigger reason than the one recorded.
  4. The remaining open question is not about these two creatives. It is whether the account can buy at $500/day without the CTR collapse that $875 produced, which is a delivery-side question and therefore answerable in a night.

10Reproduce

Everything above is one command. The unit is (ad, delivery hour); every table is a selection and a grouping over it (ad_ops/asset_estimator/cohorts.py, views in views.py).

cd C:/Projects/VideoGenAI_Related
L=_runs/2026-08-01_account_recheck

python -m ad_ops.asset_estimator hourly     --meta ad_ops/_data/b1_meta_exports/B1-Ads-by-hour-Jul-30.csv            ad_ops/_data/b1_meta_exports/B1-Ads-Jul-31-2026-Jul-31-2026.csv            $L/ads_hourly_aug1.csv     --lum $L/lum --fx $L/fx_rates.json     --ads "ww_bikini_10s,ww_bodysuit_10s,ww_pool5" --horizon 12

It refuses rather than guesses in two places: an export with no Time of day column raises (folding a day into one cell is the error the unit exists to remove), and a cohort younger than the horizon is excluded and its spend reported as missing coverage rather than truncated into the total.

The scaling test (§3), the matched-clock-hour budget contrast (§6) and the pool5 cohort reconstruction (§7) remain one-off probes in _runs/2026-08-01_regime_matched/; they answered a mechanism question once and are not part of a cycle.

Data gaps this ran into, in the order they bind:

Internal — noindex. Not for distribution outside the team.