The problem

An e-commerce site runs on many systems at once. Search ranks the products, retail media places the ads, and merchandising, personalization, pricing, and inventory each make their own calls. Every one of these systems is best-of-breed, and every one does its own job well.

They're also well integrated — the data moves between them, and they fire in the right order. But one thing is missing: no layer decides who wins when two systems make conflicting calls on the same page. Integration moves the data. Orchestration sequences the workflow. Coordination — the layer that arbitrates the moment two correct decisions contradict each other in front of the customer — doesn't exist in any retailer's stack today. (We unpack that distinction in Coordination Is Not Orchestration.)

When that arbitration is missing, you get a coordination failure. A promotion surfaces on an item that is out of stock. Three discounts stack on one product with no rule for which should win. The same product shows up twice on one page. Each system reports success, and the failure appears only on the rendered page the customer sees, which no single system is watching.

In our State of Retail Coordination scan of 15+ e-commerce surfaces, ten failure patterns recurred often enough to name, and the most common appeared on more than half the sites scanned. Across billions of sessions, the cost compounds into an estimated 2–5% of GMV — $20M to $50M a year for a $1B retailer. That figure is an operating estimate that sizes the market; the measurement below replaces it with a number for a specific retailer.

Our approach is three steps; this article is about the second

Illustration 1 — Three-step engagement, Step 2 highlighted STEP 1 Diagnostic Surface coordination failures by reading the rendered page. No integration required. THIS ARTICLE STEP 2 Measurement Put a dollar figure and a range on each failure, from the retailer's own session logs. STEP 3 Resolution Test the highest-value fixes in controlled experiments and verify the impact. Illustration 1 — The engagement in three steps; this article covers Step 2.

Our engagement runs in three steps, none of which requires integration:

1. Diagnostic. Our scanner reads a retailer's live site the way a customer sees it and surfaces the specific coordination failures.

2. Measurement. We put a defensible dollar figure on each failure, using the retailer's own data. (This article.)

3. Resolution. We design the tests, apply the fixes, and measure the lift.

Step 1 shows the failures exist. Step 2 turns "here is a problem" into "here is the dollar figure."

One failure, followed to a dollar figure

Take one common failure: a parity conflict. The same product appears in both organic search and a sponsored ad on the same page. Search says it is the most relevant result, and the ad system says it is the highest bid. Both are right. The customer now sees the same product twice instead of discovering a second one, and the retailer loses a slot it could have used to surface something else the shopper might have bought.

The fix is not to kill the ad — the retailer earns real retail-media revenue on it. The fix is coordination: keep the sponsored placement, and swap the organic duplicate for the next most relevant product. Same ad revenue, one more product discovered, a better page. The measurement question is the annual dollar value of that recovered slot.

Why the number is hard to get right

The obvious way to measure it is to compare sessions where the failure appears against sessions where it does not. That comparison will mislead you — and not by a little.

These failures do not appear at random. They cluster exactly where demand is already unusual: retail media chases the most popular searches, and out-of-stock promotions pile up on cold, slow-moving items. So a comparison of "sessions with the failure" against "sessions without" is not measuring the failure — it is measuring the demand that attracted it. The naive comparison confounds1 the two, and it can hand you a number with the wrong sign, telling you a failure helps when it hurts.

1 Confounding is when something you did not account for drives both things you are comparing, so you credit the wrong cause. A simple example: app users convert better than non-app users, but that does not mean the app causes the lift. Loyal customers are the ones who install the app, and they would buy more anyway. Loyalty drives both app use and higher conversion — so the app looks more powerful than it is until you separate the two.
Illustration 2 — The confounding trap The raw comparison every session with the failure, against every session without Sessions where the failure appears 4.0% Sessions where it does not 3.4% +0.6 points Read literally, the failure looks helpful. The sponsored placement chases the most-searched products, so those sessions were always going to convert well. The same-search comparison the same search, with the failure and without it Same search, failure present 4.0% Same search, failure absent 4.5% −0.5 points Holding the search fixed removes the demand that drew the ad there. The failure costs half a point of conversion on the pages it touches. Illustrative arithmetic, for the parity example. Figures chosen to show the mechanism, not measured. Illustration 2 — The confounding trap: the raw and same-search comparisons disagree about the sign.

Our approach: an evidence ladder

We estimate each failure's impact several ways, and we rank those methods by how closely each approximates a randomized experiment.

Illustration 3 — The evidence ladder CLOSER TO A RANDOMIZED EXPERIMENT ↑ bar length = strength of the evidence Randomized assignment Experiment A randomized A/B test: customers are assigned to one version or the other at random. A platform change outside our control Natural experiment Before-and-after a real platform change, against a comparison group. The page's own search context Same-search comparison The same search, with the failure and without it, so demand cancels out. THE CRITICAL LINE Above the line: demand cancels by construction, including demand that was never measured. Below it: only measured demand is removed, and some demand intensity is always unobserved. Adjustment for measured differences Look-alike matching & ML Statistically adjust for the measurable differences between sessions. Not a measurement of this site Published-research benchmarks — sanity band What independent studies find. We check the answer against it; we do not estimate from it. Illustration 3 — The evidence ladder: the methods we use, ranked by how closely each approximates a randomized experiment.

At the top is a randomized A/B test, the gold standard. Below it, a natural experiment: a platform change we did not cause. Below that, comparing the same search with and without the failure. At the bottom, statistical adjustment for the differences we can measure.

The line across the middle of the ladder is the one to read closely. Above the line, the demand that attracted the failure cancels out by construction. Below the line, we can only remove the demand we did measure and put into a model. We therefore anchor rather than average: the headline number comes from the strongest rung the data supports, and the weaker rungs corroborate and bound it.

The same-search comparison rests on one assumption we can actually check: that which sessions show the failure is determined by the ad auction and the platform, not by targeting the individual shopper. We test the assumption before we report a number, and if it fails, we demote the estimate and widen the range.

Every figure we give is a range, not a single point: wide where the evidence is weaker, narrow where it is stronger.

From effect to dollars

In plain terms: count how often the failure happens, multiply by what it costs each time, add back the loyalty revenue recovered by not driving customers off, then subtract any ad revenue given up by fixing it.

Net Value (annual) = (β × N × AOV) + ΔLTV − ΔAd revenue
β — the per-occurrence effect on conversion; in Illustration 2 it is the 0.5 percentage-point drop measured by the same-search comparison
N — how many sessions the failure touches in a year
AOV — average revenue per converted session
ΔLTV — the loyalty revenue recovered by not driving customers away
ΔAd revenue — the retail-media dollars given up if the fix removes a paid placement

Every term is stated in revenue, so the whole figure reads directly against GMV. The last term is easy to miss. For a parity conflict the fix changes the organic slot, so the ad revenue stays. Where a fix does remove a paid placement, we net that lost revenue out explicitly — we do not count money the retailer did not make.

We report two versions of the number: the conversion piece alone (our most defensible figure), and the full figure with loyalty and the ad trade-off included. That way you can see the floor and the fuller estimate separately, rather than one blended total.

Illustration 4 — Net Value waterfall $60M +$10M −$15M $55M Conversion + ΔLTV − ΔAd Net Value Illustrative figures only. Illustration 4 — Net Value: conversion recovered, plus retention, minus the ad-revenue trade-off.

We report each failure type on its own, so the output reads as a ranked menu of opportunities, each with a dollar figure and a range, rather than one opaque total. The biggest, best-evidenced line items are where the conversation starts.

Illustration 5 — Impact by failure type FAILURE TYPE ANNUAL IMPACT ($) % OF REVENUE PREVALENCE 90% RANGE Parity violation $30M 1.0% 11% $15M – $40M Duplication $5M 0.15% 8% $0M – $15M Out-of-stock promotion $20M 0.65% 5% $10M – $30M Ad density (to 0.10) $15M 0.45% 19% $0M – $40M Every failure is measured the same way. Illustration 5 — Impact by failure type; four of the many failure types we detect. Illustrative figures only.

From a number to a decision

Coordination failures are real, common, and costly — and the cost can be measured, defensibly, without integrating anything. The number holds up because of how it is built: a comparison that cancels out demand, an assumption that is tested rather than assumed, a range instead of a false-precise point, and validation that includes attempts to break the result.

The output is a ranked list of failures, each with a dollar figure and a confidence range. That list is where a retailer's coordination conversation can start.

Back to Insights