Back to blog
Insights / Research / Cover study
Research · Cover study

How we built the Content Conversion Standard.

Three years. 100M+ images. 10T+ retail performance data points. Inside the benchmark every CPG catalog now scores against — what we measured, what we threw out, and why a single number decides whether content ships.

EC
JR
Elena Choi & Jordan Reyes Head of Research · Principal Researcher
12 min read Mar 04, 2026 Vol. 04 · No. 09

01 · The premise What "ready" means, measurably.

For most of digital commerce's history, "is this image good?" has been a question answered by the loudest person in the room. Brand managers had taste. Creative directors had instincts. Compliance had checklists. Nobody had a number — and so nothing shipped against a standard. It shipped against an opinion.

When we started this project in early 2023, the premise was simple: convertibility is observable. Shoppers behave in patterns. Those patterns leave signal across hundreds of measurable dimensions — pack-front clarity, claim placement, visual weight, foreground contrast, the geometry of attention. Add enough of those signals together, weighted against actual conversion outcomes, and you get a number that means something the same way a credit score means something.

The harder question was what to call ready. Ready for what? Not ready to win a design award. Ready to convert at first impression against the population of shoppers who will see it on a phone, in two-tenths of a second, on a retailer page they didn't choose. That definition was the wall we kept hitting back to as we built every part of the standard.

100M+ Images analyzed
10T+ retail performance data points trained on
20+ Countries observed

We built it the way you'd build a benchmark: collect everything, kill what doesn't predict, repeat for years. The corpus that survived is what now powers every Vizit Score.

02 · The signals The signals we kept — and the ones we threw out.

We started with 1,140 candidate signals. Some were obvious: where the product sits in the frame, how much of the pack you can read at 320px wide, how many visual elements compete for attention. Others were more speculative — color temperature, claim density, the angle of the product against the cropline. Every signal had to earn its place by predicting outcomes on a held-out test set across categories the model hadn't seen.

By year two we were down to 184. By the time we shipped, 87. The cuts were sometimes counterintuitive. A few examples of what survived, and what didn't:

  • Survived: Pack-front legibility at thumb size. Product centrality. Foreground-to-background contrast ratio. Claim hierarchy. Cropline geometry.
  • Cut: Brand color saturation. Logo size as a percentage of frame. Photographic style. Image resolution above 1080px. Background "premium-ness."

The cuts were the more interesting result. Photographic style — the dimension creative teams argue about most — turned out to be a near-zero predictor of conversion once pack-front legibility was controlled for. The shopper, it turns out, does not care whether the bag is lit moodily. The shopper cares whether they can tell what's in it.

87 Fig. 1 · Final signal set

Fig. 1 The 87 signals that survived to v1, grouped by dimension. Pack-front clarity carries the heaviest weight; photographic style carries none.

Once you let outcomes do the talking, the things you argued about for a decade stop mattering. The things you'd never thought to argue about turn out to be everything.

— Jordan Reyes, Principal Researcher

The model that came out of this is calibrated against actual shopper behavior — not the opinion of any one judge, not the conventions of any one brand. That's what makes the score portable across categories, retailers, and geographies. Pack-front clarity in pet food and pack-front clarity in personal care are measured the same way, because shoppers process them the same way.

03 · The verdict Why a single score wins over a dashboard.

Most of the conversion-intelligence work we saw in the wild before Vizit was dashboard-shaped. A page would show you twelve metrics, ten meters, three trend lines, and a colored heatmap, and leave you to figure out which one mattered. We tried this. It is worse than nothing. People who had to make a ship/no-ship decision in fifteen minutes ignored the dashboard entirely.

So we collapsed it. The Vizit Score is a 0–100 number on each asset and a 0–100 number on the page that contains it. Above 80 is conversion-ready. Below 60 is leaving revenue on the table. Between is the optimization zone. That's the entire decision surface. Underneath, the score decomposes — you can open any single asset and see exactly which signals dragged it down — but the headline is a number, and the number is what ships.

0–100 Fig. 2 · The score in production

Fig. 2 A single number compresses 87 signals into one decision surface. Diagnostics live below the headline, not in place of it.

The result, in production, is that teams stop arguing about whether something is good. They argue about how to get it above 80, which is a much more productive argument. The work moves from taste to optimization, and once you've made that move, every part of the catalog gets faster.

We'll write more in the coming months about specific signal dimensions — pack-front clarity gets its own piece next month, with the dataset attached. For now, the headline: the standard exists, it's portable, and it scores against shopper behavior rather than anyone's instincts. That's the work.

See where your catalog sits against the standard.

Score your content →