← Back to Marketing Analytics Consultants

Case study · Catalog data and marketplace listings

Three hundred thousand products nobody could describe

Two decades of catalog across a website and three marketplaces. Every product had a title and almost none had a date, a place or a subject anyone could search on. The information was not lost — it had never been connected.

1. Three hundred thousand products nobody could describe

A catalog of roughly three hundred thousand items, built up over two decades, listed across a website and three marketplaces. Every product had a title. Almost none of them had a date, a place, or a subject that a customer could search on.

The information had not been lost. It had never been connected.

299,922products in the catalog
96.9%matched to their source records
286customer-facing categories built
2.1%of titles met the written spec

2. The join that unlocked everything

Each product carried an image file. Each image filename carried an identifier from the archive the picture originally came from. That identifier was the key to a public catalog record holding the date, the place, the subject and the creator — everything the product listing lacked.

Figure 1

The data was already there

Recovering what the listings had lost Product listing a title and an image file The filename carries an archive ID Source record date, place, subject 273,807 products matched to their source records. 265,215 of them — 96.9% — resolved. That recovered a date for 184,362 items, a US state for 36,897, and a country for 14,174. None of it was new data. All of it was already there, in the filename, unread for fifteen years.

Two different filename conventions were in use across the catalog. One embedded a direct archive identifier; the other embedded a sequence number that turned out to be the row position in that collection's index. Both were decoded and verified against the source records.

What a 96.9% match buys

Geography for tens of thousands of products, dates for the majority of the catalog, and a subject vocabulary in plain English rather than archival jargon. From that, 286 customer-facing categories — named the way a customer would type them, not the way a cataloger filed them.

3. Measuring what had actually shipped

A written specification for product titles already existed and it was sensible: subject, then year or era, then item type, and never internal codes. We measured the live catalog against it. Two point one percent complied.

Figure 2

The distribution that gave it away

How often each of eight generic title suffixes appeared across 299,922 products 27,183 26,793 26,687 26,655 26,579 26,555 26,513 26,223 Every one within half a percent of the others. Editorial decisions do not distribute like that.

Suffix frequency across a complete 299,922-row catalog export. Together these eight phrases account for 88.8% of all titles.

What had shipped was the original archive record with a generic phrase attached. Eight different phrases, each appearing between 26,223 and 27,183 times. That is a loop, not editing.

The rest of the damage, counted

Over 60 characters: 166,819 titles. Over 100 characters: 75,947. Truncated mid-word: 23,914. Still carrying an internal prefix: 29,062. Size baked into the title: 29,205. Character-encoding corruption: 5,903. Each of those is a specific, countable defect, and each was fixed by rule rather than by hand.

Rebuilt titles were generated for the curated tier following the original specification. Eighty-four percent carry a year, against 7.3% in the catalog as it stood.

4. The marketplace revision, and the number that was wrong

Alongside the catalog work, the marketplace listings were repositioned — keywords, bullets and metadata retargeted from obscure archival wording toward the language buyers actually use. At the same time, prices moved.

The early read on the price change was that it had worked. It had not.

Figure 3

Same change, opposite conclusion

The same price change, read two ways From the notification emails Orders up 149% Revenue up 39% Verdict: it is working From the profit log, with cost Framed volume up 35%, price down 24% Profit per framed order $86.90 to $32.20 Verdict: minus $2,475 a year The emails were accurate. They simply carried no cost of goods, so they could only ever report the half of the story that looked good.

The first reading came from marketplace order notifications, which report units and revenue. The second came from a complete order-and-profit log covering 317 real orders with cost of goods and profit recorded per order.

Why this is the most important finding on this page

The optimistic reading was ours, and it was wrong, and we said so. Marketplace notifications carry revenue and no cost. A discount that lifts volume will always look like success in a revenue-only view. The recommendation that followed — restore the framed price, keep the unframed tier — was the opposite of what the first analysis implied.

5. And one more correction

An early estimate of catalog composition was built from a sample of a few thousand products pulled from across the catalog. When a complete export arrived, the sample turned out to have clustered badly.

SegmentEstimated from the sampleActualError
Curated tier14,164155−99%
Poster collection23,3001,095−95%
Photochrom collection25,7004,455−83%
Survey collection25,66049,458+93%
What we did about it

Withdrew the estimates and every conclusion that rested on them, including a strategic argument that had been built on the curated tier being roughly the size of a competitor’s entire catalog. Paged API sampling returns records in clusters, not at random, and a sample drawn that way is not a sample. The complete export is now the only source used for any count.

6. What this case is meant to show

The usual approachWhat was done here
Rewrite listings by hand or by templateRecover the real metadata first, then write from it
Trust that the spec was followedMeasure compliance across the full catalog and count the failures
Read revenue reportsRead cost with them, because revenue alone always flatters a discount
Estimate from a sampleGet the complete export, and withdraw the estimate when it disagrees
The transferable idea

Most catalogs of any age contain more information than they expose. It is usually sitting in a filename, a legacy field, a supplier feed or an export nobody has opened. Recovering it is cheaper than creating it, and it is almost always the first thing worth doing.

Marketing Analytics Consultants

Sitting on a catalog nobody can search?

We find out what data you already have, recover what has been lost between systems, measure what actually shipped against what was specified, and tell you which of it is worth money.

Prefer email? info@marketinganalyticsconsultants.com