Case study · Catalog data and marketplace listings
Two decades of catalog across a website and three marketplaces. Every product had a title and almost none had a date, a place or a subject anyone could search on. The information was not lost — it had never been connected.
A catalog of roughly three hundred thousand items, built up over two decades, listed across a website and three marketplaces. Every product had a title. Almost none of them had a date, a place, or a subject that a customer could search on.
The information had not been lost. It had never been connected.
Each product carried an image file. Each image filename carried an identifier from the archive the picture originally came from. That identifier was the key to a public catalog record holding the date, the place, the subject and the creator — everything the product listing lacked.
The data was already there
Two different filename conventions were in use across the catalog. One embedded a direct archive identifier; the other embedded a sequence number that turned out to be the row position in that collection's index. Both were decoded and verified against the source records.
Geography for tens of thousands of products, dates for the majority of the catalog, and a subject vocabulary in plain English rather than archival jargon. From that, 286 customer-facing categories — named the way a customer would type them, not the way a cataloger filed them.
A written specification for product titles already existed and it was sensible: subject, then year or era, then item type, and never internal codes. We measured the live catalog against it. Two point one percent complied.
The distribution that gave it away
Suffix frequency across a complete 299,922-row catalog export. Together these eight phrases account for 88.8% of all titles.
What had shipped was the original archive record with a generic phrase attached. Eight different phrases, each appearing between 26,223 and 27,183 times. That is a loop, not editing.
Over 60 characters: 166,819 titles. Over 100 characters: 75,947. Truncated mid-word: 23,914. Still carrying an internal prefix: 29,062. Size baked into the title: 29,205. Character-encoding corruption: 5,903. Each of those is a specific, countable defect, and each was fixed by rule rather than by hand.
Rebuilt titles were generated for the curated tier following the original specification. Eighty-four percent carry a year, against 7.3% in the catalog as it stood.
Alongside the catalog work, the marketplace listings were repositioned — keywords, bullets and metadata retargeted from obscure archival wording toward the language buyers actually use. At the same time, prices moved.
The early read on the price change was that it had worked. It had not.
Same change, opposite conclusion
The first reading came from marketplace order notifications, which report units and revenue. The second came from a complete order-and-profit log covering 317 real orders with cost of goods and profit recorded per order.
The optimistic reading was ours, and it was wrong, and we said so. Marketplace notifications carry revenue and no cost. A discount that lifts volume will always look like success in a revenue-only view. The recommendation that followed — restore the framed price, keep the unframed tier — was the opposite of what the first analysis implied.
An early estimate of catalog composition was built from a sample of a few thousand products pulled from across the catalog. When a complete export arrived, the sample turned out to have clustered badly.
| Segment | Estimated from the sample | Actual | Error |
|---|---|---|---|
| Curated tier | 14,164 | 155 | −99% |
| Poster collection | 23,300 | 1,095 | −95% |
| Photochrom collection | 25,700 | 4,455 | −83% |
| Survey collection | 25,660 | 49,458 | +93% |
Withdrew the estimates and every conclusion that rested on them, including a strategic argument that had been built on the curated tier being roughly the size of a competitor’s entire catalog. Paged API sampling returns records in clusters, not at random, and a sample drawn that way is not a sample. The complete export is now the only source used for any count.
| The usual approach | What was done here |
|---|---|
| Rewrite listings by hand or by template | Recover the real metadata first, then write from it |
| Trust that the spec was followed | Measure compliance across the full catalog and count the failures |
| Read revenue reports | Read cost with them, because revenue alone always flatters a discount |
| Estimate from a sample | Get the complete export, and withdraw the estimate when it disagrees |
Most catalogs of any age contain more information than they expose. It is usually sitting in a filename, a legacy field, a supplier feed or an export nobody has opened. Recovering it is cheaper than creating it, and it is almost always the first thing worth doing.
Marketing Analytics Consultants
We find out what data you already have, recover what has been lost between systems, measure what actually shipped against what was specified, and tell you which of it is worth money.