Most of what we do is practical and finishes in weeks. This page is about the other part — the questions worth answering properly, and the methods behind them.
Why this is a discipline
Astronomy, climate science and genome research all became quantitative disciplines for one reason: the data outgrew what a person could hold in their head. Consumer behavior crossed that line about a decade ago, and almost nobody has treated it that way since.
One zettabyte, written out
1,000,000,000,000,000,000,000 bytes
A one followed by twenty-one zeros. The world now creates 181 of these every year.
Put together, those four machines hold about 274 kilobytes — roughly a quarter of a megabyte. They landed people on the Moon, left the solar system, and started personal computing. The world now produces more data every second than all four could have stored in several trillion lifetimes.
Sources: Apollo Guidance Computer, 2,048 words of erasable core memory and 36,864 words of core rope memory — approximately 4 KB RAM and 72 KB ROM (MIT Instrumentation Laboratory; NASA). Voyager 1 and 2 combined onboard memory, NASA Jet Propulsion Laboratory. Macintosh 128K, 128 KB RAM, Apple Computer, January 1984. Global data figure: IDC Global DataSphere via Statista, 2025.
Analytics platforms, search consoles, crawl tools, marketplace search censuses. The measurement problem is not a lack of instruments.
Three in four marketers never measure return using their own data. In physics that would be an unpublished experiment.
Not better tools. A written instrument, a baseline recorded before anything changes, and a log of what else moved.
The scale nobody can eyeball
Marketing became a measurement problem because nobody’s instincts operate at this size.
| What | How much | |
|---|---|---|
| Search | ||
| Google searches, per year | 5+ trillion | |
| Google searches, per day | ~13.7 billion | |
| Share of searches never seen before | ~15% | |
| Searches ending without a click | ~58% | |
| Video | ||
| Hours of YouTube watched, per day | 1 billion | |
| Hours of YouTube watched, per year | ~365 billion | |
| Hours of video uploaded to YouTube, per minute | 500+ | |
| Videos on YouTube | ~800 million | |
| TikTok videos watched, per minute | ~167 million | |
| Social | ||
| Social media users worldwide | ~5.2 billion | |
| Hours spent on social media, per day, worldwide | 14+ billion | |
| Average time on social, per person, per day | 2 hr 31 min | |
| Instagram Stories created, per minute | ~695,000 | |
| WhatsApp messages sent, per minute | ~41 million | |
| Messages and data | ||
| Emails sent, per day | ~361 billion | |
| Data created, captured or copied, per day | ~403 million terabytes | |
| Data created, per year | ~181 zettabytes | |
| Data moving across the internet, per minute | ~1.74 million gigabytes | |
| Share of the world’s data created in the last two years | ~90% | |
Twenty years ago a good marketer could hold the whole picture in their head and be roughly right. That is no longer possible, and pretending otherwise is how money gets wasted. The question is not whether your instincts are good — mine are, after twenty-five years. It is whether anything at this scale can be judged by instinct at all. It cannot. It has to be measured, and the measurement has to be done properly.
Figures compiled from Google’s own disclosures, DataReportal, Statista and platform reporting, 2025–2026. Where sources disagree, the more conservative figure is shown.
Our engagements are operational. We look at what a business’s marketing is doing, explain it plainly, and repair what is broken. Weeks, not years.
But the work generates data, and some of the questions it raises deserve a longer answer than an engagement allows. Those become research — done alongside graduate work, on our own data and on client data with permission, over the sort of timescale real research takes.
Worth being plain about this. We do not sell research studies. A proper attribution study or a causal analysis is a year or more of work and belongs in a journal, not in a consulting proposal. What the training does is make the everyday work correct — it is why we know when a number is wrong, when a comparison is invalid, and when somebody is fooling themselves with their own data.
Sprout Social surveys social marketers every year — a thousand or more each time. One question asks what they actually do with the data their own campaigns produce. The answer has barely moved.
| Sprout Social Index | Use social data to measure return |
|---|---|
| 2020 | 23% |
| 2021 | 15% |
| 2023 | 23% |
Three measurement cycles across six years, and the share never rises above a quarter. One year could be a bad sample. A flat line is structural.
The same surveys found 56% using social data to understand their audience. So the data is being collected, and it is being read. It stops just before the question of what came back.
The more recent research names the obstacle plainly. More than half of marketing leaders say poor integration between their social tools and the rest of their systems is the single biggest reason they cannot see what social contributes to the business. Around one organization in ten can turn an insight into action inside a few hours.
That is not a discipline problem, and it is not a talent problem. It is a plumbing problem — which is the encouraging part, because plumbing can be fixed.
The pattern shows up in ordinary technical detail long before anyone opens an analytics account. Two tag-manager containers loading on the same page, so every session is counted twice. Links shared from a site with no tracking parameters on them, so the traffic arrives labeled “direct” and the credit disappears. A tracking token still loading years after the service behind it shut down. None of it is exotic, and all of it is visible from the public page source.
Sprout Social Index 2020, 2021 and 2023, each surveying more than 1,000 social marketers; Sprout Social return-on-investment research, 2026.
These are published, peer-reviewed and checkable. We include them because they show what the difference between reporting and analysis is actually worth — and because the first one is the reason we take measurement seriously rather than taking a dashboard at face value.
| Company | The question | What the analysis found |
|---|---|---|
| eBayLarge-scale field experiment | Is our paid search advertising actually generating sales, or is it buying customers who were coming anyway? | Brand-keyword ads had no measurable short-term benefit. Across the wider program, returns were a fraction of what the non-experimental reporting showed, and average returns on paid search came out negative.1 |
| NetflixMatrix factorization | How do we recommend accurately when most customers have rated almost nothing? | Latent-factor models outperformed the nearest-neighbor methods the industry had been using, and could absorb implicit behavior and changes in taste over time.2 |
| AmazonItem-to-item collaborative filtering | How do we recommend across an enormous catalog without the computation growing with the customer base? | Restructuring the problem around item similarity rather than user similarity made recommendation scale with the catalog instead of with the number of shoppers.3 |
| eBay, againControlled shutoff | Why did the reporting say something different? | Because clicks and purchase intent are correlated. People who were already going to buy click the ad on the way. Attribution credits the ad; the experiment shows it changed nothing.1 |
The eBay result is the one to sit with. A company with excellent analysts, complete data and a mature attribution system was measuring a return that the experiment showed was not there. Nothing was broken and nobody was careless. The reporting was answering “what happened” and the question was “what did we cause” — and those need different methods.
1. Blake, T., Nosko, C., & Tadelis, S. (2015). Consumer Heterogeneity and Paid Search
Effectiveness: A Large-Scale Field Experiment. Econometrica, 83(1), 155–174.
doi:10.3982/ECTA12423
2. Koren, Y., Bell, R., & Volinsky, C. (2009). Matrix Factorization Techniques for
Recommender Systems. Computer, 42(8), 30–37. doi:10.1109/MC.2009.263
3. Linden, G., Smith, B., & York, J. (2003). Amazon.com Recommendations: Item-to-Item
Collaborative Filtering. IEEE Internet Computing, 7(1), 76–80.
We work from a written decision map rather than from habit. It runs on four things: what kind of question is being asked, what the outcome variable actually looks like, how the data is structured, and how interpretable the answer has to be. A few of the more common cases:
| What a business actually asks | What we reach for |
|---|---|
| What is a customer worth over their lifetime? | Probabilistic buying models — Pareto/NBD or BG/NBD for repeat rate and dropout, Gamma-Gamma for spend |
| Who are our real customer segments? | Latent class analysis or finite mixture models. Not k-means, which imposes a geometry survey data doesn’t have |
| Which channel actually earned the sale? | Markov chain or Shapley-value attribution rather than last-click. Bayesian structural time series for media mix |
| Did our spend cause the lift? | Geo experiments, instrumental variables or regression discontinuity. A predictive model cannot answer this |
| Who is about to leave, and when? | Survival analysis — Cox proportional hazards or Kaplan-Meier, because the answer is a time, not a yes or no |
| How many will happen next month? | Poisson or negative binomial for counts. Ordinary regression on a count outcome is a common and costly error |
| What will demand look like? | ARIMA or Prophet for seasonal series, gradient boosting or LSTM where there are many drivers |
| Which factors are driving this? | Regularized regression, or tree ensembles read through SHAP values when the relationship isn’t linear |
| What are people saying at volume? | Topic modeling and text classification — and graph analysis where the structure between things is the point |
| What should we actually do? | Constrained optimization. Budget, inventory and capacity problems have an answer, not just a description |
And when the choice isn’t obvious, we ask. Some questions sit on a genuine methodological fault line — small samples, nested data, competing causal identification strategies. We draft the approach, then confirm it with statistics and analytics advisors before a client relies on it. Being certain is easy. Being right is worth the extra call.
Most analysis begins with somebody opening a spreadsheet and looking around. This is the alternative — a documented process where each step has questions that have to be answered before moving to the next one. It is what we work from.
| Step | What actually happens | |
|---|---|---|
| 1 | Translate the business problem | Turn a sentence like “our marketing is not working” into something measurable. Two questions first: how will the answer be used, and how will it be delivered |
| 2 | Select the data | Build a wish list, then find out what actually exists. Expect different systems, different formats, missing fields and no documentation |
| 3 | Get to know the data | Examine the distributions. Compare real values against what the documentation claims. Validate assumptions. This is the step everyone rushes and where every later problem starts |
| 4 | Create the model set | One row per thing being studied. A balanced sample. Multiple timeframes. Split into training, validation and test |
| 5 | Fix the problems | Outliers, skewed distributions, missing values, categories with too many levels, and identifiers whose meaning changed at some point in the past |
| 6 | Transform the data | Reduce variables, capture relationships that are not linear, turn counts into proportions, handle inputs that move together |
| 7 | Build the model | The step everyone imagines is the whole job. It is one of eleven, and usually the fastest |
| 8 | Assess it | Accuracy, yes. But also: how comprehensible is it, and how stable. A model nobody understands does not get used |
| 9 | Deploy it | Move it into the place it will actually run, and fit it to how people work |
| 10 | Assess the results | Measure what happened. Not what the model predicted — what happened |
| 11 | Begin again | The process is iterative. Going back two steps is normal |
And the rule underneath all of it. Hold something back — a test set for a model, a control group for a campaign. Without it there is no way to know whether anything worked, only that something happened.
The first decision, and it determines everything after it.
| Directed | Undirected | |
|---|---|---|
| Is there a target? | Yes — something specific we are trying to predict | No — we are looking for structure that is already there |
| The question sounds like | Which customers will leave in the next hundred days? | What natural groups exist among our customers? |
| Typical methods | Decision trees, logistic and multiple regression | Cluster analysis, association rules |
| Who decides if it worked | The data. Test it against cases held back | You do. Are the groups meaningful to the business? |
That last row is the one that matters. A clustering will always produce clusters. Whether they correspond to anything real is a judgment, and only somebody who understands the business can make it.
These are genuine, unfinished, and stated here rather than promised to anyone.
| Question | What it asks | Data, and where it stands |
|---|---|---|
| The regional assessment study | What is the state of digital practice among organizations in one region, measured against a documented instrument | 187 organizations already assessed, 93 verified findings. Ready to publish |
| How archival content behaves over time | Every content plan assumes a post is finished within days. Old posts on a historical-image page keep collecting comments from new people in months when nothing was published | Ten years of post-level engagement. Already held. Needs a comparison group |
| Technical health across a sector | Which technical characteristics actually predict poor real-world performance, learned from the population rather than asserted from experience | Public web performance corpora, millions of pages, free. Not yet started |
| How an industry advertises | How systematically firms in one sector advertise, and whether it relates to anything else observable about them | The public advertising archive. Entirely open, and almost nobody uses it for commercial research |
Nothing, unless you want it to. The work you commission is the work you get, on the timescale quoted.
Occasionally an engagement raises a question worth answering properly. If that happens we will ask — in writing, at the outset — whether we may use the findings in anonymized, aggregated form. Your name would never appear, and you would see anything before it was published. If you would rather not, that is completely fine and the work is unaffected.
We’ll give you a free AI analysis and tell you the three things we’d fix first — in writing, for nothing.