Measuring Things: Why We Have So Much Digital Data
We do not have oceans of digital data because it is all meaningful. We have it because digital systems can measure cheaply, and because whatever can be measured can be monetized. The volume is exhaust.
By Shailin Dhar
The standard story about big data is that we drown in it because the world is finally revealing its secrets. That is backwards. We have so much digital data for a much duller reason: in a digital system, measurement is nearly free. Every click, scroll, hover, and load leaves a log entry whether anyone wanted it or not.
Two forces then compound the volume. First, whatever can be measured can be monetized: a counted event can be sold, billed, or reported. Second, whatever can be measured can be optimized: an automated system will happily chase any number you expose to it. So we instrument everything, because instrumenting is cheap and the readings might be worth something later.
The data lake is not a reservoir of insight. It is a puddle of exhaust that happened to be cheap to collect.
Volume is a byproduct, not a signal
Most of what accumulates is exhaust, the residue of a system running, not evidence that the running meant anything. A server that logs a billion requests has not learned a billion things. It has learned that logging is easy. Confusing the two is how organizations end up data-rich and insight-poor, warehousing terabytes while still unable to answer a simple question about who is real.
The discipline is to invert the reflex. Do not ask what the data says; ask why this particular data exists at all. Usually the answer is that it was measurable and someone hoped to bill against it, not that it carries meaning. Treat volume as a cost to be justified, not a virtue to be celebrated, and most dashboards get quieter and more honest.
Keep Reading
AI-assisted recommendations of related work elsewhere on the internet.