Snowflakes in an Avalanche: RTB and Data Purges
Every bid request is a unique snowflake. They arrive as an avalanche and melt on contact. The system generates data too specific to keep and too voluminous to audit.
By Shailin Dhar
In real-time bidding, every bid request is a snowflake. Device, location, page, user segment, timestamp, floor price. No two are quite identical. Each one is a tiny, hyper-specific description of a moment. And each one is genuinely unique, which sounds valuable until you see how many arrive.
They arrive as an avalanche. Billions a day, tens of billions, faster than any human process can watch. And like an avalanche, the individual snowflakes are irrelevant the moment they are part of the mass. No one inspects a snowflake mid-slide.
Impermanence at volume
Here is the mechanical reality: the system produces more detail than it can afford to store. Retaining full bid-stream data at that scale is ruinously expensive, so most of it is purged almost immediately: sampled, aggregated, or dropped. The specificity that would let you prove what happened is gone before anyone asks.
Accountability melts on contact. By the time you go looking for the snowflake, it is water, and the water is gone.
Shailin Dhar
This is why "we can't reproduce that" is such a common answer in ad tech disputes. It is often true. The data existed for milliseconds and was never meant to survive. The industry built a machine that emits perfect evidence and then discards it as a cost of doing business. Impermanence is not a bug in the design. It is the design.
If you want accountability in a system like this, you do not ask it to remember every snowflake. You put gauges on the slope: durable, sampled checkpoints that capture enough of the flow to reconstruct what fell. You cannot audit an avalanche after it has settled. You measure it while it moves.
Keep Reading
AI-assisted recommendations of related work elsewhere on the internet.