
Key takeaways
- Video is 30–45% of everything data centers store – streaming libraries, social clips and 24/7 surveillance footage.
- Enterprise data and backups take 20–25%, and a striking share is duplicates – the same databases copied 3–5 times.
- AI data (training, models, inference logs) is 10–15% and the fastest-growing slice; cold “dark data” nobody will ever open again is another 10–15%.
- About 90% of all data ever created was created in the last two years – and by 2030 data centers may store more machine-made than human-made data.
Rumors vs Reality: what actually fills the world’s data centers? 🗄️
Everyone “knows” the internet is mostly cat videos. The research says… they’re not entirely wrong.
We pulled 425 research signals from market studies, traffic reports and industry statistics to answer a simple question nobody seems to ask: what is all that storage actually holding?
Best-evidence breakdown (estimates – ranges reflect source disagreement):
📹 Video: 30–45% – by far the heaviest category. Streaming libraries, social clips, and the quiet giant: surveillance footage recording 24/7 around the world.
🏢 Enterprise data + backups: 20–25% – and a striking share of it is duplicates: the same databases copied 3–5 times for backup and compliance.
💬 Chats + text: 15–20% – hundreds of billions of messages a day, yet text is so light that all of it weighs less than one big video platform.
📷 Photos: 10–15% – trillions of photos, most viewed exactly once.
🤖 AI (training data, models, inference logs): 10–15% – the fastest-growing slice by far.
🧊 Cold/dark data: 10–15% – stored, paid for, and never accessed again. Ever.
🧩 Everything else: 5–10% – science, gaming, blockchain, the long tail.
Three facts that stopped me:
1️⃣ ~90% of all data ever created was created in the last two years. Human digital history before 2024 is a rounding error.
2️⃣ A meaningful share of everything we store is data nobody will ever open again – we are building warehouses for digital amnesia.
3️⃣ AI inference is starting to out-generate training, and synthetic data may soon out-volume human-made content.
Which leads to the real headline: the data centers of 2030 will store more machine-made data than human-made. We are becoming the minority author of our own archive.
Full breakdown with all 425 sources in the attached one-pager.
What share surprised you most? 👇
#DataCenters #BigData #AI #TechTrends
Related articles:








