Inside the Global Data Cloud: What Our Data Centers Really Store

What fills the world's data centers - breakdown of global storage by content type

Key takeaways

  • Video is 30–45% of everything data centers store – streaming libraries, social clips and 24/7 surveillance footage.
  • Enterprise data and backups take 20–25%, and a striking share is duplicates – the same databases copied 3–5 times.
  • AI data (training, models, inference logs) is 10–15% and the fastest-growing slice; cold “dark data” nobody will ever open again is another 10–15%.
  • About 90% of all data ever created was created in the last two years – and by 2030 data centers may store more machine-made than human-made data.

Rumors vs Reality: what actually fills the world’s data centers? 🗄️

Everyone “knows” the internet is mostly cat videos. The research says… they’re not entirely wrong.

We pulled 425 research signals from market studies, traffic reports and industry statistics to answer a simple question nobody seems to ask: what is all that storage actually holding?

Best-evidence breakdown (estimates – ranges reflect source disagreement):

📹 Video: 30–45% – by far the heaviest category. Streaming libraries, social clips, and the quiet giant: surveillance footage recording 24/7 around the world.
🏢 Enterprise data + backups: 20–25% – and a striking share of it is duplicates: the same databases copied 3–5 times for backup and compliance.
💬 Chats + text: 15–20% – hundreds of billions of messages a day, yet text is so light that all of it weighs less than one big video platform.
📷 Photos: 10–15% – trillions of photos, most viewed exactly once.
🤖 AI (training data, models, inference logs): 10–15% – the fastest-growing slice by far.
🧊 Cold/dark data: 10–15% – stored, paid for, and never accessed again. Ever.
🧩 Everything else: 5–10% – science, gaming, blockchain, the long tail.

Three facts that stopped me:

1️⃣ ~90% of all data ever created was created in the last two years. Human digital history before 2024 is a rounding error.
2️⃣ A meaningful share of everything we store is data nobody will ever open again – we are building warehouses for digital amnesia.
3️⃣ AI inference is starting to out-generate training, and synthetic data may soon out-volume human-made content.

Which leads to the real headline: the data centers of 2030 will store more machine-made data than human-made. We are becoming the minority author of our own archive.

Full breakdown with all 425 sources in the attached one-pager.

What share surprised you most? 👇

#DataCenters #BigData #AI #TechTrends

https://www.linkedin.com/posts/andrisgailitis_zettabytes-unpacked-the-real-contents-of-ugcPost-7492302135053946880-8U6D/

Related articles:

Mind the Gap: Closing the Expectation Divide in Cloud & Data Center Services

Key takeaways

  • 99.95% uptime sounds excellent but still allows almost 4.5 hours of downtime a year – and customers only notice the failures.
  • Standard cloud and bare-metal contracts cover infrastructure access, not data protection: if backups are not in the contract, there is nothing to restore.
  • A backup stored on the same server is not a backup – real resilience requires offsite storage or separate physical infrastructure.

Global demand in cloud computing industry and data centers is growing faster than ever. There is an explosion of hyperscalers as well as AI workloads that provide unprecedented growth impetus; companies at every level depend on providers to maintain high-performance networks. Yet amid all its innovation and growth, though, one thing is unchanged: the difference between what service agreements promise and what customers expect.

Uptime: The One-Way Street of Gratitude

The majority of professional hosting and cloud agreements make an uptime priority minimum (generally 99.95% and up) for essential network and infrastructure. 99.95% is perfect for the average joe, but in practice it allows for almost 4½ hours of downtime a year. Here’s the paradox: if a provider provides flawless service for years, no one writes a thank-you note. The silence on success is just “business as usual.” But then for 5 minutes when a blip happens … still well under 99.95% of the promise … customer support lines light up and legal clauses get quoted back to the provider. It’s not the case of customers being ungrateful, the lesson is that reliability has been rendered invisible. Uptime is required, and any deviation, no matter how slight or contractually permissible, is regrettable.

Backups: The Unpaid — and Often Misplaced — Safety Net

And the other consistent rub is backup accountability. Many customers of the cloud and bare-metal world think data can be automatically backed up when it resides at a professional data center. In practice, much of the standard agreement does not provide protection for data but access to the essential infrastructure. When a virtual machine fails or a dedicated server’s disk dies, infrequent but inevitable events, customers without a backup plan often require that the provider “just recover it.” And unless backups were part of the contract (or bought as an add-on), the provider can’t magically restore lost data. Another common yet sometimes ignored rule: You don’t have backup and recovery if you don’t pay for them. Customers can and should be told and are supposed to be educated by providers, but the responsibility of protecting data integrity falls to the data owner.

Backups on the Same Server: A Concealable Catch

Even customers who maintain backups can fall into the trap of storing those backups on the same VM or dedicated server they’re trying to protect. When the underlying hardware fails, it means that both the live data and the “backup” could disappear in a single stroke. Real resilience is holding backups offsite or at least on different physical infrastructure — in another availability zone, on another storage platform or through a managed backup service. A backup that shares the same failure domain isn’t a backup at all; it is simply yet another copy waiting to fail.

Planned Maintenance: No Good Deed Goes Unpunished

Even infrastructure most reliably established requires care. Hardware firmware ought to be patched, network gadgets upgraded, and security equipment put to the latest security updates. Nearly every service agreement specifies the timing of scheduled maintenance windows, and providers generally work on those days in the dead of night with ample notice given. Yet maintenance notices regularly provoke resistance. Some customers need zero disruption at any cost, including when the work is needed to prevent future outages. Ironically, the clients who value stability can be hostile to the very processes needed to preserve it.

Bridging the Expectation Gap

So how do providers and customers come together in the middle?

Crystal-Clear SLAs

Service Level Agreements need to be written in plain language, specifying uptime objectives, response times, and — crucially — what is not included. Define roles for backups, recovery and data retention.

Proactive Education

Providers should communicate the reality of uptime %, needs for maintenance, and responsibilities for backups during the sales process, not after the fact.

Shared Responsibility Models

When you hear the term shared responsibility, public cloud behemoths such as AWS and Azure made it famous. The former way, (whether that be infrastructure-as-a-service (IaaS) or colocation), is that the provider maintains the platform, while the customer secures and backs up their data.

Celebrate Reliability

It might seem a little self-obsessed, but frequently appearing as reports of “X days of uninterrupted service” help remind subscribers of what they’re getting back — and can help soften feelings when an unavoidable event plays out.

Not a Transaction, a Partnership

A data-center / cloud agreement is a partnership in its simplest form. Providers agree to world-class uptime, redundancy, and security; clients agree to gauge the extent of those services and plan. And when each side sees the contract as a living document and not fine print, there’s less room for surprise and fewer panicking calls when the inevitable hiccup occurs.

Takeaway: That is, nothing about the world-defining infrastructure is ever “set and forget.” Transparency is the key to successful customer relationships: explicit SLAs, contracts of mutual responsibilities, and an understanding that when it comes to maintenance, backups (carried out in their own locations) and periodic downtime, the system is better for it. Finally, a strong provider is not someone who never does need to worry about a problem, but one who talks things over openly, keeps promises and works with customers to navigate the times when the lights go out.

#CloudComputing #DataCenters #SLA #Uptime #Downtime #CloudServices #Infrastructure #DevOps #ITOperations #ServiceLevelAgreement #HighAvailability #CloudReliability #CloudBackup #PlannedMaintenance #BusinessContinuity

https://www.linkedin.com/pulse/mind-gap-closing-expectation-divide-cloud-data-center-andris-gailitis-wo94f

Related articles:

AWS, Azure, and Google Cloud vs. Everyone Else: What Truly Makes Them Different

AWS, Azure, and Google Cloud vs. Everyone Else: What Truly Makes Them Different

Key takeaways

  • AWS, Azure and Google Cloud control over two-thirds of the global market – but hyperscalers are nearly always more expensive than European managed hosting or bare-metal models.
  • Under the US CLOUD Act, American providers can be compelled to hand over data even when it is stored in Europe – a growing sovereignty risk for EU enterprises.
  • The real trap is lock-in: proprietary services, egress fees and multi-year contracts make leaving far harder than joining.
  • If your business is truly global, hyperscalers win; if it is regional, a specialist provider may be more effective.

When people speak of “the cloud,” they tend to refer to the big three: Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP). Together, they control more than two-thirds of the global market. But they are not the whole picture. And smaller providers — Oracle, IBM, and Alibaba to regional and niche players like OVHcloud, Hetzner, Scaleway, DigitalOcean, and Wasabi — are continuing to expand by providing something different.

So what actually distinguishes the hyperscalers from others? And what might lead organizations — predominantly in Europe — to hesitate to double down on the big three?


1. Scale and Global Reach

  • AWS, Azure, GCP:
  • Other providers:

👉 Tip: If your business is truly global, hyperscalers prevail. If you are region-specific, a specialist provider may be more effective.


2. Breadth of Services vs. Complexity

  • AWS, Azure, GCP:
  • Other providers:

👉 Lesson: Big clouds do mean you can be innovative but can also be dependent. Smaller providers give you focus and flexibility.


3. Pricing and the Illusion of Cheap

  • AWS, Azure, GCP:
  • Other providers:

💡 Key Takeaway: Hyperscalers are nearly always more expensive than managed hosting or hardware-for-rent models in Europe. Renting bare-metal servers with managed services can deliver similar performance at a lower TCO — without paying for unused features.


4. EU Regulation and Sovereignty

Now this gets political.

  • The EU’s Data Act and AI Act aim at guaranteeing digital sovereignty. However, most European enterprises still host sensitive workloads on U.S.-controlled hyperscalers.
  • Under the U.S. CLOUD Act, American companies can be compelled to hand over data, even if it’s stored in Europe. This creates legal uncertainty around banks, governments, and healthcare providers in the EU.
  • So European regulators and CIOs increasingly view reliance on U.S. clouds as a sovereignty risk.

Smaller European providers (OVHcloud, Scaleway, Deutsche Telekom’s Open Telekom Cloud) are picking up on this by ensuring data stays in Europe and is governed by European law.


5. The Lock-In Problem

It’s simple. You can get onto a hyperscaler with migration tools, free credits, and onboarding teams.

But getting off? That’s where the trap lies.

  • Proprietary databases, AI frameworks, serverless functions, and APIs don’t necessarily move.
  • Egress fees (paying to get your data out of the cloud) create financial barriers.
  • Complicated contracts and enterprise agreements bind customers for years.

👉 Once you’re deep into AWS, Azure, or Google Cloud, re-engineering workloads toward on-prem or moving to another provider is almost impossible.


6. Compliance and Industry Fit

  • Hyperscalers:
  • Other providers:

7. Innovation vs. Specialization

  • AWS, Azure, GCP:
  • Other providers:

The Strategic Choice

So what do decision-makers need to think about?

  • If you need global scale and advanced services, hyperscalers are unmatched.
  • If you want predictability, sovereignty, and lower costs, regional providers or managed hosting often win.
  • For some, the answer is a multi-cloud or hybrid approach: run AI workloads on a hyperscaler, hold sensitive or critical data with a sovereign provider, and use managed hosting for cost-sensitive workloads.

Final Thoughts

The big three clouds are powerful — but they come with strings attached: higher costs, lock-in, and sovereignty risks. Smaller and regional providers offer simpler pricing, local compliance, and more freedom.

In Europe in particular, the debate isn’t merely technical — it is political. Depending entirely on U.S. hyperscalers may solve today’s scaling problems, but it poses long-term risks to sovereignty and independence.

💡 Takeaway for business leaders: Don’t just ask “Which cloud is the biggest?” Ask “Which cloud best aligns with my strategy, compliance, and sovereignty needs?” The best choice might be a well-balanced one.

Subscribe & Share now if you are building, operating, and investing in the digital infrastructure of tomorrow.

#Cloud #AWS #Azure #GoogleCloud #MultiCloud #DataSovereignty #EUAIAct #CloudAct #DataCenters #ManagedHosting #DigitalSovereignty #LockIn #FinOps #AI #Infrastructure

https://www.linkedin.com/pulse/aws-azure-google-cloud-vs-everyone-else-what-truly-makes-gailitis-kfukf

Related articles:

Why Colocation and Private Infrastructure Are Making a Comeback—and Why Cloud Hype Is Wearing Thin

Colo-Coolocation

Key takeaways

  • 83% of enterprise CIOs planned to repatriate at least some workloads in 2024 (Barclays), up from 43% in 2020 – but only 8–9% plan full repatriation.
  • The drivers are unpredictable cloud billing, compliance burden, performance and control.
  • Colocation delivers predictable costs, data residency and direct hardware control.
  • The trend is not cloud vs colo – hybrid is the smarter default.

The Myth of Cloud-First—And the Reality of Repatriation.

For nearly a decade, businesses have been sold the idea of “cloud-first” as a golden ticket—unlimited scale, lower costs, effortless agility. But let’s be frank: that narrative wore thin a while ago. Now we’re seeing a smarter reality take shape—cloud repatriation: organizations moving workloads back from public cloud to colocation, private cloud, or on-prem infrastructure.

These Numbers Are Real—and Humbling

Still, let’s be clear: only about 8–9% of companies are planning a full repatriation. Most are just selectively bringing back specific workloads—not abandoning the cloud entirely. (https://newsletter.cote.io/p/that-which-never-moved-can-never)

Why Colo and On-Prem Are Winning Minds

Here’s where the ideology meets reality:

1. Predictable Cost Over Hyperscaler Surprise Billing

Public cloud is flexible—but also notorious for runaway bills. Unplanned spikes, data transfer fees, idle provisioning—it all adds up. Colo or owned servers require upfront investment, sure—but deliver stable, predictable costs. Barclays noted that spending on private cloud is leveling or even increasing in areas like storage and communications (https://www.channelnomics.com/insights/breaking-down-the-83-public-cloud-repatriation-number and https://8198920.fs1.hubspotusercontent-na1.net/hubfs/8198920/Barclays_Cio_Survey_2024-1.pdf).

2. Performance, Control, Sovereignty

Sensitive workloads—especially in finance, healthcare, or regulated industries—need tighter oversight. Colocation gives firms direct control over hardware, data residency, and networking. Latency-sensitive applications perform better when they’re not six hops away in someone else’s cloud (https://www.hcltech.com/blogs/the-rise-of-cloud-repatriation-is-the-cloud-losing-its-shine and https://thinkon.com/resources/the-cloud-repatriation-shift).

3. Hybrid Is the Smarter Default

The trend isn’t cloud vs. colo. It’s cloud + colo + private infrastructure—choosing the right tool for the workload. That’s been the path of Dropbox, 37signals, Ahrefs, Backblaze, and others (https://www.unbyte.de/en/2025/05/15/cloud-repatriation-2025-why-more-and-more-companies-are-going-back-to-their-own-data-center).

Case Studies That Talk Dollars

Let’s Be Brutally Honest: Public Cloud Isn’t a Unicorn Factory Anymore

Remember those “cloud-first unicorn” fantasies? They’re wearing off fast. Here’s the cold truth:

  • Cloud costs remain opaque and can bite hard.
  • Security controls and compliance on public clouds are increasingly murky and expensive.
  • Vendor lock-in and lack of control can stifle agility, not enhance it.
  • Real innovation—especially at scale—often comes from owning your infrastructure, not renting someone else’s.

What’s Your Infrastructure Strategy, Really?

Here’s a practical playbook:

  1. Question the hype. Challenge claims about mythical cloud savings.
  2. Audit actual workloads. Which ones are predictable? Latency-sensitive? Sensitive data?
  3. Favor colo for the dependable, crucial, predictable. Use public cloud for seasonal, experimental, or bursty workloads.
  4. Lock down governance. Owning hardware helps you own data control.
  5. Watch your margins. Infra doesn’t have to be sexy—it just needs to pay off.

The Final Thought

Cloud repatriation is real—and overdue. And that’s not a sign of retreat; it’s a sign of maturity. Forward-thinking companies are ditching dreamy catchphrases like “cloud unicorns” and opting for rational hybrids—colocation, private infrastructure, and only selective cloud. It may not be glamorous, but it’s strategic, sovereign, and smart.

Subscribe & Share now if you are building, operating, and investing in the digital infrastructure of tomorrow.

#CloudRepatriation #HybridCloud #DataCenters #Colocation #PrivateCloud #CloudStrategy #CloudCosts #Infrastructure #ITStrategy #DigitalSovereignty #CloudEconomics #ServerRentals #EdgeComputing #TechLeadership #CloudMigration #OnPrem #MultiCloud #ITInfrastructure #CloudSecurity #CloudReality

https://www.linkedin.com/pulse/why-colocation-private-infrastructure-making-cloud-hype-gailitis-bcguf

Related articles:

Proudly powered by WordPress | Theme: Baskerville 2 by Anders Noren.

Up ↑