Skip to content

The Data Scientist

The Invisible Barriers Holding Back Modern Data Teams and How to Fix Them

An interview with YouTube’s Sean McCarthy and Hydrolix COO Tony Falco 

Written by Jackson Martin

CDN logs are the nervous system of the digital experience. When every image, video, ad impression, playback, and user engagement matters and needs to be tracked, the processing system can be overwhelmed and therefore under-function. Think about this: A typical streaming service can utilize 8 TB per day and record millions of log lines per second during major live events. On a global scale, performance degradation can affect millions of users and potentially cause a system crash. 

But, there’s no need to stress. Tony Falco, COO of Hydrolix and Sean McCarthy, Head of OTT Live Engineering at YouTube, sat down to ease our collective nerves. Together, we break down the five critical obstacles to high-cardinality, high-volume data and the breakthroughs required to overcome them.

1. Fragmented CDN Data Prevents a Unified View into the Customer Experience

Most enterprises operating in a multi-CDN environment will say that every vendor records, delays CDN logs, and structures their fields differently. Unifying them can feel impossible, especially when teams prioritize and organize the data for their needs. As an example, Sean describes this scenario: “We often ran into significant problems getting a single, coherent view of the data.This was the central operational challenge of running a multi-CDN environment.”

Before real-time logging existed, his team relied on client analytics or waited for vendor outage alerts—approaches too slow for live streaming. “What unified, sub-second visibility looks like now versus the past is the difference between a blurry, historical photograph and an immediate, 4K live feed.”

For Tony, this fragmentation is precisely why Hydrolix was founded. Its roots trace back to the founding team’s experience at Cedexis, where they collected petabytes of data across CDNs. Tony says, “Every CDN request generates dozens of log events that tell you at each step how things are functioning. We were processing billions of transactions a day, and while it scaled, the cost of BigQuery was approaching the cost of headcount. The value of the data is only valuable if you can get the insights out of it. We set out to solve that foundational problem.”

The message is clear from both Sean and Tony’s perspectives, fragmentation is the enemy of observability, and solving that problem is the foundation for everything else.

2. Today’s Data Volumes Are Massive and Legacy Systems Can’t Keep Up

Even once unified, the sheer scale of CDN logs pushes traditional systems past their breaking point. “The massive amount of data that multi-CDN sources generate is hard to fathom,” Sean says. “It’s not enough to simply collect this massive amount of data; you need the ability to query and observe it as it is ingested.” 

Put bluntly, it’s probable that early connected solutions and the companies that built them couldn’t predict that the amount of managed data would be so immense.  To explain, Tony shares why older data lakes fail and why Hydrolix doesn’t, “It comes down to two major innovations: cloud object storage like S3 and Kubernetes. Together, they make a decoupled architecture where ingest and query scale independently. You can go from 10 pods to 100 pods and back down without over-provisioning.”

So what makes Hydrolix different? Confidently, Tony explains, “In a lot of systems, you can’t scale back down once you scale up. Legacy systems ship, compute, ingest, and store as one rigid unit, but Hydrolix breaks it apart. With Kubernetes, the whole front end is elastic.”

Scalability is a problem if it can’t be equally flexible up and down. This elasticity is exactly what enables companies like Fox to scale up massively for the Super Bowl, then scale right back the next day.

3. To Retain or Not to Retain—High Cost vs. Required Repository

Historically, a company generating 10 TB of daily logs for more than 24-90 days was a financial non-starter, often costing $600,000 a month. Yet, sampling is a blindfold. “Performance issues are often highly specific events masked by aggregation,” Sean explains. “Full-fidelity logs ensure you capture every unique error.”

Hydrolix changes the unit economics through 25–50× compression on commodity object storage, making long-term retention viable. This granularity is critical for isolating:

  • “Needle in the haystack” QoE anomalies
  • Device-specific failures and regional congestion
  • SLA violations and A/B test deltas

And no one knows the ins-and-outs of these issues better than Tony. “We use the most cost-effective hot storage and then we layer on our own compression and partition the data,” he says. “It retrieves data even from slower storage. All of that adds up to a highly performant, durable database at a fraction of the cost.”

This solution is well-known in the data community. Hydrolix is often referenced and referred to in data-centric, user-generated-content platforms like Reddit. Comments like:

We moved to Hydrolix. 15 months retention means we can actually do some analysisand it’s about 25% the cost of Splunk. (Username Pik000)

To sum it up, Sean summarizes the stakes clearly, “Addressing the problems Hydrolix solves is non-negotiable.” 

4. Business-Level Insights Are Often Missing, Causing Expensive Mistakes

On the other side of the user experience lies the in-house issues caused when data is fragmented, sampled, or delayed–Companies lose the ability to make business-level decisions. Tony reacts to a real-world example of how visibility changes behaviors and budgets. 

Many companies rely on WAFs (web application firewalls) to block malicious traffic, but Hydrolix focuses on exposing failures. “We found that a huge percentage of traffic that’s supposed to be blocked is not blocked.” Tony contends, “We’re seeing as much as 60% of traffic on major brands coming from bots and bypassing their firewall.” 

Why does this matter? Because the consequences impact day-to-day operations like:

  • Overage fees from CDNs and ISPs
  • Performance degradation
  • Origin overload
  • Security exposure

Hydrolix’s visibility gives companies the clarity to act quickly and effectively. The decisions leaders need to make are relatively simple, unless they aren’t getting the information needed to make those simple decisions.” Tony clarifies that “it can take teams two or three days to resolve one alert. Backlogs mount. People burn out. Being able to classify events in real time and act immediately is the key.”

Customers, like Sean, all see that once visibility is clear, teams have a single, normalized view of the data and take real-time actions to prevent–or extinguish–issues before they fully materialize.”

5. Most Teams Don’t Have the Expertise to Build Real-Time Data Pipelines

Designing a multi-CDN real-time pipeline is a massive undertaking. System architects need a deep understanding of the source data, a thoughtful data model, and a system capable of providing mappings and transforms. Sounds easy enough, right?

“This often requires a dedicated team,” Sean says. “Unless your core business is real-time analytics, it likely does not make sense to build it yourself.” Tony echoed the sentiment, adding, “Very few companies can build something this complex and obtainable on their own. Companies like Lyft, Uber, and Nielsen scaled because their entire business demanded it. Most companies try, run into intricacies, fail, and all the while spending far more money than they would with a specialized vendor.”

Working with Hydrolix empowers in-house data teams by providing the right architecture to limit the Mean Time to Resolution (MTTR). “It’s our superpower,” says Tony. “If you summarize every testimonial, case study, and call, the pain point we fix is the same: Find and fix problems before your customers—or your boss—see them.”

One example comes from the Nordic electronics retailer Elkjøp. In 2024, they saw the start of a DDoS attack during Black Friday weeks, used the Hydrolix product TrafficPeak (which is part of Akamai’s cloud and observability services), and were able to stop it instantly. When Elkjøp Team Lead, eCommerce Jonas Petersson was asked to clarify “instant” he responded, “The entire event from spotting to stopping the attack was instant. No sites went out of service and none of our customers experienced any impact whatsoever.”

Building a bespoke multi-CDN pipeline is a resource trap. Many times spiraling, creating instead of solving MTTR.

Conclusion: Data Insights Aren’t a Luxury, They’re Non-Negotiable

Streaming companies no longer have the luxury of slow analytics, partial data, or fragmented visibility. Fragmentation, volume, cost, and complexity all conspire to hide the truth just long enough for customers to feel the pain.

Tony sums up why these five problems separate the organizations that grow and organizations that put the “lie” in liability, “Reduce the time between a problem appearing and a human fixing it. When you do that, everything else—cost, retention, performance—falls into place.