On 12 June 2025, for two hours and 28 minutes, Cloudflare's TURN service failed almost every request. Its SFU could not create new sessions, although calls already running carried on. Traffic fell to 20% of normal. The cause, in Cloudflare's own postmortem, was a failure in a third-party storage service underneath one of its internal systems.

Now picture that night three ways: you built your video stack, you bought a cloud API, or you embedded a commercial platform on your own infrastructure. Every path fails sometimes. What the choice actually decides is who gets woken up, whether they can fix it, and what recovery waits on. The architecture differences are in the technical decision guide; this post is about the pager.

Video infrastructure failure modes, in three kinds

Most video incidents fall into one of three groups, and each path owns a different share of them:

  • Service failures: the media servers, relays or signalling stop working.
  • Dependency failures: something underneath fails, such as the cloud region, a storage service or DNS.
  • Change failures: a patch, a security fix or a browser release breaks something that worked yesterday.

Buy (cloud API)

Build (open source, self-run)

Embed (commercial platform, your infra)

Who is paged

You, to watch and communicate

You, to diagnose and fix

You for infrastructure; the vendor for software

Who can fix a service failure

The provider

Your team

Split: you restart and scale, the vendor ships fixes

Dependency failures

The provider's cloud, invisible to you

Your cloud, visible to you

Your cloud, visible to you

Change failures

The provider's releases and browser updates

Every patch, CVE and browser change

Your upgrades of the vendor's releases

Recovery waits on

The provider's timeline

Your team's depth

Your team for infra, the vendor for bugs

Buy: you get paged, they fix it

When a cloud video API fails, your on-call engineer is paged by your own alerts or your customers, and then mostly waits. They can confirm it is the provider, update your status page, and tell customers. They cannot fix it.

The Cloudflare incident shows the shape of a typical failure. Existing calls survived; new ones could not start; the relay layer failed outright for users who needed it. A dashboard that only counts active sessions would have looked half-fine for the first hour. Alerting on join success rate is what catches this kind of outage early.

Then there is the contract. Twilio's API service level agreement commits to 99.95% availability for standard services, which allows about 21.6 minutes of downtime in a 30-day month. Only unavailability of five or more continuous minutes counts, the remedy is a 10% credit on that service's monthly fee, and credits are described as the "sole and exclusive remedy". The document does not name Programmable Video, so check what your own video product is covered by before relying on it.

Buying hands the fixing to people who do it all day, which is usually worth a lot. What it does not hand over is the night itself: your customers still call you.

Build: you get paged, you fix it

Run the stack yourself and every failure is yours to diagnose and fix, including the ones that are not really about video.

  • Your cloud fails. From 11:48 PM PDT on 19 October to 2:20 PM PDT on 20 October 2025, about 14.5 hours, AWS us-east-1 suffered a disruption that started with an empty DNS record for DynamoDB and spread to services including EC2 and Network Load Balancer. A video stack built in that region inherited it.
  • Your relays need patching. coturn, the standard open-source TURN server, had a published access-control bypass in versions 4.5.1.x, announced 11 January 2021 and fixed in 4.5.2. Finding, testing and rolling out that patch across every relay is your job.
  • Browsers change underneath you. Chrome removed the legacy callback-based getStats API in version 117. Any monitoring built on it stopped returning data unless someone on your team had migrated it first.

The upside is real: you can see everything, change anything, and fail over on your own terms. Recovery time is set by your team's depth, which is the part that has to be staffed around the clock.

Embed: the pager is split

Running a commercial platform on your own infrastructure divides the work. You own the infrastructure, so cloud failures, capacity and networking are yours to see and act on immediately. The vendor owns the software, so a bug in the media server or SDK waits on their fix, delivered as a release you then roll out.

This is reasoning about responsibility, not a measured result, and the details depend on the support contract. The questions to ask a vendor are concrete: who is on call for software defects, what the response time is at 2am in your time zone, and how emergency fixes reach infrastructure they do not operate.

What to watch, whichever path you take

Some failure modes exist on every path: client networks, blocked ports, bad Wi-Fi. Monitoring should separate those from platform failures. A starting set, offered as design choices rather than standards:

  • Join success rate. The fastest signal for outages like Cloudflare's, where existing calls hide the problem.
  • Signalling connected versus media connected. A session that joins but never gets media points at the media path or relay, not at login.
  • Relay share. A sudden rise usually means something changed in the network path, not in your users.
  • Client-side quality. Round-trip time, loss, freezes and why the sender reduced quality, collected from the browser across real sessions.

The honest counterweight

For most teams, buying is the more reliable option in practice. A provider running video for thousands of customers has more people, more regions and more incident practice than a small team can staff. Handing them the pager is often exactly right.

What you give up is control of the timeline. When the provider is down, you wait; when their priorities change, you adapt. A failure you cannot fix is one of the few good reasons to leave a provider, and it is worth knowing in advance what you would rewrite to leave.

The Bottom Line

Video infrastructure failure modes are the same on every path: services fail, dependencies fail, and changes break things. Build, buy and embed decide who is paged, who can fix it and what recovery waits on. Buying hands over the fixing but not the night; building gives you every lever and every pager; embedding splits them at the line between infrastructure and software. Alert on join success, read the SLA's remedy clause, and decide which of those trades your team can actually staff.

What's Next

Every incident starts with evidence from real sessions. Collecting getStats evidence in production covers how to gather it from the browser before you need it.

Frequently Asked Questions

What are the main video infrastructure failure modes?

Three kinds: service failures in media servers, relays or signalling; dependency failures in the cloud, storage or DNS underneath; and change failures from patches, security fixes and browser releases. Client network problems sit on top of all three and exist on every path.

What happens to my app when my video API provider has an outage?

Usually you cannot fix it, only detect it and communicate. In Cloudflare's 12 June 2025 incident, calls already running continued while new sessions failed, so monitoring join success rather than active sessions is what shows the outage early.

What does a 99.95% SLA actually allow?

About 21.6 minutes of unavailability in a 30-day month. Under Twilio's API SLA, only outages of five or more continuous minutes count, and the remedy is a 10% service credit described as the sole and exclusive remedy.

Is self-hosted video less reliable than a cloud API?

Not inherently, but reliability depends on your team. You inherit your cloud's failures, every security patch and every browser change, and recovery time is set by how many people can respond at 2am.

Who is responsible for incidents with an embedded video platform?

Usually split: you own the infrastructure it runs on, the vendor owns software defects. With Samvyo, which is based on SFU architecture and resilient by design, the media path, TURN and recording run on infrastructure you control, so infrastructure incidents are visible to your team directly; agree software support terms before go-live.

What should a video on-call engineer monitor?

Join success rate, signalling versus media connection, relay share, and client-side quality metrics such as round-trip time, loss and freezes. Thresholds should come from your own baseline rather than a generic number.