If you arrived here from the business version, you are in the right place. That piece covered what self-hosting costs and where the break-even sits. This is the companion "what": the components you actually run, and which of them decide where your data goes.
Install an SFU and you have a demo. The first production incident usually teaches you that the SFU was one of seven things you had to run.
A self-hosted video architecture is the whole set of services that turn a media server into something people can depend on: signalling, the media servers themselves, relays for users behind restrictive networks, recording, storage, capacity management and the network edge. This post walks each one — what it does, what breaks without it — and ends on the three that matter most for control.
The self-hosted video architecture at a glance
Component | What it does | What breaks without it | Carries media? |
|---|---|---|---|
Signalling + auth | Exchanges offers/answers, issues room tokens | Nobody can join; or anyone can | No |
Media servers (SFU) | Receives and forwards audio/video | No call at all | Yes |
TURN relays | Relays media for users who can't connect directly | Calls fail behind strict NATs and firewalls | Yes |
Recording | Captures tracks or composites sessions | No record, or a record you can't defend | Yes |
Storage | Keeps recordings under a retention policy | Recordings lost, or kept forever | At rest |
Capacity + placement | Decides which server a room lands on; handles failure | Hot servers, dropped calls on failure | No |
The edge | TLS, DNS, certificates, firewall ports, public IPs | Browsers refuse to connect | No |
The last column is the one this self-hosted video architecture walkthrough keeps returning to. Start with the component everyone forgets is a component.
Signalling: the part WebRTC leaves to you
WebRTC standardises how media flows. It deliberately does not standardise how two endpoints agree to start. RFC 8829, the JavaScript Session Establishment Protocol, says call setup was designed "to focus on controlling the media plane, leaving signaling-plane behavior up to the application as much as possible," and that transport of offers and answers is "left entirely up to the application."
Media servers follow the same philosophy. mediasoup describes itself as "not a standalone server but an unopinionated Node.js module" to be integrated into a larger application. So a self-hosted deployment always includes a signalling service you write or adopt — usually a WebSocket server — plus the authentication around it: who may create a room, who may join, and with what permissions. That token service is where access control lives, and it is easy to under-build.
Signalling only introduces the parties. In a self-hosted video architecture, the media goes somewhere else.
Media servers: where the calls physically run
The SFU is the component that receives each participant's streams and forwards them to everyone else — the reason an SFU beats an MCU or mesh at meeting sizes. Self-hosting it brings three requirements that a managed service hides:
- A public address that browsers can reach. mediasoup documents an announced address for exactly this, "useful when running mediasoup behind NAT with private IP." Get it wrong and connections negotiate successfully, then carry no media.
- A range of open UDP ports. mediasoup's defaults span ports 10000–59999 (API reference); LiveKit's sample Helm configuration uses 50000–60000 plus TCP 7881 (server-sample.yaml). Those ports have to be open end to end, through cloud security groups and any corporate firewall in between.
- Host-level networking. The same LiveKit configuration notes that "you can run only one instance of LiveKit per physical node" because of those port requirements. Media servers do not sit comfortably behind the load balancers and container networking most web services use.
Capacity per server is well documented — mediasoup cites "over ~500 consumers" per worker on one core in its scalability notes — and growing past one server means clustering or cascading, which moves the bottleneck in ways covered in SFU scaling.
The SFU reaches most users directly. Not all of them.
TURN: the relay you can't skip
Some users sit behind networks that block direct UDP — symmetric NATs, corporate proxies, guest Wi-Fi. For them, media goes through a TURN relay, and TURN over TLS on port 443 is often the only path that works.
Two things make TURN an architecture decision rather than a checkbox. It carries media, so wherever the relay runs is somewhere your users' audio and video pass through. And it carries a bill that scales with the share of users who need it, which is the subject of The TURN Bill. A self-hosted deployment that uses a third-party TURN service has quietly moved part of its media path back outside its own boundary.
Recording and storage: media that stays
Live media passes through and is gone. Recorded media stays, which makes recording the component with the longest-lasting consequences.
Composite recorders render the call in a browser — Jibri launches "a Chrome instance rendered in a virtual framebuffer," and "only one recording at a time is supported on a single jibri" — so capacity is one recorder per concurrent session. Per-track capture sits next to the SFU and scales far better; the choice between them also decides whether you can redact later. Either way, recordings need a storage tier with a retention policy, integrity controls and an access log, each of which is its own design decision.
Those five components carry the load of a self-hosted video architecture. Two more keep them upright.
Capacity, placement and failure
With more than one media server, something has to decide which server each room lands on, notice when a server is overloaded or unhealthy, and move new rooms elsewhere. Without it, one server runs hot while another idles, and a single failure drops every call on the machine that failed.
Media server state lives in memory, so a server that dies takes its rooms with it; what the user experiences depends on whether clients reconnect quickly to a healthy server. In a self-built stack, you design that behaviour: health checks, draining a server before maintenance, reconnection logic in the client, and the monitoring that tells you which of these is failing. It is the least visible component and the one that most often turns a stable demo into an unstable service — the gap described in why the demo takes an afternoon and production takes months.
The edge: TLS, DNS and certificates
Browsers will not give a page camera or microphone access unless the page is secure. MDN is explicit that getUserMedia is "available only in secure contexts (HTTPS)". That makes valid certificates, DNS records for every public endpoint, and renewal automation hard requirements — for the web app, the signalling endpoint and TURN over TLS alike. An expired certificate does not degrade a video service; it stops it.
That is the full self-hosted video architecture inventory. What turns it into a decision is the question of which parts you need to own.
The three parts of a self-hosted video architecture that decide where data goes
Look back at the last column of the first table. Three components carry media: the SFU, TURN and recording. Everything else — signalling, placement, the edge — handles metadata, control and configuration.
That split is the practical core of self-hosting. If the goal is to control where audio and video go, those three have to run on infrastructure you control, and a deployment that self-hosts the SFU but uses a third-party TURN service or a hosted recorder has not achieved it. The other components still matter for reliability and access control, but they are not where the media is. How that plays out across on-prem, managed-cloud and SaaS deployments — and what each lets you prove — is the subject of where the media flows in each deployment.
Build, buy, or deploy
Build on open source
mediasoup, Janus, Pion, LiveKit and coturn cover the media, relay and recording components well. You build signalling, authentication, placement, failure handling, the recording pipeline and the operational tooling around them. This fits teams with WebRTC depth who want every layer under their control.
Buy a hosted service
Every component above is run for you, which removes most of the operational work and all of the control over where media flows. For many products that is the right trade.
Deploy a platform on your own infrastructure
Samvyo is one option between the two: based on SFU architecture, shipping embeddable SDKs, with the media path, TURN and recording — the three components that carry media — running on infrastructure you control, and resilient by design. Where it does not fit: if you have no requirement about where media flows, a hosted service is simpler; and if you want to own and modify every layer yourself, building on open source gives you that.
The Bottom Line
A self-hosted video architecture is seven components, not one media server: signalling, SFUs, TURN, recording, storage, capacity management and the edge. Three of them — the SFU, TURN and recording — carry media, and those three are what decide whether self-hosting actually keeps your data where you intended.
Inventory all seven before you commit, and draw the boundary around the three.
What's Next
For what running these components costs and where self-hosting breaks even, read the business companion on self-hosted video cost. For how the media path differs across deployment models, see where the media flows in each deployment.
Frequently Asked Questions
What does a self-hosted video architecture include?
Seven components: a signalling and authentication service, one or more SFU media servers, TURN relays, a recording pipeline, storage with retention, capacity and placement management, and the network edge — TLS, DNS, certificates and firewall ports. The media server alone is a demo, not a deployment.
Does WebRTC include signalling?
No. The JSEP specification leaves signalling to the application, and media servers such as mediasoup are libraries without a built-in signalling layer. Self-hosted deployments run their own signalling service, usually over WebSockets, plus the token-based authentication around it.
Which ports does a self-hosted video architecture need open?
A range of UDP ports for media plus a few fixed ports for signalling and TURN. mediasoup's defaults span 10000–59999; LiveKit's sample configuration uses UDP 50000–60000 and TCP 7881. TURN over TLS on 443 is usually added for users behind strict firewalls.
Do I need my own TURN server if I self-host?
If you want the media path to stay on your infrastructure, yes. TURN relays carry audio and video for users who cannot connect directly, so a third-party TURN service moves part of your media outside your boundary. It also carries its own bandwidth bill.
What happens to a call if a self-hosted media server fails?
Rooms on that server drop, because media server state is held in memory. What users experience depends on whether clients reconnect quickly to a healthy server, which in a self-built stack you design with health checks, draining and client reconnection logic.
Which parts of the stack control where my data goes?
The three that carry media: the SFU, TURN relays and recording. Samvyo keeps exactly those three on infrastructure you control and is resilient by design; signalling, placement and the edge handle control and configuration rather than audio and video.