If you arrived here from the business version, you are in the right place. That piece compared what on-prem, managed cloud and SaaS let you claim about your data. This is the companion "why": where the media actually flows in each, and what you can prove about it.
Tell a customer that their video never leaves your infrastructure and the next question, from anyone serious, is: prove it.
The answer lives in the video deployment architecture — not the logo on the contract, but which component each packet passes through, which of those components can read it, and what record that leaves behind. Every WebRTC call has four distinct flows, and they do not all go to the same place.
The four flows in every video deployment architecture
Flow | What it carries | Protection in transit | Who can read the content |
|---|---|---|---|
Signalling | Room joins, offers/answers, tokens, participant lists | TLS (WSS/HTTPS) | Whoever runs the signalling server |
Media | Audio and video between client and SFU | DTLS-SRTP, mandatory | The client and the SFU |
Relay | The same media, via TURN, for users who can't connect directly | DTLS-SRTP inside the relayed packets | Not the relay — only the endpoints |
Recording | Media captured and written to storage | Depends on the recorder and store | Whoever runs the recorder and holds the storage keys |
The third column looks reassuring — everything is encrypted. The fourth column is the one that matters, and it starts with a property of every SFU.
Encrypted on every hop, readable at the SFU
WebRTC encryption is mandatory — unencrypted media is not allowed. The WebRTC security architecture, RFC 8827, is flat about it: "All media channels MUST be secured via SRTP," and "media traffic MUST NOT be sent over plain (unencrypted) RTP or RTCP."
But that encryption is hop by hop. Each client negotiates DTLS-SRTP with the media server it connects to, not with the other participants, because the SFU has to receive, inspect and re-send each stream — that is what an SFU is for. The SFU decrypts what arrives and re-encrypts what it forwards. Wherever the SFU runs is where your media exists in the clear, in memory, for as long as the call lasts.
That single fact does most of the work in any video deployment architecture. "Encrypted in transit" is true of every WebRTC service. "Nobody outside our organisation can see the media" is only true if the SFU is yours — or if the media is encrypted a second time, end to end, which comes later. How the handshake reaches this point is walked through in the WebRTC architecture overview.
The relay is different, and more benign than it looks.
TURN sees traffic, not content
When a user sits behind a network that blocks direct paths — the situation NAT traversal is about — their media goes through a TURN relay. RFC 8656 describes a TURN server as one that "relays data between a TURN client and its peer(s)," and is explicit that "applications that want end-to-end security should encrypt the data sent between the client and a peer."
WebRTC does exactly that: the DTLS-SRTP session runs between the client and the SFU, through the relay. The relay forwards packets it cannot decrypt. What it does see is metadata — which IP addresses talked, when, and how much — and that metadata is itself sensitive. A third-party TURN service cannot watch your calls, but it can see that they happened, from where and for how long. Where the relay runs is therefore a metadata question and a cost question; the cost side is The TURN Bill.
Live media is transient. Recorded media is the part of a video deployment architecture that lasts.
Recording: where media becomes data at rest
A recorder is a subscriber: it receives decrypted media from the SFU and writes it somewhere. Two locations matter — where the recorder runs, because it holds cleartext media, and where the files land, because that is the copy that outlives the call and can be demanded later. Whoever holds the storage encryption keys decides who can read it.
This is also where the long-term evidence obligations sit: completeness, integrity, access logs and export, covered in the recording architecture audit.
Tracing the video deployment architecture in each model
Put the four flows against the deployment models from the business companion and the differences become concrete.
Flow | SaaS | Managed cloud | On-prem | Hybrid |
|---|---|---|---|---|
Signalling | Vendor | Your cloud account | Your network | Often vendor |
Media (cleartext at SFU) | Vendor | Your cloud account | Your network | Your infrastructure |
Relay (metadata) | Vendor or its TURN supplier | Your cloud account | Your network | Your infrastructure |
Recording (at rest) | Vendor's storage | Your bucket, provider's custody | Your storage | Your storage |
Two readings of that table are worth drawing out. In managed cloud, media and recordings are yours to configure but physically in a provider's custody — which is why the business companion separates residency from who can be compelled to disclose. And in a hybrid, the media, relay and recording rows are yours while signalling may not be, so the precise question is what the signalling side carries: room identifiers, participant identities and timings are metadata, not media, but they are not nothing.
Knowing where the flows go is the video deployment architecture. Showing it is the evidence.
What you can prove, and how
Every claim about a flow in your video deployment architecture can be backed by a record the system already produces, if you keep it.
Claim | Evidence |
|---|---|
"This call used our SFU and our relay" | Client-side WebRTC statistics. The W3C stats spec exposes the selected ICE candidate pair and each candidate's type — host, srflx, prflx or relay — so each session can record which path and which relay address it used. |
"Media traffic stayed inside our network" | Network flow records at the media servers. AWS describes VPC Flow Logs as capturing "information about the IP traffic going to and from network interfaces"; on-prem, the firewall and switch logs do the same job. |
"Only our relays were used" | TURN server allocation logs, matched against the relay candidates reported in client statistics. |
"Recordings are stored here, readable only by us" | Storage location configuration, key ownership, and the access log for every read. |
"The vendor cannot see our media" | Either the SFU is outside the vendor's control — shown by the rows above — or the media is end-to-end encrypted. |
The last row points at the one technique that changes the fourth column of the first table.
End-to-end encryption: the strongest claim, and what it costs
End-to-end encryption adds a second layer that the SFU cannot remove. SFrame, RFC 9605, describes a mechanism "in which central media servers (Selective Forwarding Units or SFUs) can access the media metadata needed to make forwarding decisions without having access to the actual media." With it, even a vendor-operated SFU forwards media it cannot read.
The cost is real and usually underestimated. Anything that needs to see the media on the server side stops working unless it holds the keys: server-side recording, transcription, live captioning, real-time AI features and compliance review all depend on cleartext. For many regulated uses — where a recording must exist and be retrievable — end-to-end encryption conflicts directly with the requirement.
That leaves two honest ways to make the strongest claim. Run the SFU, relay and recorder yourself, so the point of cleartext is inside your boundary; or use end-to-end encryption and accept that the server can do nothing with the media. Which one fits depends on whether you need the server to see it.
Build, buy, or deploy
Building on open source puts every flow on infrastructure you choose — the full component list is the price. Buying SaaS puts every flow with the vendor, which is fine where the claims you need are the vendor's to make. Samvyo sits between: based on SFU architecture, with the media path, TURN and recording — the flows that carry media — running on infrastructure you control, resilient by design. Where it does not fit: if you need the server to never see media at all, end-to-end encryption is the requirement, not deployment location; and if no claim about media location matters to your customers, a hosted service is simpler.
The Bottom Line
WebRTC encrypts every hop, but the SFU decrypts media to forward it, so the SFU's location is where your media exists in the clear. TURN relays see traffic, not content; recorders and their storage hold the copy that lasts.
Trace the four flows in your video deployment architecture, keep the evidence each one produces, and make only the claims those records support.
What's Next
For which claims each deployment model supports at the business level, read the business companion on video deployment models. For the full component inventory behind these flows, see what you actually run to self-host real-time video.
Frequently Asked Questions
What is a video deployment architecture?
It is the map of where each part of a video call runs and flows: signalling, media through the SFU, relayed media through TURN, and recordings into storage. It determines which component can read the media and which records exist to prove where it went.
Is WebRTC media encrypted?
Always in transit. RFC 8827 requires all media channels to be secured with SRTP and forbids unencrypted RTP. But the encryption is hop by hop: the SFU decrypts media to forward it, so it can read the content.
Can a TURN server see my video?
No. TURN relays the encrypted DTLS-SRTP packets between the client and the media server without the keys to decrypt them. It does see metadata — IP addresses, timing and traffic volume — which is why relay location still matters.
How can I prove what my video deployment architecture actually did on a call?
Record client-side WebRTC statistics for the selected ICE candidate pair, which show the path and candidate type, including relay addresses. Match them against TURN allocation logs and network flow logs at your media servers.
Does end-to-end encryption work with an SFU?
Yes. SFrame, RFC 9605, lets SFUs forward media using metadata without access to the content. The trade-off is that server-side recording, transcription and other processing cannot work unless they hold the keys.
Where does media exist in the clear in a hybrid deployment?
At the SFU and the recorder. In a hybrid where those run on your infrastructure, cleartext media stays inside your boundary even if some management runs elsewhere. Samvyo keeps the media path, TURN and recording on infrastructure you control.