Someone who was on a recorded call asks for their copy. The law says you have to give it to them — and that you have to protect everyone else who was on it.
Whether that takes ten minutes or ten days was decided long before the request arrived, on the day someone picked a recording mode. Composite vs per-track recording is usually framed as a question of convenience and cost. It is really a question of reversibility: you can always mix separate tracks into one file later, and you can never pull a mixed file back apart into tracks.
This post argues composite vs per-track recording from that one-way door.
Composite vs per-track recording, in one paragraph
A composite recording renders the call the way a participant saw it — a layout of tiles, one mixed audio track — and encodes that into a single file. Jibri is the classic example: it launches "a Chrome instance rendered in a virtual framebuffer" and encodes the output. A per-track recording (Agora calls it individual recording) writes each participant's audio and video as separate streams. Janus is explicit about how literal that is: "each mjr file contains a single media stream," so an audio/video publisher produces two files.
The mainstream platforms offer both, which is the first sign that neither is simply better. Agora recommends individual mode "if you want more flexibility in processing the recorded files" and composite mode to store everyone "in one file, with no need to combine them after recording." Twilio records every track separately — seven participants with camera and microphone produce "a total of 14 media tracks" — and builds playable compositions from them afterwards. How each model behaves at scale is covered in the recording pipeline at scale. What matters here is what each one lets you do on the day someone asks for a copy.
The request that tests the choice
Under the GDPR, a person can ask for a copy of their personal data, and Article 15(3) says "the controller shall provide a copy." Article 15(4) adds the constraint: that right "shall not adversely affect the rights or freedoms of others."
For video, the European Data Protection Board has spelled out what that means. Its guidelines on video devices say the controller "should not in some cases hand out video footage where other data subjects can be identified" — but that protecting third parties "should however not be used as an excuse to prevent legitimate claims of access," so the controller should "implement technical measures to fulfil the access request (for example, image-editing such as masking or scrambling)."
So the realistic request is this: one participant on a five-person, 45-minute call wants their copy, and the other four have to be masked. Here is what that asks of each recording mode.
Why a composite recording resists redaction
In a composite file, every person on the call has been fused into the same frames and the same waveform. Removing one of them is no longer an edit to a file — it is an edit to every frame of it.
- The video is a moving target. Four other faces sit in tiles across 81,000 frames (45 minutes at 30 fps). Speaker-focused layouts move and resize those tiles every time someone new talks, so a mask drawn for minute three is in the wrong place by minute four. Automated face-tracking helps, and every miss is a disclosure.
- The audio is already mixed. Twilio describes its own composition audio as recordings "mixed through a linear adder." Once voices are summed into one signal, there is no clean way to subtract one. The practical options are muting whole stretches — which also silences the person who asked, whenever anyone talks over them — or source-separation tools that approximate the split.
- The screen share came along too. Whatever was on a shared screen — a customer list, someone else's document — is rendered into the same frames and needs the same treatment.
None of this is impossible. It is manual, slow, and imperfect in exactly the places where imperfection matters — the first real cost in composite vs per-track recording. Per-track recording turns the same request into a selection.
Why per-track recording makes it a selection
With separate tracks, the requester's camera and microphone are already their own files. Answering the request starts by including those and leaving everyone else's out — the decision a composition step makes anyway. Twilio's description of that step is the whole point: developers "can select the specific audio recordings to be included" and the video recordings to be laid out. Exclusion is a parameter, not an edit.
The same property handles the harder variants. If the requester needs to see the conversation in context, you can compose their tracks with the others masked or replaced by a placeholder, working from clean sources rather than a mixed frame. If one participant has to be removed from an archive entirely, you delete their tracks and keep everyone else's intact.
It is not a complete answer, and the edges are worth knowing:
- A track is a microphone, not a person. Two people sharing a meeting-room device share one track, and a microphone picks up whoever is near it. Per-track gets you to the person in most cases, not all.
- The requester's own track can still contain others. Their screen share, or a colleague speaking audibly behind them, still needs review.
- Tracks have to be re-aligned. RFC 3550 notes that timestamps "from different media streams may advance at different rates and usually have independent, random offsets," so a recorder that discards the timing data leaves tracks that cannot be put back in sync. Keeping it is part of the recording architecture audit.
Those are real limits. They are also the reason per-track is a better starting point, not a reason to prefer composite — which brings us to what composite genuinely does better.
Where composite recording is the right call
Composite has honest advantages, and in many composite vs per-track recording decisions they win.
Consideration | Composite | Per-track |
|---|---|---|
Playable straight away | Yes — one standard file | No — Twilio notes raw track recordings are "incompatible with common media players"; Janus files need its post-processor |
What it shows | What participants saw, layout included | What each participant sent |
Storage | One encoded stream per session | Every published stream, kept separately |
Processing | Encoding happens live, per session | Composition happens later, only when needed |
Redaction / access requests | Frame-by-frame masking; mixed audio | Select or exclude tracks; mask from clean sources |
Reversible | No — cannot recover tracks | Yes — a composite can be built at any time |
The storage row deserves a number, with its assumption stated: if each of five cameras publishes at 1 Mbps and a composite is encoded at 2 Mbps, per-track video costs about 2.5 times as much to keep. Retention multiplies that difference for as long as you hold the footage. And the recorder has to ask for the right quality — an SFU forwards whichever simulcast layer a subscriber requests, so a track recorder that accepts a thumbnail layer archives a thumbnail.
The "what it shows" row matters too. A composite is a record of what was on screen; if the question is ever "what did the participants see?", that is precisely the evidence you want. A per-track archive can reconstruct it, but only if the layout logic is reproducible.
So composite vs per-track recording is less binary than it looks.
How to choose between composite vs per-track recording
Ask one question first: will anyone ever need to separate one person from this recording? That covers access requests, redaction for disclosure, removing a participant, and reviewing one speaker's contribution. If the honest answer is "never" — an internal all-hands, a published webinar — composite is simpler and cheaper, and the one-way door costs you nothing.
If the answer is "possibly," capture per-track and compose on demand. This is the pattern the large platforms already use, and it fits the way media servers work: an SFU forwards streams without decoding them, the property that makes it scale, so capturing tracks is close to free for it, while composing requires decoding and re-encoding, the work an MCU does live. Doing that work later, only for the recordings that need it, is usually cheaper than doing it for every session.
For regulated recordings — banking verification calls, disputes, anything that may be disclosed — the case for tracks is strongest, and what each regime demands makes it stronger.
The Bottom Line
Composite vs per-track recording is a choice about reversibility. Tracks can always become a composite; a composite never becomes tracks. When a request for a copy arrives and everyone else on the call has to be masked, that asymmetry is the difference between selecting files and editing every frame.
In composite vs per-track recording, record per-track wherever someone might one day need to be separated from the recording, and compose only what you need to play.
What's Next
For the evidence a recording pipeline has to produce beyond the recording mode — completeness, timing, chained hashes and export — read the recording architecture audit. For how composite and track recorders behave under load, see what runs behind the recording pipeline at scale.
Frequently Asked Questions
Composite vs per-track recording: what is the difference?
A composite recording mixes everyone on a call into one file, with a video layout and a single mixed audio track. A per-track recording keeps each participant's audio and video as separate streams. Per-track can be turned into a composite later; a composite cannot be split back into tracks.
Is individual recording the same as per-track recording?
Yes. Agora calls it individual recording mode; others call it track recording, track egress or raw capture. In every case each participant's media is written separately rather than mixed into one file.
Why is it hard to redact one person from a composite recording?
Because their face and voice are fused into every frame and the mixed audio. Masking means tracking their tile across a layout that changes whenever someone speaks, and mixed audio cannot be cleanly un-mixed. With per-track recording you exclude or mask their tracks instead.
Does GDPR require redacting other people from a video access request?
Article 15(4) says the right to a copy must not adversely affect the rights of others, and EDPB guidance on video says controllers should use technical measures such as masking or scrambling rather than refuse the request. This is how the texts read, not legal advice for a specific case.
Does per-track recording cost more storage?
Usually yes, because every published stream is kept separately. As an illustration, five cameras at 1 Mbps each against a 2 Mbps composite is about 2.5 times the storage. The trade is that composition work happens only for the recordings that need it.
Can I record per-track and still get a normal video file?
Yes. That is the common pattern: capture tracks, then compose a playable file on demand with the layout and participants you choose. Keep the tracks as the source of record and treat each composite as a derived copy.
Can recordings and redacted copies stay on our own infrastructure?
Yes, if the platform keeps recording processing and storage on infrastructure you control. Samvyo works this way: it is based on SFU architecture, so recordings and any redacted copies made from them stay inside your own boundary rather than a vendor's. Which recording mode fits each workflow is still a per-use-case decision.