EP05: How Simulcast works and what it's for¶
| Recap | EP04 "Transport showdown: RTMP / SRT / WHIP" covered how the three ingest protocols handle loss |
|---|---|
| Next | EP06 "Publish/play security: signing, tokens and anti-hotlinking" |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Explain Simulcast's underlying mechanism: why "one publish, multiple quality layers" needs no server transcode, and who encodes the layers, where, and how.
- Explain why Simulcast, not SVC — not an arbitrary choice but a trade-off set by the codec ecosystem today.
1. Opening hook (script notes)¶
EP02 said "no transcoding on the network side". EP04 covered the ABR degradation logic "lower bitrate first, then resolution". This episode fills the missing link between them: how multiple layers are produced at the sender, and how the server switches quality layers without doing pixel-domain transcoding.
2. What Simulcast is: one publish, multiple independent quality layers at once¶
Simulcast is not "one encode that the server crops into layers" — the sender produces multiple independently decodable encodings at the same time. For example, a 1080p input with 3 simulcast layers can produce three independent encodings (1080p, 720p, 360p) sent as separate RTP streams. A product may implement this with multiple encoder instances, or with a hardware/software encode pipeline that supports multiple outputs; "you must run several independent encoder instances" is not a standard requirement.
Compare the traditional way to get multiple layers without Simulcast: server-side transcoding — the server fully decodes the incoming stream to raw pictures, then re-encodes it once per target layer. Decode + multiple re-encodes is a heavy compute load and naturally adds processing latency.
Simulcast moves that work from "decode then encode on the server" to "encode several copies at the publisher", so each layer the server receives is already a complete picture it can forward directly to viewers — no decode, no re-encode, just byte forwarding. That's the technical root of EP02's "no transcoding on the network side": the transcoding work didn't vanish; it moved to the publisher's encoder, trading publisher encode cycles for server transcode cycles.
3. Why Simulcast, not SVC¶
Someone who knows codecs may ask: isn't Scalable Video Coding (SVC) more bitrate-efficient? Why not SVC?
- SVC's idea: the encoded output contains a base layer plus one or more temporal/spatial/ quality enhancement layers, with dependencies between layers. A forwarder can select a subset to forward according to the dependency structure. It is usually more bitrate-efficient than multiple fully independent encodes, but encoding complexity, actual gains and sender compute cannot be generalized.
- The cost is complexity and compatibility: the SFU/Edge must understand and correctly handle layer identifiers, dependencies and switch points, and the receiver must support the matching codec/profile/scalability mode. VP9 and AV1 commonly use SVC modes in WebRTC; whether H264/HEVC is usable depends on browser, system decoders, negotiation and product implementation — a product's current state is not a standard prohibition.
- Simulcast's encodings are independently decodable: WebRTC commonly uses RID, SSRC and SDP/RTP mapping to distinguish encodings; RID does not necessarily exist in every protocol and container. The forwarder still has to maintain RTP sequence, timestamps, RTCP feedback and layer-switch state — it's not unconditionally "just byte forwarding".
This project currently centers on H264/HEVC and chooses Simulcast based on target browsers and existing sender/receiver capabilities. This is PPCDN's current product compatibility trade-off, not the H264/HEVC standard excluding SVC, nor the only choice for every WebRTC product. For selection, go by your target device matrix and interoperability tests.
4. How to use it: layering, ABR-safe switching, and multitrack as a separate dimension¶
- Layer cap: PPCDN currently limits H264 to 4 layers and HEVC to 3 (see
docs/design/whip-hevc-h264-multitrack-simulcast-design.zh-CN.mdfor the design: "HEVC simulcast's total layer count maxes out at 3 — it must stay below 4; H264 keeps its existing layer-count cap"); this is not a universal cap defined by Simulcast, RTP or the codec standards, but a product trade-off under this project's current encoder pipeline and device matrix. Each encoding needs a distinguishable RTP identifier and independent state; whether that's RID, SSRC or another mapping depends on negotiation and implementation. - How ABR switches layers: the server can step up/down based on receiver feedback and available bandwidth estimates. Switching to another independent encoding usually waits for the target layer's random-access point; for this project's H264/HEVC streams, the common approach is to request and wait for an IDR/keyframe, then handle RTP timestamp and sequence continuity. A keyframe may need to be requested via PLI/FIR; generating one costs bitrate and has wait time, so a correct switch avoids broken reference chains and artifacts but does not guarantee zero wait, zero stutter or that it is always imperceptible.
- Multitrack (H264/HEVC dual streams) is a separate dimension: multitrack answers "give browsers with different decode capability H264 or HEVC respectively", while Simulcast answers "give different network/quality needs different layers". They can run together: H264 and HEVC each maintain their own simulcast layer count, RID sequence and encoder, independently (e.g. H264 with 4 layers while HEVC independently runs 3).
5. Diagrams¶
5.1 Server transcoding vs Simulcast: where the transcode work went¶
[Traditional: server transcoding]
Publisher ──single encoded stream──▶ Server
│ decode to raw pictures
▼
re-encode per target layer (1080p / 720p / 360p)
│ CPU-heavy, adds processing latency
▼
deliver to different viewers
[Simulcast]
┌─ encoder A: 1080p ──┐
Publisher ────┼─ encoder B: 720p ──┼──▶ Server (byte forwarding; no decode/encode) ──▶ deliver to viewers
└─ encoder C: 360p ──┘
(three encodings run at the publisher, independent and directly decodable)
5.2 Simulcast layer structure¶
Publisher encoder
│
├── main encoder (e.g. 1080p) ──▶ RTP stream, RID = "high"
├── scaled-layer encoder (720p)──▶ RTP stream, RID = "mid"
└── scaled-layer encoder (360p)──▶ RTP stream, RID = "low"
The three streams are independent and fully decodable; the server identifies them by RID and forwards the chosen layer
5.3 ABR switches safely at keyframe boundaries¶
Viewer currently plays the "mid" layer (720p)
│
▼
Server sees rising loss, decides to drop to "low" (360p)
│
▼
Request or wait for a usable keyframe (IDR here) on the "low" layer
│
▼
From that IDR on, forward the "low" layer's data
│
▼
Goal: avoid artifacts; whether it's imperceptible depends on keyframe wait, buffering and RTP continuity handling
6. Wrap-up and next episode (script notes)¶
This episode broke "one publish, multiple layers, no server transcoding" down to the mechanism: Simulcast has the sender produce multiple independent encodings; the server needn't do pixel-domain transcoding but still handles RTP/RTCP and switch state. PPCDN chose it over SVC as a product trade-off for today's device matrix. Switching to an independent layer should start from a usable random-access point and handle feedback and RTP continuity correctly; a keyframe boundary alone is not unconditional seamlessness.
Next episode, we move from "how to encode" to "how to defend" — publish/play signing, token mechanisms and anti-hotlinking.