EP01: The interactive-video dilemma¶
0. Goals of this episode¶
By the end, the viewer should be able to answer two questions:
- Is my interactive live-streaming / interactive-game project also stuck in the latency, legacy-protocol and single-point/sovereignty traps?
- If my pipeline still uses RTMP or HTTP-FLV, what happens on a weak network, and why is that not an exaggeration?
This episode only lays out problems and verifiable facts — no cure; the cure is EP02–EP20.
1. Opening hook (script notes)¶
If you've built interactive live streaming — real-person video, online auctions, co-streamed games, or games with barrage feedback — you may have hit this: the host has already moved to the next step, but the viewer's picture is still on the previous one.
That's usually not just "the network happened to hiccup". It's the combined result of capture, encoding, transport, jitter resistance and decode/render. Today we break these problems into several real dilemmas and separate protocol mechanics, configuration and measured data.
2. Three real dilemmas¶
2.1 Dilemma 1: the latency paradox — mature protocols are the opposite of real-time business needs¶
HLS/DASH is the most mature way to deliver internet video. Based on "segment then download", it inherently carries seconds of latency. That barely matters for one-way VOD, but it is a structural conflict for highly interactive scenarios:
| Scenario | Why latency-sensitive | Consequence of traditional segmentation |
|---|---|---|
| Live video / real-time trading | The interaction window is measured in seconds; the picture must match the business state | Users see a stale picture and can't participate |
| Online auctions / flash sales | Bid order decides fairness; the picture can't lag the business state | Ordering confusion, disputes |
| Interactive classrooms | Teacher-student Q&A and exercises need instant feedback | Interaction degrades into one-way playback |
| Social co-streaming | Conversation depends on natural turn-taking | Talking over each other, latency stacking — unusable |
Do real-time protocols, usually UDP-based (WebRTC/SRT), fix it automatically? No. End-to-end latency is the combined result of capture, encoding, network queueing and propagation, loss recovery, receiver jitter buffer, decode and render. Some waits are bounded by parameters, some vary dynamically with network and implementation; you can't just add a few defaults together and call it a measurement, and you certainly can't treat one buffer parameter as a strict mathematical upper bound for an entire protocol.
The project's confirmed, same-caliber measurements are below. They are PPCDN case data, not a guarantee of WebRTC, SRT, or any codec in all environments:
| Scenario | Measured end-to-end latency | Notes |
|---|---|---|
| P2P direct | ~70ms | Real link, media bypasses Origin/Edge forwarding |
| WHIP ingest (via Edge) | ~108ms | Measured on the current actual pull path |
| SRT ingest (via Edge) | ~386ms | Measured on the current actual pull path |
The WHIP-vs-SRT gap mainly comes from the ingest protocol itself — SRT's receive window (TSBPD) inherently queues longer than WebRTC ingest does. That's a protocol difference, not a codec difference: HEVC currently doesn't even participate in P2P direct connect, it only goes through Edge delivery, so the encoding format isn't the variable driving these numbers. P2P skipping a server hop usually helps latency, but the result still depends on the network path, NAT traversal outcome, encoder and player. A protocol name only tells you which mechanisms are available; whether it meets a real-time goal must be measured with unified instrumentation, endpoints and clock caliber.
2.2 Dilemma 2: a protocol being retired — RTMP/HTTP-FLV on weak networks¶
2.2.1 An industry migration that already happened¶
- Adobe formally ended Flash Player support on 2020-12-31, and mainstream browsers then removed the Flash runtime. RTMP (Real-Time Messaging Protocol) was originally designed for Flash Player playback in the browser — that "native browser playback" path no longer exists.
- What people call "RTMP live" in a browser today doesn't actually use RTMP for playback:
it's re-wrapped as HTTP-FLV (same TCP transport, but the browser still needs a JavaScript
library like
flv.jsto de-mux client-side), or transcoded to HLS/DASH (segment download buys playback compatibility at the cost of real-time — see Dilemma 1). - WHIP (WebRTC-HTTP ingestion) and WHEP (WebRTC-HTTP egress) matured during IETF standardization and have been adopted by mainstream publishers including OBS Studio and several real-time services as the next-generation ingest/playback protocols — precisely to replace RTMP at the "ingest" position.
- This isn't one vendor's preference but the joint direction of browser capability, codec standards and protocol standardization over recent years: the playback-side RTMP/Flash combo is gone, and ingest-side RTMP is being replaced by WHIP/SRT.
2.2.2 Root cause: why it's a disaster on weak networks, not hyperbole¶
RTMP and HTTP-FLV are both built on TCP. TCP delivers an "ordered, complete" byte stream: if any packet is lost, later data that already arrived must wait for its retransmission before reaching the application — this is Head-of-Line Blocking, a design property of TCP itself, not a defect of some implementation.
On a weak network (with some loss rate), this cascades:
- A packet is lost → retransmit, waiting at least one more round-trip (RTT);
- During retransmission, all later-arrived audio/video is blocked in the buffer and can't play;
- Sustained loss or repeated retransmits → blocking can accumulate from tens of ms to seconds or more, and has no theoretical upper bound;
- The player buffer fills or drains → stutter, catch-up, and in the worst case a reconnect.
This differs from common WebRTC/SRT trade-offs: an application can bound how long it waits for retransmission based on a packet's timeliness and drop data once it's stale, avoiding endless accumulation for the sake of completeness. SRT does this with ARQ, TSBPD and stale-packet drop; WebRTC does it with RTP/RTCP feedback, congestion control, jitter buffer and decoder frame-drop policy. Neither offers a strict mathematical latency bound independent of network, configuration and implementation: congestion, queueing, outages, reconnects or implementation policy can still push latency well up.
In one sentence: TCP byte streams prioritize reliable, ordered delivery, so loss causes head-of-line blocking; a real-time media stack can drop stale data by timeliness, which usually makes latency easier to control — but it is not an unconditional upper bound.
2.2.3 Protocol comparison¶
| Dimension | RTMP / HTTP-FLV | SRT | WebRTC (WHIP / WHEP) |
|---|---|---|---|
| Transport | TCP | UDP + ARQ | Usually UDP + RTP/RTCP; can fall back to TURN/TCP/TLS |
| Loss handling | TCP must deliver reliably in order; possible head-of-line blocking | ARQ requests retransmit; TSBPD delivers by time; configurable stale-drop can discard late data | Whether NACK/FEC is used depends on negotiation/implementation; the receiver can drop stale frames |
| Weak-network latency | Queueing/retransmit can accumulate to seconds or kill the stream | Engineering bounds via latency etc., but no unconditional math bound |
Jitter buffer and congestion control adjust dynamically; no unconditional math bound |
| Native browser playback | No longer supported (relies on the discontinued Flash Player) | Not supported (usually needs protocol conversion server-side) | Native, standard Web APIs |
| Current role | Legacy protocol exiting the playback path | The modern choice for weak uplinks | The mainstream real-time playback protocol |
2.2.4 A concrete architecture choice¶
This isn't abstract — the project's public technical notes say it plainly: ingest supports both WHIP and SRT (the SRT ingest measurements in §2.1 are exactly that path), playback is uniformly WHEP, and HTTP-FLV playback endpoints are deliberately not provided. In other words, "no HTTP-FLV" isn't an omission but a constraint locked in at protocol-selection time; SRT is not contradictory — it addresses weak-uplink robustness, while playback still goes only through WHEP. Such constraints are increasingly common in mature real-time systems and are the concrete form of the "industry migration" discussed here.
2.3 Dilemma 3: the single-point and sovereignty paradox — the easier it is, the more you hand over your lifeline¶
Beyond protocols, real projects carry two equally important structural risks that don't show up in latency numbers:
- Single point of failure: an Origin is usually a single copy — the sole entry for all publishing. When it fails, the result is not "degradation" but a whole-path outage: every viewer loses the picture at once. This isn't theory — the architecture design doc lists "Origin publish disconnect" as a known failure mode, and explicitly not an automatic, seamless failover: the publisher must re-publish to a new Origin before any viewer can recover. This is an objective risk point in many real interactive-live architectures today, and something to weigh carefully later: whether a single point is worth eliminating is an engineering calculation, not the dogma "a single point must be eliminated".
- Content sovereignty: to cut engineering effort, some teams outsource distribution to a single third party, even treating it as a "backup fallback". It looks like double insurance, but actually hands business continuity to a third party you can't control — their rate limiting, service changes or cross-border compliance shifts can all directly affect your availability, and you have no say in those decisions.
These two are still blank spots in most discussions about interactive-live projects — few systematically assess whether a single point is worth eliminating, and few discuss how to weigh content sovereignty (EP03 expands on it). This episode just points them out.
3. Diagrams¶
3.1 End-to-end latency composition: defaults can't replace measurement¶
Capture → Encode → Send queue → Network propagation/queueing/loss recovery
→ Receiver reorder & jitter buffer → Decode → Render
These stages overlap or vary dynamically; configuration values can't be mechanically added.
PPCDN same-caliber measurements:
P2P direct ~70ms
WHIP ingest (via Edge) ~108ms, SRT ingest (via Edge) ~386ms
(the gap comes from the ingest protocol's receive window, not the codec —
HEVC currently doesn't participate in P2P direct connect)
Diagram points (narration cues): unify the measurement endpoints and clock first, then
look at per-stage instrumentation. Network, encoder and buffer policy can each become the
bottleneck; you can't infer end-to-end latency from a single RTP packet count, an SRT
latency parameter, or a player's target buffer value.
3.2 The protocol triangle: latency, weak-network robustness, ecosystem compatibility¶
Lowest latency
╱ ╲
╱ WebRTC ╲
╱ (WHIP/WHEP) ╲
╱________________╲
Best weak-net robustness ──── Best ecosystem compatibility
(SRT, trading retransmit (RTMP once unified the world
window for weak-net via the Flash Player;
availability) after Flash ended, this
corner collapsed)
Diagram points (narration cues): RTMP's historical edge was exactly the "ecosystem compatibility" corner — almost every browser could play it via Flash. After Flash ended, that corner no longer holds for RTMP, and it never had an upper bound on the "weak-network robustness" corner either (see the head-of-line analysis in §2.2.2). That's the graphical explanation of "RTMP/HTTP-FLV is being retired": it wasn't beaten by a competitor; its one advantage collapsed on its own.
3.3 Weak-network loss handling: head-of-line blocking vs dropping by timeliness¶
[RTMP / HTTP-FLV (TCP)]
frame1 frame2 frame3(lost) frame4 frame5 ...
│
▼
TCP demands "ordered + complete" delivery
│
▼
frame4, frame5 must queue until frame3 is retransmitted
│
▼
Retransmit waits ≥ 1 RTT; on a weak network, repeats stack to seconds, no theoretical bound
│
▼
Player buffer fills → stutter / catch-up / reconnect
[SRT / WebRTC (real-time media policy, typically UDP)]
frame1 frame2 frame3(lost) frame4 frame5 ...
│
▼
Try retransmit, FEC, or wait for reorder per protocol/implementation policy
│
▼
Once stale, recovery can be abandoned to avoid further backlog
│
▼
Viewer experience: usually trades localized quality loss for more controllable latency
3.4 A real project's architecture today: single point and third-party dependency¶
[Viewers / Internet]
│ egress
┌─────────────────┼─────────────────┐
│ │ │
[Edge 1] [Edge 2] ... [Edge N]
└─────────────────┼─────────────────┘
internal forwarding
│
[Origin] ← sole entry, single point, failure = whole-path outage
│
ingest
│
[publish site · long online]
[Origin] ──backup route──> [third-party live service] ← continuity depends on a third party's decisions
Diagram points (narration cues): this is the distribution architecture of one real interactive-live project today, not a hypothetical. Both risks of Dilemma 3 map onto it — the Origin is the sole entry (single-point risk), and the backup route hangs off a third-party cloud (sovereignty risk). This is why the course starts from "dilemmas" rather than jumping straight to solutions.
4. Wrap-up and next episode (script notes)¶
The dilemmas seen here reduce to one line: a traditional live path optimized for one-way viewing cannot be assumed to meet real-time interaction goals. Dynamic or configured buffering, TCP head-of-line blocking, architecture single points and third-party dependencies all need scenario-specific measurement and trade-offs — not conclusions from a protocol name alone.
Next episode, we start taking apart how PPCDN addresses these architecturally — trying P2P first and falling back to Edge in sequence on failure, controlling weak-network backlog by media timeliness, and keeping the critical path in your own hands.