Skip to content

EP04: Transport showdown — RTMP / SRT / WHIP

Open on YouTube

Recap EP03 "Content sovereignty: control it yourself, the ultimate moat" wrapped up Module 0 (breaking out)
Next EP05 "How Simulcast works and what it's for"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Explain the concrete loss/latency trade-off mechanisms of RTMP, SRT and WHIP, rather than the coarse "TCP bad, UDP good" conclusion.
  2. Understand why this project keeps WHIP and SRT side by side instead of picking one and dropping the other — they solve different problems.

1. Opening hook (script notes)

Module 0 covered three dilemmas and the overall architecture. From this episode we enter Module 1, back to the basics — ingest protocols. EP01 already covered RTMP/HTTP-FLV's head-of-line blocking on weak networks; this episode digs one level deeper: how do RTMP, SRT and WHIP actually handle "loss"? And why does this project support both SRT and WHIP rather than keeping just one?


2. RTMP: a retiring protocol still worth understanding

RTMP (Real-Time Messaging Protocol) usually runs over a persistent TCP connection; after the handshake, messages are split into chunks multiplexed over several chunk streams. AMF is mainly used to serialize objects for command and data messages; audio/video messages carry the respective encoded payload, so it's wrong to say flatly that "media is AMF-encoded". RTMPS carries RTMP over TLS, giving the link confidentiality and integrity; it does not remove TCP head-of-line blocking.

This design made sense in its era: TCP's reliable delivery saved the protocol from handling retransmission itself. But the cost is exactly the head-of-line blocking from EP01 — if any TCP segment is lost, later data that already arrived must queue until it is retransmitted. That wasn't a problem when RTMP was designed (typical scenarios were stable wired uplinks), but on today's mobile/weak networks it's a handicap.

RTMP's place in the ecosystem is thin now: Flash's end removed it from the playback side (EP01), leaving only "ingest entry" — some old encoders/publishing tools still speak only RTMP. Even that is being replaced: SRT and WHIP both offer ingest better suited to modern networks, which is why this project's public ingest protocol list has no RTMP.


3. SRT: a modern choice designed for weak networks

SRT (Secure Reliable Transport) sits on UDP and targets low-latency reliable transport over unstable networks. Understand it by separating three mechanisms:

  • ARQ: the receiver detects missing packets by sequence number and sends NAK; the sender retransmits while data is still timely. It solves "how to recover loss".
  • TSBPD: schedules delivery based on send timestamps and the negotiated latency, absorbing jitter and restoring send timing. It solves "when to deliver" and is not a synonym for the ARQ window.
  • TLPKTDROP / stale-packet drop: when enabled and its conditions are met, abandons data that can no longer be delivered on schedule, letting both ends skip stale packets. It solves "when to stop remediating", and should not be written as "TSBPD drops packets automatically".

The latency parameter gives retransmission and jitter absorption a time budget: larger usually improves recovery chances and adds delivery wait; smaller does the opposite. But it is not a strict mathematical latency bound for the whole SRT session: handshake, network queueing, application buffering, configuration differences, disconnects and reconnects all lie outside it. For this project, use end-to-end instrumentation as the source of truth rather than summing config values to derive total latency.

SRT can be configured with AES-based payload encryption. The exact mode and key length depend on protocol version and implementation configuration; it protects the transport link and is not automatically end-to-end encryption across every intermediate processing node.

SRT's core question is "can a weak uplink still deliver the picture intact" — trading latency for robustness is its design intent, not a defect.


4. WHIP: standardizing WebRTC into an ingest protocol

WHIP (WebRTC-HTTP Ingestion Protocol) is not a brand-new transport but a standardized "publish signaling procedure" added to WebRTC (using HTTP for the SDP offer/answer exchange); media still travels the native WebRTC stack:

  • ICE opens the network path (including NAT-traversal address negotiation).
  • DTLS-SRTP encrypts the media on the WebRTC peer connection. If media is processed or re-encrypted after SRTP terminates at an SFU/Edge, it's hop-by-hop protection, not E2EE only decryptable by the communicating parties; true E2EE needs additional frame-level media encryption and key distribution.
  • RTP carries media, with NACK (selective retransmission requests) and optional FEC handling loss, while congestion control (e.g. GCC) adjusts the send bitrate.

WHIP has no SRT-style TSBPD. RTP implementations usually keep a receive/reorder history counted in packets or bytes to detect reordering, generate NACKs or hold retransmittable data; e.g. this project's WebRTCInboundRTPBufferSize is a packet capacity. A packet cache is not a time window: the same 512 packets cover different durations at different bitrates, packet sizes and frame rates, and it can't be used to derive playback wait directly. Actual waiting is also determined by the jitter buffer, NACK policy, decode deadline and congestion control.

WHIP was published as RFC 9725 (2025), no longer an IETF draft. It standardizes WebRTC ingestion's HTTP signaling and resource lifecycle; media security, congestion control and loss recovery still come from the WebRTC/RTP system.


5. Side-by-side comparison

Dimension RTMP SRT WHIP
Transport TCP UDP + ARQ UDP + RTP (NACK / FEC / congestion control)
Loss handling TCP must retransmit and deliver in order; possible head-of-line blocking ARQ recovery; TSBPD timed delivery; configured stale-drop can skip late data NACK/FEC used per negotiation and implementation; jitter buffer and decode deadline decide waiting
Latency bound No strict math bound latency is an engineering budget, not an unconditional end-to-end bound Dynamic buffering and congestion control; no unconditional math bound
Encryption RTMP itself unencrypted; RTMPS over TLS Configurable AES payload encryption Mandatory DTLS-SRTP link encryption; not automatically E2EE
Native browser support No longer (relies on the discontinued Flash) No Yes (standard WebRTC APIs)
Typical role Legacy, exiting The modern weak-uplink choice Mainstream real-time publish/play protocol
Supported here? No Yes (ingest) Yes (ingest + playback all via WHEP)

6. How this project chooses: unify the signal, differentiate the handling

WHIP and SRT are not "who replaces whom" here but provided together for different ingest scenarios: prefer WHIP when the network is clean and you want the lowest latency; use SRT and trade a little latency when the uplink is unstable and you need more robustness.

The two receive mechanisms are not equivalent: SRT's TSBPD schedules delivery by timestamp, while an RTP packet cache usually keeps history by packet/byte count. This project can normalize both sides into an Unrecoverable Loss Rate (ULR) for product policy, but that is not a standard metric jointly defined by SRT and WebRTC. The computation must also unify sampling window, denominator, retransmission de-duplication and counter-reset semantics — you can't compare directly just because field names look similar.

This unified signal already backs a single degrade state machine shared by WHIP and SRT in this project (see docs/design/publish-degrade-protocol.zh-CN.md for the design): the raise threshold degradeRaisePct defaults to 0.8%, the recovery threshold degradeLowerPct defaults to 0.3%, the sampling window degradeSampleSec defaults to 6 seconds, and the cooldown after each state transition, degradeObservationSec, defaults to 60 seconds. These are configuration defaults, not an SLA guaranteed network-wide — production thresholds still need calibration against real link measurements.

Degradation is two steps, cost increasing:

  1. Lower the encoding bitrate first: takes effect in real time, no interruption.
  2. Only if bitrate is at the floor and still short, lower resolution (drop simulcast layers): requires restarting the publish, about 1 second of interruption.

The logic is straightforward: bitrate changes are free and can trigger often; layer changes cost an interruption and are used only when "even the floor isn't enough". The "about 1 second of interruption" figure today is mainly an SRT-side production observation — mmx carries a dedicated log marker, SRT publish resumed on path X after <gap> of ingest interruption, used to measure actual reconnect time — not a value promised across every network environment; the WHIP side still lacks a production measurement at the same granularity for its restart cost.


7. Diagrams

7.1 Loss handling across the three protocols

[RTMP / RTMPS (TCP)]
loss → TCP must retransmit → in-order delivery → later bytes queue → wait can grow unbounded

[SRT (UDP)]
ARQ: detect a gap → NAK → retransmit while still timely
TSBPD: schedule delivery by timestamp + latency
TLPKTDROP: when enabled and a packet is stale, skip data that can't be delivered in time

[WHIP (WebRTC/RTP, typically UDP)]
sequence numbers and RTCP feedback: detect loss → NACK/FEC can attempt recovery
jitter buffer: handle jitter and playout timing against a dynamic target
decode deadline policy: drop frames and keep playing once recovery is worthless
Note: a packet-count RTP cache is not a millisecond window

7.2 Unify the decision semantics, not the raw counters

SRT native stats                        RTP/WebRTC native stats
loss / retrans / drop / too-late        sequence gap / late / RTX / FEC
         │                                      │
         └──────────────┬───────────────────────┘
                        ▼
  Derive a product metric from window, original-packet identity, recovery result and playout deadline
                        │
          ┌─────────────┴─────────────┐
          ▼                           ▼
  Metrics keep degrading →         Metrics recover steadily →
  lower bitrate / drop a layer     cautiously step back up

Both sides can share the product decision semantics of "when to degrade, when to recover", but the different protocols' cumulative counters cannot be dropped into one formula. Formal thresholds also need calibration per protocol, direction, encoding layer and real network samples.


8. Wrap-up and next episode (script notes)

This episode dug one level below EP01's "head-of-line blocking": RTMP/RTMPS inherits TCP's reliable-ordered byte-stream semantics; SRT combines ARQ, TSBPD and optional stale-drop; WHIP reuses WebRTC/RTP feedback, congestion control and dynamic buffering. None offers a math latency bound independent of implementation and network conditions. This project keeps WHIP and SRT together and normalizes their stats for one product degradation policy.

Next episode: Simulcast — how one publish produces multiple quality layers in one pass, and why the server needn't transcode.