Skip to content

EP08: Video instrumentation — measuring real end-to-end latency

Recap EP07 "How to strengthen resilience on weak networks"
Next EP09 "Latency stability matters more than average latency"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Understand why "end-to-end latency" cannot simply be "now minus timestamp", and where measuring latency in a browser is genuinely hard.
  2. Know the three other playback-experience instrumentation metrics besides latency: video load time, connection success rate and rebuffer rate — and how each is reported and aggregated.

1. Opening hook (script notes)

The current measured baseline: P2P really works at around 70ms; for ingest measured on the Edge-distribution path, WHIP comes in around 108ms and SRT around 386ms (the gap comes from the ingest protocol's receive window, not the codec). How are these numbers obtained? Stamping a timestamp at the publisher isn't enough — the player must know which frame that timestamp belongs to, when that frame actually rendered, and handle the clock error between the two ends. And latency alone doesn't describe the experience; first-frame, connection success and rebuffer caliber have to fill it out.


2. The problem: this subtraction can't be done correctly in a browser

2.1 How timestamps get into the picture

ppobs stamps every video frame (not just keyframes): H.264 wraps it in a user_data_unregistered SEI, HEVC in a Suffix SEI, and AV1 in a metadata OBU. The payload is a fixed 16-byte UUID (the ASCII encoding of "OBS-ABSTS-SEI-v1", directly recognizable in a packet capture) plus an 8-byte big-endian timestamp — 24 bytes total, and it carries only a timestamp, no explicit frame number. The reason: this data is physically embedded in that frame's own bitstream and travels with it, so the instant the player decodes the SEI it inherently knows "this timestamp belongs to the frame I'm decoding right now" — no extra ID is needed to correlate them.

WHIP's "obs-timestamp" DataChannel is a fully independent side channel; by the time it reaches the player it's already detached from the video frame, so this message has to carry its own identity fields to be matched back up: {"frame_no": <RTP sequence number>, "timestamp": <ms>, "rid": "<simulcast layer id>"}. frame_no is not a globally incrementing "frame count" — it's the sequence number of that frame's first RTP packet within that simulcast layer's own sequence-number space; rid distinguishes which layer, since each layer counts independently and frame_no alone isn't unique across layers. Both paths share the same difficulty — reordering, loss and staleness — it's just that SEI is naturally immune to identity mismatches because it's embedded in the frame, while the DataChannel solves the same problem with this set of explicit fields.

2.2 The player's dilemma: no NTP, only a local clock of unknown offset

A latency sample should be generated at the callback where the target frame is actually submitted for display, not when a packet arrives or decode completes. Conceptually it's render_time(frame_id) - capture_time(frame_id). A browser usually only has the OS wall clock and a monotonic clock; even if both ends did clock sync, there remain offset, drift, sync accuracy and network-asymmetry errors.

On the ppobs side, "did clock sync" specifically means: at startup it spins up a background thread that polls public NTP servers (time.cloudflare.com, time.windows.com, pool.ntp.org), takes up to 4 samples per server and keeps the one with the lowest RTT, ending sampling early once an RTT under 30ms is seen; after that it recalibrates every 30 minutes. As long as it has synced successfully even once, it keeps using that last-calibrated offset even if every server later becomes unreachable, rather than falling back to the raw system clock — only if it has never synced successfully at all (e.g. the network blocks UDP port 123) does it fall back to an uncalibrated local wall clock. A browser can't do any of this (it has no access to UDP 123), which is exactly the asymmetry the rest of this section has to work around.

That offset directly breaks the computation: if the player's local clock lags ppobs's and the true latency is small enough, subtracting directly yields a negative number — and latency can't be negative, which shows the subtraction used the wrong baseline. That's the core problem this episode solves.


3. The fix: an application-layer "NTP-style" offset calibration

3.1 The four-timestamp algorithm

You don't need to implement real NTP, just borrow NTP's "four timestamps compute an offset" algorithm over an existing application-layer connection:

T1 = player's local time when it sends the request
T2 = server time when it receives the request
T3 = server time when it sends the reply
T4 = player's local time when it receives the reply

offset = ((T2 − T1) + (T3 − T4)) / 2   ≈ server clock − player clock
rtt    = (T4 − T1) − (T3 − T2)          network round-trip, excluding server processing

An easy trap: the sign of offset is "server clock minus player clock", so correcting local time means adding offset, not subtracting — the sign is easy to flip; the right approach is to verify with concrete numbers before shipping (e.g. set a known case "local clock is 100ms ahead of the server", verify the computed offset is −100, and only then does correction cancel that 100ms). After calibration, latency becomes:

corrected_now  = local reading + offset
one_way_delay  = corrected_now − embedded timestamp

3.2 Why round-trip against ppcenter, not ppobs

The first intuition is "the player should align to the publisher", but that assumption is structurally invalid: P2P quota is at most 3 per publisher (EP02), so the vast majority of viewers on Edge aren't connected to the publisher at all and can't do this round-trip with it.

What you actually estimate is the offset between the capture end's and player's clock domains. Using a clock-synced server as a UTC approximation is one engineering implementation, but neither the server nor the publisher is an error-free UTC: their own sync errors add into the result, and the four-timestamp algorithm also assumes roughly symmetric up/down delay. So record the minimum-RTT sample, offset age and estimation uncertainty alongside latency, rather than treating the calibration as truth. High-precision scenarios can have the capture end and player each sync to the same controlled time source and monitor drift.

3.3 How to sample reliably

  • Run one round of sampling immediately after the connection is established, 5–8 probes per round, and take the offset of the minimum-RTT probe — a smaller RTT means that round-trip was less disturbed by queueing and jitter. This is the standard NTP/SNTP client selection method, far more reliable than a single probe.
  • Report no latency metric before the first calibration completes, to avoid dirty data from an untrusted offset — the same "no trusted data, no output" principle recurring throughout this course.
  • Re-calibrate periodically (on the order of tens of seconds to a minute), because device clocks drift over time and a one-time calibration goes stale on long streams.

4. How timestamps reach the player: two connections, each handling a part

One easily confused point: "calibrating the offset" and "delivering timestamps" go over two fully independent connections, each handling a different job.

  • Players supporting Insertable Streams can parse the SEI/OBU timestamp directly from the bitstream — no extra channel needed.
  • Cross-browser fallback: some browsers (e.g. iOS Safari without Insertable Streams) can't read in-bitstream timestamps. In that case the Origin node parses the publisher's "obs-timestamp" DataChannel, re-wraps the timestamp into a message, and broadcasts it over a WebSocket control connection the player already has open — a connection originally established for layer switching (the ABR layer selection from EP05). It now carries one more duty — forwarding timestamps — instead of opening another DataChannel. One fewer connection, reusing an existing channel, is the direct reason for the implementation choice here.

This mechanism is now extended to the Origin→Edge forwarding segment (viewers on Edge also receive the timestamp broadcast), but limited by the lack of a real multi-node environment, it's only verified that "each independent stage works as designed and degrades gracefully" — a full real multi-node end-to-end integration hasn't been run. Worth stating honestly here rather than dressing it up as "fully verified end to end".


5. Handling residual noise: invalid samples must not enter the official SLA

Negative latency is physically invalid, meaning at least one of clock calibration, frame matching or sampling timing is unreliable. It must not be clamped to 0 and mixed into the official SLA, nor participate as a signed raw value in official percentiles — either would artificially improve the distribution. The right approach is to keep it in a separate data-quality stream, tag the reason and trigger recalibration; the official SLA accepts only non-negative samples with successful frame matching, an unexpired offset and uncertainty within budget, while disclosing the invalid-sample rate. Persistent negatives need an alarm; occasional negatives count toward measurement quality, not business latency.


6. Complementary, not substitutes: distinguish ICE RTT from RTCP RTT

In WebRTC, at least two quantities are casually called RTT and must not be mixed. RTCIceCandidatePairStats.currentRoundTripTime comes from the current candidate pair's STUN connectivity check and describes the ICE transport path; RTCRemoteInboundRtpStreamStats.roundTripTime and other RTP stats are based on RTCP messages and describe the media-control feedback for the corresponding SSRC. Field availability depends on browser and session state.

Metric Coverage
SEI/OBU timestamp one-way latency (this episode's approach) Full path from publisher capture to player render, including encode/decode, jitter buffer, queueing
ICE candidate-pair RTT STUN round-trip on the current ICE candidate pair; excludes encode/decode and playout buffering
RTCP RTT Control-message round-trip for the corresponding RTP/SSRC; not equal to the ICE check RTT, and excludes the full encode/decode path

Both signals should be collected together: currentRoundTripTime is a free direct-connection quality signal, but only this episode's timestamp approach yields latency data covering the full path.


7. Beyond latency: three equally important playback instrumentation metrics

Knowing "what the latency is" isn't enough — a play request that takes forever to connect, connects but takes long to show a picture, or shows a picture but stutters throughout: none of those three bad experiences is captured by a single "end-to-end latency" number. The system maintains three separate instrumentation metrics, each answering one question.

7.1 First-frame time: from playback intent to the first actually-rendered frame

First-frame must unify its endpoints: start is the monotonic time when the player accepts a valid play request, end is when the first non-placeholder video frame is actually submitted for render. Merely receiving RTP packets, a complete frame, or finishing decode does not mean the user has seen a picture. Reporting should also note whether it included auth, signaling and autoplay waiting; if the business needs a separate "first frame after connect", name it as a different metric.

This field is used in two reports: the one-shot connection-result report carries it (to measure "how was this connection"), and the play-session start report carries it (to compute first-frame P50/P95/P99).

The two paths don't currently implement this at the same precision: the P2P path uses the browser's video.requestVideoFrameCallback — a callback the browser fires only when a frame is actually submitted for composited display, which maps exactly onto the "first non-placeholder frame actually submitted for render" definition above. The Edge path currently approximates "decoding has started" by polling getStats and checking for fps > 0, with a 1-second poll interval — coarser granularity than the P2P path. This isn't a doc/code mismatch; it's an engineering trade-off driven by differing browser capabilities — requestVideoFrameCallback is more precise, but today it's only used on the P2P leg.

7.2 Connection success rate: a one-shot report, not just a success/fail boolean

After each play decision settles (P2P or Edge chosen), the player sends a one-shot "connection result" report with these core fields:

  • path: whether this actually used p2p or edge.
  • isOK: whether the connection succeeded.
  • reason: on failure, the specific reason — e.g. ICE checks didn't select a usable direct candidate pair, P2P quota full, or a network/permission error; even on success it can carry the candidate type and path-selection result, for grouping by network condition. Don't judge failure by a "symmetric NAT" label alone.

Why separate "no live stream is publishing" as a failure: if a viewer reports "connection failed" but the stream is actually offline (the host just ended it that second), this failure doesn't indicate a problem with the system's connection capability — the server checks whether the stream is really online; if not, the failure record counts only toward "total play attempts" (used for stats like P2P hit rate) and not toward "connection failures", so it doesn't lower the success rate — avoiding a normal event like a host ending the stream being miscounted as a service fault.

7.3 Rebuffer rate: unify the event and the denominator first

Rebuffering and connection success are not the same class of problem. Define rebuffering as: after the first frame, while the user expects playback and the page is in the counted state, render progress stalls beyond the product threshold; one event lasts until rendering stably resumes. Active pause, background suspension, seek, stream switch and playback end don't count. The threshold is a product parameter recorded per platform and version — pplayer's current default is 2000ms (DEFAULT_LAG_THRESHOLD_MS in lag-tracker.mjs), meaning a <video> element's waiting event has to persist past 2 seconds before it counts as a rebuffer; shorter waits below that threshold (e.g. a brief waiting triggered by an ABR layer switch) don't count. Heartbeat and final reports carry two cumulative fields:

  • lagCount: how many rebuffers since playback started (the player judges "does this count as a rebuffer" against a threshold).
  • lagDurationMs: how many milliseconds those rebuffers totaled.

The time-based rebuffer rate aggregates as "total rebuffer duration ÷ effective watch time"; you can also report rebuffers per minute. Effective watch time excludes pre-startup, active pause and background suspension. Numerator, denominator and event-merge rules must be fixed, or different players' data aren't comparable.

An easily-missed implementation detail: the playback-end report must be sent with the browser's sendBeacon, because at the moment the page closes/refreshes, a normal async request may be terminated before it finishes; sendBeacon is the API browsers designed for "still send a request as the page unloads". But sendBeacon can't set custom request headers, so auth info can't go in the Authorization header like other APIs — it must go in the body. So here it reuses the txTime/txSecret short-lived playback token from EP06, passed directly as a body field, and the player never holds appSecret.

7.4 Why three separate metrics, not one big event

Video load time, connection success rate and rebuffer rate all look like parts of "playback experience", but they weren't crammed into one universal reporting API, because their timing and lifecycle are completely different: connection success is a one-shot judgment at the instant playback is established, load time is a one-shot process from request to first frame, and rebuffering is a state that accumulates throughout playback. Designing to each lifecycle as "one-shot result report" and "session heartbeat" is easier to maintain than one oversized, semantically mixed event, and lets the server aggregate independently along the three dimensions "connection quality", "startup speed" and "playback smoothness" without first deconstructing a big event to compute each dimension's numbers.


8. Diagrams

8.1 The full path of a timestamp from marking to the player

ppobs (accurate UTC after NTP sync)
   │  stamp an absolute timestamp per frame (SEI for H.264/HEVC, metadata OBU for AV1)
   │  WHIP publish additionally broadcasts the same timestamp to the "obs-timestamp" DataChannel
   ▼
Origin (no decode, no transcode; passes through as-is)
   │
   ├─ Path A: browsers with Insertable Streams ──▶ parse SEI/OBU directly from the bitstream
   │
   └─ Path B: cross-browser fallback
         Origin parses the "obs-timestamp" DataChannel
                │
                ▼
         wrap into a message, broadcast over the player's already-open ABR control WebSocket
                │
                ▼
         the player receives the timestamp on the same connection it uses for layer switching

8.2 Offset calibration: four timestamps for one trusted offset

Player                                   Server (ppcenter)
  │  T1: send calibration request ──────────▶ │  T2: receive request
  │                                            │
  │ ◀──────────────────────────────────────   │  T3: send reply
  │  T4: receive reply                         │

offset = ((T2 − T1) + (T3 − T4)) / 2   (server clock − player clock)
rtt    = (T4 − T1) − (T3 − T2)         (network round-trip, excluding server processing)

Repeat 5~8 times, take the offset of the minimum-rtt probe as this round's calibration

8.3 The full instrumentation lifecycle of one playback

Viewer clicks play
   │
   ▼
Play decision settles (P2P or Edge chosen)
   │
   ├──▶ one-shot "connection result" report: path, isOK, reason, costTime
   │      (measures: did this connection succeed, and why it failed)
   ▼
First frame received
   │
   ├──▶ session "start" report: costTime, path, qualityLevel
   │      (measures: request to first frame, how long startup took)
   ▼
Playing continuously
   │
   ├──▶ periodic "heartbeat": cumulative lagCount / lagDurationMs, p2pDelayMs
   │      (measures: how many rebuffers during playback, total duration, live P2P link latency)
   ▼
Playback end / page close
   │
   └──▶ "stop" report (sendBeacon to guarantee delivery): endReason, final lagCount / lagDurationMs
         (auth info in the body, since sendBeacon can't set custom headers)

9. Wrap-up and next episode (script notes)

This episode explained why the end-to-end latency subtraction is error-prone in a browser, and the fix: choose the right round-trip target (UTC, not the publisher), then reuse an existing connection for offset calibration. But latency is just one dimension of playback experience — video load time measures how fast startup is, connection success rate measures whether you can connect, and rebuffer rate measures how stable it is once connected. Together, the four metrics are the complete data behind the words "playback experience".

Next episode, why "average latency" alone isn't enough — latency stability often matters more than the average.