EP08: Video instrumentation — measuring real end-to-end latency¶
| Recap | EP07 "How to strengthen resilience on weak networks" |
|---|---|
| Next | EP09 "Latency stability matters more than average latency" |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Understand why "end-to-end latency" cannot simply be "now minus timestamp", and where measuring latency in a browser is genuinely hard.
- Know the three other playback-experience instrumentation metrics besides latency: video load time, connection success rate and rebuffer rate — and how each is reported and aggregated.
1. Opening hook (script notes)¶
The current measured baseline: P2P really works at around 70ms; for ingest measured on the Edge-distribution path, WHIP comes in around 108ms and SRT around 386ms (the gap comes from the ingest protocol's receive window, not the codec). How are these numbers obtained? Stamping a timestamp at the publisher isn't enough — the player must know which frame that timestamp belongs to, when that frame actually rendered, and handle the clock error between the two ends. And latency alone doesn't describe the experience; first-frame, connection success and rebuffer caliber have to fill it out.
2. The problem: this subtraction can't be done correctly in a browser¶
2.1 How timestamps get into the picture¶
ppobs stamps every video frame (not just keyframes): H.264 wraps it in a user_data_unregistered
SEI, HEVC in a Suffix SEI, and AV1 in a metadata OBU. The payload is a fixed 16-byte UUID (the
ASCII encoding of "OBS-ABSTS-SEI-v1", directly recognizable in a packet capture) plus an 8-byte
big-endian timestamp — 24 bytes total, and it carries only a timestamp, no explicit frame
number. The reason: this data is physically embedded in that frame's own bitstream and travels
with it, so the instant the player decodes the SEI it inherently knows "this timestamp belongs to
the frame I'm decoding right now" — no extra ID is needed to correlate them.
WHIP's "obs-timestamp" DataChannel is a fully independent side channel; by the time it reaches
the player it's already detached from the video frame, so this message has to carry its own
identity fields to be matched back up: {"frame_no": <RTP sequence number>, "timestamp": <ms>,
"rid": "<simulcast layer id>"}. frame_no is not a globally incrementing "frame count" — it's
the sequence number of that frame's first RTP packet within that simulcast layer's own
sequence-number space; rid distinguishes which layer, since each layer counts independently and
frame_no alone isn't unique across layers. Both paths share the same difficulty — reordering,
loss and staleness — it's just that SEI is naturally immune to identity mismatches because it's
embedded in the frame, while the DataChannel solves the same problem with this set of explicit
fields.
2.2 The player's dilemma: no NTP, only a local clock of unknown offset¶
A latency sample should be generated at the callback where the target frame is actually submitted for
display, not when a packet arrives or decode completes. Conceptually it's
render_time(frame_id) - capture_time(frame_id). A browser usually only has the OS wall clock and a
monotonic clock; even if both ends did clock sync, there remain offset, drift, sync accuracy and
network-asymmetry errors.
On the ppobs side, "did clock sync" specifically means: at startup it spins up a background
thread that polls public NTP servers (time.cloudflare.com, time.windows.com, pool.ntp.org),
takes up to 4 samples per server and keeps the one with the lowest RTT, ending sampling early once
an RTT under 30ms is seen; after that it recalibrates every 30 minutes. As long as it has synced
successfully even once, it keeps using that last-calibrated offset even if every server later
becomes unreachable, rather than falling back to the raw system clock — only if it has never synced
successfully at all (e.g. the network blocks UDP port 123) does it fall back to an uncalibrated
local wall clock. A browser can't do any of this (it has no access to UDP 123), which is exactly
the asymmetry the rest of this section has to work around.
That offset directly breaks the computation: if the player's local clock lags ppobs's and the
true latency is small enough, subtracting directly yields a negative number — and latency can't
be negative, which shows the subtraction used the wrong baseline. That's the core problem this
episode solves.
3. The fix: an application-layer "NTP-style" offset calibration¶
3.1 The four-timestamp algorithm¶
You don't need to implement real NTP, just borrow NTP's "four timestamps compute an offset" algorithm over an existing application-layer connection:
T1 = player's local time when it sends the request
T2 = server time when it receives the request
T3 = server time when it sends the reply
T4 = player's local time when it receives the reply
offset = ((T2 − T1) + (T3 − T4)) / 2 ≈ server clock − player clock
rtt = (T4 − T1) − (T3 − T2) network round-trip, excluding server processing
An easy trap: the sign of offset is "server clock minus player clock", so correcting local time
means adding offset, not subtracting — the sign is easy to flip; the right approach is to verify
with concrete numbers before shipping (e.g. set a known case "local clock is 100ms ahead of the
server", verify the computed offset is −100, and only then does correction cancel that 100ms). After
calibration, latency becomes:
corrected_now = local reading + offset
one_way_delay = corrected_now − embedded timestamp
3.2 Why round-trip against ppcenter, not ppobs¶
The first intuition is "the player should align to the publisher", but that assumption is structurally invalid: P2P quota is at most 3 per publisher (EP02), so the vast majority of viewers on Edge aren't connected to the publisher at all and can't do this round-trip with it.
What you actually estimate is the offset between the capture end's and player's clock domains. Using a clock-synced server as a UTC approximation is one engineering implementation, but neither the server nor the publisher is an error-free UTC: their own sync errors add into the result, and the four-timestamp algorithm also assumes roughly symmetric up/down delay. So record the minimum-RTT sample, offset age and estimation uncertainty alongside latency, rather than treating the calibration as truth. High-precision scenarios can have the capture end and player each sync to the same controlled time source and monitor drift.
3.3 How to sample reliably¶
- Run one round of sampling immediately after the connection is established, 5–8 probes per round, and take the offset of the minimum-RTT probe — a smaller RTT means that round-trip was less disturbed by queueing and jitter. This is the standard NTP/SNTP client selection method, far more reliable than a single probe.
- Report no latency metric before the first calibration completes, to avoid dirty data from an untrusted offset — the same "no trusted data, no output" principle recurring throughout this course.
- Re-calibrate periodically (on the order of tens of seconds to a minute), because device clocks drift over time and a one-time calibration goes stale on long streams.
4. How timestamps reach the player: two connections, each handling a part¶
One easily confused point: "calibrating the offset" and "delivering timestamps" go over two fully independent connections, each handling a different job.
- Players supporting Insertable Streams can parse the SEI/OBU timestamp directly from the bitstream — no extra channel needed.
- Cross-browser fallback: some browsers (e.g. iOS Safari without Insertable Streams) can't read
in-bitstream timestamps. In that case the Origin node parses the publisher's
"obs-timestamp"DataChannel, re-wraps the timestamp into a message, and broadcasts it over a WebSocket control connection the player already has open — a connection originally established for layer switching (the ABR layer selection from EP05). It now carries one more duty — forwarding timestamps — instead of opening another DataChannel. One fewer connection, reusing an existing channel, is the direct reason for the implementation choice here.
This mechanism is now extended to the Origin→Edge forwarding segment (viewers on Edge also receive the timestamp broadcast), but limited by the lack of a real multi-node environment, it's only verified that "each independent stage works as designed and degrades gracefully" — a full real multi-node end-to-end integration hasn't been run. Worth stating honestly here rather than dressing it up as "fully verified end to end".
5. Handling residual noise: invalid samples must not enter the official SLA¶
Negative latency is physically invalid, meaning at least one of clock calibration, frame matching or sampling timing is unreliable. It must not be clamped to 0 and mixed into the official SLA, nor participate as a signed raw value in official percentiles — either would artificially improve the distribution. The right approach is to keep it in a separate data-quality stream, tag the reason and trigger recalibration; the official SLA accepts only non-negative samples with successful frame matching, an unexpired offset and uncertainty within budget, while disclosing the invalid-sample rate. Persistent negatives need an alarm; occasional negatives count toward measurement quality, not business latency.
6. Complementary, not substitutes: distinguish ICE RTT from RTCP RTT¶
In WebRTC, at least two quantities are casually called RTT and must not be mixed.
RTCIceCandidatePairStats.currentRoundTripTime comes from the current candidate pair's STUN
connectivity check and describes the ICE transport path;
RTCRemoteInboundRtpStreamStats.roundTripTime and other RTP stats are based on RTCP messages and
describe the media-control feedback for the corresponding SSRC. Field availability depends on browser
and session state.
| Metric | Coverage |
|---|---|
| SEI/OBU timestamp one-way latency (this episode's approach) | Full path from publisher capture to player render, including encode/decode, jitter buffer, queueing |
| ICE candidate-pair RTT | STUN round-trip on the current ICE candidate pair; excludes encode/decode and playout buffering |
| RTCP RTT | Control-message round-trip for the corresponding RTP/SSRC; not equal to the ICE check RTT, and excludes the full encode/decode path |
Both signals should be collected together: currentRoundTripTime is a free direct-connection quality
signal, but only this episode's timestamp approach yields latency data covering the full path.
7. Beyond latency: three equally important playback instrumentation metrics¶
Knowing "what the latency is" isn't enough — a play request that takes forever to connect, connects but takes long to show a picture, or shows a picture but stutters throughout: none of those three bad experiences is captured by a single "end-to-end latency" number. The system maintains three separate instrumentation metrics, each answering one question.
7.1 First-frame time: from playback intent to the first actually-rendered frame¶
First-frame must unify its endpoints: start is the monotonic time when the player accepts a valid play request, end is when the first non-placeholder video frame is actually submitted for render. Merely receiving RTP packets, a complete frame, or finishing decode does not mean the user has seen a picture. Reporting should also note whether it included auth, signaling and autoplay waiting; if the business needs a separate "first frame after connect", name it as a different metric.
This field is used in two reports: the one-shot connection-result report carries it (to measure "how was this connection"), and the play-session start report carries it (to compute first-frame P50/P95/P99).
The two paths don't currently implement this at the same precision: the P2P path uses the
browser's video.requestVideoFrameCallback — a callback the browser fires only when a frame is
actually submitted for composited display, which maps exactly onto the "first non-placeholder frame
actually submitted for render" definition above. The Edge path currently approximates "decoding has
started" by polling getStats and checking for fps > 0, with a 1-second poll interval — coarser
granularity than the P2P path. This isn't a doc/code mismatch; it's an engineering trade-off driven
by differing browser capabilities — requestVideoFrameCallback is more precise, but today it's only
used on the P2P leg.
7.2 Connection success rate: a one-shot report, not just a success/fail boolean¶
After each play decision settles (P2P or Edge chosen), the player sends a one-shot "connection result" report with these core fields:
path: whether this actually usedp2poredge.isOK: whether the connection succeeded.reason: on failure, the specific reason — e.g. ICE checks didn't select a usable direct candidate pair, P2P quota full, or a network/permission error; even on success it can carry the candidate type and path-selection result, for grouping by network condition. Don't judge failure by a "symmetric NAT" label alone.
Why separate "no live stream is publishing" as a failure: if a viewer reports "connection failed" but the stream is actually offline (the host just ended it that second), this failure doesn't indicate a problem with the system's connection capability — the server checks whether the stream is really online; if not, the failure record counts only toward "total play attempts" (used for stats like P2P hit rate) and not toward "connection failures", so it doesn't lower the success rate — avoiding a normal event like a host ending the stream being miscounted as a service fault.
7.3 Rebuffer rate: unify the event and the denominator first¶
Rebuffering and connection success are not the same class of problem. Define rebuffering as: after the
first frame, while the user expects playback and the page is in the counted state, render progress
stalls beyond the product threshold; one event lasts until rendering stably resumes. Active pause,
background suspension, seek, stream switch and playback end don't count. The threshold is a product
parameter recorded per platform and version — pplayer's current default is 2000ms
(DEFAULT_LAG_THRESHOLD_MS in lag-tracker.mjs), meaning a <video> element's waiting event has
to persist past 2 seconds before it counts as a rebuffer; shorter waits below that threshold (e.g. a
brief waiting triggered by an ABR layer switch) don't count. Heartbeat and final reports carry two
cumulative fields:
lagCount: how many rebuffers since playback started (the player judges "does this count as a rebuffer" against a threshold).lagDurationMs: how many milliseconds those rebuffers totaled.
The time-based rebuffer rate aggregates as "total rebuffer duration ÷ effective watch time"; you can also report rebuffers per minute. Effective watch time excludes pre-startup, active pause and background suspension. Numerator, denominator and event-merge rules must be fixed, or different players' data aren't comparable.
An easily-missed implementation detail: the playback-end report must be sent with the browser's
sendBeacon, because at the moment the page closes/refreshes, a normal async request may be
terminated before it finishes; sendBeacon is the API browsers designed for "still send a request as
the page unloads". But sendBeacon can't set custom request headers, so auth info can't go in the
Authorization header like other APIs — it must go in the body. So here it reuses the txTime/txSecret
short-lived playback token from EP06, passed directly as a body field, and the player never holds
appSecret.
7.4 Why three separate metrics, not one big event¶
Video load time, connection success rate and rebuffer rate all look like parts of "playback experience", but they weren't crammed into one universal reporting API, because their timing and lifecycle are completely different: connection success is a one-shot judgment at the instant playback is established, load time is a one-shot process from request to first frame, and rebuffering is a state that accumulates throughout playback. Designing to each lifecycle as "one-shot result report" and "session heartbeat" is easier to maintain than one oversized, semantically mixed event, and lets the server aggregate independently along the three dimensions "connection quality", "startup speed" and "playback smoothness" without first deconstructing a big event to compute each dimension's numbers.
8. Diagrams¶
8.1 The full path of a timestamp from marking to the player¶
ppobs (accurate UTC after NTP sync)
│ stamp an absolute timestamp per frame (SEI for H.264/HEVC, metadata OBU for AV1)
│ WHIP publish additionally broadcasts the same timestamp to the "obs-timestamp" DataChannel
▼
Origin (no decode, no transcode; passes through as-is)
│
├─ Path A: browsers with Insertable Streams ──▶ parse SEI/OBU directly from the bitstream
│
└─ Path B: cross-browser fallback
Origin parses the "obs-timestamp" DataChannel
│
▼
wrap into a message, broadcast over the player's already-open ABR control WebSocket
│
▼
the player receives the timestamp on the same connection it uses for layer switching
8.2 Offset calibration: four timestamps for one trusted offset¶
Player Server (ppcenter)
│ T1: send calibration request ──────────▶ │ T2: receive request
│ │
│ ◀────────────────────────────────────── │ T3: send reply
│ T4: receive reply │
offset = ((T2 − T1) + (T3 − T4)) / 2 (server clock − player clock)
rtt = (T4 − T1) − (T3 − T2) (network round-trip, excluding server processing)
Repeat 5~8 times, take the offset of the minimum-rtt probe as this round's calibration
8.3 The full instrumentation lifecycle of one playback¶
Viewer clicks play
│
▼
Play decision settles (P2P or Edge chosen)
│
├──▶ one-shot "connection result" report: path, isOK, reason, costTime
│ (measures: did this connection succeed, and why it failed)
▼
First frame received
│
├──▶ session "start" report: costTime, path, qualityLevel
│ (measures: request to first frame, how long startup took)
▼
Playing continuously
│
├──▶ periodic "heartbeat": cumulative lagCount / lagDurationMs, p2pDelayMs
│ (measures: how many rebuffers during playback, total duration, live P2P link latency)
▼
Playback end / page close
│
└──▶ "stop" report (sendBeacon to guarantee delivery): endReason, final lagCount / lagDurationMs
(auth info in the body, since sendBeacon can't set custom headers)
9. Wrap-up and next episode (script notes)¶
This episode explained why the end-to-end latency subtraction is error-prone in a browser, and the fix: choose the right round-trip target (UTC, not the publisher), then reuse an existing connection for offset calibration. But latency is just one dimension of playback experience — video load time measures how fast startup is, connection success rate measures whether you can connect, and rebuffer rate measures how stable it is once connected. Together, the four metrics are the complete data behind the words "playback experience".
Next episode, why "average latency" alone isn't enough — latency stability often matters more than the average.