Skip to content

EP10: How to quantify video quality

Recap EP09 "Latency stability matters more than average latency"
Next EP11 "What Netflix's seamless switching teaches us"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Understand why VMAF/PSNR/SSIM don't belong in the real-time forwarding hot path, yet remain usable for offline or side-channel sampling evaluation.
  2. Separate bitstream/signaling-verifiable attributes, encoder self-reported config, and pixel-level perceptual quality, and know what each kind of evidence can answer.

1. Opening hook (script notes)

The last two episodes were about whether latency is stable; this one changes dimension: is the picture any good. Pixel-level comparison does require decoding, but "the real-time forwarding path doesn't decode" doesn't mean "the system can never assess quality". The key is to separate the online hot path, side-channel sampling, and offline experiments.


2. How the industry measures quality, and why it shouldn't block real-time forwarding

2.1 Three mainstream objective quality metrics

Video quality is usually measured with three kinds of objective metrics:

  • PSNR (peak signal-to-noise ratio): pixel-by-pixel numerical difference between the compressed and original picture; the oldest, but weakly correlated with human perception — two pictures that are numerically very close may look very different to the eye, and vice versa.
  • SSIM (structural similarity): not per-pixel differences but whether local structure, brightness and contrast are similar; closer to how the eye perceives "picture structure", but still a pure mathematical model.
  • VMAF (Video Multi-method Assessment Fusion): a perceptual-quality fusion model led by Netflix, combining several sub-metrics with machine-learned weights into a score highly correlated with human subjective ratings — currently the objective metric most accepted in the industry as close to real perception.

2.2 All three need pixels, but can run offline or on a side channel

Whichever you use, the common precondition is getting both the original and compressed pixel matrices frame by frame and comparing them. That means some stage must decode the compressed stream back to pictures to compute these metrics.

Keeping the Origin→Edge real-time media hot path pass-through and without transcoding is a sound trade-off; decoding frame by frame and waiting on VMAF in that synchronous link would add compute, latency and failure surface. But the evaluation task can be an asynchronous side channel: copy a small amount of the stream to an analysis service per policy, or save both the reference source and the encoded output in a test environment, then decode, align and compute offline. It doesn't block the production media forwarding.

So the accurate statement is: don't compute full pixel metrics in the real-time forwarding hot path, but you can run VMAF/PSNR/SSIM via offline baselines, regression tests and side-channel sampling. Full-reference VMAF also requires the reference source and distorted stream to be correctly paired and frame-aligned; without a reference, choose a no-reference/reduced-reference model or subjective sampling — don't pass off a proxy metric as VMAF.


3. Online observation: layer verifiable attributes and proxy metrics

When the online hot path doesn't do pixel comparison, you can combine three kinds of evidence to detect anomalies, but the conclusion should be called "config/transport/playback health", not equated with a perceptual quality score.

3.1 Attributes verifiable from signaling or bitstream

Some attributes can be verified without pixel decoding: negotiated signaling reveals codec, profile, payload type and RID/simulcast declarations; RTP headers and receive stats give actual throughput, packet rate and timestamp cadence; parsing H.264/HEVC parameter sets and frame headers yields encoding resolution, partial color/layer info and keyframe identification. Whether a given item is readable depends on the container, encryption boundary and protocol implementation, but "the server knows nothing" isn't accurate.

Target bitrate, encoding preset, quality parameter and the encoder's internal complexity still usually need sender reporting. Self-reported values should be cross-checked against bitstream-observable values, clearly separating "declared config" from "actual output".

This project has already turned that cross-check into a shipped feature: ppobs reports the encoder's declared codec, resolution, target bitrate, keyframe interval and simulcast layer count as summary fields to ppcenter (POST /v1/encoders/report), and the AI live-health report then compares that declared value against the same stream's measured SRT bitrate — when the declared bitrate is far above the measured bitrate, it docks points on the "uplink" dimension, while a keyframe interval over 2 seconds only triggers a notice, not a deduction. When the encoder hasn't reported, or is running too old a version, that dimension is marked as missing data and the report still generates normally — missing one category of evidence doesn't fail the whole report. This is a real example of recording "declared config" and "actual output" separately and cross-checking them.

An encoder declaring "target bitrate 3Mbps" doesn't mean the actual output is constantly 3Mbps, especially with VBR and low-complexity scenes. Judge from actual send bitrate, available send bitrate, queue, protocol-native loss/retransmit stats and publish-layer state — don't assert degradation from "measured 1Mbps" alone. Content complexity and rate-control mode are also explanatory variables.

Production has a real-world case that illustrates this: while troubleshooting a single H.264 push stream, raising the encoder's target bitrate from 1500kbps to 3000kbps pushed the actual uplink from 2.87Mbps to 5.35Mbps, and SRT ingest packet loss at Origin jumped from 0% to 3.6%-4.7%, after which downstream nodes' RTP fragment reassembly started throwing errors — while at the very same time, internal forwarding from Origin to the downstream nodes stayed at zero loss. Only by putting all three kinds of signal together could the real bottleneck be pinned down to the uplink hop from the push-stream source to Origin, rather than internal forwarding or the protocol itself; looking only at the encoder's self-reported target bitrate, or only at downstream alerts, wouldn't get you there. The specific ceiling on that uplink bandwidth was still an open item, not yet independently measured at the time, and doesn't represent a fixed value across all deployment environments.

3.3 Playback experience: no matter how good the quality, it's meaningless if viewers don't get it

The pull success rate, first-frame P95, rebuffer rate and end-to-end latency P95 from EP08/EP09 aren't strictly "quality" itself, but they determine whether viewers actually receive and smoothly watch the quality the encoder was supposed to provide. A stream configured at 1080p with a healthy uplink but a high rebuffer rate on the player still feels bad — judging quality only from the encoder and link, without the playback leg, gives an incomplete conclusion.


4. An example: same "1080P" label, wildly different experiences

Looking at the "resolution" label alone is not enough. Two streams both configured at 1080p: one with a healthy uplink, stable protocol-native loss/retransmit metrics and zero rebuffering on the player; the other with a persistently backed-up send queue, frequent publish-layer adjustments and a non-trivial rebuffer rate — even with identical encoder-reported resolution, the perceived quality experienced is not on the same level. This is why you combine the three signals rather than staring at one isolated number (e.g. only the "nominal resolution" or only the "current loss rate") — a single signal easily tells a one-sided or misleading story, the same lesson as EP09's "the average lies": the one-sidedness of a single metric is often filled in only by cross-checking multiple dimensions.


5. Diagrams

5.1 The two quality-evaluation paths

[Offline/side-channel: pixel-level objective metrics]
Reference source + sampled stream ──▶ decode and frame-align ──▶ PSNR / SSIM / VMAF
                          │
                          └── runs asynchronously, doesn't block the real-time forwarding hot path

[This project's approach: infer from a combination of proxy metrics]
Encoder self-reported config ──┐
                                ├──▶ three signals cross-check ──▶ a judgment "this stream's quality is probably normal"
Uplink health               ──┤
                                │
Playback experience         ──┘

(online observation does no pixel decoding; offline/side-channel evaluation supplies perceptual-quality evidence)

5.2 The three pieces of the proxy-metric puzzle

┌─────────────────┐   ┌─────────────────┐   ┌─────────────────┐
│ signaling/bitstream │ │   uplink health  │ │  playback experience │
│ codec/resolution/  │ │ protocol-native  │ │ pull success rate/  │
│ layers/keyframe    │ │ stats/ degrade & │ │ first-frame P95/    │
│ interval/simulcast │ │ recovery events/ │ │ rebuffer rate/      │
│ layer count        │ │ config vs actual │ │ E2E latency P95     │
└────────┬─────────┘   └────────┬─────────┘   └────────┬─────────┘
         │                      │                      │
         └──────────────────────┼──────────────────────┘
                                 ▼
                    combined judgment "is quality a problem"
                 (any one piece alone can be one-sided;
                  cross-checking approaches real experience)

6. Wrap-up and next episode (script notes)

This episode separated online from offline evaluation: the real-time forwarding hot path shouldn't be blocked by pixel-level analysis, but offline baselines and side-channel sampling can absolutely compute VMAF/PSNR/SSIM. Online, distinguish signaling/bitstream-verifiable attributes, encoder declarations, link health and playback experience; these signals are good for finding anomalies but must not masquerade as a perceptual quality score.

Next episode, let's talk about Netflix — what its seamless-switching experience can teach this system.