EP19: Load-testing methodology — from a single box to global load¶
| Recap | EP18 "Recording & playback: adding VOD to live streaming" |
|---|---|
| Next | EP20 "Is a 30% bandwidth saving realistic? A real-data review" (series finale) |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Understand testing is layered: unit tests, black-box integration tests (local one-shot instance / production regression against the live deployment), single-box single-stream latency measurement, single-box capacity load testing, real-network acceptance, global capacity load testing — each layer answers a different question and none can substitute for another.
- See how "end-to-end latency" is measured: instrumentation, clock calibration, cross-browser, sample size and error, and the difference between "one measurement" and a network-wide SLA.
- Understand why some things can't be mocked: NAT diversity, real racing timing, failure fallback, real provider network throttling — only testable on real networks.
- Know where this system stands on the test ladder: latency measurement is complete; black-box
testing now has two tiers (a local one-shot binary covering the recording path, and production
regression against the live
api.pp-cdn.org/record.pp-cdn.orgcovering the Play/NAT/records/split-rec HTTP contract); the self-hosted standalone node's single-box capacity now has two real stress-test reports. But global capacity load testing still hasn't been done, and the platform's own Origin/Record/EdgemmxNodeCapacity(EP15) remains a pending calibration — these two are not the same thing.
1. Opening hook (script notes)¶
The last eighteen episodes took the system from architecture to deployment to recording. This episode does the most-often-skipped and most failure-prone thing: testing. Many people think "tests passed" means "
go test ./...is green", but in the real world there are several steps between green unit tests and "surviving production traffic". This episode climbs those steps bottom-up — from single-box single- stream to the (not yet reached) global load test. I'll try to make clear what each level answers and what it doesn't, because the biggest testing trap is using one level's pass to prove another level's conclusion.
2. Testing is a layered climb¶
A table of what each layer tests and doesn't:
| Level | Typical practice | Answers | Doesn't answer |
|---|---|---|---|
| Unit tests | go test ./..., go vet ./... |
Is the logic right, are edges handled | Does it work on a real network |
| Black-box integration tests (local, one-shot) | integrationTest/record: compiles the real mmx binary, real ffmpeg publish |
Is the protocol/interface contract right | Performance ceiling, capacity |
| Black-box production regression tests | integrationTest/api: real HTTPS requests against the live deployment |
Do production's auth/validation/routing contracts still match the docs | Whether media actually establishes, playback quality |
| Single-box single-stream latency measurement | One real stream, a limited network scenario | End-to-end latency magnitude and composition | Concurrency scale, network-wide SLA |
| Single-box capacity load test | Concurrency-ladder load against one real node | Where this machine's concurrency ceiling is, whether degradation is linear or a cliff | Multi-box/multi-region capacity |
| Real-network acceptance | Two genuinely independent NAT environments | NAT traversal, racing, failure fallback | At-scale capacity |
| Global capacity test | Multi-region, high concurrency, peak traffic | System capacity ceiling, where the bottleneck is | (not done yet) |
Core principle: test at a level and you only get that level's conclusion. Green unit tests prove "the code logic is self-consistent", not "production can take it" — repeated in many projects but still easy to forget when actually doing it.
3. How to measure "end-to-end latency": caliber and principle¶
Latency is the through-line of this course; how is it measured? The system has an established reproducible, publishable method:
① Stamp an absolute timestamp in the bitstream. ppobs writes an absolute UTC timestamp into the stream during capture (H.264/H.265 SEI); Origin/Edge pass it through hop by hop without decode or transcode — so the instrumentation itself doesn't break the "no transcoding on the network side" principle (EP05/EP10).
② The player does application-layer clock calibration. ppplayer has no system-level NTP, so it performs one application-layer clock calibration against ppcenter (NTP-style offset estimation), corrects the local clock, then subtracts the embedded timestamp:
corrected_now = local_clock + offset
one_way_delay = corrected_now - embedded_timestamp
③ Don't hide the unobservable segment with a fixed constant. If the span between stamping and display can't be measured directly, instrument or report a "measurement boundary" separately, and validate with high-speed photography/external reference. Mechanically adding one compensation to every reading introduces systematic error, and different encoders, devices and render paths don't share one constant.
④ Cross-browser measurable. iOS Safari lacks WebCodecs Insertable Streams and can't read the in-bitstream
SEI — so to measure those devices too, the Edge node parses the ingest stream's SEI and delivers
OBS_TIMESTAMP to the player over the ABR control WebSocket (reusing the connection the player already holds
for layer switching).
⑤ From a single reading to a distribution with sample info. Player reports are grouped by scenario,
codec, device and network, reporting sample size n, test window, P50/P80/P95/P99 and, where possible, a
confidence interval or repeat-trial dispersion. A percentile without sample size can't be used to judge
stability, and mixing scenarios into one aggregate misleads.
The results this method produces (after tuning; see the Live Streaming Latency Performance Test Report,
ppcenter/docs/test/live-latency-performance-test-report.zh-CN.md):
| Path | Measured latency |
|---|---|
| P2P direct | ~70ms |
| WHIP ingest (Edge delivery) | ~108ms |
| SRT ingest (Edge delivery) | ~386ms |
Under these real test conditions, P2P direct has the shortest path and the lowest latency; when delivery goes through Edge, the ingest protocol is the main variable — WHIP ingest has no fixed receive window, while SRT deliberately widens its ingest-side TSBPD receive window to retransmit around packet loss on weak networks. That's a protocol-level trade-off between "loss robustness" and "latency", not an implementation defect. Note carefully that the variable being compared here is the ingest protocol (WHIP vs SRT), not the codec (HEVC vs H264) — isolating the codec's own effect on latency would require holding the ingest protocol fixed and only switching codec in a controlled experiment, which this report doesn't cover. You can't use a protocol-difference number to answer a codec-difference question — this is exactly the "don't mix two different axes" point the next episode keeps hammering. These numbers themselves are still approximate point estimates; three numbers alone aren't enough to assert that the protocol difference accounts for the entire observed gap. The report must also give each group's sample size, sampling window, device, encoding parameters, network conditions, clock-sync error and percentiles; with insufficient samples it should be called "this observation" only, not a network-wide SLA.
4. Black-box integration tests: test against the real artifact, in two tiers¶
Before single-box measurement, there's a layer more "real" than unit tests but more controllable than a real deployment: black-box integration tests. In this system it actually comes in two tiers, differing in how "real" they are.
4.1 Testing against a local, one-shot real binary — integrationTest/record¶
Taking the recording path (EP18) as the example, integrationTest/record:
- Compiles the real mmx binary (not an in-process mock), starting a real process;
- Uses real ffmpeg RTMP publish to trigger fMP4 segments actually landing on disk;
- Covers
/v3/recordings/*and the split-rec two-phase protocol and auth-failure branches over real HTTP; - Cross-checks the API responses against the files really written on disk.
Its value: an in-process unit test can never prove that "real binary + real encoder + real network stack" is connected. A black-box test truly assembles and runs the whole pipeline.
It has a few iron rules worth borrowing for any integration test:
- One-shot, isolated: each run compiles, starts the process and uses a temp dir itself; artifacts and secrets are regenerated each time.
- Never touch production:
SPLIT_REC_SECRETis randomly generated each time, and thedeletesegmentcase only deletes files it created. Tests may really write and delete, but only in their own one-shot sandbox. - Know your "not covered" list: this suite covers only the record role, and without local object storage it only verifies "upload doesn't block the response", not real upload success. Writing down "what's not covered" matters more than a few extra cases.
4.2 Testing against the real live deployment — integrationTest/api¶
This system has a second, even more "real" tier of black-box testing: no compiling, no local process — it
sends real HTTPS requests directly against the live production deployment at https://api.pp-cdn.org
(ppcenter) and https://record.pp-cdn.org (mmx-recorder), covering the Play API (/v1/play/requests, the
Edge/P2P routing entry point from EP13), NAT probing (/v1/nat/probe), recording round/manifest queries
(/v1/records/*), and the split-rec two-phase protocol. It uses a dedicated registered test account
(credentials live in a machine-local .env, never committed to source), covering negative cases — auth
failure, field validation, expired tokens, cross-app token misuse — and the positive cases are written
honestly too: for example the Play API's happy-path test accepts any of three outcomes — 200 (a healthy
Edge happened to be available), 503 no_edge_available (none happened to be available), or
402 account_in_arrears (the test account happened to be in arrears) — because this suite doesn't control
whether an Edge is actually available at that moment, and an honest "no capacity" still proves the
auth+validation layer works; it shouldn't count as a test failure.
This tier isn't decorative: the production deployment log (docs/deployment/ppcdn生产部署总结.md) shows it
being run as a real regression case after every ppmmx/ppcenter release, with recent runs at
35 PASS / 0 FAIL / 5 SKIP (the skips are all environment-dependent, such as the manifest backend being
temporarily unreachable). It isn't just there to confirm "nothing broke", either — running it for real has
caught leftover test-code drift after a protocol upgrade (the split-rec signing-key source — the test code
never followed through on the "stop using the deployment-level shared secret" upgrade) and one real
production nginx routing gap; both the findings and the fixes are recorded in the deployment log, another
example of a problem only a real environment surfaces.
But it still only tests the HTTP control-plane contract: the Play API returns the URL and mode
(edge-only/p2p-connect) telling the player which Edge to WHEP into, and this suite only verifies that
response is correct — it never actually opens a WebRTC/WHEP connection to check whether video shows up.
Automated end-to-end testing of the real origin/edge media path is still empty (see Section 7).
5. Real-network acceptance: what can't be mocked¶
Some things no amount of good unit or integration testing can mock; only real networks work. P2P has one real proven run of ~70ms, but one success can't cover NAT diversity and the failure distribution, so systematic acceptance remains:
- Prove baseline connectivity first: with the easiest-to-admit combination (one end on a public IP), confirm offer/answer/ICE/media is connected, ruling out "never even reached P2P" issues.
- Then build real NAT combinations: both ends each behind a real carrier/router NAT, and the two must not be the same NAT device. This is a permanent blind spot of mocks — you can mock the type "symmetric NAT", but not a real NAT device's behavior on a real carrier network.
- Verify "rejection must also be imperceptible": after deciding against TURN (EP13), "is the
should-reject combination cleanly rejected and completely imperceptible to the user" becomes the key to
shipping. This validates not "can it identify CGNAT" (the
/v1/nat/proberequest/response contract already has black-box production regression coverage from §4.2), but "after identifying it, does the viewer feel nothing" — a UX dimension the contract test can't reach. - Independent media evidence: look at this stream's session bytes on the Edge node — after P2P is established it should trend to zero. This is independent evidence that "media really didn't go through Edge", rather than "ppcenter reported p2p-connect but it's actually still falling back to Edge".
- Separate exploratory scenarios from pass/fail: cases like "the real elapsed-time distribution of falling back to Edge in sequence after P2P fails" (today the server branches by codec up front and falls back sequentially — it's not parallel racing; see EP13/EP09) have the goal of getting a set of real data, not a pass/fail verdict — because that data is used to calibrate parameters (e.g., whether the 5-second handshake timeout and the first-frame grace period are set right), not to produce a verdict.
- Keep the evidence: each run keeps the
chrome://webrtc-internalsscreenshot, the full/v1/play/requestsresponse, and ppcenter's admission reason that time. A "tested" without evidence is not tested.
In one line: real-network acceptance is slow, human-driven, and irreplaceable.
6. Single-box capacity load testing: same method, two machines, degradation is always a cliff, never a gradient¶
The previous sections tested "is the protocol/interface right" and "latency magnitude" — both single-stream views. One question remains unanswered: how many concurrent streams can one real machine actually hold, and what does it look like when the ceiling collapses?
This is related to — but not the same as — the pending item EP15 left behind ("the single-box capacity
value looks optimistic and needs load-test calibration"): EP15 was about the platform's own
Origin/Record/Edge three-role node pool, whose mmxNodeCapacity still hasn't been calibrated by a real load
test to this day. What's discussed here is a different, already-running load-test effort aimed at
self-hosted standalone nodes (the single-machine, single-role deployment licenseCode self-hosting
customers use, covered in EP17). The method is the same: hit one real node with whep-loadgen in a
concurrency ladder, watching CPU/memory/disconnects/packet loss as the load climbs. There are now two
reports, and because the machine specs differ, the numbers can't be mixed:
2 vCPU / 4GB (LightNode, Manila, first round, single-track SRT H264,
benchmark/report/ppmmx-standalone-stress-test.zh-CN.md): 12→25→50 concurrent streams all held stable
(CPU averaging 26%→63%), then 75 failed outright: CPU spiked to 134%/177% on the 2-core box, B2 segment
uploads timed out, and the server dropped frames for slow readers. But this report's own §7 partly overturns
that conclusion: switching to a different cloud provider (AWS Lightsail, same 2vCPU/2GB spec) and re-running
the same ladder, it cleanly reached 250 concurrent streams, only failing at 300 due to memory exhaustion.
Comparing the two revealed that the first provider's real measured egress bandwidth was only about
50-106Mbps (the advertised 1Gbps was never reached) — that, not the machine's CPU ceiling, was the real
bottleneck behind the 50→75 crash.
1 vCPU / 2GB (LightNode, a different instance, same method, only the spec changed, 2026-10-08,
docs/test/ppmmx-standalone-1vcpu-stress-test.zh-CN.md): stable throughout up to 18 concurrent streams
(CPU ≤53%), then at stream 19, CPU jumped from ~40% to 88-103% within about 15 seconds and stayed there,
with SRT ingest RTT spiking from ~4ms to 100-500ms — this time the load generator and the target machine
were in roughly the same data center (RTT ~4ms), so this really did measure the machine's own ceiling. This
time it hit CPU first, not memory (unlike the earlier finding that smaller specs hit memory first —
because this test cut away an entire core).
Put the two reports side by side and three lessons emerge, unrelated to unit testing but just as important:
- Degradation is a cliff, not a gradient. Both machine types went from "healthy" to "overloaded" within seconds to the low tens of seconds, with almost no transition band — CPU climbed gently from 10→18 streams (~2-3%/stream), then jumped straight to 88-103% between 18→19 streams, when extrapolating the earlier slope would predict only ~48-55%.
- You can't linearly extrapolate across machine types. Halving the 2vCPU machine's slope ("CPU% ≈ 15 + 0.97% × concurrency") to estimate the 1vCPU machine's capacity gives an optimistic ~35 streams, but the real measurement was only 18 — a single-core box has no second core to fall back on for GC, network polling and bursts, so the cliff arrives much earlier and much narrower.
- "Is it the machine's fault" can only be confirmed by changing environments. The same 2vCPU machine, moved to a cloud provider that actually delivers the network bandwidth, jumped from holding 50 to holding 250 — proof that the original "crashed at 75" measurement wasn't really testing CPU at all, it was testing the provider's hidden egress throttling. Without re-testing in a different environment, a single load test's failure conclusion can easily mistake "the provider's hidden rate limit" for "the machine's real ceiling."
This round of testing also had an unplanned payoff, echoing Section 4's theme — some bugs only surface in
a real environment: during the 1vCPU test, two real bugs affecting every licenseCode self-hosted
deployment customer were incidentally found and fixed — NodeRegisterAck had never echoed nodeSecret back
to nodes registering with only a licenseCode, so those nodes could never get credentials, their publish
whitelist stayed permanently empty, and no appId could ever push a stream through (now fixed and deployed to
production); and under the default bridge-network deployment, the WebRTC ICE candidate address collected
was the container's internal IP rather than its public IP, breaking WHEP playback (recorded as a known
issue, not yet code-fixed — worked around that time with a manual config option). These two bugs weren't the
"data points" the load test set out to collect, but they were its most valuable output — once again
confirming Sections 4 and 5's point: some things only surface when you run real binaries under real network
conditions; unit tests and small-scale mocks never force them out.
7. Honest status: the ladder is only half climbed¶
This episode's stance: honestly mark where each rung has been reached, and which ones are still empty.
- Done: unit tests (
go test/go vetgreen); single-box single-stream latency measurement (the report is complete and reproducible); black-box integration tests covering the recording path (local one-shot binary, §4.1); black-box production regression tests covering the Play/NAT/records/split-rec HTTP contract (against the live deployment, §4.2); single-box capacity load tests of self-hosted standalone nodes (two real reports, §6). - Half done: P2P really established a connection and was observed at ~70ms, but sample size and error analysis across carriers, NAT types and devices are still insufficient to form a hit rate or SLA.
- Still empty:
- No automated end-to-end test for the real origin/edge media path (WebRTC connection setup + playback quality) — the §4.2 production regression only verifies that the Play API's returned URL/mode is correct; it never actually takes that URL and opens a WHEP connection to look at the picture;
- No global capacity load test;
- The platform's own Origin/Record/Edge
mmxNodeCapacity(EP15) still awaits load-test calibration — this is a different thing from the §6 standalone-node load tests, and the latter's results can't be used to claim the former is calibrated.
The conclusion must be blunt: production capacity cannot be derived from unit test results alone. Not humility but engineering fact — capacity is beaten out of real streams, not derived from code.
8. Take-away testing principles¶
Seven principles:
- Test the right question at the right level. Don't use green unit tests to prove production capacity.
- Test against the real artifact. Compiling a real binary and running a real publish is far more real than an in-process mock.
- Isolated, one-shot, never touch production. Tests may really write and delete, but only in their sandbox.
- Write down the "not covered" list. Knowing what you can't test matters more than a few extra cases.
- Separate "pass/fail" from "collect data". For parameter-calibration tests, the goal is the distribution, not a binary verdict.
- Load tests must approach reality. NAT diversity, geographic distribution, peak shape — the more perfect the mock, the further from reality.
- Doubt your own test environment. The same "failure" conclusion can turn out completely wrong on a different cloud provider or a different machine — before concluding anything, ask: "am I measuring the machine itself, or a hidden limit of the environment?"
9. Diagrams¶
9.1 The test ladder: higher is more real, slower and costlier¶
Global capacity test ← not done (multi-region / high concurrency / peak)
▲
Real-network acceptance ← plan complete, lacks real-environment execution (NAT diversity, racing, fallback)
▲
Single-box capacity test ← done (standalone nodes, 1vCPU/2vCPU reports; cliff-edge degradation)
▲
Single-box latency test ← P2P ~70 / WHIP ingest ~108 / SRT ingest ~386ms; samples & error to be expanded
▲
Black-box production regr. ← done (real HTTPS against prod API; contract only, not real media)
▲
Black-box integration (local) ← done (real binary + real publish; covers only record)
▲
Unit tests ← done (logic correctness; can't derive capacity)
9.2 The end-to-end latency measurement path¶
ppobs capture ── write an absolute UTC timestamp (SEI)
│
▼
Origin / Edge ── pass through hop by hop, no decode, no transcode
│ │
│ └─ Edge parses SEI, delivers OBS_TIMESTAMP over the ABR WS (to browsers without WebCodecs)
▼
ppplayer ── application-layer clock calibration (aligned to ppcenter) → subtract
│
▼
report to ppcenter ── group by scenario, record n, P50/P80/P95/P99 and error info
10. Wrap-up and next episode (script notes)¶
This episode climbed the test ladder bottom-up: real P2P direct connections run at ~70ms; when delivery goes through Edge, WHIP ingest runs ~108ms and SRT ingest ~386ms (the gap comes from the ingest protocol, not the codec). Black-box testing now has two tiers — a local one-shot binary (covering the recording path) and production regression against the live deployment (covering the Play/NAT/records/split-rec contract, but unable to touch the real media path). Self-hosted standalone nodes now also have two real capacity load-test reports, with a shared lesson: degradation is a cliff, not a gradient, and you can't linearly extrapolate across machine types. Green unit tests don't mean production capacity, and one real successful run doesn't mean a network-wide hit rate or SLA; global capacity load testing is still blank, and the platform's own Origin/Edge
capacitystill hasn't been calibrated by load testing.Next episode is the last one in this course. We return to the original business question: does the "save 30% bandwidth" claim have real data behind it? We'll lay out, item by item, what's measured, what's modeled, and what's still unverified.