EP16: Handling sudden bursts of high concurrency¶
| Recap | EP15 "The relationship between horizontal scaling and low latency" |
|---|---|
| Next | EP17 "Deploy from scratch: run your own CDN in 15 minutes" |
0. Goals of this episode¶
By the end, the viewer should be able to:
- Separate "heavy load" from "load that arrives suddenly": the hard part of a burst isn't the total but the peak arriving faster than you can react — too fast to scale manually.
- Explain the boundaries of three common tools: P2P offloads only a fixed small number of views per publisher, on-demand pull reduces idle upstream cost, publisher-side degradation protects the publish link; viewer burst capacity still relies on pre-provisioned Edge, admission and scaling.
- Know what happens when elasticity can't hold: soft-threshold alarms + manual scaling, and their real limits in a burst (no auto-scaler, each channel's 30-second cooldown swallowing alarms).
- Correctly understand cloud vendor quotas: a monthly egress allowance is usually a billing threshold that turns into pay-as-you-go beyond it, not an inherent cut-off; only a product with a hard cap or arrears policy affects service.
1. Opening hook (script notes)¶
Last episode we said horizontal scaling buys capacity. But real traffic often isn't a slow climb letting you add machines calmly — it's an event starting on the dot and, within seconds, hundreds or thousands clicking in at once. That's a sudden burst of high concurrency. The question this episode answers: when the load isn't "heavy" but "urgent", what holds? Honestly, it has no "auto-scaler" to add machines instantly. It relies on several layers of elasticity, plus a manual process. This episode makes those layers clear and honestly brings out where they can't hold.
2. Why a burst is hard: not "big total", but "fast arrival"¶
Distinguish two loads first:
- Uniform high load: user count climbs slowly, you see capacity alarms, and you have time to add machines step by step. That's EP15's scenario.
- Sudden burst: traffic hits a peak in a very short time, far above the mean. Typical sources: the opening instant of an event, popular content being shared out, all viewers flooding in the moment a stream starts.
The real difficulty of the latter isn't the total but the time scale:
- Unpredictable: you don't know how many will flood in.
- Too late to scale manually: scaling is a manual process (EP15's 7 steps) taking minutes; a burst is second-scale. By the time machines are ready, the surge may be over — or may have already crushed one layer.
- Amplifies every single point: a single Origin / single Edge that holds under normal load saturates first at peak. A burst essentially exposes the thinnest link.
So handling a burst isn't about "being big enough" but having things that auto-adapt the instant load changes. This system has three.
3. Auxiliary offload: P2P provides a fixed-cap offload¶
P2P direct (EP13) can cut some Edge egress, but first look at who supplies it. Here, media is served by the publisher directly to a few viewers, not viewer forwarding to viewer:
- The publisher uses spare uplink bandwidth for direct viewers; viewers don't become new distribution nodes.
- Each publisher serves at most
maxP2PSessionsconcurrent P2P connections — the platform default is 3, but a superadmin can raise it (up to 50) via/admin/play-settings;/v1/publish/requestsreturns this number to ppobs, so ppobs no longer hardcodes 3 on the client the way earlier versions did. So whether a single stream grows from 10 viewers to 100,000, P2P offloads that same fixed N Edge views and doesn't grow with viewer count. - For a business with many publishers going live at once, total offload grows with publisher count; but that still isn't a burst-capacity defense for a single hot stream.
Of course it's not unlimited; two boundaries worth stating:
- Each publisher serves at most
maxP2PSessionsP2P direct connections at once (default 3; EP13's atomic lease atomically checks "occupied slots ≥ cap" at allocation time); anyone beyond that cap must go through Edge. The cap is now a platform-level adjustable parameter, not a constant hardcoded into the binary — but it still doesn't grow on its own just because one stream's viewer count spikes, so P2P remains a capped cost optimization, not auto-scaling for a viewer burst. - P2P hit rate depends on real network behavior. Ideal NAT engineering would judge mapping and filtering behavior separately; but PPCDN's current server-side pre-filter still works at the label level — hitting a symmetric or CGNAT label gets a flat rejection before ICE is even attempted. That's an honest current limitation, not the fine-grained judgment call it ought to be (see EP13).
- The fallback is a sequential model where the server branches by codec up front, not "P2P racing Edge in parallel": H264 clients try P2P alone first and only switch to the Edge address (delivered alongside the decision) on failure or timeout; HEVC clients go straight to Edge and never attempt P2P. The exact timeout/grace period should be calibrated from real network samples, not hardcoded.
In one line: P2P can reduce a little Edge egress, but each publisher defaults to at most 3 (superadmin-adjustable, and it never grows automatically with a stream's viewer count); burst capacity planning can't count it as a resource that expands automatically.
4. Second line: on-demand pull + pre-provisioned idle Edge¶
The idea of the second layer: make idle machines cost near zero, so you can pre-provision many on standby.
- On-demand pull (EP15): before viewers arrive, an Edge has no pipeline to Origin; an Edge with no viewers adds almost 0 load to Origin.
- So pre-provisioning idle Edge is worthwhile: you can start a batch of edge machines across regions in advance; when nobody's watching they consume almost no Origin bandwidth, only their own running cost.
- When a burst arrives and viewers land, these Edge each pull on demand and start distributing. Capacity is "dynamically enabled by where the viewers actually land", not all occupying upstream from the start.
The cost remains honest: on-demand pull has cold-start link setup and first-frame wait. Mitigate with parallel attempts on other usable paths and reasonable timeouts; on one Edge, viewers of the same stream share the pull link, triggered by the first viewer and reused by the rest.
5. Publish-link degradation: protects the publisher, doesn't replace viewer capacity¶
The goal of publisher-side adaptive degradation is to protect the main publish link when the publisher's uplink or ingress is congested. It may lower the per-stream bitrate for later distribution, but the trigger and the controlled object are on the publish side; it's not a capacity controller for a viewer concurrency burst.
This system's degradation is layered and ordered (EP07/EP11 covered the mechanism; restated here from the "burst" angle):
- Triggers come from protocol-native stats and product-derived metrics: SRT and RTP/WebRTC counters have different semantics; decisions need the statistical window, recovery result and playout deadline, not renaming one counter as a unified ULR.
- Lower bitrate first, then resolution: step one only lowers the encoder bitrate
(
obs_encoder_update, real-time, no interruption), 100% → 80% → 60%; only when bitrate is at the floor and still short does it enter the resolution stage, dropping the highest simulcast layer (~1s interruption). - Recovery is slow and steady: quality metrics must fall back into the recovery range and hold through an observation window before stepping back up, and each added layer in the resolution stage needs reconfirmation to avoid oscillation. Thresholds and observation times are version-specific product parameters, to be calibrated per protocol and real network samples.
After bitrate drops, Edge's per-stream egress bytes may fall — a side benefit; but if Edge's bottleneck is connection count, CPU, memory or scheduling slots rather than bandwidth, publish degradation won't add viewer capacity. A viewer surge still needs pre-provisioned capacity, load balancing, admission/queueing, rate limiting and automatic or manual scaling.
State the boundary clearly: publish degradation protects the publisher's main link, not viewer-burst capacity. It can't fix insufficient Edge connections or machines, and shouldn't wait for a viewer surge to trigger.
6. When elasticity can't hold: soft-threshold alarms + manual scaling¶
If the peak exceeds what the three layers hold, what's left is monitoring + humans.
There are three signal types:
| Signal | Trigger | Meaning |
|---|---|---|
capacity_high |
single node used/capacity ≥ 80% (warning), ≥ 95% (critical) | "this one is nearly full, prepare to add machines" |
pool_exhausted |
the app's resolved main pool has no available candidate and fallback is disabled | scheduling fails outright |
pool_fallback |
main pool has no candidate, fallback allowed, and the shared pool has one | lands on the shared pool (log only by default, no alarm) |
These signals' design intent is "scaling trigger": at 80% capacity, act. But in a burst they have limits that must be stated:
- No auto-scaling executor. Alarms just notify a human, who then runs the manual 7-step process. For a seconds-scale burst that chain is too slow — scaling chases the surge after the fact.
- Each alarm channel enforces its own global 30-second cooldown: the generic webhook and
Telegram (production is currently wired to a Telegram ops group) each maintain their own cooldown
lock, but for a given channel, once anything has gone out in the last 30 seconds, every subsequent
alarm (regardless of node or type) is silently dropped. In a burst, multiple nodes often cross their
threshold at the same moment, so only the first notification to arrive actually gets sent. This is a
real gap — the moment you most need dense alarming is exactly the moment alarms get swallowed
hardest. Production already hit a real bug in this pipeline (fixed 2026-09-15): four check points
including
CheckTrafficQuota/CheckCapacityused to unconditionally resend a "resolved" notification any time a metric dipped below the resolve threshold, which flooded the Telegram group with an hourly repeat oftraffic_quota ... 0% of quotastarting 2026-09-14 17:32 UTC — it only stopped once the fix made it notify only when an actually-active alarm was cleared. That incident is itself evidence the alarm pipeline really runs in production and really does get stressed by bursts or anomalous states — it's not just a design on paper. pool_exhausted/pool_fallbackaren't persisted; without a configured webhook they're visible only in server logs, not the console.
One unavoidable trade-off in a burst: isolation (fail-closed) and availability (fallback) conflict. If an app's resolved main pool has no candidate and is configured not to fall back, scheduling fails outright; a pool can be app-exclusive or tenant/user-shared. Whether to temporarily borrow the default shared pool is a business-priority question, and the shared pool itself must have available candidates.
7. The fourth signal: cloud vendor egress quota — now auto-alarmed, but still an approximation¶
There's one more metric that earlier versions couldn't see at all, and that's now been partly filled in: the cloud vendor's monthly egress allowance.
- In-plan egress is usually a billing threshold (e.g. AWS Lightsail's monthly traffic allowance). A burst consumes it fast; going over it usually just incurs extra charges, not an automatic cut-off. Whether there's throttling or suspension depends on the specific cloud product, region and account policy.
ppcenter's concurrency-capacity monitoring (capacity/pipelines) still isn't byte-level traffic — that part hasn't changed. But the system now has aTrafficQuotaChecker(internal/manager/traffic_quota_checker.go, committed 2026-09-13, confirmed live in production 2026-09-15): every hour it sums each node's downstream bytes for the current month (TrafficUsageNodeDaily, produced by the same usage-reporting pipeline the Edge nodes already feed for billing — the same underlying data), and compares that againstalarm.trafficQuotaBytes(default 3TB, matching the AWS Lightsail$12/motier used uniformly across the three nodes). At 70% (alarm.trafficQuotaWarning) it automatically raisesAlarmTrafficQuota; falling back below 60% (alarm.trafficQuotaResolve) auto-clears it — running through the exact same alarm pipeline ascapacity_high, with the same webhook/Telegram configurability and the same 30-second cooldown from the previous section. Production's Telegram ops group has been genuinely receivingtraffic_quota origin-node-1 ... monthly traffic N% of quotanotifications since 2026-09-14.- But this still isn't "wired into the cloud vendor's billing API":
TrafficQuotaCheckercompares each node's self-reported byte count against a fixed number written intoppcenter's config, not AWS's actual billing usage. If the cloud plan, region or billing tier changes (a different instance size, a different traffic tier),trafficQuotaBytesdoesn't update itself — someone on ops has to keep that constant in sync by hand. And it's a per-node byte approximation, not an account-level dollar cost. The truly accurate monthly bill still has to be checked by logging into the cloud console. - Net result: the old problem — "concurrency alarms can't see how much of the traffic budget is spent at all" — is now half-solved by a self-reported-bytes approximation. But whether that quota number still matches the real cloud bill is still something a human has to keep watching as the cloud vendor's plans and prices change.
In other words: a burst can push the cost curve past the in-plan allowance before nodes even saturate — there's now at least one layer of automatic early warning based on self-reported bytes (alarm at 70%, clear at 60%), a real improvement over the old "purely log into the console by hand" world. But the quota number itself is still a human-maintained approximation and shouldn't be treated as being as precise as the cloud vendor's actual bill. This is first and foremost a budget problem, not one that should be described as a guaranteed availability outage; if a vendor genuinely enforces a hard cap, that calls for a separate service-risk alarm.
8. Honest status and this episode's stance¶
Summarizing:
- Today's main capability for viewer bursts = pre-provisioned Edge + on-demand pull + scheduling/ admission + manual scaling. There's no auto-elastic scaling.
- P2P defaults to offloading at most 3 per publisher (superadmin-adjustable, but still doesn't grow automatically with viewer count); publish degradation protects the publisher link. Neither replaces Edge capacity for a hot stream.
- So the realistic burst strategy is:
- Pre-provision idle Edge (they eat almost no upstream bandwidth when idle), so a burst has somewhere to land;
- Treat P2P's default cap of 3 as a fixed offload benefit, not capacity that grows with viewers;
- Script the scaling process, compressing the manual 7 steps to under a minute;
- Calibrate capacity against real measurements (per EP15: origin/record's
mmxNodeCapacityare still clearly too optimistic, while edge is already close to the economics-derived value), so soft thresholds don't protect an inflated denominator; - Keep
TrafficQuotaChecker's quota constant in sync with the actual cloud plan — byte-level monitoring is now automated, buttrafficQuotaBytesis a fixed number sitting in config; when the cloud plan changes, someone has to go update it, or this layer of alarming quietly drifts out of sync. - Final stance: a burst amplifies every single point. Single-region deployment, a single Origin, a single Edge may suffice under uniform load but are the first to break in a burst. At low load, the value of redundancy (N+1) often exceeds "a bit more capacity" — the note EP15 ended on.
9. Diagrams¶
9.1 Three layers of elasticity for a burst¶
Load spikes instantly
│
▼
┌───────────────────────────────────────────────┐
│ Auxiliary offload: P2P direct │
│ defaults to 3 per publisher (superadmin-adjustable); doesn't grow with viewers │
│ tried before Edge, sequential fallback on failure, timeout calibrated from real samples │
└───────────────────────────────────────────────┘
│ traffic still lands on Edge
▼
┌───────────────────────────────────────────────┐
│ Layer 2: on-demand pull + pre-provisioned idle Edge │
│ idle Edge load on Origin ≈ 0 → pre-provision many │
│ pulls only when viewers land; same-stream viewers share one pull link │
└───────────────────────────────────────────────┘
│ nodes/links start congesting
▼
┌───────────────────────────────────────────────┐
│ Publish-link protection: adaptive degradation │
│ ULR over limit → lower bitrate (no interruption) → drop the top layer │
│ protects the publisher; doesn't fix insufficient Edge connections/machines │
└───────────────────────────────────────────────┘
│ all three can't hold
▼
┌───────────────────────────────────────────────┐
│ Fallback: capacity_high alarm → manual scaling (7-step process) │
│ ⚠ no auto-scaling executor; each channel's 30s cooldown swallows alarms │
│ ⚠ traffic quota now has TrafficQuotaChecker auto-alarming, │
│ but the quota number still needs manual sync with the cloud bill │
└───────────────────────────────────────────────┘
9.2 Isolation promise vs burst availability¶
The app's main pool has no available candidate
│
▼
allowFallbackToDefault ?
│
┌────┴──────┐
false true
│ │
▼ ▼
scheduling fall back to the default shared pool
fails (preserve availability,
(keep the but the isolation promise temporarily fails
isolation and by default no alarm)
promise,
sacrifice
availability)
10. Wrap-up and next episode (script notes)¶
A burst ultimately relies on available Edge capacity, scheduling admission and fast enough scaling. On-demand pull makes pre-provisioned Edge's upstream idle cost low; P2P defaults to offloading at most 3 per publisher (superadmin-adjustable, but it never grows automatically with viewer count); publish degradation protects the main publish link, not viewer capacity. The cloud vendor's monthly allowance mainly changes overage cost rather than causing an automatic cut-off — the system now has a
TrafficQuotaCheckerthat auto-alarms off each node's self-reported bytes, but that's only an approximation, and the quota number itself still needs to be kept in sync with the cloud bill by hand.That wraps up Module 3 "Advanced architecture". Next episode begins Module 4 "Practice & verification" — the first thing is to deploy the whole system from scratch and run these mechanisms by hand.