Skip to content

EP15: The relationship between horizontal scaling and low latency

Recap EP14 "Node-pool management: putting resource boundaries in the scheduler"
Next EP16 "Handling sudden bursts of high concurrency"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Debunk the most common misconception: when nodes are idle and the path and media parameters don't change, horizontal scaling generally doesn't reduce single-stream latency; it mainly buys capacity, and it prevents overload from degrading latency.
  2. Explain how "capacity" is measured and scheduled here: capacity/pipelines, scheduling by most remaining capacity first, and why the capacity threshold is soft admission, not a hardware guarantee.
  3. Understand scaling's upstream impact: Origin egress should sum per actually-pulling Edge; pre-push fanout and on-demand pull costs can't be mixed.
  4. Know the current boundary: capacity alarms are available, but the auto-scaling executor isn't implemented — scaling is currently a manual process.

1. Opening hook (script notes)

A common question: will buying more machines make a single stream faster? State the conditions fully first: when nodes aren't overloaded and the network path, protocol buffers and media parameters are unchanged, adding machines usually doesn't shorten a single stream's pipeline; when the original node is already queuing, dropping or CPU-saturated, offloading can of course restore latency. Horizontal scaling mainly changes capacity and high-load stability, not the latency of a healthy, idle path.


2. Busting a misconception: the latency floor isn't "not enough compute"

End-to-end latency is composed of capture and encoding, network propagation, protocol buffering, queueing, decode and render. Horizontal scaling directly addresses only resource contention and queueing; it doesn't automatically shorten propagation or shrink existing jitter buffers or GOPs. Bigger buffers are usually steadier but cost higher latency; the exact numbers must be measured per protocol, config and load, not assumed as one cross-scenario fixed floor.

So the conclusion is direct:

  • Under healthy idle conditions, adding Edge or Origin won't shrink the established buffers automatically. One pipeline's serial overhead doesn't shorten just because you spread more nodes in parallel.
  • When already overloaded or badly routed, scaling and rescheduling can reduce queueing, loss and retransmission, or move users closer, so observed latency may drop. What improves is congestion or the path, not the node count itself.
  • PPCDN's proven P2P run is about 70ms; that's a case of changing the media path, not an effect of horizontal scaling, and can't be extrapolated into a network-wide SLA.

In one line: horizontal scaling solves "can we hold more people", not "can we be faster". The two are often conflated.


3. What horizontal scaling really buys: capacity, and "holding latency under load"

If adding machines can't lower latency, what does it buy? Capacity — and an easily-missed indirect benefit: keeping latency from degrading as load rises.

  • One Edge can serve a finite number of viewers/concurrent connections. Two roughly double the cap, and one overloading doesn't drag down the rest — the capacity version of EP02's "the Edge tier isn't a single point".
  • The scheduler picks the node with the most remaining capacity, steering new traffic to the idler one. As long as nodes aren't pushed past the overload line, latency won't degrade from queueing, loss and retransmission.
  • In other words: horizontal scaling doesn't push the latency curve down; it flattens the latency curve at higher load. This extends EP09's "latency stability matters more than average latency" into the capacity dimension — what you really protect is "don't get worse when crowded".

4. How capacity is measured: capacity / pipelines

The node-side capacity model is simple, two numbers:

  • capacity: the configured concurrent-slot ceiling (e.g. an Edge with mmxNodeCapacity: 12).
  • pipelines: the currently occupied media pipelines. It's an implementation-defined scheduling count, not equal to viewer count; one shared origin-pull pipeline can serve several viewers, and different protocols, codec tracks or forwarding targets may each occupy a pipeline.
  • Remaining capacity free = capacity − pipelines: the scheduler picks in the pool by this value, descending (EP14 covered "filter by pool first"; here's the in-pool ranking dimension).

Two realities that must be stated:

First, capacity is a scheduling admission threshold, not a hardware guarantee. The scheduler can exclude a node from new-request candidates when pipelines >= capacity, but that only means "no new pipelines", not that the node is healthy below the threshold. Over-configuration, a momentary spike in existing sessions, or CPU/NIC bottlenecks arriving first can still cause loss or anomalies; under-configuration wastes resources. So it must be calibrated by load testing and used alongside CPU, memory, bandwidth and loss monitoring.

Second, alarms are for humans, not machines. So the system uses three soft thresholds to give advance warning for scaling:

Threshold Default Meaning
alarm.capacityWarning 0.80 at 80% → warning, "start preparing to add nodes"
alarm.capacityCritical 0.95 at 95% → critical, "this is dangerous"
alarm.capacityResolve 0.70 falling back below 70% → auto-clear the alarm

Pipeline count is an integer, so thresholds should be integer comparisons. For example with ceiling rounding, capacity=12 gives a warning at ceil(12×0.80)=10 pipelines and critical at ceil(12×0.95)=12 — not at 9.6 or 11.4 pipelines. The rounding rule must be consistent between implementation and monitoring queries. The scheduler steers new requests to other nodes with headroom, but it won't add machines for you.


5. The most counter-intuitive lesson: more Edge makes Origin busier

By now "more machines = more capacity" seems obvious. But this architecture has an easy trap: scaling Edge pushes load onto Origin.

The reason depends on the distribution mode and can't be mechanically multiplied by the total deployed Edge count:

Origin egress = Σ (the bitrate each stream actually sends to each downstream target)

Under push fanout (pre-push), every configured downstream keeps receiving the stream, so a new Edge immediately adds Origin egress. Under on-demand pull, only an Edge actually serving that stream's viewers pulls from Origin; a new but idle Edge produces no media egress for that stream. And several viewers on one Edge sharing an origin pull also can't be counted per viewer.

So when a new Edge actually starts pulling, Origin adds:

  • a new forwarding link (egress bandwidth),
  • a new pre-decode media-forwarding session (session count, memory),
  • these costs stack per actually-pulling downstream target (Edge/Recorder etc.), not necessarily with the deployed Edge count.

Engineering lesson: don't treat capacity as "how many streams this machine is guaranteed to run", and don't replace traffic measurement with "streams × fixed targets". Sum per stream's actual bitrate and actual downstream connections, find the CPU/memory/NIC/loss knee points in load tests, then leave headroom for the admission threshold. This isn't hypothetical — a 2026-09-20 production capacity assessment (a modeled estimate, not a live measurement) ran exactly this formula against Origin's then-configured mmxNodeCapacity=100: 100 streams × roughly 5Mbps each × 2 forwarding targets (record + 1 edge) ≈ 1-1.5Gbps of aggregate egress — two orders of magnitude above the single verified "comfort zone" anchor in the deployment plan (15Mbps per stream, three-tier simulcast, "a rounding error on 2 vCPUs"). The conclusion was to conservatively lower mmxNodeCapacity first, rather than add more Edge and see whether Origin can take it.

The takeaway: horizontal scaling isn't "add where it's short"; follow the data flow and see — will each machine you add shift pressure to its upstream. Here, Edge's upstream is Origin.


6. The mechanism that makes edge scaling cheap: on-demand origin pull

Since Origin is the bottleneck, does every added Edge necessarily add pressure? Not necessarily — it depends on whether Edge pre-pulls.

This system uses on-demand pull (sourceOnDemand): before viewers arrive, an Edge has no pipeline to Origin. The flow:

  1. A viewer starts a WHEP play request;
  2. This Edge pulls the stream from Origin;
  3. Several viewers of the same stream on one Edge share one origin-pull link.

Two direct benefits:

  • An Edge with no viewers adds almost 0 load to Origin. You can pre-provision many edge machines "on standby" that consume no Origin bandwidth until viewers actually land.
  • Edge scaling doesn't touch the publish side. Adding Edge just adds a potential origin-pull point; the publisher is oblivious.

Honest costs and boundaries: on-demand pull means cold start adds link setup and first-frame wait. Today's fallback isn't "P2P and Edge racing in parallel" — it's a sequential model where the server branches by codec up front. H264 clients try P2P alone first, and only fall back to the Edge address (delivered alongside the decision) on failure or timeout; HEVC clients go straight to Edge from the start and never attempt P2P. The exact timeout/grace period should come from real network-distribution calibration, not a fixed constant hardcoded across environments (see EP13).


7. Honest status

  • Capacity alarms are persisted and can send a webhook, but the auto-scaling executor isn't implemented. Scaling is currently a manual 7-step process: create machine → deploy mmx → configure node id and capacity → open firewall, configure TLS → append the new node's private IP to Origin's forwardMmxTargets and reload → update region mapping if any → confirm the node is online in ppcenter and verify with a WHEP pull.
  • Once node count grows, the manual process becomes the bottleneck. That's why the business plan lists "scripted scaling / node onboarding" as the next focus — marginal ops cost eats gross margin and is the most realistic challenge for a self-hosted CDN.
  • Starting out with one Edge isn't a latency problem but an availability problem. One Edge failing means all viewers are cut. At low load, "add one more for N+1 redundancy" actually outranks "add a bit more capacity".
  • Single-machine capacity is currently optimistic and pending load-test calibration. Following the 2026-09-20 capacity assessment (a modeled estimate, not a live measurement), production's three-node mmxNodeCapacity was lowered from origin=100/record=50 to origin=48, record=48, edge=12 (edge not yet re-verified against real hardware). edge=12 lines up with the quota-economics math (derived from a 3TB/month traffic budget) and is the closest of the three to being realistic. But origin=48 still works out to roughly 720Mbps of aggregate throughput by the egress formula above — 48 times the verified safety zone. record=48 doesn't actually fix anything either: on a 60GB disk under the 7-day retention policy, 48 concurrent full recordings can only sustain about 24-30 minutes, barely different from the old value of 50. This is still an explicit documented TODO: find a controlled window, push real streams to the knee point, and backfill a capacity value close to what's actually measured — not just pick a number that merely looks appropriately large.

8. Diagrams

8.1 Single-stream latency vs capacity ceiling: two different things

     Latency (per stream)                    Capacity (node count)
─────────────────────────        ─────────────────────────────
 capture/network/buffer/decode     ┌────┐ ┌────┐ ┌────┐
 with healthy idle and unchanged   │Edge│ │Edge│ │Edge│   ← add machines
 path, adding nodes usually        └────┘ └────┘ └────┘
 doesn't shorten these serial        │      │      │
 costs                                └──────┼──────┘
     │                                       ▼
     └─ if the original node is              raise concurrency and
        overloaded, offloading                redundancy
        can reduce queueing

8.2 Origin egress grows with the number of Edge

              ppobs ──(one stream)──▶ Origin
                                   │
                    ┌──────────────┼──────────────┬──▶ record
                    ▼              ▼              ▼
                 Edge-1         Edge-2         Edge-3
                 (each with its own origin-pull link)

 Origin egress = sum per actually-pulling Edge/Recorder target
             → push fanout counts configured targets; on-demand pull counts actively-pulling Edge

8.3 On-demand pull: an Edge nobody watches eats no Origin bandwidth

Before viewers:           After viewers:
  Edge ──✗── Origin         Edge ──▶ Origin ──▶ stream
  (no pipeline)              ↑
                          several viewers share one Edge→Origin pull link

  idle Edge adds ≈ 0 load to Origin;
  cost: cold start adds link setup/first-frame wait → P2P tried first, sequential fallback to Edge on
  failure/timeout (not a parallel race)

9. Wrap-up and next episode (script notes)

This episode separated "adding machines" from "lowering latency": under healthy idle with an unchanged path, horizontal scaling mainly adds capacity; under overload it can restore latency by reducing queueing and loss. capacity is an admission threshold to be calibrated by load testing, and pipeline isn't viewer count. Computing Origin egress must distinguish pre-push fanout from on-demand pull and sum per actual downstream connection.

Next episode, we push it one step further: when traffic isn't a slow climb but a sudden burst — sudden high concurrency, with only soft-threshold alarms and a manual scaling process — what happens and what to do.