Skip to content

EP14: Node-pool management — putting resource boundaries in the scheduler

Recap EP13 "P2P direct connect: faster than CDN for the last mile"
Next EP15 "The relationship between horizontal scaling and low latency"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Explain what "node pool" actually solves — it adds a resource-isolation boundary beyond role and region; an app can own a pool or share one, and binding must not be blanket-explained as physical exclusivity.
  2. Explain how one scheduling request flows: first get available candidates by pool, then rank by remaining capacity; fallback is only discussed when the main pool has no available candidate.
  3. Separate the general model from the PPCDN case: a general system can build pools by tenant, user or app; PPCDN currently buckets apps by their owner's pool by default and also allows explicit app binding.

1. Opening hook (script notes)

A node pool isn't a simple machine label but the candidate-set boundary of the scheduler. First ask whom the boundary belongs to: a tenant pool is usually shared by several apps of the same tenant, a user pool may carry all of that user's apps, and only a pool dedicated to one app is "app-exclusive". PPCDN is one implementation: an app can bind a pool explicitly, and when it doesn't, it usually resolves to its owner's pool. Below we cover both the general scheduling principles and this case's product trade-offs.


2. The problem: why "role + region" isn't enough

The status quo is plain: Origin/Edge nodes have only two filtering dimensions — role (Origin or Edge) and region (region/city). All online nodes of the same role are treated as one big resource pool for ranking and selection, with no idea who the machines belong to or serve.

That model breaks in these scenarios:

  • A big customer / high-priority app needs a dedicated set of nodes, or several apps of one tenant need to share reserved resources, to avoid capacity being squeezed by others.
  • Nodes with different hardware specs / encoding capability must be scheduled separately — e.g. nodes that support higher-resolution encoding only serve a specific plan.
  • A new version or a freshly provisioned node must first go into a "canary pool" for validation, not immediately into production scheduling.
  • Ops want to split node management authority and visibility by team or contract boundary.

The common thread: on top of "role + region", you need one more dimension — node pool — as the resource-isolation boundary. It isn't a "label" ranking machines, but a hard constraint on who can be seen.


3. Mechanism design

3.1 The pool is the first scheduling filter

An Origin/Edge scheduling request is handled in this order:

  1. Resolve the request's pool: from an explicit app binding, or from a tenant/user default pool or the shared pool.
  2. Compute available candidates in that pool: role-matching, online, healthy, not draining and satisfying admission conditions. "There are nodes in the pool" alone doesn't mean there are candidates.
  3. Rank available candidates by most remaining capacity (most free; workerId ascending as the tiebreak) and take the best.

The key: pool filtering happens first, and no matter how the ranking inside the pool changes, it can't alter the fact that "nodes in other pools aren't even candidates". That's where isolation lands.

3.2 Data model: nodes declare ownership, apps bind to a pool

Two entities, each with its job:

NodePool(poolId, name)                                     // the pool's identity
AppPoolBinding(appId, poolId, allowFallbackToDefault)      // which pool an app binds to, and whether it may fall back when full
  • A node carries a PoolId field, written at registration — but the value written is one the control plane resolves from the node's credential record, not something the node reports about itself (details in §3.3). If there's no matching credential, or the credential's PoolId is empty, it's normalized to the built-in default shared pool — existing nodes upgrade with zero changes: same behavior before and after.
  • App binding and user pool are different concepts: an explicit app binding decides the app's pool; only absent an explicit record does PPCDN resolve to the owner's pool. The former gives app-level isolation, the latter a user-level boundary shared by several apps.
  • A node's pool, like region, is a "statically fixed at deployment" attribute, not dynamic per-request data.

Implementations: ppcenter/internal/models/node_pool.go, ppcenter/internal/store/node_pool_store.go, ppcenter/internal/store/app_pool_binding_store.go.

3.3 Who decides pool ownership: from "node self-reports" to "control plane assigns authoritatively"

There's an evolution here worth walking through — this wasn't a one-shot, settled design.

The original design really was "the node self-reports": ppmmx had a static deployment item, mmxNodePoolId, and reported it to ppcenter at registration through the poolId field (proto field number 11) of the then-current WorkerIndication message — decided and reported by the node itself, same as region and capacity.

That's no longer how it works. After ppcenter introduced "create node in console + nodeSecret/licenseCode registration auth" on 2026-09-09, the registration protocol was renamed to NodeRegister (see ppcenter/pb3/mmx.proto) — whose fields are only msgType/nodeSecret/nodeType/diskSerial/version/ region/webrtcBaseUrl/publishUrl/abrNegotiationAddr/licenseCode, ten in total, with no poolId field at all. Pool ownership became something the control plane assigns authoritatively, with the node having no say in it whatsoever:

  • When a node is created (a superadmin creates it in the console, or a self-hosted node completes its first registration with a licenseCode), a credential record is generated in the mmx_nodes table, whose PoolID defaults to pool-<userId> for the node's owning user (NodePoolIDForUser(userID) in store/mmx_node_store.go).
  • Every time a node registers with its nodeSecret/licenseCode, ppcenter first looks up that credential record, then assembles, server-side, the internal WorkerIndication struct used downstream, filling in cred.PoolID (the SetNodeRegisterHandler callback in ppcenter/main.go: if cred.PoolID != "" { ind.PoolId = &cred.PoolID }), and passes it to NodeManager.RegisterIndication.
  • "The node makes no pool-related decisions locally" is still true, but the reason has shifted from "the control plane's scheduling logic is simply a better fit for this" to a harder constraint — the node no longer even has the ability to self-report its pool.

This wasn't a casual change; it's a security tightening. Once self-hosted nodes are allowed to register with a licenseCode (a lower trust tier than managed nodes), you can no longer let a node report "which pool I belong to" on its own word — otherwise a self-hosted node could in principle claim to belong to someone else's pool at registration time, sidestepping the "one user, one pool" isolation boundary. Pulling pool ownership back into an authoritative control-plane record — one the node can't reach or alter — is what keeps the isolation promise from being bypassed on the node side.

Implementations: ppcenter/main.go (SetNodeRegisterHandler, assembling cred.PoolID into ind.PoolId), ppcenter/internal/store/mmx_node_store.go (NodePoolIDForUser, writing PoolID at node creation), ppcenter/internal/manager/node_manager.go (applyIndication writing it onto the online node), ppcenter/pb3/mmx.proto (the NodeRegister message). The ppmmx side historically did have an mmxNodePoolId config item, but the WorkerIndication.poolId field it fed no longer exists in the current registration protocol — even if that config item is still sitting in the ppmmx codebase, ppcenter no longer reads it.

3.4 Ranking within a pool: only remaining capacity (region ranking retired)

How to rank inside a pool also went through a revision. The early version ranked by "city priority + health + remaining capacity"; after 2026-09-20, Origin/Edge in-pool ranking collapsed to only most remaining capacity, with region/city no longer involved.

This wasn't a casual deletion but a re-division of responsibilities: "grouping nodes" is now done by the pool. If the pool already expresses the machine set you want, stacking another regional priority only muddies the semantics of "best by capacity within pool" — should it favor region or idleness? After removing region ranking, the in-pool rule is clean enough to be one sentence: whoever is idlest wins. The RegionRoute config and admin UI remain, but no longer affect Origin/Edge selection.

3.5 Fallback semantics: fail-closed, never silently break isolation

This is the most important part to explain and the easiest to get wrong.

Suppose an app binds a pool but the pool has no available candidate. "No available candidate" includes no nodes, nodes offline or unhealthy, draining, or no node meeting capacity admission — not just "zero nodes" or a colloquial "pool full". Two approaches differ completely in semantics:

  • Silently fall back to the default shared pool: scheduling "never fails", but the isolation promise is quietly broken — the business thinks it's on dedicated resources but is actually on the public pool, and cost accounting and isolation expectations are all off.
  • Explicit config + explicit alarm: fallback is a switch that must be actively turned on (allowFallbackToDefault), and it is off by default.

This system chose the latter, with a fail-closed default:

  • An app with an explicit binding has allowFallbackToDefault false by default.
  • Main pool has no candidate and fallback not allowed → scheduling fails outright, and a pool_exhausted alarm is emitted.
  • Main pool has no candidate and fallback allowed → recompute candidates in the default pool; fallback succeeds only if there are candidates, otherwise scheduling still fails. A successful fallback produces a distinguishable pool_fallback record, not a silent pass.
  • An app with no binding record at all → at the level of the "general mechanism" filter itself, the default semantics are equivalent to binding default, same behavior as before the upgrade. But PPCDN's product layer overrides that with a more specific rule — see §5: absent an explicit binding, an app actually resolves to its owning user's pool, not literally default. The 2026-09-13 production incident happened precisely because these two layers of semantics weren't aligned.

The philosophy is simple: isolation is a promise, and silent fallback means the promise failed. So it must be explicitly enabled by the business, and either outcome must leave a trace — either failure alarms or fallback is recorded; never "looks all normal".

3.6 unassigned and the canary pool: two "invisible" uses

Two very practical edge uses hide in the pool mechanism, resting on the same principle: as long as no app resolves to a pool, its nodes are never selected by any scheduling:

  • unassigned (removed from scheduling): a dedicated sentinel pool id. No app resolves to it (an app resolves either to its own pool-<userId> or to default), so a node set to unassigned is cleanly "pulled out of scheduling" while still running, still connected, and movable back at any time. This is what the superadmin UI's "remove node" button does.
  • Canary pool: register a new node into a pool no app is bound to, and it naturally "hides" from production scheduling — not a candidate, but heartbeat, capacity reporting and health checks proceed as usual. Once validated, move it into the real production pool. No separate canary deployment flow needed.

3.7 Four call sites, a small blast radius

Pool filtering was added to four scheduling entry points: SelectOriginForPublish (WHIP publish → Origin), SelectOriginForStream (per-stream → Origin), SelectOriginForSRTPublish (SRT publish → Origin), and FindEdgesForStream/SelectEdgeForStream (→ Edge). These call sites already held appId before the change, so it was just passing one more argument and wrapping the candidate set with pool filtering — no new global state or inter-service dependency.


4. Who manages it, and how

Pool config and node ownership are superadmin capabilities, following the same pattern as other admin config:

  • Management APIs: what actually shipped is much leaner than the original design — just a read-only GET /admin/node-pools (list existing pools) and PUT /admin/mmx-nodes/{id}/pool (move a node into/out of a pool). The original design's two "create pool / bind app" surfaces, PUT/DELETE /admin/node-pools and the entire GET/PUT/DELETE /admin/app-pool-bindings surface, were removed outright in v1.0.19 (2026-09-21) — see §5 for why: it's the direct consequence of the product collapsing "create pools freely" down to "one pool per user." AdminHandler doesn't even keep a bindingStore field anymore. What remains is superadmin-only, through the same auth middleware and audit log.
  • Moving a node takes effect immediately: when changing a node's pool, the control plane writes the new pool to the persisted node record (mmx_nodes.pool_id, surviving the next node re-registration) and mirrors the change onto the current in-memory node object, so the next scheduling request uses the new ownership immediately without waiting for a node reconnect.
  • Admin UI: a node-pools page aggregates nodes by pool, supports "add node to a pool" and removal (i.e. set to unassigned). default is always shown; unassigned is pinned last.
  • Alarms: pool_exhausted carries role + poolId + appId as de-duplication dimensions to avoid one binding spamming; pool_fallback is "a normal path in the shared-pool case", so it only logs, no alarm — otherwise every app without its own nodes would fire a notification on each scheduling. The code comment (ppcenter/internal/manager/alarm_notifier.go) is blunt about it: "OnPoolFallback is now the normal scheduling path...Deliberately log-only - it must not page anyone."

One more detail worth mentioning: display paths use a "quiet" filter. Paths polled every time, like the Active Streams list and Edge preview, would repeatedly trip alarms if they reused the alarming filterByPool; so they go through poolFilterQuiet — same filter logic, no pool_exhausted/ pool_fallback alarms. Read paths and scheduling paths share semantics but not side effects.

Implementations: ppcenter/internal/apis/admin_mmx_node.go, internal/apis/admin.go, internal/apis/route.go, internal/manager/alarm_notifier.go, web-admin-react/src/pages/NodePoolsPage.tsx.


5. A real product simplification: from "arbitrary pools" to "one pool per user"

The general scheduling model can support arbitrary pools per role, candidate-pool priorities and active/draining/disabled states, or one-pool-per-node vs multi-pool labels. These are optional product capabilities, not inherent requirements of the node-pool mechanism.

In the actual product, this flexibility was largely collapsed:

  • Pools are fixed: one pool per user, id like pool-<userId> (NodePoolIDForUser). No entry to let users freely create, name and order pools.
  • App binding can be derived automatically: if an app has no explicit binding, it resolves to its owner's pool-<userId>; when that user hasn't deployed nodes, allowFallbackToDefault is true to use the shared pool as a stopgap rather than blocking the request.
  • The superadmin action is simplified from "create pool + bind app" to "move node": the UI no longer creates/updates pools or binds apps to pools, only "add node to pool / remove from pool".

This "auto-resolve" rule wasn't designed ahead of need — it was forced out by a real production incident. On the evening of 2026-09-13, /v1/publish/requests started returning 503 no_healthy_origin for every app. The superadmin console showed all three nodes' Status/Capacity/heartbeat as perfectly healthy, yet scheduling couldn't pick a single one. The root cause: all three nodes were actually registered under pool-1 (i.e., pool-<userId>), while the filtering logic at the time matched apps with no explicit pool binding against nodes in pool=default — the two sides never lined up, so every app without an explicit binding, including the official demo account, matched zero nodes, with no connection whatsoever to node health, mmx version, or heartbeat — it was purely that the app→pool step had simply never been configured. Once "one pool per user" was confirmed as the settled design, the correct fix wasn't to move the nodes back to default (that would violate the design) but to make apps with no explicit binding fall back to "their owning user's pool" — exactly the behavior described in the second bullet above. The fix was verified and redeployed the same day; see the 2026-09-13 (evening) deployment log entry for details.

Why collapse it? Because "named pool + binding-priority list" is too heavy for current product users. PPCDN reduces the common need to "this user's apps prefer this user's nodes", with explicit app binding for the rare app-level need. A user pool doesn't mean each app inside it owns machines exclusively; exclusivity depends on how many subjects actually bind the pool and on fallback policy.

The honest cost: the more flexible shapes in the design (arbitrary named pools, pool-priority lists, node M:N, draining/disabled states) are not all exposed to the product; they stay in the design for when they're truly needed. "Designed capability" and "exposed capability" are two different things, and this episode is an example.


6. Honest boundaries

  • It doesn't depend on active-active. Node pools and the ppcenter active-active redesign come from the same design doc (the latter's HA part), but the two tracks are independent: the HA plan was reviewed and not funded, while node pools advanced independently and are live. The "per-app node pool isolation" mentioned in EP02/EP03 is this shipped capability.
  • It currently covers only the Origin and Edge roles. Pool filtering is added to the SelectOriginFor*/FindEdgesForStream/SelectEdgeForStream function family; Recorder scheduling doesn't go through this path at all (Recorder was once assumed to fold into Origin and not need its own pool, but the product later split it into an independent role without ever catching it up on pool isolation). The Forward (cross-region bridging) and Standalone (self-hosted single-box) roles added on 2026-10-06 are, by design, not part of Origin/Edge scheduling either, and likewise fall outside this isolation mechanism's reach — their node records do carry a PoolId field (reported/normalized at registration just like the other roles), but that value is never read by any candidate-filtering logic; it's simply persisted statically alongside the node, with no actual effect today.
  • It's "scheduling-level isolation", not inherent app exclusivity or physical isolation. The pool constrains whom the control plane assigns a request to; within a tenant/user pool, several apps may still share nodes, and network and process boundaries aren't isolated automatically.
  • Flexibility leaves blanks. Whether a node may belong to several pools at once, whether per-pool capacity alarm thresholds are configured separately, and how the fallback default is set in the product are explicitly marked "to review" in the design; the current implementation takes the simple, conservative values.
  • The canary pool lacks an explicit test case. Logically "a pool no app binds doesn't participate in scheduling" holds, but it's currently covered mainly by isolation unit tests; the canary pool itself has no dedicated test case.

7. Diagrams

7.1 Scheduling: filter by pool first, then rank by remaining capacity within it

Scheduling request (role=origin/edge, appId)
        │
        ▼
Resolve the app's candidate pool
(explicit binding first; else → the user's pool-<userId>; else → default)
        │
        ▼
Filter by pool: keep only online nodes belonging to the candidate pool
(nodes outside the pool are already not candidates here)
        │
        ▼
Rank within the pool: most remaining capacity first, workerId ascending tiebreak
(region / city no longer participate since 2026-09-20)
        │
        ▼
Return the best node; main pool has no candidate → go to §7.2 fallback decision

7.2 Fallback semantics: fail-closed, refuse silent crossing of the isolation boundary

The app's resolved main pool has no available candidate
        │
        ▼
allowFallbackToDefault = ?
        │
   ┌────┴─────┐
  false       true
   │           │
   ▼           ▼
scheduling     recompute available candidates in the default shared pool
fails          │
+ pool_exhausted alarm   ├─ has candidates: fallback succeeds + pool_fallback record
                         └─ no candidates: scheduling fails

7.3 Node-pool ownership and management actions

   ppmmx Origin / Edge node
   (registers with only nodeSecret/licenseCode + nodeType + region, etc. —
    the NodeRegister message has no poolId field at all; the node's own word doesn't count)
              │ registers with nodeSecret / licenseCode
              ▼
        ppcenter control plane
   ┌───────────────────────────────────┐
   │ looks up the mmx_nodes credential   │
   │ record for cred.PoolID              │
   │ (defaults to pool-<userId> when the │
   │  node is created)                   │
   │ assembles it server-side into the   │
   │ internal WorkerIndication           │
   │ node's pool: pool-<userId> /         │
   │              default / unassigned    │
   │ app binding: AppPoolBinding          │
   │ scheduling: pool filter → best by capacity within pool │
   └───────────────────────────────────┘
              ▲
              │ superadmin moves node: PUT /admin/mmx-nodes/{id}/pool
              │  • writes mmx_nodes.pool_id (survives re-registration)
              │  • mirrors onto the online node (next scheduling takes effect immediately)
              │  • pool-<userId> = belongs to a user
              │  • unassigned    = pulled out of scheduling (node still running)

8. Wrap-up and next episode (script notes)

This episode took node pools down to the scheduling code: nodes declare ownership, requests resolve to a pool by app/user/tenant rule, and scheduling only ranks within the available candidates; when the main pool has none, whether to fall back must be explicitly defined. PPCDN's "user pool + optional explicit app binding" is just a case, and a user pool sharing resources shouldn't be absolutized into per-app exclusivity.

Next episode, we pull the lens back: what's the relationship between horizontal scaling and low latency — adding machines looks simple, but it has costs and trade-offs too.