Skip to content

EP13: P2P direct connect — faster than CDN for the last mile

Recap EP12 "Benchmarking low-latency live streaming and the current baseline", end of Module 2
Next EP14 "Node-pool management: putting resource boundaries in the scheduler"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Understand the gap between the ideal NAT/ICE engineering model — separately evaluating mapping/filtering behavior plus connectivity checks — and what PPCDN actually implements: a hard pre-filter on the classic NAT-type label. Be able to explain why that gap is a reasonable engineering trade-off, not a mistake.
  2. Understand that TURN, Edge fallback and the direct-connect quota are all product trade-offs — PPCDN chose "self-hosted STUN + sequential Edge fallback" over a TURN relay, and the cost of that choice is that symmetric-NAT/CGNAT users currently have no fallback path at all.

1. Opening hook (script notes)

Module 3 starts here. The current P2P end-to-end latency that really works is about 70ms, but this result can't be simplified into a guarantee from some "NAT type label". What actually establishes the connection is ICE: gather candidates, exchange candidates, check candidate pairs and select a usable path; prediction rules can only optimize whether it's worth trying, not replace the actual checks.


2. Prediction vs. measurement: a NAT-label hard pre-filter gates ICE, which decides whether direct connect actually works

2.1 The gap between the better theoretical model and what PPCDN actually implements

The classic four-way "full cone / restricted / symmetric NAT" taxonomy is considered an oversimplification in modern ICE practice: the more accurate approach is to observe address/port mapping behavior and inbound-packet filtering behavior separately, then factor in hairpinning support, mapping lifetime, IPv6, UDP availability, and whether the network just switched. CGNAT only means address translation happens inside the carrier's network — it doesn't necessarily mean hole-punching is impossible; endpoint-dependent mapping lowers the success rate, but in theory shouldn't be a death sentence from the label alone.

That's the more advanced industry approach, but honesty requires saying: PPCDN's current server-side decision (EvaluateDirectEligibility in ppcenter/internal/p2p/eligibility.go) doesn't go that far yet. It consumes a single classic label reported by ppobs/pplayer after their own probing — one of public/cone/restricted/symmetric/cgnat (parsed by parseProbeNATType in nat_probe_v1.go, whose own comment admits that "browser-side probing can only really distinguish public from restricted — it can't even tell full-cone from port-restricted apart"). The only field in the eligibility logic that gestures at "behavior" is a boolean, MappingStable, but that's itself just a direct mapping of "does the NAT type fall into one of the three good labels, {public, cone, restricted}" — not an independent multi-port/multi-destination behavior probe. More importantly: the moment either side reports symmetric or cgnat, EvaluateDirectEligibility returns a rejection immediately (EligibilityReason of symmetric_nat/cgnat), without ever reaching the later UDP, address-family, capability or cooldown checks — and without ever getting a chance to attempt ICE. That is exactly the "death sentence from a label alone" the previous paragraph warned about.

That doesn't make the current approach wrong: symmetric-NAT/CGNAT combinations really do have a low hole-punching success rate, and rejecting early saves a round of signaling and ICE overhead that was nearly doomed anyway — that's a reasonable engineering trade-off. But it shouldn't be dressed up as "the system already does fine-grained judgment based on mapping/filtering behavior" — the honest framing is: the ideal ICE/NAT engineering practice tiers decisions on observed behavior; PPCDN's current implementation is a hard pre-filter on the classic type label, a simplified version that's still evolving. The finer-grained signals this course often mentions — historical success rate, whether the network just switched — also aren't wired into this decision yet: nat_observation_store.go only keeps each participant's most recent probe and expires it, with no success-rate accounting; P2P hit-rate is currently visible only after the fact, in the GET /admin/p2p-penetration report — an analytics view that never feeds back into the real-time decision.

2.2 Two gates in sequence: the server-side pre-filter first, the ICE connectivity check after

Past the pre-filter, the flow finally reaches real ICE: gather host, server-reflexive (and relay if needed) candidates; exchange credentials and candidates; form candidate pairs; then verify bidirectional reachability and nominate a path through authenticated STUN connectivity checks. But the pre-filter and the ICE check aren't two independent pieces of evidence that corroborate each other — they're a strict two-gate sequence. EvaluateDirectEligibility runs a single, one-shot hard filter against each side's most recent probe snapshot: the probe must be fresh and match on identity and streamPath; the mapping must be stable (not symmetric/cgnat); both sides must have UDP available; there must be a shared public address family; ICE/H.264/Opus capability must be compatible; neither side can be in a failure cooldown; and at least one side must have a server-validated public UDP endpoint the peer can reach. Only when every condition holds does the server hand the client a p2p-connect decision and let it actually start ICE; fail any single one and the client gets edge-only back directly — ICE is never even attempted.

So "eligibility passed" (eligibility.Eligible == true) precisely means "all pre-filter rules are satisfied, so it's worth letting the client go run a real ICE attempt" — not "these two ends are proven to connect". What actually decides whether direct connect succeeds is always the client's own ICE candidate-pair connectivity check in that moment; the pre-filter's job is only to turn away combinations that are clearly hopeless, saving the cost — it never predicts "will this specific attempt succeed."

(The design doc's DIRECT_ELIGIBLE "validation matrix" isn't, in the current implementation, a table that pre-scores every NAT combination — the real code is a sequence of conditions evaluated in order, rejecting on the first hit. Both describe the same design intent, but the actual implementation is plainer; this course no longer uses the name DIRECT_ELIGIBLE, since it never actually appears in the code.)

2.3 Honest boundary: the network world has no mathematical 100%

This must be said plainly: NAT pre-filtering cannot provide a mathematically 100% success guarantee on the public internet. Even after passing every pre-filter rule in §2.2, the actual connection attempt can still fail — a firewall policy, a network switch, a mapping that happens to expire right at that moment. The pre-filter is only an admission condition that "reduces wasted attempts"; it's never a guarantee that "this one will definitely connect."

That's also why "ICE might fail" has to come with a fallback path — not an optional nice-to-have. But the specific fallback mechanism changed on 2026-10-04, and it's worth explaining carefully: it is no longer "P2P and Edge race in parallel, whichever produces a frame first wins"; it's now a sequential model — the client tries P2P first, and on failure or timeout switches straight to an Edge address it was already handed. PPCDN's playback-decision API splits by decoding capability first: clients that support HEVC go straight to Edge and never attempt P2P; H264 clients that pass the §2.2 pre-filter get a p2p-connect decision that already carries a pre-signed Edge WHEP fallback address — no second request needed (see ppcenter/internal/apis/player_v1.go; the server-side comment states outright, "the client follows the decision - no client-side race"). Whether a client failed the pre-filter, failed to establish P2P, or the publisher's quota is full, it switches to that same Edge address that was delivered alongside the decision — this is not two links racing for the first frame, it's one primary path plus one backup address that's always ready. This replaces the earlier "P2P/Edge race in a 500ms window" design; if a source you've seen still describes parallel racing, it's describing the state before this change.


3. TURN, Edge and direct: trade-offs across product paths

TURN is ICE's standard relay-candidate source, not "a fake P2P success". Its value is raising the connection success rate on restricted networks, behind firewalls, or where UDP is unavailable, and providing uniform session connectivity semantics; the cost is relay bandwidth, deployment/ops and possibly added path latency.

If a product already has a mature Edge media path, it can choose to fall back to Edge after a direct failure instead of deploying TURN. That's a product trade-off among cost, protocol consistency, coverage and ops complexity — it doesn't mean TURN is inferior, nor is it valid to pre-reject CGNAT or endpoint-dependent-mapping users. The safer strategy is to let ICE checks try usable host/srflx candidates within a controlled budget and quickly pick Edge on failure; if the business values connection success rate or general WebRTC interop more, provide TURN.

PPCDN today has chosen the former: it self-hosts exactly one STUN server (ppcenter/internal/stun/server.go, listening on :3478 by default, its address delivered to the client alongside the p2p-connect decision), and there is no TURN relay implementation or config option anywhere in the repo. The only fallback on a direct-connect failure is the Edge WHEP fallback described above — not a media relay. That means combinations the pre-filter currently rejects outright, like symmetric NAT and CGNAT, have zero "still connect via relay" path — Edge is the only option. That's an honest statement of the current state, not a verdict that "a worse experience for CGNAT users is fine" — if PPCDN ever needs to cover those users, deploying TURN is a step it can't skip.


4. Quota allocation: atomic leases to prevent over-issue

4.1 Each publisher serves at most 3 direct connections

P2P direct connect consumes the publisher's own uplink (EP07: main publish first, P2P uses the spare), so you can't let every viewer try to connect directly to the same publisher without limit — each publisher serves at most 3 P2P direct connections at any time, and the 4th viewer must go through Edge. This cap protects two things: the publisher's own uplink won't be crushed by "one-drag-many", and the already-established direct connections won't degrade because new ones joined.

That "3" isn't a hardcoded magic number — it's a platform default (defaultMaxP2PSessions = 3 in ppcenter/internal/apis/play_settings.go): a superadmin can adjust it within (0, 50] (maxP2PSessionsCeiling = 50), the change takes effect on p2p.Coordinator immediately with no restart needed, and /v1/publish/requests passes the currently-effective value through to ppobs — the publisher side no longer hardcodes 3 either. That detail is itself a concrete example of the "quota allocation is a product trade-off" principle from earlier.

4.2 Why "atomic lease", not simply "count how many are active"

There's a classic concurrency problem here: if several viewers almost simultaneously ask "can I connect directly P2P to this publisher", simply "query the active count, agree if under 3" races — two requests both see "2 active", both judge "a slot is free", both get approved, and 4 connections are actually established, exceeding the cap.

The fix is to merge "check if a slot is free + occupy one" into one indivisible atomic operation (atomic lease): on arrival, try to atomically decrement a slot; only success lets the request continue down the P2P flow, and failure (the slot was taken by another request) immediately disqualifies it back to Edge. However many requests arrive at once, at most 3 obtain a slot, and over-issue can't happen. This "resource lease" pattern isn't P2P-specific; any "finite resource, high-concurrency applications" scenario faces the same problem and uses the same class of solution.

Inside PPCDN, this "atomic" operation is implemented as a single global mutex: the Coordinator in ppcenter/internal/p2p/coordinator.go wraps the entire "check eligibility + allocate a slot via allocateLocked" block (the AllocateEligible function) in c.mu.Lock() — not a CPU-level lock-free atomic instruction, but equivalent in effect: nothing can cut in between the check and the occupy. The lock's granularity is "per coordinator instance", not "per publisher," but since this code path only does in-memory comparisons and map operations with no network I/O, the critical section is extremely short and doesn't become a bottleneck under high concurrency.


5. How signaling flows: the control plane only forwards and validates, never touches media

Before the P2P connection is established, the Offer, Answer and ICE candidate signaling messages all pass through the server — but the server only forwards and validates identity here, not participating in media itself (echoing EP02's most basic principle, "the control plane doesn't touch media"). The core of validation is confirming that the two parties' identities and the session actually match, preventing anyone from injecting signaling into another stream or another user's session — the last gate against cross-stream, cross-user signaling forgery. After signaling passes, media travels the direct publisher↔ viewer channel, entirely bypassing the server. Code location: ppcenter/internal/ws/p2p_signal.go, with the identity/session validation logic in validateInbound.


6. Honest status: connectivity is now confirmed; systematic acceptance is still the blank

The ~70ms P2P direct-connect figure is an end-to-end result from a real 2026-09-29 link test (see the latency test report) under specific experimental conditions — not a theoretical number, and not a promise for every network.

Connection reliability itself has made real progress recently: the earlier stretch where "P2P almost always fell back to Edge" has been through a round of fixes (completing the STUN config, fixing signaling identity matching, fixing STUN server address-family selection — several deployment-log entries around 2026-09-22 are this round of fixes), and P2P direct connect has now tested successfully across multiple real network environments — it's no longer that earlier "structural, near-total failure" state. But that only answers "it connects on the network combinations that have been tested" — it's not the same as completing a systematic acceptance pass covering every mapping/filtering behavior, CGNAT, multiple carriers, IPv4/IPv6, UDP restrictions, and network switching. Of the seven acceptance scenarios laid out in docs/test/p2p-real-environment-acceptance-plan.zh-CN.md, most are still marked "exploratory" rather than completed hard acceptance — the plan calls for recording per-ICE-candidate-type success rates, connection setup time, fallback rate and end-to-end latency distribution, not just one or two successful connections. Design complete, single-connection verification, and large-scale systematic acceptance are three different stages; right now PPCDN is at the second one.


7. Diagrams

7.1 P2P eligibility decision flow

Play request arrives (carrying NAT probe results)
        │
        ▼
Supports HEVC? ──yes──▶ return edge-only directly (never tries P2P)
        │no
        ▼
Passes the server-side pre-filter (EvaluateDirectEligibility, all §2.2 hard conditions)?
        │
   ┌────┴────┐
   no         yes
   │          ▼
   │    Does the publisher still have a free P2P slot (< maxP2PSessions)?
   │          │
   │     ┌────┴────┐
   │     no         yes
   ▼     ▼          ▼
return edge-only (already carries a signed Edge address)   return p2p-connect
(reason = whichever pre-filter check failed / at_capacity) (also carries the Edge fallback address, no second request needed)
                                             │
                                             ▼
                              client starts ICE: gather/exchange candidates + connectivity check
                                             │
                                        ┌────┴────┐
                                   fail/timeout      success
                                        │             │
                                        ▼             ▼
                              switch to the Edge address already in hand   direct P2P media
                             (sequential fallback, not a parallel race)

7.2 Atomic lease of P2P quota

Many play requests arrive almost at once (all want P2P to the same publisher)
        │
        ▼
Try the atomic operation one by one: merge "check slot < 3" + "occupy a slot" into one step
        │
   ┌────┴────┐
   decrement ok   decrement failed (slots full)
   │              │
   ▼              ▼
continue the    immediately disqualify, return Edge
P2P flow
   │
   ▼
at most 3 leases exist at once; no over-issue

7.3 Signaling forwarding: the control plane validates identity, not media

Publisher                    Control plane (ppcenter)              Player
  │  Offer ──────────────────▶  validate: which session does
  │                              this signaling belong to? do the
  │                              two identities match this stream?
  │                                    │
  │                              ┌─────┴─────┐
  │                              pass        reject
  │                              │            │
  │  ◀───────────────────────  forward       reject, log anomaly
  │                              │
  │  Answer / ICE candidate  ──▶ same identity check, then forward ──▶

   After signaling passes, media travels the direct publisher ↔ player channel,
   no longer passing the control plane

8. Wrap-up and next episode (script notes)

This episode took P2P direct connect down to the ICE layer: the ideal NAT engineering practice would describe mapping/filtering behavior separately rather than relying on the classic four-way label, but PPCDN's current server-side pre-filter still lives at the label layer — hit a symmetric/cgnat label and it rejects outright, without ICE ever being attempted. That's an honest gap, not fine-grained judgment that's already been achieved. Once a request clears the pre-filter, whether it actually connects is still decided by that attempt's own ICE connectivity check; the pre-filter's only job is to skip attempts that are already hopeless. TURN is the standard relay capability; PPCDN currently self-hosts only STUN and hasn't deployed TURN, choosing Edge fallback over relay — and the fallback mechanism itself just changed from "P2P/Edge racing in parallel" to "try P2P first, fall back in sequence to the Edge address already in hand on failure." The current ~70ms is an observed baseline from one real link; connection reliability has improved recently, but systematic acceptance covering every NAT combination and recording ICE candidate-type success-rate distributions is still a blank.

Next episode, node-pool management: why "role + region" isn't enough, how to build resource boundaries by user, tenant or app, and why there's no fallback by default when the main pool has no candidate.