Skip to content

EP17: Deploy from scratch — run your own CDN in 15 minutes

Recap EP16 "Handling sudden bursts of high concurrency", end of Module 3
Next EP18 "Recording & playback: adding VOD to live streaming"

0. Goals of this episode

By the end, the viewer should be able to:

  1. Review and run the repo's docker compose as the current deployment template, and state each component's place in the topology, its ports, and the prerequisite fixes.
  2. See which config keys in this local orchestration map exactly to the mechanisms covered earlier (node pool / capacity / on-demand pull / ABR / node auth).
  3. Separate "the environment came up" from "end-to-end verified" — the smoke test checks control-plane health, not a real publish-to-play path.

1. Opening hook (script notes)

The first sixteen episodes covered general media-system mechanics; this one uses the PPCDN repo's compose as a deployment case. State the status first: it's the current deployment template, not a turnkey release; running it as-is still lacks database config, the mmx image and a TLS-mount consistency fix. The commands below always use the same set of compose files, but it's only a runnable flow after those prerequisites. "Containers are up" and "media runs end to end" remain two different things.


2. First, see which components must come up

This system isn't "one process"; the local orchestration brings up four things at once, exactly the topology diagram recurring in earlier episodes:

Component Container Ports Role
nginx ppcdn-nginx 80 / 443 Front door, unified entry: reverse proxy for /origin/, /edge/, /v1/, /ws/
ppcenter ppcdn-ppcenter 8090 Control plane: node registration, scheduling, auth, alarms, admin UI
mmx-origin ppcdn-mmx-origin 1935 (RTMP in), 8889 (WebRTC), 8189/udp Media plane: the sole ingest entry
mmx-edge ppcdn-mmx-edge 8890 (WebRTC), 8190/udp Media plane: pulls from Origin, distributes to viewers

Plus two dependencies: MySQL (ppcenter's persistence, required) and Redis (off by default; the control plane runs read-only without Redis).

Three prerequisites to know up front, so "15 minutes" doesn't mislead:

  1. The mmx image doesn't come from this repo. ppmmx is a separate repo (this monorepo only has its config files); compose uses the ppcdn/mmx:latest image, which you must build or obtain from the ppmmx repo first.
  2. ppcenter hard-depends on MySQL. deploy/config/ppcenter-local.json currently has empty database fields; you must at least set host=mysql, user, password and dbName=philcdn, and keep them consistent with the MySQL override.
  3. nginx's TLS mount is currently incomplete. deploy/nginx.conf references /etc/nginx/tls.crt and /etc/nginx/tls.key, but compose doesn't mount them. Fix the template before running: for dev, remove the 443 SSL listener; when WebRTC/public access is needed, generate certs and mount the cert and key to those paths.

So "15 minutes" only applies once the image is ready, the database config is filled, and the TLS template is fixed. Without the prerequisites, this page doesn't promise the commands run as-is.

This isn't the only self-hosting path

This episode is about running the entire control plane yourself (ppcenter + mmx-origin + mmx-edge + nginx). If your goal is just to add a node you're paying for to someone else's already-running PPCDN control plane (say you just want to contribute a VPS, pay by license, and skip running scheduling/auth/billing yourself), there's a shorter path: buy a ppmmx license (a license code), and in the separate ppmmxDocker repo run deploy.sh --license-code <code>. The node auto-detects its role, registers, and gets billed by license as a NODE_ROLE_STANDALONE node (a single instance that both ingests the publish and serves WHEP viewers directly, without peering with other nodes) — none of this episode's steps are needed. This path is already live, with real self-hosted customers using it (see docs/deployment/ppcdn生产部署总结.md, entries from 2026-10-06 onward). On 2026-10-08 it surfaced and fixed a combined defect: on licenseCode registration, ppcenter never echoed the resolved nodeSecret back to the node, so the node had no credential, its whitelist sync failed permanently, and no appId could publish; after the fix the node reports its own secret and the console no longer shows nodeSecret. The two paths solve different problems — don't conflate them.


3. The 15-minute steps

Step 1: Get the code

git clone <ppcdn repo>
cd ppcdn

Step 2: Bring up dependencies (MySQL / Redis)

The default compose has no database; layer it on with the override:

First define the same file set used by all commands in this section:

COMPOSE="docker compose -f deploy/docker-compose.yml -f deploy/docker-compose.mysql.yml"
$COMPOSE up -d mysql redis

This starts MySQL 8.0 (db philcdn) and Redis 7 on ports 3306 and 6379.

Step 3: Fill local config

Edit deploy/config/ppcenter-local.json — it's mounted as /app/conf/ppcenter.json in the container. Three key items:

  • mysql.host/user/password/dbName: point at the MySQL just started (host mysql). The current override only sets a root password, so the template should at least set user=root, password=ppcdn_local, dbName=philcdn; production should use a dedicated user and secret.
  • nodeAuth.bearerToken: the platform-level node ingest token, local default local-dev-token-change-in-prod, must be changed in production (EP06's boundary).
  • redis.enabled: set true and fill the address to enable Redis; without it the control plane runs read-only.

Step 4: Prepare images and start

First prepare ppcdn/mmx:latest from the separate ppmmx repo and complete the TLS fix above, then in this repo:

docker build -t ppcdn/ppcenter:latest -f deploy/ppcenter.Dockerfile ppcenter
$COMPOSE up -d --build

Don't use make docker-up here, because it loads only the base compose and misses the MySQL/Redis override this section needs.

Step 5: Wait for health checks

The ppcenter container defines a healthcheck pinging http://localhost:8090/health every 10s. Wait for healthy:

$COMPOSE ps
curl -f http://localhost:8090/health

Step 6: Access

  • ppcenter admin: http://localhost:8090 (the superadmin SPA, baked in at build time).
  • Unified entry: http://localhost/ (nginx), with /origin/ and /edge/ proxied to the two mmx nodes.
  • Swagger: ppcenter/bin/docs/swagger.yaml in the repo.

Step 7: Clean up

$COMPOSE down -v

4. The current smoke template and a runnable alternative

The repo has make docker-smoke, but the current version can't be the runnable one-shot entry for the template above. deploy/smoke-test.sh's COMPOSE_FILE concatenates deploy/docker-compose.yml again from the deploy/ dir (inconsistent path); the script also doesn't load the MySQL override, and the nginx TLS mount remains unresolved. Before fixing these, use the same $COMPOSE commands as the previous section and run these control-plane checks by hand:

$COMPOSE up -d --build
curl -f http://localhost:8090/health
curl -i http://localhost:8090/admin/nodes
curl -i -X POST http://localhost:8090/v1/play/requests -H 'Content-Type: application/json' -d '{"streamName":"test","clientId":"smoke","requestRegion":"local","capabilities":["whep"]}'
$COMPOSE down -v

These verify ppcenter starts and its control-plane endpoints are reachable. They don't publish, play, or verify P2P, and don't prove the mmx media path works. HTTP status codes are also affected by auth and node registration state; judge with the response body, not by treating several status codes as success.


5. Three config items, read against earlier episodes

deploy/config/mmx-origin.yml and mmx-edge.yml have keys that put earlier mechanisms into concrete values:

① Node identity and capacity (EP14 pool / EP15 capacity / EP16 alarms)

mmxControl:
  nodeRole: origin          # or edge
  nodeRegion: local
  heartbeatIntervalSeconds: 10
  nodeCapacity: 30          # origin 30, edge 50

Nodes register their role, region and capacity with ppcenter via this config (EP14's poolId is the same kind of "statically declared at deployment" attribute). nodeCapacity is the denominator of EP15's "soft admission threshold" and the trigger basis for EP16's capacity_high alarm.

② ABR adaptive bitrate (EP11)

# mmx-edge.yml
webrtcABREnable: true
webrtcABRWSPath: /ws/control
webrtcABRSwitchCooldown: 3000

Edge has ABR on and switches layers over the /ws/control WebSocket — EP11's "seamless switching" lands on this switch and cooldown.

③ Pull mode: Publisher vs on-demand pull (EP15)

# mmx-origin.yml          # mmx-edge.yml
paths:                     paths:
  all:                       all:
    source: publisher          source: redirect
    sourceOnDemand: false      sourceOnDemand: true

Origin is publisher (waits for the publisher), Edge is redirect + sourceOnDemand: true — exactly EP15's "an Edge with no viewers doesn't consume Origin bandwidth" in config form.


6. From local to production: same components, three shapes

This local compose is only the minimal shape. The repo also gives two others for the same four components:

  • K8s manifests: deploy/k8s-*.yaml (ppcenter, mmx-origin, mmx-edge, nginx, infra), using Service DNS instead of hard-coded IPs, config via ConfigMap, and /health probes. For cluster scenarios.
  • Production deployment: systemd brings up mmx on each machine, the front door is nginx with real TLS certs, and MySQL/Redis are separate instances; the accompanying deploy/remote-deploy-*.sh scripts stop the service, back up the old binary, swap, restart, check status and clean old backups. Scaling is still EP15's manual 7 steps — after adding a machine, append the new Edge's private IP to Origin's forwardMmxTargets and reload.

local compose → K8s → systemd production can reuse components and protocols, but the differences are far more than orchestration and secrets. Network exposure, load balancing, persistent volumes, upgrade strategy, failure domains, observability and certificate automation all change production behavior.

Kubernetes especially needs three prerequisite designs:

  • Public UDP: WebRTC media ports can't rely on in-cluster Service DNS alone; you need a LoadBalancer, NodePort, host networking or a UDP-capable entry, and open cloud security groups and node firewalls.
  • ICE reachability: candidate addresses must be client-reachable public IPs/ports; handle Service NAT, external traffic policy and multi-NIC address advertisement; for complex NAT, evaluate STUN/TURN rather than assuming Pod IPs are reachable.
  • TLS and domains: the browser WebRTC/auth entry needs a trusted cert, correct SNI and WebSocket/HTTP upgrade config; use Ingress/Gateway with cert-manager or equivalent, and make clear that UDP media is not the same as an HTTPS Ingress.

6.1 PPCDN's own production shape

The above is generic; PPCDN's actual production fleet (full plan in docs/deployment/ppcdn生产部署方案.md, what actually happened in the sibling ppcdn生产部署总结.md) turns it into a few concrete trade-offs, which are the "local to production" landing of this episode:

  • Media plane in one Region, one AZ: Origin / Record / Edge all sit in the same availability zone of the media Region (e.g. AWS Singapore), so Origin→Record/Edge private-IP traffic is free; crossing AZs bills roughly $0.01/GB each way. This must be explicitly pinned at instance creation and verified per-machine — it isn't a default guarantee.
  • Control plane and data layer can split, even across Regions: ppcenter and ppcenter-db (MySQL8 + Redis) may live in a different Region from the media plane — nodes all dial outward and ppcenter never reaches back into media nodes, so no media data path is needed between them; only private/secure connectivity from control plane to data layer (cross-Region traffic and latency must be accounted for separately). Splitting the data layer into its own instance buys memory isolation, independent scaling and independent snapshots for +$12/mo.
  • Capacity is two unrelated lines: a single $12/mo (2vCPU/2GB) Lightsail Edge has two ceilings — instant concurrency (mmxNodeCapacity, 12 in production) and the Lightsail monthly egress quota (3TB ≈ 5.5 average concurrent viewers). They must be set and monitored separately: watching only the concurrency alarm and not monthly traffic gets you a bill far above the instance itself at month's end. It's the same trap EP15/EP16 keep flagging as "calibrate the denominator from real measurements."
  • Derive node count from the Region's vCPU quota: how much vCPU the media Region is granted directly caps the Edge count — the current line is 64 vCPU, minus 2 vCPU each for Origin/Record, allowing 30 $12 Edges, about 360 peak concurrency / ~165 average; for 500 peak concurrency, at 12 concurrent per instance you need about 88 vCPU, and you'd actually request 96–100 for headroom. These are usage-cap estimates, not an SLA; real capacity is further bounded by CPU/memory, public egress, P2P hit rate, TLS/DNS and Origin's forwarding-target scale.
  • Scaling is still a manual 7-step process: the capacity alarm can fire a webhook, but no autoscaler exists yet; every new Edge must have its private IP appended to Origin's forwardMmxTargets and reloaded.

7. Honest boundaries

  • Smoke ≠ end to end. The fixed make docker-smoke or the manual commands here only verify the control plane. Truly verifying a stream needs ppobs publish + ppplayer play, handling real NAT/network conditions.
  • This repo's integration tests only cover the record role's recording surface (integrationTest/record, compiling a real mmx binary and using real ffmpeg publishing to verify segment writes). The origin/edge media paths currently have no automated end-to-end test.
  • All local defaults must be changed: nodeAuth.bearerToken is a dev value, TLS is self-signed, MySQL passwords are local. None can go to production as-is.
  • Move this compose as-is onto a public cloud host and WHEP playback will most likely fail to connect (a confirmed, still-unfixed real bug). The local config's webrtcAdditionalHosts: [localhost, 127.0.0.1, mmx-origin/mmx-edge] only covers "accessing it via localhost on the same machine." Once you deploy the same template on a public cloud host (still docker bridge networking + port mapping), the default webrtcIPsFromInterfaces: true picks up the docker bridge's internal IP (e.g. 172.18.0.2) instead of the public IP, so the WebRTC ICE candidate address points somewhere the outside world can never reach — publishing (RTMP/SRT) is unaffected, but playback gets stuck on deadline exceeded while waiting connection. This is a real defect confirmed in production on 2026-10-08 while stress-testing a self-hosted node (see the corresponding entry in docs/deployment/ppcdn生产部署总结.md), and as of now it is still unfixed. Before deploying publicly you must manually add the public IP to webrtcAdditionalHosts, or switch the network mode to host.
  • The mmx image isn't built in this repo. Local compose uses a prebuilt ppcdn/mmx:latest, produced from the separate ppmmx repo.
  • compose/smoke is the current template, not a verified one-shot release. Fix the DB config, smoke path/override and TLS mount first, then talk about timing; "15 minutes" is the goal for a fully prepared environment, not a guarantee.
  • There's now a measured benchmark for roughly how much load this class of machine can take. On 2026-10-08, a 1 vCPU / 2GB self-hosted node (ppmmx-standalone-2, running the standalone deployment from the separate ppmmxDocker repo, not this episode's two-container origin+edge compose) went through a stepped load test — pure H.264 video, 1280×720@30, roughly 2.5 Mbps, SRT publish, with the load generator in roughly the same datacenter as the machine under test (RTT ~4ms). Result: stable throughout up to 18 concurrent WHEP viewers (CPU peaking at 53%), and the 19th viewer pushed CPU from ~40% to 88-103% within about 15 seconds and started dropping frames — there's almost no transition band between 18 and 19 viewers; once a single-core machine nears saturation, congestion arrives as a cliff, not a gradual decline. This isn't a direct test of this episode's compose topology (here origin and edge split the load across two containers, whereas standalone carries everything in one process), so treat it only as a capacity reference point for "a similarly small machine running the same mmx media core" — your own hardware and bitrate still need their own measurement and calibration. Full method and data: docs/test/ppmmx-standalone-1vcpu-stress-test.zh-CN.md.

8. Diagrams

8.1 Local compose components and ports

                       ┌───────────────────────────┐
   browser / client ──▶│        nginx   :80/:443     │
                       │  /origin/  /edge/  /v1/  /ws/│
                       └───────┬───────────┬────────┘
                               │           │
              ┌────────────────┘           └────────────────┐
              ▼                                             ▼
      ┌───────────────┐                             ┌───────────────┐
      │  mmx-origin   │◀── pull (only after viewers)─│   mmx-edge    │
      │ :1935 RTMP in │                             │  :8890 WebRTC │
      │ :8889 WebRTC  │                             │  :8190/udp    │
      │ role=origin   │                             │  role=edge    │
      │ cap=30        │                             │  cap=50, ABR  │
      └───────┬───────┘                             └───────┬───────┘
              │        node registration / heartbeat / commands │
              └───────────────┬─────────────────────────────┘
                              ▼
                     ┌─────────────────┐        ┌──────────────┐
                     │   ppcenter      │───────▶│  MySQL(req.) │
                     │   :8090 control │        └──────────────┘
                     │  /health /admin │        ┌──────────────┐
                     └─────────────────┘───────▶│ Redis(opt.)  │
                                                └──────────────┘

8.2 The journey of one command

Current manual smoke (same compose files)
      │
      ▼
$COMPOSE up -d --build
      │
      ├─ /health                        control-plane health
      ├─ check /admin/nodes, /v1/play/requests
      └─ $COMPOSE down -v              clean up

Verifies "orchestration + control plane" works; not "publish → play" works

9. Wrap-up and next episode (script notes)

This episode reviewed the current compose as a deployment template: fill in the mmx image, database config and TLS mount first, then start, check and clean up with the same compose files; the existing make docker-smoke still needs path and dependency fixes. A control-plane smoke isn't media end to end. On Kubernetes, you must separately solve public UDP, ICE-reachable addresses and trusted TLS — production differences can't be reduced to "just a different orchestrator".

Next episode, we add a new capability to this running live system: recording & playback — turning a live stream into per-session replayable VOD.