Skip to content

EP20: Is a 30% bandwidth saving realistic? A real-data review

Recap EP19 "Load-testing methodology: from a single box to global load"
Next Series finale (end of Module 5)

0. Goals of this episode

By the end, the viewer should be able to:

  1. Rewrite "save 30% bandwidth" as a testable hypothesis: HEVC's saving vs H264 depends on encoder, preset, content and quality metric; P2P direct saves part of Edge egress — the two can't be conflated.
  2. Separate three kinds of numbers: measured, modeled and unverified. This episode lays all three out.
  3. Understand unit economics: why every saved TB is gross margin, and which variables erode it.
  4. Take away an honest conclusion: any saving ratio must be verified with a same-quality target, real device mix, P2P hit rate and one unified cost model — industry rules of thumb can't be treated as operational results.

1. Opening hook (script notes)

The very first episode raised cost: traditional CDN cost grows linearly with viewers, and egress is the big item. In this final episode we return to that business question — does the specific-sounding claim "save 30% bandwidth" have real data behind it? I'll be blunt: behind that phrase are two different things — one a codec-layer estimate, one a distribution-structure capability — and part of it is not yet verified by real operating data. No number-pumping here, just an honest review.


2. Split "save 30% bandwidth" into two things first

The same phrase "save bandwidth" maps to two completely different technical paths here:

Path A: HEVC saves traffic Path B: P2P saves egress cost
Mechanism HEVC has higher coding efficiency than H264 at equal quality On a successful direct connection, media bypasses Edge
Layer Codec layer Distribution-structure layer
Saves The bitrate/traffic each viewer occupies The Edge node's egress bandwidth
Magnitude Set an interval hypothesis first, then measure per encoder/preset/content/quality metric Depends on hit rate, and at most 3 per publisher
Precondition Playback device supports HEVC Meets P2P direct admission (EP13)

Separating the two is the precondition for this episode's clarity. Lumping them into "save 30%" is a simplification unfair to both mechanisms.


3. Path A: HEVC savings are an interval to be measured, not a fixed 30%

The approach is multitrack (EP05): the publisher sends two independent streams, H264 and HEVC, and the player picks by the browser's decode capability — HEVC-capable devices pull HEVC, others fall back to H264.

"HEVC saves 30% at equal quality" can serve as the central value of an interval hypothesis for capacity planning, not a universal fact. The actual gap varies with encoder implementation and version, preset/ latency profile, resolution and frame rate, content motion and noise, bitrate range and quality metric; and "equal quality" by subjective rating, VMAF, SSIM or PSNR may differ. Do bitrate-ladder tests on representative content, plot BD-rate or same-quality bitrate difference, and report an interval.

But three preconditions must be stated:

  • It only applies on HEVC-capable devices. Unsupported users still go H264. The overall saving can't be simplified to 30% × device share but must be weighted by each codec's actual viewing bytes, measured bitrate gap and fallback rate.
  • It adds publisher-side cost: two independent streams clearly increase the publisher's encoding resources and uplink bandwidth (the system accounts capacity, bitrate and alarms per codec session, so EP15/EP16's capacity model must count two sessions).
  • It saves "traffic", not "count": the viewer count is unchanged, the path is unchanged (still via Edge), only each copy of the stream is smaller.

In one line: 30% is only a hypothesis to verify. Fix the quality target and low-latency constraint first, then measure the saving interval across encoders, presets, content sets and device compatibility.


4. Path B: P2P savings — clear per connection, bounded in total

The second path is where this system is more imaginative and needs more verification: P2P direct connect.

Its cost-saving logic is direct (EP13):

  • Every successful direct view has zero Edge egress — data flows from the publisher straight to the viewer, entirely bypassing the edge server.
  • Each publisher serves at most 3 direct connections at once by default (maxP2PSessions, default 3, adjustable from the superadmin console), strictly controlled by an atomic lease on the central side.
  • The main publish link always has priority: P2P uses only the safely-measured spare of the publisher's uplink, and on detected uplink congestion the system will immediately sacrifice P2P connections to protect the main publish — so enabling P2P never risks "saving money at the cost of the host's quality".
  • P2P is an enhancement layered on a stable Edge network, not essential: when a stream can't go direct at all, viewer experience equals a traditional CDN and the cost advantage is pure upside.

A single P2P direct connection does remove that viewer's Edge media egress, but you can't claim the overall saving necessarily exceeds some codec ratio. At most 3 per publisher, so a hot stream's P2P offload ratio falls as viewer count grows. The overall saving must be computed from the actual P2P byte share, Edge's marginal billing and P2P signaling/ops cost.

And the hit rate — precisely where real data is most lacking.


5. A real-data review: laying out three kinds of numbers

The core of this episode is to divide all relevant numbers into three classes. An honest review never dresses the second and third classes up as the first.

① Measured (reported, reproducible)

  • End-to-end latency: P2P direct runs at ~70ms; when delivery goes through Edge, the ingest protocol is the main variable — WHIP ingest ~108ms, SRT ingest ~386ms (the gap mainly comes from SRT's fixed TSBPD receive window on the ingest side, a protocol-level trade-off between "loss robustness" and "latency", not a codec difference — the variable in this comparison is the ingest protocol, WHIP vs SRT, not the codec, HEVC vs H264; don't mix the two axes when citing this, see the Live Streaming Latency Performance Test Report for detail).
  • The various NAT-admission rejection reasons, with unit-test coverage; the Play/NAT/records/split-rec HTTP contract also has a layer of black-box production regression testing against the live deployment (EP19 §4), though it only verifies the interface contract, not whether real media actually connects.
  • But note: these are approximate observations under limited scenarios. They must be published with each group's sample size, percentiles, device/network/encoding parameters and measurement error; don't treat a point estimate as an SLA, nor infer causality from a difference alone.

② Modeled (a formula and caliber, but not measured)

To avoid mixing the "per-node amortization" and "per-GB cloud billing" models, this episode uses one monthly full-cost model:

Monthly total cost = node fixed fee + actual egress overage fee + cross-region/origin-pull fee
                   + storage & request fee + control-plane/monitoring fee + ops cost
Effective delivered traffic = Edge delivered bytes + P2P delivered bytes
Unit cost = monthly total cost / effective delivered traffic
  • Traffic included in a plan's nodes counts into the node fixed fee; only bytes beyond the plan incur overage — you can't book the same traffic both at $4/TB and $0.0122/GB.
  • "Weighted-average bitrate 1.75 Mbps" and the viewing-resolution distribution can feed traffic prediction, but must be updated with actual watch time, codec share and observed bitrate.
  • P2P billing is modeled as "duration × average throughput", and in the current implementation "average throughput" is a fixed constant of 1Mbps — it does not collect each session's real throughput (docs/design/ppcdn-billing-and-traffic-monitoring.zh-CN.md §2.7). So even if this number is later used to back out "how much traffic P2P saved you", that would still be a modeled number, not a measured one.
  • Price and cost are different dimensions. With a target gross margin m, price = unit cost / (1-m); "cost × 2" corresponds to a 50% margin or 100% markup — keep the terms straight.

③ Unverified (gaps)

  • P2P's real hit rate and actual byte-offload share: P2P works, but there aren't enough production samples to quantify the cost benefit; a working latency run doesn't mean the hit rate is verified.
  • HEVC's measured bitrate interval and actual usage share: needs per-content and quality-target experiments, not fixing at 30% upfront.
  • Cross-AZ traffic, cloud quota overage, ops headcount: variables that erode gross margin and lie outside system monitoring (EP16 covered the cloud quota blind spot).

The most important sentence of this episode: the link working doesn't mean the saving ratio is verified. HEVC needs same-quality encoding experiments, P2P needs actual delivered-byte-share measurement, and then everything goes into one monthly full-cost model.


6. Unit economics: compute the saving with one model

Don't directly subtract an "Edge marginal cost" from another model from the price. In one unified model, run the baseline and optimized scenarios separately:

baseline unit cost  = baseline monthly total cost / baseline effective delivered traffic
optimized unit cost = optimized monthly total cost / optimized effective delivered traffic
monthly saving      = baseline monthly total cost - optimized monthly total cost

If Edge is usage-billed, P2P offload may directly lower marginal egress; if Edge is a monthly node with included traffic still within the plan, the short-term bill may not change at all, only raising sellable headroom. A confirmable cost saving exists only when it avoids new nodes, reduces overage, or lowers the actual bill. Whether revenue is unchanged also depends on the billing contract; don't automatically equate offloaded bytes with gross margin.

"Is revenue unchanged" isn't just a rhetorical hedge — this system's own billing reality is a ready-made example of it. Production currently actually runs a flat $1/day per active account, completely blind to bytes (the hardcoded dailyFeeUSD constant in internal/store/billing_store.go): which means every GB of egress saved today falls 100% into the operator's own gross margin and changes nothing on any customer's bill, because the bill never reads a traffic field at all. The new usage-based scheme ($0.0244/GB downstream + $0.02/GB-month of recording storage, charged in real time per minute, no monthly minimum, with any day that has usage but meters below $0.0001 bumped up to $0.0001) is already code-complete (ProcessRealtimeUsageBills/BillingManager), but it's gated by the BillingConfig.UsageBasedEnabled config switch — BillingManager.Start() picks one of two loops to run based on that switch, either the old flat-daily-fee loop or the new per-minute real-time billing loop — and there is no known production deployment record proving that switch has ever been flipped on: a true value shown in the local repo's config file isn't evidence (remote config has been hand-edited out of sync with the local copy before), and the deployment log itself contains no entry saying usage-based billing has gone live. Even P2P's "duration × 1Mbps" billing accumulator hangs off the same switch-gated real-time billing loop (chargeRealtimeP2PUsage is a sub-function called from inside ProcessRealtimeUsageBills, not an independent standing loop) — the design doc's claim that P2P billing is "independent of Edge downstream billing" refers to independent accounting structure (its own accumulator, its own bill line item), not independently active: when the switch is off, this billing logic doesn't run at all, exactly like bandwidth billing. This is yet another place, one this episode keeps returning to, where "implemented in code" and "live in production" are easy to conflate. Put differently: if that switch is ever actually flipped on, the traffic HEVC/P2P save would for the first time show up directly on a customer's bill; but as of now, saving bandwidth has no confirmed effect on the revenue side at all — every confirmed effect is on the cost side.

But gross margin is eroded by four variables, all of which must be watched:

  1. P2P hit rate — real proven samples exist, but production samples are insufficient to estimate a stable hit rate (§5's gap).
  2. HEVC usage share — more capable devices, more traffic saved.
  3. Plan allowance and overage price — the allowance is usually a billing threshold, not a cut-off line; the marginal cost inside and outside differs.
  4. Ops headcount — scaling, certs and node onboarding are still largely manual (EP15), the most realistic marginal cost of a self-hosted CDN.

One easy-to-assume claim needs clearing up here: "HEVC and P2P can stack" is not currently true — playback decisions branch by codec up front (see EP13): HEVC-capable clients go straight to Edge and never attempt P2P at all from the start; only H264 clients ever get a chance at P2P direct-connect eligibility. In other words, these two cost-reduction paths are mutually exclusive today: a given stream gets either HEVC's codec-layer traffic savings or P2P's distribution-layer egress savings, never both at once. Getting both requires finishing an engineering task first — bringing HEVC into P2P eligibility — which hasn't been done yet; it isn't a capability that's already stacking in production.


7. Honest conclusion

  • "HEVC saves traffic": direction holds, the ratio isn't a constant. Measure an interval across encoders, presets, content, low-latency constraints and quality metrics, then weight by real device/viewing share.
  • "P2P saves bandwidth cost": P2P works at ~70ms, and a direct connection reduces the corresponding Edge egress; but at most 3 per publisher, and the production hit rate, byte-offload share and bill saving remain to be verified.
  • Turning "cost reduction" from an architectural capability into a stably reproducible business metric is the first priority of the next stage — which is exactly the methodology from EP19: real-network acceptance + reproducible evidence.
  • So the honest answer to "is a 30% bandwidth saving realistic" is: treat 30% as an experimental hypothesis, not a conclusion; the final number must come from same-quality encoding tests, real traffic distribution and one unified cost model.

8. Diagrams

8.1 Two cost-reduction paths, don't conflate them

        "save 30% bandwidth"
              │
      ┌───────┴────────┐
      ▼                ▼
  HEVC (codec layer)   P2P (distribution layer)
  measured saving      direct success → that view's Edge media egress = 0
  interval
  saves "each copy     saves "a whole egress stream"
   is smaller"
      │                │
  precondition:         precondition: H264 (HEVC currently excluded
   device supports        from P2P eligibility)
   HEVC                 variable: hit rate × egress unit price
      │                │
      └───────┬────────┘
              ▼
   Mutually exclusive today, not stackable: HEVC clients go straight to
   Edge and never attempt P2P; only H264 clients enter P2P eligibility
   (sequential fallback, not racing)

8.2 Three kinds of numbers, don't treat the latter two as the first

① Measured (reported)
   · latency: P2P direct ~70 / WHIP ingest ~108 / SRT ingest ~386 ms (protocol difference, not codec)
   · boundary: report sample size, percentiles, scenario and error; not an SLA

② Modeled (a formula)
   · weighted bitrate 1.75 Mbps
   · P2P billing caliber: duration × a flat 1Mbps (not measured throughput)
   · monthly total cost / effective delivered traffic; same formula for baseline and optimized
   · boundary: marginal cost differs inside/outside the plan allowance; cost and price are separate;
     saved traffic ≠ saved bill (production runs a flat $1/day today; the usage-based switch is
     not confirmed to be on)

③ Unverified (gaps)
   · P2P production hit rate / byte-offload share
   · HEVC same-quality bitrate interval / actual device share
   · cross-AZ / cloud quota overage / ops headcount

9. Series wrap-up

That's the whole course. We started from episode one's "interactive-video dilemma" and went through protocol choices, weak networks, observability, quality, P2P, node pools, scaling, bursts, deployment and recording, finally returning to the business model itself.

If one sentence sums up the methodology: every number must state its measurement boundary, sample size and model. The ~70ms P2P direct connection and the ~108ms/386ms WHIP/SRT ingest numbers are both real-link observations, and the gap between them comes from the ingest protocol, not the codec; "save 30%" is only a hypothesis awaiting verification across encoders, presets, content and quality metrics; and even if P2P genuinely saves bandwidth, under the current flat-$1/day billing model that saving today only affects the operator's own gross margin — there's no confirmed evidence yet that it shows up on any customer's bill. Separating measured, modeled and unknown matters more than a pretty single-point number.

Thanks to everyone who read this far. The system's code is open source — go deploy it, load-test it, and fill in the "not yet verified" — that's the real sequel to this course.