Performance methodology
Published Performance numbers are generated on a single reference host. A Python
harness runs curated dnsperf scenarios against a Conduit binary, records
same-host metrics into JSON, and — after validation — that JSON is promoted
into the repository under perf/results/references/. Operator-docs charts, CSV
downloads, and study evidence are then regenerated from that committed
reference only; docs CI does not re-run load against a live binary.
The sections below explain how to interpret those published cells: suites and load shapes, how load is applied, answer-quality gates, drain vocabulary, and how to read charts. Commands to replay suites locally, refresh the publish set, or wire annotations are under Reproduce against a binary (Maintainer publish for the promote path).
Suites
| Suite | What it measures | Primary metrics |
|---|---|---|
scale |
Runtime models (sync / split_io) under paired load shapes |
Achieved QPS, latency |
shutdown_drain |
Drain policies under load (especially slow upstream) | Drain duration, client loss during stop |
feature_tax |
Cost of metrics / dnstap / OTLP / logging / tracing relative to observability off (including collect-vs-emit, combined surfaces, split_io scrape, and scrape-hammer cadence) |
QPS / latency delta vs baseline |
lifecycle |
Cold start and config apply | Cold-start ms, apply latency |
lossless_upgrade |
Zero-downtime upgrade handoff (when available) | Gap / loss, child READY, drain duration |
Maintainer Layer A microbenchmarks (make performance) are separate and are not the
published contract.
Load shapes
Answer-source shapes used by forward suites:
| Shape | Meaning |
|---|---|
forward_fast |
Stub upstream answers immediately |
forward_slow |
Stub upstream holds every answer for a fixed 50 ms |
cache_hit |
Lookup cache enabled and warmed |
cache_churn |
Lookup cache under matched high turnover (short TTL, small max_entries, diverse queries; no warm plateau) |
Warm cache_hit cells (memory or LMDB) are read-mostly after warm: the harness
probes a small set of answers so the load window is near-100% hits with rare
inserts. That shape measures lookup/serve cost, not fill/evict churn, capacity
pressure, or hit-rate under turnover.
High-churn cache_churn cells intentionally keep the cache under continuous
fill/evict pressure (large query file, stub TTLs long enough that entry-cap
turnover dominates, cap below query cardinality). Published takeaways for that
pole lead with relative QPS and hit/miss (or hit-rate), not QPS alone.
Average latency under the elevated outstanding window still tracks roughly
outstanding ÷ QPS (Little's Law)
and must not be read as raw disk service time. The published memory-vs-LMDB
churn study uses eight sync ingress workers (and an explicit LMDB
shard_count of 16) so the compare is not starved by a two-thread sync path.
A thin-ingress (two-worker) companion pair remains in the catalog for
continuity. When record_cache_metrics is set, the harness also scrapes mean
fill and capacity-eviction path durations (conduit_cache_fill_duration_seconds,
conduit_cache_eviction_duration_seconds) into run secondary so eject/populate
cost can be compared without reading dnsperf queue latency. The published
memory-vs-LMDB churn study surfaces those means on the QPS evidence table and a
dedicated fill-duration figure. Lazy TTL expiry is
wall-clock deterministic on read; what is arbitrary under LMDB capacity
pressure is when_full victim selection (evict_one / sample), not the
expiry clock.
LMDB cache cells
Published LMDB performance cells use a real-disk environment path (for
example under /var/tmp/… on a disk-backed mount). Do not promote LMDB
numbers from a tmpfs path — that understates durable-backend I/O cost.
Published warm LMDB cells use lmdb.sync: full (default); sync mode is
not a warm-pole matrix. High-churn comparative cells annotate each LMDB member’s
explicit lmdb.sync value (full, no_meta, periodic, or none) as
first-class peers. A periodic cell also records its lmdb.sync_interval
because its durability window and write cost depend on that setting. Absolute LMDB
QPS is lab-, disk-, and sync-mode-dependent; prefer
relative claims on the same host. Revisit absolute LMDB churn QPS after
changing sync mode or storage before treating older numbers as current.
The stub upstream is never the constraint
Forward shapes answer from a stub responder, not a real resolver, and the stub is built so that it cannot be what a chart is measuring:
- Replies are served by several worker processes sharing one port, so the responder scales past a single CPU. Loaded directly, with no Conduit in the path, it answers roughly 480k QPS on the fast shape — about three times the highest rate any published cell drives through it.
- The delayed responder queues its replies rather than sleeping on them, so
forward_slowmodels a backend that is slow but still accepts new queries — the delay never serializes into one answer per 50 ms. On that shape the harness tops out near 39k QPS, which is the load generator's own in-flight window rather than the responder; the responder loses no queries there. - On a lab host with mixed fast and efficient CPU cores, Conduit runs on the fast cores and every harness process — load generator, stub responder, and telemetry receivers — is confined to the others, so the measurement apparatus cannot take CPU away from the thing being measured.
A stub that saturates does not simply cap a chart; it changes what the chart measures, because Conduit then starts refusing queries it cannot forward. See Only successful answers count for how the harness detects that.
Runtime compares must present paired shapes — especially forward_slow for
sync vs
split_io — so charts do not favor only the shape that helps one model.
See Sync vs split_io.
Reading forward_slow cells — saturation against a 50 ms backend
forward_slow cells offer far more
concurrency than a runtime that blocks on upstream latency can absorb, which is
the point: they show what happens to each runtime model when the backend is slow
and the client keeps asking. Read both columns together. A runtime that
multiplexes in-flight queries answers near the upstream delay itself; a runtime
that occupies a worker for the whole round trip reports a small fraction of that
throughput and an average latency of seconds, because queries wait in Conduit
rather than at the backend. Every published cell here is measured on a stub
upstream that stays well clear of saturation and is checked for successful
answers, so the numbers are Conduit's behavior, not the harness reaching its
limit. See
How load is applied
and Only successful answers count.
Drain policy vocabulary
Shared by shutdown_drain (and future lossless upgrade):
| Policy | Intent |
|---|---|
drain_complete |
Wait until idle (or a very long budget) before exit |
drain_budgeted |
Configured drain_timeout_ms budget |
drain_minimal |
Near-zero wait / drain: false cutover |
Config fields: Shutdown. Comparative evidence: Drain policy under slow upstream.
Lab profiles
Every run records a lab host profile (id, CPU, cores, OS, binary identity, loadgen).
Published charts come from a single reference host (maintainer-ws-1). They
compare scenarios on that same host; they are not cross-host capacity claims.
See the disclaimer on the performance hub.
Publish-quality labs require CPU frequency governors to be performance when the
host offers that governor (powersave / schedutil / mixed governors introduce large
QPS swings). Boards that never expose performance are not blocked. How to set
the governor (or intentionally bypass the gate for a noisy probe) is under
Reproduce — Maintainer publish.
Publish hosts also need a large enough UDP receive memory ceiling
(net.core.rmem_max, typically ≥ 4 MiB) so fixture
listeners.rcvbuf can take effect.
With the OS default (~208 KiB), elevated same-host dnsperf can overflow
Conduit's ingress socket: the kernel increments RcvbufErrors, dnsperf reports
Queries lost, and Conduit metrics still show equal query and response counts
for every datagram that arrived. That pattern is host socket buffering, not
dnsperf soft-stop cutoff and not Conduit refusing to answer received queries.
Remediation (sysctl + harness gate) is under
Reproduce — Maintainer publish
and the in-tree perf/README.md.
Primary load generator
Published QPS / latency cells use DNS-OARC dnsperf. The default obtain-and-run path
is a pinned Docker image (dnsconduit-dnsperf:2.14.0 from perf/fixtures/dnsperf/).
A native dnsperf on $PATH is an optional override. Exact flags and query files live
in perf/README.md in a source checkout.
How load is applied (not a fixed offered QPS)
Published scale / feature-tax cells are not “sustain X queries per second” runs. dnsperf sends as fast as its outstanding-queries window allows, and that window's size is the single most important knob for whether a chart measures Conduit's own capacity or just restates 1/latency:
dnsperf … -c <clients> -T <threads> -l <seconds>
| Knob | Meaning |
|---|---|
-l |
Timed window (CLI --time; publish-quality omits short smoke overrides) |
-c / -T |
Parallel client sockets and send/receive pairs |
-Q |
Not set for any published cell — no offered-QPS cap |
-q |
Max outstanding queries; sender pauses when this many are in flight until replies or timeouts free slots |
By Little's Law, achieved QPS × average latency ≈ outstanding queries. With a
thin window (dnsperf default -q ≈ 100, -c 4/-T 2), a fast round trip
(sub-millisecond) means QPS is almost entirely a restatement of
100 / latency — two configurations that both answer quickly can trade chart
rank on nothing but a fraction of a millisecond of host noise, even though
neither is actually CPU- or capacity-bound. That thin window is only a
trustworthy throughput signal when the round trip is slow enough, or
deliberately used as an artificial constraint, that the outstanding window
itself is the intended object of study.
Recipe by load shape:
| Load shape | Recipe | Why |
|---|---|---|
forward_fast, cache_hit, cache_churn, forward_slow (scale, feature_tax) |
Elevated: clients 16 / dnsperf_threads 8 / max_outstanding 2000 |
One recipe across published comparative cells. The elevated window lets achieved QPS reflect Conduit's own capacity instead of restating 1/latency, and on the slow shape it keeps a multiplexing runtime from being capped by the loadgen rather than by Conduit. |
shutdown_drain, lossless_upgrade |
Thin: dnsperf default (-q ≈ 100), -c 4/-T 2 |
These cells measure drain timing and client loss during a stop window, not steady-state throughput. |
forward_slow cells run a 30 second window instead of 10. A runtime that
blocks on a 50 ms upstream completes very few queries, and a short window would
leave too small a sample to mean anything.
Published scale and feature_tax cells are each the median of 3
independent rounds (same scenario, same recipe, rerun end to end) rather than
a single draw, so a one-off scheduling hiccup on the shared lab host cannot flip
a ranking. The observed per-round range is recorded on each cell's
quality.notes in the reference JSON. A full curated publish-set lab refresh
uses the same N=3 → merge-median bar for every member in that bag (including
shutdown_drain when it is in the publish-set). lifecycle cells outside that
refresh path remain single-shot.
N=3 is the default publish bar. Maintainers may remeasure a noisy subset
at N=5 (then merge-median) when a cell's per-round min–max span in
quality.notes is large relative to the median (roughly ≳20–25%) — for example
an elevated split_io pole that swung by half in one round. Do not raise the
default bag-wide N to 5; prefer subset remesure when one axis is the problem.
Only successful answers count
A forwarding proxy that cannot forward still answers: it returns SERVFAIL, and it does so in microseconds. A load generator counts those refusals as completed queries, so an overwhelmed configuration can report a higher QPS and a lower average latency than a healthy one. Published throughput must never be built on that.
Every load-bearing cell therefore records the loadgen's response-code
breakdown, and the harness compares it against the answer the scenario is meant
to produce (NOERROR for forward and cache shapes):
- At least 99% successful answers: the cell is a valid measurement.
- Below that: the cell is recorded as invalid, and promoting a reference that contains it fails outright. There is no override for published data.
Cells where failed answers are the subject of the measurement — the
shutdown_drain stop window, for example — opt out of the check explicitly in
the scenario catalog.
Both parts of the check appear in the reference JSON as
metrics.response_codes and metrics.answer_ok_percent.
Changing the outstanding window (or other recipe knobs) for a local probe is fine for exploration, but those results are not interchangeable with published cells unless the promoted reference is refreshed under that recipe. See Reproduce against a binary.
Separate from dnsperf: Conduit fixtures also set
forward.outstanding_per_backend and orchestrator.txn_table_capacity. Those
are server-side in-flight limits, not the loadgen -q window. Lab fixtures
size them well above the offered window (8192 and 65536) so that a published
chart measures the runtime rather than a fixture ceiling; production values
should be chosen for your own backends, not copied from the lab.
Studies (comparative evidence)
Tuning evidence (studies) group published scenario cells into operator questions (runtime choice, metrics tax, drain policy, and similar). Prefer findings and studies for decision narrative; use reference results only when you need the dense chart warehouse. Study pages embed the same promoted evidence as the warehouse.
Charts, CSV, and tables
Operator-docs charts are static SVG derived from promoted JSON. Paired load shapes
that differ by orders of magnitude (for example forward_fast vs forward_slow) use
separate charts so each Y-axis stays readable. Each chart has a CSV download of
the values shown in that figure. Tables include richer columns (sent / completed /
lost, workers, latency) when recorded, and scenario labels link to
scenario descriptions. Missing series are omitted or
marked unavailable — numbers are never fabricated. Callouts next to some figures
are stable footnotes from the harness catalog; how they are attached at promote
time is under
Reproduce — Maintainer publish.