Performance
Directional performance evidence for DNS Conduit: runtime model tradeoffs, worker sizing, observability tax, and shutdown drain under load. Published figures are same-host comparisons on a single reference host, not service-level objectives.
Findings
Short directional takeaways from published studies. Each bullet is a same-host relative result on a single reference host — not a portable capacity target or SLO. Absolute QPS on study pages is lab detail only; remeasure on your hardware before sizing.
- Runtime model — Under a fast upstream,
split_ioreaches about 1.9× the QPS ofsync(~141k vs ~74k). → Sync vs split_io - Sync ingress sizing — More ingress workers raise throughput under fast forward (~38k → 74k → 146k → 240k from 1→2→4→8); gains stay large at the top of that range on this lab. → Ingress concurrency (sync)
- Answer cache — When nearly every query hits, a warm cache path is about 3.5× forward-fast QPS on this lab — a ceiling, not a forecast of your hit rate. → Cache hit vs forward
- Memory vs LMDB (warm) — Warm LMDB costs about 6% QPS versus warm memory on this lab (~311k vs ~329k) under sync cache_hit — read-mostly after warm, not churn. → Memory vs LMDB warm cache_hit
- Memory vs LMDB (churn) — Under matched high churn (ingress-8), LMDB
sync: fullcosts about 99% QPS versus memory (~3k vs ~232k; roughly 83.9×).no_metais only about 1.6×full(~4k);periodic(sync_interval1s) is about 43.8×no_meta(~192k, about 17% versus memory);noneis nearby (~177k, about 24% versus memory). Hit rates stay similar; pick sync mode from the durability decision tree, not the QPS chart alone. → Memory vs LMDB high-churn cache - Metrics scrape — Versus observability off, standard scrape costs about 9% QPS on the sync ladder in this median refresh; collect carries most metrics cost, with emit a thin band. → Metrics scrape tax, Collect vs emit
- Logging and tracing — Debug logging costs about 5% versus warn on this median; full pipeline tracing (~100% sample) costs about 37% versus observability off. Keep tracing for diagnosis windows, not standing production. → Logging verbosity tax, Pipeline tracing tax
- Dnstap and OTLP — Sampled dnstap about 6%, fuller emit about 10%; scrape+dnstap together about 15%. OTLP push about 8% — same band as standard scrape. → Dnstap emit tax, Combined metrics + dnstap, OTLP tax under load
- Shutdown drain — Under slow upstream, complete drain is ~114 ms; budgeted ~164 ms; minimal ~63 ms, with complete recording the most in-flight client failures. Pick a policy for your restart window on purpose. → Drain policy under slow
Full comparisons, charts, and omitted-member notes live under Tuning evidence (studies).
Start here if deploying
Three steps for first sizing on your hardware. Published figures stay same-host comparisons on a single reference host — not capacity SLOs.
- Pick a runtime — read the runtime model finding (~1.9×
split_iovssyncunder a fast upstream), then the Sync vs split_io study. - Size metrics scrape — read the metrics scrape finding (~9% standard vs off on the sync ladder), then Metrics scrape tax and Operator metrics bases.
- Remeasure one study — replay a study against your binary with
Reproduce against a binary (
--study …) before locking in worker counts or observability posture.
How to use this section
- Decide — skim Findings, then open the matching study for charts and takeaways.
- Interpret — Methodology explains load shapes, how load is applied, and how to read published numbers.
- Remeasure — Reproduce against a binary on your hardware before sizing decisions.
- Look up raw numbers (optional) — When you need the full charts or CSV rows, open Reference results. To see what a specific table row means, follow its link into Scenarios.
Already know which tradeoff you are weighing? Jump to the decision map on the studies hub.
Disclaimer
Published reference results were measured on a single reference host
(maintainer-ws-1). Charts and studies are same-host comparisons — each
figure contrasts configurations against baselines taken on that host under the
same load recipe. They are not capacity guarantees or portable cross-host
SLOs. Reproduce with the harness against your Conduit binary and hardware
before making local sizing decisions. Do not treat absolute QPS as transferable
across machines.
In this section
- Findings — short directional takeaways
- Start here if deploying — runtime → metrics tax → remeasure
- Tuning evidence (studies) — decision evidence by category
- Methodology — how to interpret published numbers
- Reproduce against a binary — remeasure on your hardware
- Reference results — dense chart/CSV warehouse (optional)
- Scenarios — row-level glossary (deep links from tables)