Troubleshooting
This hub collects symptom-oriented pointers for common operator issues. Each section links to the canonical topic page — it does not duplicate full configuration reference.
| Section |
Covers |
| DNS and forwarding |
Startup bind failures, no client reply, SERVFAIL, upstream timeouts |
| Backend health |
Probes, passive fast-trip, fail-open, drain/freeze, health metrics |
| Dataplane runtime and concurrency |
Slot-pool exhaustion, split_io runtime and worker misconfiguration |
| Control plane |
conduitctl connectivity, rejected reload/apply, overlay surprises, restart-pending changes |
| Observability |
Metrics scrape, OTEL push, dnstap, tracing, logging |
DNS and forwarding
Conduit exits at startup or never serves DNS
| Symptom |
Likely cause |
What to check |
| Process exits immediately; errors on stderr |
Invalid YAML, validation, or snapshot compile at first startup |
conduitctl validate --file PATH; fix messages (script '…', rule '…', data source '…', field validation) — Config file — Validation |
Address already in use / bind error in logs |
Listener port taken, or threads > 1 on UDP without reuse_port: true |
Reference: listeners (reuse_port, threads); ss -ulnp \| grep PORT |
Permission denied binding port 53 (or other privileged port) |
Process lacks bind capability |
Run as root, use CAP_NET_BIND_SERVICE, or bind a high port (for example 15353) — Install and run |
| Process runs but clients get no answer |
Empty listeners.listeners or no pool/backends |
At least one listener and one pool with a backend — Config file — What makes a config runnable |
conduitctl validate does not bind sockets — bind failures appear at startup or after restart, not during validate alone.
Confirm listeners after a successful start:
ss -ulnp | grep conduit
# or match your listener port from config
Client gets no response (timeout or silence)
Quick path check (adjust ports):
dig @127.0.0.1 -p 15353 +time=3 +tries=1 example.com A
ss -ulnp | grep -E '15353|5300'
SERVFAIL or other error RCODE
| Symptom |
Likely cause |
What to check |
SERVFAIL on most queries |
Upstream unreachable, wrong pool/backend, or forward failure |
Backend address correct and reachable? Upstream resolver running? — First query |
SERVFAIL after retries exhausted |
max_attempts, pool exhausted, or max_txn_duration_ms |
Retries and transactions; conduit_retries_total |
SERVFAIL immediately (no upstream wait) |
Missing pool, empty backends, or route failure |
Pool name from rules matches pools:; at least one backend — Pools and backends |
NXDOMAIN / other RCODE |
Upstream answer or policy set_rcode |
Expected upstream behavior vs response-hook policy — Rules and actions |
When metrics are enabled:
curl -sS "http://127.0.0.1:9090/metrics" | grep conduit_forward_errors
See conduit_forward_errors_total (pool, backend, reason = timeout, send_error, etc.).
Upstream timeouts and slow responses
| Symptom |
Likely cause |
What to check |
Answers often SERVFAIL after ~2s (default) |
forward.timeout_ms exceeded |
Default 2000 ms — Reference: forward; increase timeout or fix upstream latency |
Counter reason="timeout" on forward errors |
Upstream slow or down |
conduit_forward_errors_total; test backend with dig @BACKEND_IP -p PORT |
reason="table_full" on forward errors |
Too many in-flight queries to one backend |
Lower load or raise forward.outstanding_per_backend (default 100) — Reference: forward |
| Wrong egress path (multi-homed host) |
Source bind or pool sources_* |
Dual-stack forwarding |
Response-hook retry can run after timeout — Conduit still reaches Response rules so policy can fail over to another pool.
Backend health
For probes, passive fast-trip, eligibility, and operator controls, see Backend health. Health is opt-in (pools[].health.enabled: true).
| Symptom |
Likely cause |
What to check |
| Traffic still hits a dead backend |
Health disabled for the pool (default) |
Enable health: on the pool — Backend health — When to enable |
| All backends marked down; queries still succeed |
Fail-open floor or single-backend pool |
min_eligible; single-backend pools always fail open — Backend health — Route |
| Backend stays down after upstream recovers |
Waiting for probe rise; passive fast-trip cannot mark up |
Wait for rise consecutive successful probes, or conduitctl health set up / health resume — Active probes and passive fast-trip |
| Drained backend returns to rotation unexpectedly |
Scope not frozen, or health resume snapped applied to observed |
conduitctl health show; drain is health set down (implies freeze) — Operator controls |
| Applied health stale after clear/freeze sequence |
Clear-while-frozen footgun |
Prefer atomic conduitctl health resume — Clear-while-frozen |
conduitctl health fails |
No control plane at process start |
Same as conduitctl cannot connect |
| Health gauges missing from scrape |
Health category excluded, metrics disabled, or health not enabled on the pool |
metrics.enabled: true with health in the plan (base: minimal or standard include it) and at least one pool with health — Built-in metrics — Backend health, Metrics configurability |
| High probe load on upstreams |
Low interval_ms × many backends |
Raise interval; size for upstream tolerance — Probe behavior |
conduitctl health show
curl -sS "http://127.0.0.1:9090/metrics" | grep -E 'conduit_backend_health_|conduit_probe_results'
Process logs: active-probe transitions at INFO (backend health transition); passive fast-trip at WARN (passive health: forward failure, passive fast-trip: backend marked down).
Dataplane runtime and concurrency
For runtime models, worker pools, and the transaction slot pool, see Runtime and concurrency and the Dataplane runtime tuning guide. Settings under dataplane:, listeners:, and forward: are start-time — a reload or conduitctl apply updates the stored snapshot but you must restart to apply them on the wire.
Slot pool exhaustion
Slot gauges require the runtime category (base: standard); the exhaustion counter is in failures (both minimal and standard):
curl -sS "http://127.0.0.1:9090/metrics" \
| grep -E '^conduit_(slots_in_use|slots_capacity|slot_pool_exhausted_total)'
Split I/O runtime or worker misconfiguration
Control plane
conduitctl cannot connect
| Symptom |
Likely cause |
What to check |
Connection refused to 127.0.0.1:5199 |
No control: at process start |
Add control.listen_address and restart — gRPC and conduitctl — Enabling the control plane |
| Worked before; fails after config edit |
control: added via reload only — listener not started |
Restart after enabling control — snapshot updates but gRPC binds at startup |
| TLS / protocol errors |
http:// vs https:// mismatch |
Plain TCP → http://; with control.tls → https:// — gRPC and conduitctl — Connecting |
Unauthenticated / RPC denied |
control.api_keys set; missing or wrong key |
--api-key / CONDUIT_API_KEY — API keys |
| mTLS required |
control.tls.client_ca_path set |
Client cert + API key rules — mTLS |
conduitctl validate --file runs offline and does not need the control plane.
Reload or apply fails validation
| Symptom |
Likely cause |
What to check |
conduitctl reload exits non-zero; DNS still works |
Bad on-disk YAML; reload rejected |
Last-good snapshot still active — fix file, validate, reload again — Config file — Startup vs reload |
| SIGHUP sent; no config change |
Same as reload failure |
Process logs for validation/compile errors |
conduitctl apply exits non-zero |
Patch fails validate/compile |
Sparse patch only; no forbidden keys — Configuration model — overlay |
| Startup exits; DNS never worked |
First snapshot never installed |
Fix stderr from conduit or validate --file before retrying start |
After a failed reload, conduitctl export still reflects the running effective config (last good), not the rejected file.
Apply rejected or overlay surprises
| Symptom |
Likely cause |
What to check |
Apply error mentions rules, metrics, or tracing |
Those sections are file-layer only |
Edit startup YAML and reload — not overlay-eligible — Configuration model — overlay |
| Weights reverted unexpectedly |
Reload or SIGHUP clears overlay |
Expected — reload re-reads disk and drops in-memory patches — Control plane workflows |
export differs from file on disk |
Active overlay or normalized defaults |
Compare export to file; use apply --clear to drop overlay without re-reading disk — Reload and export — clear vs reload |
| Patched wrong startup file |
Reload always uses path from process start |
Edit the file Conduit was started with, not only a copy used with validate --file — Config file — Overview |
Listener, forward, or control change had no effect
Observability
Metrics scrape returns connection refused or empty
| Symptom |
Likely cause |
What to check |
curl to /metrics fails with connection refused |
metrics.prometheus not configured, metrics disabled, or Conduit not listening on that address |
Confirm metrics.enabled: true and metrics.prometheus.listen_address in the startup config; Metrics |
Connection works but no conduit_* series |
metrics: omitted or enabled: false |
Built-ins are off when the block is missing — Metrics — Enabling export |
| Scrape works but counters stay at zero |
No DNS traffic yet, or wrong listener port in dig |
Send a query to the configured listener; check conduit_queries_total |
| Changed scrape address or enabled metrics after start — still old behavior |
Export listeners bind at process start |
Restart Conduit after metrics: changes — Observability — Changing observability config |
Smoke test:
curl -sS "http://127.0.0.1:9090/metrics" | head
dig @127.0.0.1 -p 15353 +time=3 example.com A
curl -sS "http://127.0.0.1:9090/metrics" | grep conduit_queries
Adjust host, scrape port, and listener port to your config.
OTEL metrics push failures
| Symptom |
Likely cause |
What to check |
Log line failed to build OTLP metric exporter at startup |
Invalid metrics.otel.endpoint (not http:// or https://) |
Metrics — Export architecture; validate with conduitctl validate --file |
Periodic otel metrics push failed at warn |
Collector down, TLS verify failure, or network block |
Endpoint reachable; for self-signed HTTPS use allow_invalid_certs: true (lab only) or fix collector cert; local smoke: OTLP metrics push smoke with conduit-otlp-metrics-tracer |
| No push logs |
Push interval default 15s; successes log at debug only |
Set logging.level: debug briefly to see otel metrics push ok |
| Enabled OTEL after process start — no push |
OTEL task starts at process start |
Restart after adding or changing metrics.otel |
| Push rejected with 401 / 403 |
Collector requires auth |
Set metrics.otel.headers (for example Authorization: Bearer …) — Metrics — OTEL |
Bind Prometheus scrape to loopback or restrict with firewall — scrape has no built-in auth today.
Event export / dnstap gaps
| Symptom |
Likely cause |
What to check |
| No frames at collector |
Collector not running, wrong socket path, or sink filters exclude the query |
Start conduit-dnstap-tracer before Conduit; see Event export and dnstap |
conduit_events_queue_dropped_total increasing |
Collector slow or down; queue full |
Event export — Overload and metrics; fix collector throughput |
| Added a new sink via reload — no effect |
New sinks require restart |
Event export — Changing events config |
Query frames missing pool/backend on query emit |
Expected — pool/backend filters apply to response / retry only |
Event export — Filters |
Pipeline trace not found
| Symptom |
Likely cause |
What to check |
conduitctl trace N — not found |
Wrong txn_id, trace expired, or activation did not match |
Tracing — Activation; traces TTL 5 minutes, store cap 1000 |
| Control plane unavailable |
No control: at startup |
conduitctl trace needs gRPC — gRPC and conduitctl |
Unsure of txn_id |
Id is per-worker and increments |
Set logging.level: debug, send query, read txn_id from query complete — Logging |
| Enabled tracing after start — no traces |
Tracing compiled at process start |
Restart after tracing: changes — Tracing — Changing tracing config |
Logging surprises
| Symptom |
Likely cause |
What to check |
| No per-query lines at default level |
query complete is debug only |
By design — use Metrics for volume; enable debug briefly for txn_id |
RUST_LOG overrides config |
Env set at startup |
Unset RUST_LOG in production — Logging — RUST_LOG override |
Changed logging.level via reload — no effect |
Subscriber binds at process start |
Restart after logging changes — Logging — Changing logging config |
- Getting started — First query — end-to-end lab and basic
dig failures
- Guide: Backend health — probes, drain, and resume lab
- Control plane workflows — reload, apply, export, and restart
- Config file — validation, startup path, startup vs reload
- Configuration model — overlay, last-good snapshot, pending reconcile
- Runtime and concurrency — runtime models, worker pools, slot pool
- Dataplane runtime tuning —
sync vs split_io, worker sizing, slot pool
- Observability — which signal to use, OTEL naming, reload matrix
- Metrics and tracing — metrics + tracing lab
- Operator metrics bases —
minimal vs standard scrape comparison
- Event export and dnstap — dnstap lab