Skip to content

Tracing

Tracing records optional per-query pipeline traces for the dataplane: which phases ran, cumulative elapsed time, and selected pool/backend at each step. Tracing is off by default — when disabled, Conduit does not allocate trace buffers on the query path.

Use tracing when you need to debug routing, forwarding, or retries on specific queries. For aggregate volume and latency, use Metrics. For wire-level query/response export to a collector, use Event export.

Pipeline tracing is not OTEL traces

The tracing: config block controls in-process pipeline traces stored for GetTrace / conduitctl trace and optional JSON log output. It is not OpenTelemetry distributed trace export over OTLP (not implemented). OTLP metrics live under metrics.otel — see Metrics.

Enabling tracing

Add a tracing: block with enabled: true. The control plane must be running if you want to fetch traces with conduitctl trace or gRPC GetTrace.

control:
  listen_address: "127.0.0.1:5199"
tracing:
  enabled: true
  activation:
    tag: trace
    selectors:
      - type: qtype
        value: A
    sample_percent: 100
  output:
    log_json: false
SettingMeaning
tracing.enabledMust be true for activation and recording
tracing.activationOptional filters — which transactions get a trace (see Activation)
tracing.output.log_jsonWhen true, emit completed traces as JSON on the process log at info (target: conduit::trace)

Field reference: Config schema: metrics and tracing.

How tracing fits the query path

sequenceDiagram
  participant C as Client
  participant L as Listener worker
  participant O as Orchestrator
  participant S as TraceStore

  C->>L: DNS query
  L->>O: pipeline phases
  Note over O: After Request rules:<br/>activation check
  O->>O: record phase events<br/>(if activated)
  O->>C: DNS response
  O->>S: store trace by txn_id
  1. The query runs through the defined pipeline (Parse through Send).
  2. After Request rules complete, Conduit evaluates activation once per transaction. Matching queries allocate an in-memory trace buffer.
  3. Each completed pipeline phase appends a trace event (phase name, elapsed microseconds since transaction start, optional pool/backend).
  4. At transaction completion, the trace is inserted into a bounded in-memory TraceStore (1000 entries, 5 minute TTL). Optionally, log_json writes the same events to stderr/stdout.

Non-matching queries pay no trace allocation cost.

Activation

Activation decides which transactions receive a pipeline trace. All configured clauses must pass (logical AND). Evaluated after Request rules, so tags set in request policy are visible to activation.

FieldMeaning
activation.tagTransaction must have the named tag key (same semantics as tag_required on event export sinks)
activation.selectorsSelector list — same types as rules: qname_suffix, qname_exact, qtype, rcode, tag. All must match
activation.sample_percentFloat in [0, 100]; deterministic sampling (default 100). Same algorithm as event-export
activation.sample_keyOptional static salt for sample_percent (mutually exclusive with sample_key_from)
activation.sample_key_fromOptional qname — per-query-name salt for sample_percent

When activation is omitted, every transaction matches (subject to sample_percent default 100).

Example — trace only A queries that carry a debug tag:

tracing:
  enabled: true
  activation:
    tag: debug
    selectors:
      - type: qtype
        value: A
    sample_percent: 25

Tag the transaction in Request rules (or Rhai) before activation runs:

rules:
  match_mode: first_match
  rules:
    - name: tag-debug
      hook: request
      selectors:
        - type: qname_suffix
          value: "lab.example."
      actions:
        - type: set_tag
          value: debug=1

Selectors at activation time

Activation runs after Request rules, before Lookup. Selectors that depend on rcode or final pool/backend usually will not match at activation — prefer qname, qtype, and tags for trace gating.

Trace events

Each event in a stored trace has:

FieldMeaning
phaseTop-level pipeline phase — for example parse, request_rules, lookup, response_rules, send. Route, forward, and wait appear as nested events inside lookup when the forward provider runs; cache provider outcomes are nested the same way
elapsed_usMicroseconds since the transaction started (cumulative, not per-phase delta)
poolSelected pool at that phase, when applicable
backendSelected backend at that phase, when applicable — configured backend name when set, else the ip:port address (same name-when-set identity as metrics/logs/events)
cacheNamed cache instance on nested cache provider events (provider cache answered / miss / bypass)
messageOptional detail string (for example nested provider outcome text)

Retries re-enter at Lookup; you will see additional lookup (with nested forward events when applicable) and response_rules events on the same transaction trace.

Fetching traces

Traces are keyed by internal transaction id (txn_id). In lab setups the first query is often 1 (incrementing per worker). With logging.level: debug, Conduit also logs txn_id on each query complete line — use that value with the CLI or gRPC.

conduitctl trace

Requires a running control plane and matching control.listen_address (or CONDUIT_CONTROL):

conduitctl trace 1

Prints one line per event: phase, elapsed microseconds, pool, backend, cache, message. Exits non-zero if no trace was found (wrong id, TTL expired, or activation did not match).

Details: gRPC and conduitctl — trace.

gRPC GetTrace

grpcurl -plaintext \
  -d '{"txn_id":"1"}' \
  127.0.0.1:5199 \
  conduit.v1.ConduitControl/GetTrace

Returns found and an events array with the same fields as above. Traces expire after 5 minutes or when the store exceeds 1000 entries (oldest evicted).

JSON log output

When output.log_json: true, Conduit logs the completed trace as JSON at info with log target conduit::trace:

INFO conduit::trace: pipeline trace txn_id=42 events=[{"phase":"parse","elapsed_us":12,...}, ...]

Use this for ad hoc debugging or shipping traces to a log aggregator. It is separate from logging.level and from OTLP log export (not implemented).

Cost and when to enable

State Hot-path cost
tracing: omitted or enabled: false None — no trace buffer
Enabled, activation does not match Activation check only
Enabled, trace captured Per-phase append + store insert at completion

Keep tracing disabled in production unless you need it. Use activation (tag, selectors, sample_percent) to limit volume — same spirit as event export filters.

Built-in phase histograms (conduit_phase_duration_seconds, full profile) provide aggregate timing without per-query storage.

Changing tracing config

The tracing: block lives in the file layer only — overlay patches that include tracing are rejected. Edit the file on disk, then reload or send SIGHUP so validation and the snapshot reflect the change.

Tracing activation rules and log_json are compiled into the process at startup from the initial config. Turning tracing on or off, changing activation, or toggling log_json requires a process restart after updating the file (same limitation as metrics export listeners). See Configuration model — What takes effect when.

Lab smoke test

  1. Start an upstream resolver on 127.0.0.1:5300 (or adjust pool backend in your config).
  2. Start Conduit with control: and tracing.enabled: true (minimal example in Enabling tracing).
  3. Send a matching query: dig @127.0.0.1 -p 15353 +time=3 test.example.com A.
  4. Run conduitctl trace 1 (or GetTrace via grpcurl). If unsure of the id, set logging.level: debug and read txn_id from a query complete line.

Expect top-level phase lookup (with nested route/forward events when upstream runs) and send.