Skip to content

Retries and transactions

This page explains how Conduit handles retries — sending the same client transaction through Lookup again after an upstream answer or timeout — and the global limits that stop further attempts. For declarative actions, see Rules and actions. For the full query path, see Architecture and packet path.

Overview

A transaction is everything Conduit remembers for one client query from Receive through Send or drop. Retries reuse that same transaction: the original question, client address, tags, and request-side pool choice stay in place unless policy changes them on a later hook.

Only Response rules (built-in actions or Rhai on the response hook) can trigger a retry. Request rules run once at the start of the transaction; they do not run again when Conduit re-enters at Lookup.

When policy requests a retry, Conduit jumps from Response rules back to Lookup, runs the full provider chain again (subject to cache eligibility), picks an eligible backend in the target pool (see Backend selection on retries), and forwards again when the forward provider runs. When limits are reached, every eligible backend in the target pool was already tried, or policy accepts the outcome, Conduit continues to Send and replies to the client.

stateDiagram-v2
  [*] --> Lookup: first attempt
  Lookup --> ResponseRules
  ResponseRules --> Send: accept / no retry
  ResponseRules --> Lookup: retry allowed
  ResponseRules --> Drop: drop
  Send --> Reply: to client

Requesting a retry

Retry intent comes from a matching response rule or Rhai script on the response hook:

Mechanism Pool for the next attempt
retry or retry_now action (response) Uses selected_pool on the next retry Lookup — the pool from the last forward attempt on this transaction (see Pool selection lifecycle)
set_retry_pool + retry or retry_now (either hook for set_retry_pool; response for retry) Uses retry_pool on the next retry Lookup if retry occurs
txn.request_retry() or txn.request_retry_now() in Rhai (response) Same as retry / retry_now — stay in the current pool
txn.set_retry_pool("name") in Rhai Pool for retry Lookup if retry occurs; first forward ignores (both hooks) — pair with txn.request_retry() on the response hook to fail over
set_retry_source_v4 / set_retry_source_v6 + retry (request or response for source; response for retry) One-shot egress bind on the next retry forward only — see Source selection lifecycle
txn.set_retry_source_v4(addr) / txn.set_retry_source_v6(addr) in Rhai Same as set_retry_source_* — does not trigger retry; pair with txn.request_retry() on the response hook

At forward route inside Lookup, when attempt_count > 0 (retry re-entry), Conduit uses retry_pool if set (then clears it), then falls back to selected_pool, then the default pool. On the first forward (attempt_count == 0), retry_pool is ignored. Full lifecycle: Pool selection lifecycle.

Response rules run after an upstream answer or after a forward timeout (still with no stored answer). That lets you retry on SERVFAIL, NXDOMAIN, slow upstreams, and other conditions you express with selectors such as rcode.

Declarative examples

Fail over to another pool on SERVFAIL:

orchestrator:
  max_attempts: 3
  max_txn_duration_ms: 5000

pools:
  - name: primary
    backends:
      - address: "10.0.0.1:53"
  - name: secondary
    backends:
      - address: "10.0.0.2:53"

rules:
  match_mode: first_match
  rules:
    - name: use-primary
      hook: request
      selectors:
        - type: qname_suffix
          value: ".example."
      actions:
        - type: set_pool
          value: primary
    - name: servfail-retry-other-pool
      hook: response
      selectors:
        - type: rcode
          value: SERVFAIL
      actions:
        - type: set_retry_pool
          value: secondary
        - type: retry

Retry within the same pool (try another backend in that pool):

pools:
  - name: primary
    backends:
      - address: "10.0.0.1:53"
      - address: "10.0.0.2:53"
      - address: "10.0.0.3:53"

rules:
  match_mode: first_match
  rules:
    - name: servfail-retry-same-pool
      hook: response
      selectors:
        - type: rcode
          value: SERVFAIL
      actions:
        - type: retry

On SERVFAIL, Conduit re-enters at Lookup. With set_retry_pool + retry, the next forward uses that pool’s backends. With retry alone, Conduit keeps the pool from the first attempt and selects a different backend there when more than one is configured.

Rhai

On the response hook:

  • txn.request_retry() — soft retry in the current pool (same as retry in YAML).
  • txn.request_retry_now() — hard retry in the current pool (same as retry_now in YAML).
  • txn.set_retry_pool("pool-name") — pool for retry Lookup if retry occurs; first forward Route ignores. Add txn.request_retry() or txn.request_retry_now() to trigger failover.

See Transaction API — Outcomes (request_retry, request_retry_now) and Routing (set_retry_pool).

What happens on each attempt

When the forward provider reaches Route inside Lookup:

  1. Pool selection — see Pool selection lifecycle below.
  2. Backend selection — see Backend selection on retries.
  3. Attempt counter — Conduit increments the attempt count for this transaction before Forward.

conduit_queries_by_pool_total increments for each attempt that reaches Route → Forward inside Lookup, including retries.

Tags set on request rules or earlier response rules persist across retries unless a script clears them. You can branch later rules on tags (for example “already retried”).

Pool selection lifecycle

Each transaction carries two pool-related fields set by policy:

Field Set by Role
selected_pool set_pool / txn.set_pool, request rules, or default pool Primary pool for routing
retry_pool set_retry_pool / txn.set_retry_pool Optional one-shot override for the next retry Route only

At each Route, Conduit resolves the pool name, picks a backend, then updates selected_pool to the pool that attempt actually used. That update matters on later retries: after a cross-pool failover, selected_pool reflects the failover pool even though request rules originally chose another name.

First attempt (attempt_count == 0 before Route):

  • Uses selected_pool (from request policy or default). retry_pool is ignored, even if stashed on the request hook.

Retry re-entry (attempt_count > 0):

  • If retry_pool is set → use that name for this attempt, then clear it (consumed).
  • Else → use selected_pool (the pool from the last successful Route on this transaction).
  • Else → default pool.

So retry_pool is not a standing “retry target” for every subsequent attempt. You do not need to call set_retry_pool again on each response pass just to stay on a pool you already failed over to — after the one-shot stash is consumed, later retries follow selected_pool, which Route already updated to that pool.

When to set retry_pool again:

  • You want a different pool on the next retry than selected_pool would give (for example tertiary after secondary).
  • You called clear_retry_pool / txn.clear_retry_pool() and later need a new override.
  • Response policy sets a fresh stash on each response pass that requests retry (unusual; only when each retry should honor a newly chosen override).

Worked example — request rule sets set_pool: primary and set_retry_pool: secondary, response rule retries on SERVFAIL:

Step attempt_count before Route retry_pool selected_pool before Route Pool used
First Route 0 secondary (ignored) primary primary
Response: retry — secondary (unchanged) primary —
Second Route 1 secondary → cleared primary secondary
Response: retry again (no new set_retry_pool) — — secondary (updated at Route) —
Third Route 2 — secondary secondary

Further retries in secondary use another backend there when available; they do not snap back to primary unless policy changes selected_pool or stashes a new retry_pool.

Rhai reference: Transaction API — Routing. Pipeline detail: Architecture — Route.

Source selection lifecycle

Egress bind IP is separate from pool choice. Each transaction carries up to four egress-related fields:

Field Set by Role
source_override_v4 / source_override_v6 set_source_v4 / set_source_v6 or Rhai txn.set_source_* (request hook only) Standing local bind for every forward attempt (unless retry source wins on that attempt)
retry_source_override_v4 / retry_source_override_v6 set_retry_source_* / txn.set_retry_source_* (request or response hook) Optional one-shot bind for the next retry forward only

At each Forward, after attempt_count is incremented for this attempt:

First forward (attempt_count == 1 at Forward):

  • retry_source_override_* is ignored, even if stashed on the request or an earlier response pass.
  • Uses standing source_override_* if set, else pool/global round-robin among configured sources.

Retry forward (attempt_count > 1 at Forward):

  • If retry_source_override_* is set for the backend’s address family → use it once, then clear it (consumed).
  • Else → standing source_override_* if set.
  • Else → pool/global round-robin.

set_retry_source_* does not trigger retry and does not affect the first upstream attempt. Pair with retry / retry_now or Rhai txn.request_retry() when outcome-driven failover should use a different bind IP. clear_retry_source_* removes the stash without clearing standing set_source_* overrides.

Worked example — request rule sets set_source_v4: 127.0.0.1 and set_retry_source_v4: 10.0.0.5, response rule retries on SERVFAIL:

Step attempt_count at Forward retry_source_override_v4 Bind used (v4 backend)
First forward 1 10.0.0.5 (stashed) 127.0.0.1 (standing; stash ignored)
Response: retry — — —
Second forward 2 consumed → cleared 10.0.0.5 (one-shot retry source)
Third forward (if any) 3 — 127.0.0.1 (standing again)

Allowed-set enforcement at Forward is unchanged. See Dual-stack forwarding and Transaction API — Egress.

Backend selection on retries

Attempt Behavior
First (attempt_count 0 before Route) Sticky weighted pick among eligible backends in the pool — same as normal Pools and backends load balancing
Retry (attempt_count > 0) Weighted pick among eligible backends in the target pool that were not already used for that pool on this transaction

Eligible means every configured backend when pool health is off, and backends whose applied health is up when health is on (plus any fail-open treatment). Retries never prefer a backend that Route would skip on the first attempt.

On a cross-pool retry, only backends tried in the target pool are excluded — backends used in other pools do not count.

When every eligible backend in the target pool was already tried (or none are eligible), Route cannot select another backend. Conduit sets SERVFAIL and moves to Send (pool exhausted for this transaction).

A pool with only one backend cannot offer an alternate target on retry; the next retry attempt hits pool exhaustion immediately after the first forward fails.

Global limits (orchestrator)

The top-level orchestrator: block caps how long a transaction may loop and how many Lookup forward attempts are allowed. Field reference: Reference: orchestrator. When omitted, Conduit uses the same defaults as in Minimal configuration.

Field Default Meaning
max_attempts 3 Maximum forward-provider Route → Forward cycles inside Lookup for one client query
max_txn_duration_ms 5000 Wall-clock limit for the whole transaction from start to Send or drop
txn_table_capacity 1024 Capacity for tracking in-flight transactions on the dataplane (not per-query retry count)

Conduit checks max_txn_duration_ms and max_attempts before each forward attempt inside Lookup. When either limit is exceeded, Conduit sets SERVFAIL on the transaction and moves to Send instead of forwarding again.

A retry stops when any of these occurs:

  • max_attempts reached
  • max_txn_duration_ms exceeded
  • Pool exhausted — no unused eligible backend left in the target pool for this transaction
  • Policy accepts the answer (no retry intent on Response rules)

Validation: max_attempts must be ≥ 1.

What the client sees when limits hit

When retries are exhausted, the pool is exhausted, or the transaction runs too long, the client receives a synthesized SERVFAIL (unless an upstream wire answer was already stored and policy sends the pipeline to Send without another retry). Synthesized errors echo the question section from the original query. Details: Send in Architecture and packet path.

You can adjust response metadata before Send with the set_rcode action on response rules when policy accepts the answer instead of retrying.

Observability

Signal When
conduit_retries_total{pool} Response rules send the pipeline back to Lookup; pool is the target pool for the next attempt
conduit_queries_by_pool_total{pool} Each attempt that reaches Forward, including retries
Event export retry frames When sinks are configured with retry emission — see Event export