Retries and transactions
This page explains how Conduit handles retries — sending the same client transaction through Lookup again after an upstream answer or timeout — and the global limits that stop further attempts. For declarative actions, see Rules and actions. For the full query path, see Architecture and packet path.
Overview
A transaction is everything Conduit remembers for one client query from Receive through Send or drop. Retries reuse that same transaction: the original question, client address, tags, and request-side pool choice stay in place unless policy changes them on a later hook.
Only Response rules (built-in actions or Rhai on the response hook) can trigger a retry. Request rules run once at the start of the transaction; they do not run again when Conduit re-enters at Lookup.
When policy requests a retry, Conduit jumps from Response rules back to Lookup, runs the full provider chain again (subject to cache eligibility), picks an eligible backend in the target pool (see Backend selection on retries), and forwards again when the forward provider runs. When limits are reached, every eligible backend in the target pool was already tried, or policy accepts the outcome, Conduit continues to Send and replies to the client.
stateDiagram-v2
[*] --> Lookup: first attempt
Lookup --> ResponseRules
ResponseRules --> Send: accept / no retry
ResponseRules --> Lookup: retry allowed
ResponseRules --> Drop: drop
Send --> Reply: to client
Requesting a retry
Retry intent comes from a matching response rule or Rhai script on the response hook:
| Mechanism | Pool for the next attempt |
|---|---|
retry or retry_now action (response) |
Uses selected_pool on the next retry Lookup — the pool from the last forward attempt on this transaction (see Pool selection lifecycle) |
set_retry_pool + retry or retry_now (either hook for set_retry_pool; response for retry) |
Uses retry_pool on the next retry Lookup if retry occurs |
txn.request_retry() or txn.request_retry_now() in Rhai (response) |
Same as retry / retry_now — stay in the current pool |
txn.set_retry_pool("name") in Rhai |
Pool for retry Lookup if retry occurs; first forward ignores (both hooks) — pair with txn.request_retry() on the response hook to fail over |
set_retry_source_v4 / set_retry_source_v6 + retry (request or response for source; response for retry) |
One-shot egress bind on the next retry forward only — see Source selection lifecycle |
txn.set_retry_source_v4(addr) / txn.set_retry_source_v6(addr) in Rhai |
Same as set_retry_source_* — does not trigger retry; pair with txn.request_retry() on the response hook |
At forward route inside Lookup, when attempt_count > 0 (retry re-entry), Conduit uses retry_pool if set (then clears it), then falls back to selected_pool, then the default pool. On the first forward (attempt_count == 0), retry_pool is ignored. Full lifecycle: Pool selection lifecycle.
Response rules run after an upstream answer or after a forward timeout (still with no stored answer). That lets you retry on SERVFAIL, NXDOMAIN, slow upstreams, and other conditions you express with selectors such as rcode.
Declarative examples
Fail over to another pool on SERVFAIL:
orchestrator:
max_attempts: 3
max_txn_duration_ms: 5000
pools:
- name: primary
backends:
- address: "10.0.0.1:53"
- name: secondary
backends:
- address: "10.0.0.2:53"
rules:
match_mode: first_match
rules:
- name: use-primary
hook: request
selectors:
- type: qname_suffix
value: ".example."
actions:
- type: set_pool
value: primary
- name: servfail-retry-other-pool
hook: response
selectors:
- type: rcode
value: SERVFAIL
actions:
- type: set_retry_pool
value: secondary
- type: retry
Retry within the same pool (try another backend in that pool):
pools:
- name: primary
backends:
- address: "10.0.0.1:53"
- address: "10.0.0.2:53"
- address: "10.0.0.3:53"
rules:
match_mode: first_match
rules:
- name: servfail-retry-same-pool
hook: response
selectors:
- type: rcode
value: SERVFAIL
actions:
- type: retry
On SERVFAIL, Conduit re-enters at Lookup. With set_retry_pool + retry, the next forward uses that pool’s backends. With retry alone, Conduit keeps the pool from the first attempt and selects a different backend there when more than one is configured.
Rhai
On the response hook:
txn.request_retry()— soft retry in the current pool (same asretryin YAML).txn.request_retry_now()— hard retry in the current pool (same asretry_nowin YAML).txn.set_retry_pool("pool-name")— pool for retry Lookup if retry occurs; first forward Route ignores. Addtxn.request_retry()ortxn.request_retry_now()to trigger failover.
See Transaction API — Outcomes (request_retry, request_retry_now) and Routing (set_retry_pool).
What happens on each attempt
When the forward provider reaches Route inside Lookup:
- Pool selection — see Pool selection lifecycle below.
- Backend selection — see Backend selection on retries.
- Attempt counter — Conduit increments the attempt count for this transaction before Forward.
conduit_queries_by_pool_total increments for each attempt that reaches Route → Forward inside Lookup, including retries.
Tags set on request rules or earlier response rules persist across retries unless a script clears them. You can branch later rules on tags (for example “already retried”).
Pool selection lifecycle
Each transaction carries two pool-related fields set by policy:
| Field | Set by | Role |
|---|---|---|
selected_pool |
set_pool / txn.set_pool, request rules, or default pool |
Primary pool for routing |
retry_pool |
set_retry_pool / txn.set_retry_pool |
Optional one-shot override for the next retry Route only |
At each Route, Conduit resolves the pool name, picks a backend, then updates selected_pool to the pool that attempt actually used. That update matters on later retries: after a cross-pool failover, selected_pool reflects the failover pool even though request rules originally chose another name.
First attempt (attempt_count == 0 before Route):
- Uses
selected_pool(from request policy or default).retry_poolis ignored, even if stashed on the request hook.
Retry re-entry (attempt_count > 0):
- If
retry_poolis set → use that name for this attempt, then clear it (consumed). - Else → use
selected_pool(the pool from the last successful Route on this transaction). - Else → default pool.
So retry_pool is not a standing “retry target” for every subsequent attempt. You do not need to call set_retry_pool again on each response pass just to stay on a pool you already failed over to — after the one-shot stash is consumed, later retries follow selected_pool, which Route already updated to that pool.
When to set retry_pool again:
- You want a different pool on the next retry than
selected_poolwould give (for example tertiary after secondary). - You called
clear_retry_pool/txn.clear_retry_pool()and later need a new override. - Response policy sets a fresh stash on each response pass that requests retry (unusual; only when each retry should honor a newly chosen override).
Worked example — request rule sets set_pool: primary and set_retry_pool: secondary, response rule retries on SERVFAIL:
| Step | attempt_count before Route |
retry_pool |
selected_pool before Route |
Pool used |
|---|---|---|---|---|
| First Route | 0 | secondary (ignored) | primary | primary |
| Response: retry | — | secondary (unchanged) | primary | — |
| Second Route | 1 | secondary → cleared | primary | secondary |
Response: retry again (no new set_retry_pool) |
— | — | secondary (updated at Route) | — |
| Third Route | 2 | — | secondary | secondary |
Further retries in secondary use another backend there when available; they do not snap back to primary unless policy changes selected_pool or stashes a new retry_pool.
Rhai reference: Transaction API — Routing. Pipeline detail: Architecture — Route.
Source selection lifecycle
Egress bind IP is separate from pool choice. Each transaction carries up to four egress-related fields:
| Field | Set by | Role |
|---|---|---|
source_override_v4 / source_override_v6 |
set_source_v4 / set_source_v6 or Rhai txn.set_source_* (request hook only) |
Standing local bind for every forward attempt (unless retry source wins on that attempt) |
retry_source_override_v4 / retry_source_override_v6 |
set_retry_source_* / txn.set_retry_source_* (request or response hook) |
Optional one-shot bind for the next retry forward only |
At each Forward, after attempt_count is incremented for this attempt:
First forward (attempt_count == 1 at Forward):
retry_source_override_*is ignored, even if stashed on the request or an earlier response pass.- Uses standing
source_override_*if set, else pool/global round-robin among configured sources.
Retry forward (attempt_count > 1 at Forward):
- If
retry_source_override_*is set for the backend’s address family → use it once, then clear it (consumed). - Else → standing
source_override_*if set. - Else → pool/global round-robin.
set_retry_source_* does not trigger retry and does not affect the first upstream attempt. Pair with retry / retry_now or Rhai txn.request_retry() when outcome-driven failover should use a different bind IP. clear_retry_source_* removes the stash without clearing standing set_source_* overrides.
Worked example — request rule sets set_source_v4: 127.0.0.1 and set_retry_source_v4: 10.0.0.5, response rule retries on SERVFAIL:
| Step | attempt_count at Forward |
retry_source_override_v4 |
Bind used (v4 backend) |
|---|---|---|---|
| First forward | 1 | 10.0.0.5 (stashed) | 127.0.0.1 (standing; stash ignored) |
| Response: retry | — | — | — |
| Second forward | 2 | consumed → cleared | 10.0.0.5 (one-shot retry source) |
| Third forward (if any) | 3 | — | 127.0.0.1 (standing again) |
Allowed-set enforcement at Forward is unchanged. See Dual-stack forwarding and Transaction API — Egress.
Backend selection on retries
| Attempt | Behavior |
|---|---|
First (attempt_count 0 before Route) |
Sticky weighted pick among eligible backends in the pool — same as normal Pools and backends load balancing |
Retry (attempt_count > 0) |
Weighted pick among eligible backends in the target pool that were not already used for that pool on this transaction |
Eligible means every configured backend when pool health is off, and backends whose applied health is up when health is on (plus any fail-open treatment). Retries never prefer a backend that Route would skip on the first attempt.
On a cross-pool retry, only backends tried in the target pool are excluded — backends used in other pools do not count.
When every eligible backend in the target pool was already tried (or none are eligible), Route cannot select another backend. Conduit sets SERVFAIL and moves to Send (pool exhausted for this transaction).
A pool with only one backend cannot offer an alternate target on retry; the next retry attempt hits pool exhaustion immediately after the first forward fails.
Global limits (orchestrator)
The top-level orchestrator: block caps how long a transaction may loop and how many Lookup forward attempts are allowed. Field reference: Reference: orchestrator. When omitted, Conduit uses the same defaults as in Minimal configuration.
| Field | Default | Meaning |
|---|---|---|
max_attempts |
3 | Maximum forward-provider Route → Forward cycles inside Lookup for one client query |
max_txn_duration_ms |
5000 | Wall-clock limit for the whole transaction from start to Send or drop |
txn_table_capacity |
1024 | Capacity for tracking in-flight transactions on the dataplane (not per-query retry count) |
Conduit checks max_txn_duration_ms and max_attempts before each forward attempt inside Lookup. When either limit is exceeded, Conduit sets SERVFAIL on the transaction and moves to Send instead of forwarding again.
A retry stops when any of these occurs:
max_attemptsreachedmax_txn_duration_msexceeded- Pool exhausted — no unused eligible backend left in the target pool for this transaction
- Policy accepts the answer (no retry intent on Response rules)
Validation: max_attempts must be ≥ 1.
What the client sees when limits hit
When retries are exhausted, the pool is exhausted, or the transaction runs too long, the client receives a synthesized SERVFAIL (unless an upstream wire answer was already stored and policy sends the pipeline to Send without another retry). Synthesized errors echo the question section from the original query. Details: Send in Architecture and packet path.
You can adjust response metadata before Send with the set_rcode action on response rules when policy accepts the answer instead of retrying.
Observability
| Signal | When |
|---|---|
conduit_retries_total{pool} |
Response rules send the pipeline back to Lookup; pool is the target pool for the next attempt |
conduit_queries_by_pool_total{pool} |
Each attempt that reaches Forward, including retries |
Event export retry frames |
When sinks are configured with retry emission — see Event export |
Related topics
- Rules and actions —
retry,retry_now,set_retry_pool,set_rcode, response selectors - Declarative failover — end-to-end SERVFAIL / timeout failover lab
- Rule action order — soft vs hard drop/retry; request stash on first Route
- Pools and backends — pool names, weights, default pool
- Backend health — eligibility and fail-open at Route
- Architecture and packet path — Response rules, Send, timeouts
- Rhai — Transaction API (Routing) —
txn.set_pool,txn.clear_pool,txn.set_retry_pool,txn.clear_retry_pool - Rhai — Transaction API (Egress) —
txn.set_source_*,txn.set_retry_source_*,txn.clear_retry_source_* - Built-in metrics — counters, profiles, and pipeline mapping