Skip to content

Drain policy under slow upstream

How do complete, budgeted, and minimal drain policies behave under forward_slow load at stop?

Numbers are same-host comparisons on a single reference host and are not service-level objectives. See the performance hub disclaimer.

When this matters

Shutdown and restart windows trade how long Conduit waits for in-flight work against client failures during stop. See graceful drain on shutdown. Slow upstreams keep transactions outstanding longer, so shutdown.drain policy choice shows up clearly.

What we varied

  • Varied: drain policy (drain_complete, drain_budgeted, drain_minimal)
  • Held constant: forward_slow load overlapping SIGTERM, same lab recipe
  • Read the timing columns, not the throughput columns: these cells keep the thin load recipe on purpose, because the subject is what happens across the stop window. Their QPS and latency are incidental and are not a throughput ranking.

Evidence

Drain duration under forward_slow

Drain duration under forward_slow

Download CSV

Drain policy Drain duration (ms) Client failures during stop QPS Avg latency (ms)
drain_complete 113.5 299 6.5 996.6
drain_budgeted 113.6 200 10.0 909.2
drain_minimal 63.7 200 9.7 902.1

At a glance

  • Drain duration under forward_slow: drain_complete ≈ 114 ms, drain_budgeted ≈ 114 ms, drain_minimal ≈ 64 ms

Takeaway

Minimal drain stops fastest; complete leaves the most in-flight clients to fail. On this lab (median of three rounds), minimal finishes in ~64 ms; complete and budgeted both take ~114 ms. Complete records the most client failures during stop (299 vs 200 for the others) — a longer wait gives doomed slow-upstream requests more time to time out.

What to do: choose complete, budgeted, or minimal for your upgrade/restart window from your failure budget — pick a policy on purpose. Vocabulary: Performance methodology — Drain policy.

Member scenarios