ADR-007 — Per-Host Circuit Breaker
Context and Problem Statement
ADR-001 mandated a circuit breaker at the HttpClient layer to prevent
cascading failures when a target host becomes unavailable or starts
rejecting connections. CircuitOpenError was reserved as a placeholder
but never raised. Without a circuit breaker HttpClient hammers a
failing host indefinitely, consuming retry budget and delaying detection
of systematic outages.
Decision Drivers
- Prevent thundering-herd retries against a host that is already down.
- Keep per-host state so one failing domain does not affect others.
- Remain disabled by default to avoid breaking existing callers.
- Fit cleanly into the existing sync
_request()pipeline.
Considered Options
- A — Count-based per-host state machine (chosen)
- B — Rate-based (failure ratio over a sliding window)
- C — External circuit breaker (e.g.
pybreakerlibrary)
Decision Outcome
Chosen: Option A.
A count-based, per-host state machine is simple, deterministic, and requires no additional dependencies.
State machine
CLOSED ─(failures >= threshold)─► OPEN ─(recovery elapsed)─► HALF_OPEN
▲ ▲ │
└───────── success ────────────────┼──────────────────────────────┘
└──────── failure ─────────────┘
- CLOSED — normal operation; consecutive failure counter tracked.
- OPEN — all requests to this host blocked with
CircuitOpenError; timer running towardrecovery_seconds. - HALF_OPEN — one probe allowed; success → CLOSED, failure → OPEN.
Implementation
- New class
CircuitBreakerinladon.networking.circuit_breaker. HttpClientConfigfields (both off by default):circuit_breaker_failure_threshold: int | None = Nonecircuit_breaker_recovery_seconds: float = 60.0HttpClient._circuit_breakers: dict[str, CircuitBreaker]keyed bynetloc; created lazily on first request to a host.- Check at start of
_request()before_enforce_rate_limit(); raiseCircuitOpenErrorif blocked. - Record success/failure on every
_request()completion.
Async concurrency amendment (2026-08-02, Issue #160)
The original caller-enforced single-probe assumption is insufficient for
AsyncHttpClient and AsyncCurlHttpClient: several tasks can be admitted
before any one of them records an outcome. The shared async policy therefore
uses an event-loop-local admission guard with these semantics:
- CLOSED requests remain concurrent. Outcome updates are synchronous between await points, so failure counts cannot be partially mutated.
- The first request after OPEN recovery reserves the HALF_OPEN probe. Other
requests to that host receive
CircuitOpenErroruntil the probe records an outcome or is cancelled. - Every admission carries a circuit generation. A slow CLOSED request from an earlier generation cannot close or re-open a newer HALF_OPEN probe.
- Cancellation releases a HALF_OPEN reservation without counting as host success or failure.
Async rate limiting uses a separate per-host asyncio.Lock. A caller waits
and commits its next eligible start time while holding that host's lock, so
concurrent callers cannot wake as a batch. Locks are not shared across hosts,
and all event-loop-bound guard state is discarded when the client closes.
Async client instances remain single-event-loop objects and must not be shared
across threads or loops.
HTTP 5xx health-accounting amendment (2026-08-02, Issue #165)
HTTP responses with status_code >= 500 count as circuit-breaker failures;
responses below 500 that are not configured in retry_on_status, including
ordinary 4xx responses, count as successes. This origin-health accounting is
independent of ADR-002's result contract: a 5xx response can remain Ok(...)
for caller-owned status interpretation while still contributing one failure
for the logical call sequence. Statuses handled by retry_on_status, including
4xx statuses such as 429, record one failure only after retries are exhausted.
Polite retry pacing amendment (2026-08-12, Issue #167)
The default exponential-backoff base is 0.5 seconds so a retryable response
without a Retry-After header cannot immediately re-fire under default
configuration. Setting backoff_base_seconds=0.0 remains supported as an
explicit opt-out and emits UserWarning when retries are enabled.
Every retry attempt re-evaluates its host's politeness interval, including any
robots.txt Crawl-delay override. Retry-After or exponential backoff and the
remaining per-host interval are merged by taking their maximum, producing one
sleep between attempts. Async clients perform this merged wait and timestamp
reservation while holding the existing per-host lock, preserving staggered
starts for concurrent callers.
Consequences
- Good: cascading failures to a dead host are cut off quickly.
- Good: per-host tracking means unrelated domains are unaffected.
- Good: disabled by default — zero behaviour change for existing callers.
- Neutral:
thresholdcounts call sequences (oneHttpClient.get()invocation), not individual HTTP attempts. Withretries=2andthreshold=3, the circuit opens after 3 exhausted call sequences (up to 9 underlying HTTP attempts). Operators should sizethresholdwith this in mind — a value of 3 is more tolerant than it first appears. During HALF_OPEN, the single probe may itself involve up toretries + 1raw HTTP attempts; observers watching traffic may see more than one request. - Bad: no persistence across client instances — circuit state resets on every new client construction.
Rejected options
B (rate-based): Requires a sliding window, timestamps per request, and is harder to reason about in tests. Count-based state plus explicit async admission generations provides deterministic semantics without that extra model.
C (pybreaker): Adds a runtime dependency for functionality that is straightforward to implement cleanly. Keeping it in-house means the behaviour is explicit, testable, and auditable.