provider_failover¶
Circuit breaker for exhausted LLM provider credit, plus model fallback-chain construction.
The problem¶
An LLM provider can return a permanent credit or billing error (for example, "credit balance is too low"). If you keep calling that provider, every call wastes a round trip. You want to skip it and use a fallback model instead, automatically.
How this bite helps¶
It keeps a process-level registry of exhausted providers. When a provider is marked exhausted, every subsequent model build for that provider skips it immediately and promotes the first available fallback. model_with_fallbacks builds a LangChain RunnableWithFallbacks with the circuit-breaker guard baked in.
The moving parts, and where each fits:
is_provider_exhausted/mark_provider_exhausted/on_provider_exhausted— the circuit breaker. A provider key (e.g."claude"from"claude-sonnet-4-5") is marked exhausted for the lifetime of the process.on_provider_exhaustedlets model/agent caches clear themselves when a provider goes down.ExhaustedProviderError— raised by your ownmodel_builderwhen asked to build a model for an exhausted provider, so a fallback is chosen instead.ExhaustedProviderCallback— a LangChainBaseCallbackHandleryou attach to a model instance (model.callbacks = [ExhaustedProviderCallback("claude-sonnet-4-5")]). Because it subclassesBaseCallbackHandler, LangChain's callback system invokes itson_llm_errorautomatically whenever a call to that model raises an exception. The handler checks the error against known credit/billing patterns; on a match it marks the provider exhausted and logs amodel_credit_balance_exhaustedwarning.is_fallback_error— the predicate for which exceptions should trigger LangChain's fallback chain (theexceptions_to_handleyou would otherwise pass towith_fallbacks). Rate-limit (429) errors are excluded — those belong to retry middleware.model_with_fallbacks— the entry point. Given a primary model name, a comma-separated fallback list, and yourmodel_builder, it skips any exhausted provider and returns aRunnableWithFallbacks(or a plain model when there are no fallbacks). Pass the result directly tocreate_agent(model=...).
How it relates to LangChain's built-in middleware. LangChain ships ModelRetryMiddleware (retries the same model on transient errors such as 429s) and ModelFallbackMiddleware (switches models when the primary fails). This module augments rather than replaces them: ModelRetryMiddleware handles transient 429 stalls by waiting and retrying the same model; this module handles permanent credit/billing failures (HTTP 400/401/402/403) by skipping the provider and switching models. Use model_with_fallbacks where you would otherwise wire ModelFallbackMiddleware, and keep ModelRetryMiddleware for the 429 path.
What topologies it supports¶
create_agentwith multiple LLM providers, where you want automatic fallback when one provider fails.- Model factories that build models from a name, where you want a guard against exhausted providers.
- Agent or model caches that need to be cleared when a provider goes down, via the
on_provider_exhaustedcallback.
Configured vs exhausted providers¶
The module keeps a process-level set of exhausted providers — the provider keys (e.g. "claude") that have been marked as credit/billing-exhausted. It does not track "configured" providers itself; your model_with_fallbacks call supplies the configured list (primary + fallbacks), and the module filters those against the exhausted set at build time.
Where the exhausted list shows up:
- In the exception message.
ExhaustedProviderErroris raised by your ownmodel_builderwhen asked to build a model for an exhausted provider. Its message names the one provider that was requested and the model, e.g.Provider 'claude' is exhausted (credit/billing). Use a fallback model instead of 'claude-sonnet-4-5'.It names that single exhausted provider — it does not enumerate every exhausted provider. Useis_provider_exhausted(name)to test a specific provider, andmark_provider_exhausted/on_provider_exhaustedto manage the set. - In the logs. When a provider trips, the
model_credit_balance_exhaustedwarning includes anexhausted_providersfield with the full current set (sorted) — e.g.exhausted_providers=['claude', 'deepseek']— so operators can see all providers that are currently skipped. - In the fallback decision.
model_with_fallbacksfilters the configured fallback list against the exhausted set at build time. It logsmodel_provider_exhausted_skipping(the skipped primary, the promoted fallback, and the remainingfallbacks) when the primary is exhausted, andmodel_fallback_chain_configured(theprimaryand thefallbacksthat were actually built after dropping exhausted ones).
What gets logged¶
The module emits structured log events (via structlog) so operators can see provider failures and fallback decisions.
| Event | Level | When | Fields |
|---|---|---|---|
model_credit_balance_exhausted |
warning | A provider returns a credit/billing error and is marked exhausted | model, provider, exhausted_providers, error |
model_provider_exhausted_skipping |
warning | An exhausted primary is skipped and a fallback is promoted | primary, promoted_fallback, remaining_fallbacks |
model_fallback_chain_configured |
info | A model is built with its fallback chain | primary, fallbacks |
Example¶
See examples/provider_failover.py for a runnable example of this bite. Run it with:
API reference¶
langshark_bites.provider_failover
¶
Build resilient LLM model fallback chains for LangChain create_agent.
LangChain's built-in fault-tolerance middleware includes
ModelRetryMiddleware, which handles transient errors such as rate
limits (HTTP 429) by retrying the same model, and
ModelFallbackMiddleware, which switches to an alternative model when
the primary fails (internally using with_fallbacks /
RunnableWithFallbacks). This module augments that story for one
specific case the built-ins leave open: a provider returns a
permanent credit/billing error (HTTP 400/401/402/403 with messages like
"credit balance is too low", "insufficient balance", "credits exhausted").
When that happens, retrying is pointless and slows the agent. This module
marks the provider as exhausted at the process level, so every subsequent
model build for that provider skips it immediately (zero wasted round-trips)
and promotes the first available fallback to primary. It then wires a
fallback chain via LangChain's with_fallbacks/RunnableWithFallbacks.
How the pieces fit¶
is_provider_exhausted/mark_provider_exhausted/on_provider_exhausted: the process-level circuit breaker. A provider key (e.g."claude"from"claude-sonnet-4-5") is marked exhausted and remembered for the lifetime of the process.on_provider_exhaustedlets model/agent caches clear themselves when a provider goes down.ExhaustedProviderError: raised by your ownmodel_builderwhen asked to build a model for an exhausted provider, so a fallback is chosen instead.ExhaustedProviderCallback: a LangChainBaseCallbackHandleryou attach to model instances (callbacks=[...]). When the provider raises a credit/billing error, it marks the provider exhausted.is_fallback_error: predicate for which exceptions should trigger LangChain's fallback chain — i.e. theexceptions_to_handleyou would otherwise pass towith_fallbacks. Rate-limit (429) errors are excluded; those are handled by retry middleware, not by switching models.model_with_fallbacks: the entry point. Given a primary model name, a comma-separated fallback list, and yourmodel_builder, it skips any exhausted provider and returns aRunnableWithFallbacks(or a plain model when there are no fallbacks). Pass the result directly tocreate_agent.
The split is: ModelRetryMiddleware handles transient 429 stalls by
waiting and retrying the same model; this module handles permanent
credit/billing failures by skipping the provider and switching models
(like ModelFallbackMiddleware, but with a circuit-breaker that avoids
rebuilding the failed provider). Use them together.
Usage¶
from langshark_bites.provider_failover import (
ExhaustedProviderError,
ExhaustedProviderCallback,
model_with_fallbacks,
on_provider_exhausted,
)
# Guard your model factory so it refuses exhausted providers:
def create_model(model_name: str, max_tokens: int = 8192):
if is_provider_exhausted(model_name):
raise ExhaustedProviderError(model_name, model_name)
...
# Attach the callback to provider model instances:
model = ChatAnthropic(...)
model.callbacks = [ExhaustedProviderCallback("claude-sonnet-4-5")]
# Build the model with a baked-in fallback chain:
model = model_with_fallbacks(
"claude-sonnet-4-5",
"deepseek-v4-flash,gpt-4o-mini",
max_tokens=8192,
model_builder=create_model,
)
# Pass the result directly as model= to create_agent().
ExhaustedProviderCallback
¶
ExhaustedProviderCallback(model_name, extra_patterns=None)
Bases: BaseCallbackHandler
LangChain callback that detects credit/billing errors from LLM providers.
This is a BaseCallbackHandler that you attach to a model instance via
model.callbacks = [callback]. LangChain's callback system then invokes
the handler's on_llm_error automatically whenever a call to that model
raises an exception. When the error matches a known credit/billing
pattern, the provider is marked exhausted so subsequent calls skip it.
Parameters¶
model_name : str
Full model name (e.g. "claude-sonnet-4-5"). The provider prefix
is extracted as model_name.split("-")[0].
extra_patterns : list[str] | None
Additional lowercase substrings to match in the error message.
The built-in set covers Anthropic, OpenAI, and OpenRouter credit errors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
The model whose provider is marked exhausted on match. |
required |
|
list[str] | None
|
Additional lowercase substrings to match. |
None
|
on_llm_error
¶
on_llm_error(error, **kwargs)
Invoked by LangChain when a model call raises an exception.
LangChain calls this automatically through the callback list attached
to the model (the on_llm_error hook of BaseCallbackHandler).
If the error message matches a known credit/billing pattern, the
provider is marked exhausted (via :func:mark_provider_exhausted) and
a model_credit_balance_exhausted warning is logged so operators
can see which provider tripped.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
BaseException
|
The exception raised by the model call. |
required |
|
Any
|
Additional callback context (ignored). |
{}
|
ExhaustedProviderError
¶
Bases: Exception
Raised when creating a model for an exhausted provider.
is_fallback_error
¶
Return True if exc should trigger the model fallback chain.
Matches permanent billing/credit/auth errors. Rate-limit errors (429) are excluded — those are handled by retry middleware.
is_provider_exhausted
¶
Return True if provider has been marked exhausted in this process.
mark_provider_exhausted
¶
Mark provider as exhausted and fire its on_provider_exhausted callbacks.
model_with_fallbacks
¶
model_with_fallbacks(
primary_name,
fallbacks_csv,
max_tokens=8192,
model_builder=None,
)
Create a model with a baked-in fallback chain and circuit-breaker guard.
The returned object is a LangChain RunnableWithFallbacks wrapping a
primary model and ordered fallback models. Pass it directly to
create_agent (model=...) — no ModelFallbackMiddleware needed.
If the primary provider is already marked as exhausted (from a prior call anywhere in the process), the first available non-exhausted fallback is promoted to primary immediately. Only credit/billing/auth exceptions trigger fallback — rate-limit errors (429) are excluded and left to retry middleware.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
str
|
Primary model identifier (e.g. |
required |
|
str
|
Comma-separated ordered fallback model names
(e.g. |
required |
|
int
|
Maximum output tokens — applied to primary AND all fallbacks. |
8192
|
|
Callable[..., BaseChatModel] | None
|
Function |
None
|
Returns:
| Type | Description |
|---|---|
Any
|
A |
on_provider_exhausted
¶
Register a callback invoked when any provider is marked exhausted.
The callback receives the provider key (e.g. "claude"). Use this to
clear agent/model caches so the next build promotes a fallback to primary.