Skip to content

provider_failover

Circuit breaker for exhausted LLM provider credit, plus model fallback-chain construction.

The problem

An LLM provider can return a permanent credit or billing error (for example, "credit balance is too low"). If you keep calling that provider, every call wastes a round trip. You want to skip it and use a fallback model instead, automatically.

How this bite helps

It keeps a process-level registry of exhausted providers. When a provider is marked exhausted, every subsequent model build for that provider skips it immediately and promotes the first available fallback. model_with_fallbacks builds a LangChain RunnableWithFallbacks with the circuit-breaker guard baked in.

The moving parts, and where each fits:

  • is_provider_exhausted / mark_provider_exhausted / on_provider_exhausted — the circuit breaker. A provider key (e.g. "claude" from "claude-sonnet-4-5") is marked exhausted for the lifetime of the process. on_provider_exhausted lets model/agent caches clear themselves when a provider goes down.
  • ExhaustedProviderError — raised by your own model_builder when asked to build a model for an exhausted provider, so a fallback is chosen instead.
  • ExhaustedProviderCallback — a LangChain BaseCallbackHandler you attach to a model instance (model.callbacks = [ExhaustedProviderCallback("claude-sonnet-4-5")]). Because it subclasses BaseCallbackHandler, LangChain's callback system invokes its on_llm_error automatically whenever a call to that model raises an exception. The handler checks the error against known credit/billing patterns; on a match it marks the provider exhausted and logs a model_credit_balance_exhausted warning.
  • is_fallback_error — the predicate for which exceptions should trigger LangChain's fallback chain (the exceptions_to_handle you would otherwise pass to with_fallbacks). Rate-limit (429) errors are excluded — those belong to retry middleware.
  • model_with_fallbacks — the entry point. Given a primary model name, a comma-separated fallback list, and your model_builder, it skips any exhausted provider and returns a RunnableWithFallbacks (or a plain model when there are no fallbacks). Pass the result directly to create_agent(model=...).

How it relates to LangChain's built-in middleware. LangChain ships ModelRetryMiddleware (retries the same model on transient errors such as 429s) and ModelFallbackMiddleware (switches models when the primary fails). This module augments rather than replaces them: ModelRetryMiddleware handles transient 429 stalls by waiting and retrying the same model; this module handles permanent credit/billing failures (HTTP 400/401/402/403) by skipping the provider and switching models. Use model_with_fallbacks where you would otherwise wire ModelFallbackMiddleware, and keep ModelRetryMiddleware for the 429 path.

What topologies it supports

  • create_agent with multiple LLM providers, where you want automatic fallback when one provider fails.
  • Model factories that build models from a name, where you want a guard against exhausted providers.
  • Agent or model caches that need to be cleared when a provider goes down, via the on_provider_exhausted callback.

Configured vs exhausted providers

The module keeps a process-level set of exhausted providers — the provider keys (e.g. "claude") that have been marked as credit/billing-exhausted. It does not track "configured" providers itself; your model_with_fallbacks call supplies the configured list (primary + fallbacks), and the module filters those against the exhausted set at build time.

Where the exhausted list shows up:

  • In the exception message. ExhaustedProviderError is raised by your own model_builder when asked to build a model for an exhausted provider. Its message names the one provider that was requested and the model, e.g. Provider 'claude' is exhausted (credit/billing). Use a fallback model instead of 'claude-sonnet-4-5'. It names that single exhausted provider — it does not enumerate every exhausted provider. Use is_provider_exhausted(name) to test a specific provider, and mark_provider_exhausted/on_provider_exhausted to manage the set.
  • In the logs. When a provider trips, the model_credit_balance_exhausted warning includes an exhausted_providers field with the full current set (sorted) — e.g. exhausted_providers=['claude', 'deepseek'] — so operators can see all providers that are currently skipped.
  • In the fallback decision. model_with_fallbacks filters the configured fallback list against the exhausted set at build time. It logs model_provider_exhausted_skipping (the skipped primary, the promoted fallback, and the remaining fallbacks) when the primary is exhausted, and model_fallback_chain_configured (the primary and the fallbacks that were actually built after dropping exhausted ones).

What gets logged

The module emits structured log events (via structlog) so operators can see provider failures and fallback decisions.

Event Level When Fields
model_credit_balance_exhausted warning A provider returns a credit/billing error and is marked exhausted model, provider, exhausted_providers, error
model_provider_exhausted_skipping warning An exhausted primary is skipped and a fallback is promoted primary, promoted_fallback, remaining_fallbacks
model_fallback_chain_configured info A model is built with its fallback chain primary, fallbacks

Example

See examples/provider_failover.py for a runnable example of this bite. Run it with:

uv run python examples/provider_failover.py

API reference

langshark_bites.provider_failover

Build resilient LLM model fallback chains for LangChain create_agent.

LangChain's built-in fault-tolerance middleware includes ModelRetryMiddleware, which handles transient errors such as rate limits (HTTP 429) by retrying the same model, and ModelFallbackMiddleware, which switches to an alternative model when the primary fails (internally using with_fallbacks / RunnableWithFallbacks). This module augments that story for one specific case the built-ins leave open: a provider returns a permanent credit/billing error (HTTP 400/401/402/403 with messages like "credit balance is too low", "insufficient balance", "credits exhausted").

When that happens, retrying is pointless and slows the agent. This module marks the provider as exhausted at the process level, so every subsequent model build for that provider skips it immediately (zero wasted round-trips) and promotes the first available fallback to primary. It then wires a fallback chain via LangChain's with_fallbacks/RunnableWithFallbacks.

How the pieces fit

  • is_provider_exhausted / mark_provider_exhausted / on_provider_exhausted: the process-level circuit breaker. A provider key (e.g. "claude" from "claude-sonnet-4-5") is marked exhausted and remembered for the lifetime of the process. on_provider_exhausted lets model/agent caches clear themselves when a provider goes down.
  • ExhaustedProviderError: raised by your own model_builder when asked to build a model for an exhausted provider, so a fallback is chosen instead.
  • ExhaustedProviderCallback: a LangChain BaseCallbackHandler you attach to model instances (callbacks=[...]). When the provider raises a credit/billing error, it marks the provider exhausted.
  • is_fallback_error: predicate for which exceptions should trigger LangChain's fallback chain — i.e. the exceptions_to_handle you would otherwise pass to with_fallbacks. Rate-limit (429) errors are excluded; those are handled by retry middleware, not by switching models.
  • model_with_fallbacks: the entry point. Given a primary model name, a comma-separated fallback list, and your model_builder, it skips any exhausted provider and returns a RunnableWithFallbacks (or a plain model when there are no fallbacks). Pass the result directly to create_agent.

The split is: ModelRetryMiddleware handles transient 429 stalls by waiting and retrying the same model; this module handles permanent credit/billing failures by skipping the provider and switching models (like ModelFallbackMiddleware, but with a circuit-breaker that avoids rebuilding the failed provider). Use them together.

Usage

from langshark_bites.provider_failover import (
    ExhaustedProviderError,
    ExhaustedProviderCallback,
    model_with_fallbacks,
    on_provider_exhausted,
)

# Guard your model factory so it refuses exhausted providers:
def create_model(model_name: str, max_tokens: int = 8192):
    if is_provider_exhausted(model_name):
        raise ExhaustedProviderError(model_name, model_name)
    ...

# Attach the callback to provider model instances:
model = ChatAnthropic(...)
model.callbacks = [ExhaustedProviderCallback("claude-sonnet-4-5")]

# Build the model with a baked-in fallback chain:
model = model_with_fallbacks(
    "claude-sonnet-4-5",
    "deepseek-v4-flash,gpt-4o-mini",
    max_tokens=8192,
    model_builder=create_model,
)
# Pass the result directly as model= to create_agent().

ExhaustedProviderCallback

ExhaustedProviderCallback(model_name, extra_patterns=None)

Bases: BaseCallbackHandler

LangChain callback that detects credit/billing errors from LLM providers.

This is a BaseCallbackHandler that you attach to a model instance via model.callbacks = [callback]. LangChain's callback system then invokes the handler's on_llm_error automatically whenever a call to that model raises an exception. When the error matches a known credit/billing pattern, the provider is marked exhausted so subsequent calls skip it.

Parameters

model_name : str Full model name (e.g. "claude-sonnet-4-5"). The provider prefix is extracted as model_name.split("-")[0]. extra_patterns : list[str] | None Additional lowercase substrings to match in the error message. The built-in set covers Anthropic, OpenAI, and OpenRouter credit errors.

Parameters:

Name Type Description Default

model_name

str

The model whose provider is marked exhausted on match.

required

extra_patterns

list[str] | None

Additional lowercase substrings to match.

None

on_llm_error

on_llm_error(error, **kwargs)

Invoked by LangChain when a model call raises an exception.

LangChain calls this automatically through the callback list attached to the model (the on_llm_error hook of BaseCallbackHandler). If the error message matches a known credit/billing pattern, the provider is marked exhausted (via :func:mark_provider_exhausted) and a model_credit_balance_exhausted warning is logged so operators can see which provider tripped.

Parameters:

Name Type Description Default
error
BaseException

The exception raised by the model call.

required
kwargs
Any

Additional callback context (ignored).

{}

ExhaustedProviderError

ExhaustedProviderError(provider, model_name)

Bases: Exception

Raised when creating a model for an exhausted provider.

is_fallback_error

is_fallback_error(exc)

Return True if exc should trigger the model fallback chain.

Matches permanent billing/credit/auth errors. Rate-limit errors (429) are excluded — those are handled by retry middleware.

is_provider_exhausted

is_provider_exhausted(provider)

Return True if provider has been marked exhausted in this process.

mark_provider_exhausted

mark_provider_exhausted(provider)

Mark provider as exhausted and fire its on_provider_exhausted callbacks.

model_with_fallbacks

model_with_fallbacks(
    primary_name,
    fallbacks_csv,
    max_tokens=8192,
    model_builder=None,
)

Create a model with a baked-in fallback chain and circuit-breaker guard.

The returned object is a LangChain RunnableWithFallbacks wrapping a primary model and ordered fallback models. Pass it directly to create_agent (model=...) — no ModelFallbackMiddleware needed.

If the primary provider is already marked as exhausted (from a prior call anywhere in the process), the first available non-exhausted fallback is promoted to primary immediately. Only credit/billing/auth exceptions trigger fallback — rate-limit errors (429) are excluded and left to retry middleware.

Parameters:

Name Type Description Default

primary_name

str

Primary model identifier (e.g. "claude-sonnet-4-5").

required

fallbacks_csv

str

Comma-separated ordered fallback model names (e.g. "deepseek-v4-flash,gpt-4o-mini"). Empty string or whitespace-only → returns a plain model with no fallback chain.

required

max_tokens

int

Maximum output tokens — applied to primary AND all fallbacks.

8192

model_builder

Callable[..., BaseChatModel] | None

Function (model_name, max_tokens) → BaseChatModel. Usually create_model from your model factory. If None, the caller must set model_builder (required for circuit breaker).

None

Returns:

Type Description
Any

A BaseChatModel (or RunnableWithFallbacks wrapping one).

on_provider_exhausted

on_provider_exhausted(callback)

Register a callback invoked when any provider is marked exhausted.

The callback receives the provider key (e.g. "claude"). Use this to clear agent/model caches so the next build promotes a fallback to primary.