Skip to content

Advanced Configuration

This page covers advanced configuration topics that go beyond the basics in the Configuration Reference. Each section dives deeper into tuning, behavior details, and practical examples.

A model can have multiple endpoints for redundancy and load distribution. When more than one endpoint is configured, llmsoup distributes requests across them based on their weight values.

Each endpoint’s weight determines its share of traffic. The probability of an endpoint being selected is weight / total_weights.

models:
- name: gpt-5-mini
provider: openai
access_key:
env: OPENAI_API_KEY
endpoints:
- url: https://primary.openai.example.com/v1/chat/completions
weight: 70
timeout_ms: 5000
description: "Primary region (us-east)"
- url: https://secondary.openai.example.com/v1/chat/completions
weight: 30
timeout_ms: 8000
description: "Secondary region (eu-west)"

In this example, roughly 70% of requests go to the primary endpoint and 30% to the secondary. If weight is omitted, it defaults to 1 — so two unweighted endpoints split traffic 50/50.

FieldTypeDefaultDescription
urlstring(required)Full URL to the chat completions endpoint.
weightinteger (>0)1Load balancing weight. Higher values receive more traffic.
timeout_msintegerglobal defaultPer-endpoint timeout override in milliseconds.
descriptionstringHuman-readable label (appears in logs and metrics).
  • Geographic redundancy — Route to the nearest region with higher weight, fail over to others.
  • Provider diversity — Split traffic between OpenAI and Azure OpenAI for the same model.
  • Rate limit management — Distribute requests across multiple API keys or accounts.

Reasoning configuration controls how llmsoup tells upstream models to use chain-of-thought reasoning. This is configured per-model with reasoning_family and overridden per-rule with model_refs.

The reasoning_family field on a model definition tells llmsoup how to pass reasoning parameters to that model’s API. It accepts two forms:

String shorthand — for providers that use a standard parameter name:

models:
- name: gpt-5.2
reasoning_family: reasoning_effort

This tells llmsoup the model supports a reasoning_effort parameter directly in the API request.

Object form — for providers with custom parameter names:

models:
- name: deepseek-r1
reasoning_family:
type: chat_template_kwargs
parameter: thinking_mode

This tells llmsoup to pass reasoning controls via the thinking_mode parameter in chat_template_kwargs.

Within a routing rule, model_refs can override reasoning behavior per-model:

rules:
- name: deep-analysis
priority: 100
conditions:
- signal: keyword.complex_query
action:
strategy: default
primary_model: gpt-5.2
model_refs:
- model: gpt-5.2
use_reasoning: true
reasoning_effort: high
- model: gpt-5-mini
use_reasoning: false
FieldTypeDescription
modelstringModel name (must reference a configured model).
use_reasoningbooleanEnable or disable reasoning for this rule.
reasoning_effortstringEffort level: low, medium, or high.

When use_reasoning is true, llmsoup mutates the outbound request to include the reasoning parameter appropriate for the model’s reasoning_family. When false, reasoning parameters are stripped even if the model supports them.


Plugins add pre-processing and post-processing to routing rules. They are configured in the plugins array on each rule and execute in order.

rules:
- name: secure-route
priority: 100
conditions:
- signal: keyword.sensitive
action:
strategy: default
primary_model: gpt-5.2
plugins:
- type: jailbreak
configuration:
enabled: true
threshold: 0.8
- type: system_prompt
configuration:
system_prompt: "Answer carefully and factually."

Every plugin has a type field and a configuration block. The available plugin types are:

Plugin typePurpose
system_promptInject or replace system prompts
semantic-cacheCache responses by semantic similarity
jailbreakDetect prompt injection attempts
piiDetect personally identifiable information
header_mutationModify HTTP headers on requests/responses
hallucinationFlag potential hallucinations in responses
router_replayRecord routing decisions for debugging

For complete per-plugin configuration details, field descriptions, and examples, see the Plugins Reference.


llmsoup uses three independent caching layers, each with different eviction strategies and tuning knobs.

Caches full model responses keyed by request content. When a cache hit occurs, the response is returned immediately without calling the upstream model.

defaults:
model_cache_ttl_seconds: 3600
model_cache_max_capacity: 1000
SettingDefaultDescription
model_cache_ttl_seconds3600Time-to-live in seconds. Entries expire after this duration regardless of access.
model_cache_max_capacity1000Maximum number of cached model response entries. Oldest entries are evicted when capacity is reached (LRU).

Eviction behavior: Entries are evicted after the TTL expires or when the cache reaches model_cache_max_capacity entries (whichever comes first). The capacity limit uses LRU eviction — the least recently used entry is removed to make room for new ones. This prevents unbounded memory growth under heavy traffic.

When to tune:

  • Lower TTL (e.g., 60–300) for rapidly changing data or when freshness matters.
  • Higher TTL (e.g., 7200+) for stable queries where the same prompt always has the same answer.
  • Set to 0 to effectively disable model response caching.

Caches computed embedding vectors to avoid re-running BERT inference for repeated text. This primarily benefits the embedding signal evaluator.

defaults:
embedding_cache_capacity: 1000
SettingDefaultDescription
embedding_cache_capacity1000Maximum number of cached embedding entries. Oldest entries are evicted when capacity is reached.

Eviction behavior: Least Recently Used (LRU). When the cache is full, the entry that hasn’t been accessed for the longest time is evicted to make room.

When to tune:

  • Increase if your workload has many unique prompts and you see high embedding computation times in metrics.
  • Decrease if memory is constrained — each embedding entry holds a vector of floats.

Tracks per-model Time Per Output Token (TPOT) using Exponential Moving Average smoothing. This is used internally by the latency signal evaluator — it is not directly configurable via YAML.

How it works: After each model response, llmsoup computes the actual TPOT and updates the smoothed average:

smoothed_tpot = alpha × new_tpot + (1 - alpha) × previous_smoothed_tpot

The default EMA alpha is 0.3, which means:

  • 30% weight to the most recent observation
  • 70% weight to historical average

This smoothing prevents a single slow response from drastically changing the model’s latency estimate. The latency signal evaluator compares the smoothed TPOT against the max_tpot threshold configured on each latency signal.

When multiple embedding signals reference the same model (e.g., sentence-transformers/all-MiniLM-L6-v2), llmsoup automatically shares a single loaded model instance across all signals. This is transparent — no configuration is needed.

Impact: Each BERT embedding model uses ~130MB of RAM. Without sharing, six signals using the same model would consume ~780MB. With sharing, they use ~130MB total.

How it works: During startup, llmsoup groups embedding signals by their resolved model path. Signals that share a path receive a shared handle to a single model instance. Each signal still maintains its own reference text, candidates, and cache — only the underlying BERT model weights are shared.


llmsoup loads its BERT models lazily on the first request, so idle memory stays low. Under load, resident memory is driven by two things: the resident model weights and the per-request activation/scratch buffers of concurrent forward passes. The knobs below keep both bounded so the process stays within the memory budget (idle ≤ 500 MB, loaded ≤ 1 GB).

Every embedding and classifier forward pass runs on a blocking thread. Without a limit, N concurrent requests run N concurrent forwards, and peak memory scales with request concurrency. defaults.max_inference_concurrency caps how many forward passes execute at once, so peak memory scales with a small fixed limit instead.

defaults:
max_inference_concurrency: 4
  • Default: min(number_of_cores, 4), floored at 2. On Linux the core count is cgroup-aware, so a container with a CPU quota gets a sensible value automatically.
  • Override at runtime: set the LLMSOUP_MAX_INFERENCE_CONCURRENCY environment variable (used when the config value is omitted). An explicit config value always wins.
  • Validation: the value must be greater than 0 and no larger than 1024; out-of-range values are rejected at startup with a file:line:column error (the ceiling prevents an oversized value from aborting the process when the semaphore is built).
  • Backpressure, not bypass: requests wait for a permit before running a forward pass — the cap is never silently exceeded under load. Because bounding peak memory is the goal, the queue applies as hard backpressure rather than flooding the blocking pool, so a burst trades latency (not memory) for the guarantee.
  • Backward compatibility: setting the value at or above your maximum expected in-flight requests restores the prior effectively-unbounded behavior.

Because routing inference is CPU-bound and runs batch-1 with short prompts, concurrency beyond a few gives essentially no throughput gain but grows memory linearly — so a small limit is the right default. The only realistic cost of a low limit is worst-case tail latency under a hard burst (roughly ceil(burst / limit) × per-forward-ms), not steady-state throughput.

LLMSOUP_INFERENCE_THREADS caps the number of intra-op (gemm/rayon) threads used per forward pass. It is applied at startup before the first matmul by seeding RAYON_NUM_THREADS (an explicitly set RAYON_NUM_THREADS takes precedence). Lowering it reduces duplicated gemm scratch on many-core hosts without changing routing decisions.

Terminal window
LLMSOUP_INFERENCE_THREADS=2 llmsoup serve

On the glibc/Linux target, MALLOC_ARENA_MAX=2 limits the number of per-thread malloc arenas, which curbs arena retention from freed inference scratch. The production container image sets this by default. It has no effect on musl/Alpine, which is why the shipped image is glibc-based (Debian slim).

llmsoup can be built with a tuned global allocator via the jemalloc cargo feature:

Terminal window
cargo build --release --features jemalloc

This installs jemalloc as the global allocator on the Linux target (with background_thread:true,dirty_decay_ms:1000 to purge dirty pages promptly) and leaves macOS on the system allocator. On the Linux/glibc target this is the single most effective memory lever: in the container benchmark at 50-concurrency it cut loaded RSS from ~2.8 GB (glibc, MALLOC_ARENA_MAX=2) to ~1.2 GB, because glibc retains freed inference scratch in its arenas while jemalloc returns dirty pages to the OS. When jemalloc is active it supersedes MALLOC_ARENA_MAX.

Because the saving is decisive on the production target, the shipped Docker image is built with jemalloc by default, and the same feature is applied to the downloadable Linux release binaries (*-unknown-linux-gnu). Plain cargo build (e.g. local dev on macOS) stays on the system allocator; rebuild the image without it via --build-arg CARGO_FEATURES="" if you need the glibc allocator.

The shipped Dockerfile is a multi-stage, glibc (Debian slim) build that carries only the release binary, so memory behavior can be measured on and deployed to the production target rather than only macOS. It applies the memory-tuned default MALLOC_ARENA_MAX=2 and leaves inference concurrency at the cgroup-aware auto-default (so a small CPU quota yields a small limit automatically — set LLMSOUP_MAX_INFERENCE_CONCURRENCY to override), runs as a non-root user, exposes port 8080, and loads its configuration from the path in LLMSOUP_CONFIG (mount your own; no secrets are baked into the image).

Terminal window
# Build (context is the repo root; the crate lives in llmsoup/)
docker build -t llmsoup .
# Run with a mounted config and a persistent HuggingFace model cache
docker run --rm -p 8080:8080 \
-v "$PWD/config.yaml:/etc/llmsoup/config.yaml:ro" \
-v llmsoup-models:/models \
llmsoup

The same image doubles as the reproducible memory-measurement harness — run llmsoup benchmark --use-models inside the container to record idle and loaded RSS on the Linux/glibc target. (macOS numbers come from the Metal backend and are indicative only.)


llmsoup resolves secrets (API keys, authentication tokens) from external sources at startup. Secrets are never stored in the configuration file itself. The resolution methods are tried in the order they appear in the configuration.

The simplest and most common method. Reads the secret from an environment variable.

access_key:
env: OPENAI_API_KEY

Behavior: Looks up the environment variable at config load time. Fails validation if the variable is not set or is empty.

Error: Secret resolution failed: environment variable 'OPENAI_API_KEY' not set

Reads the secret from a file path. Ideal for Docker secrets and Kubernetes secret volumes.

access_key:
file: /run/secrets/openai_api_key

Behavior: Reads the entire file content and trims leading/trailing whitespace. Fails if the file does not exist or is not readable.

Error: Secret resolution failed: cannot read file '/run/secrets/openai_api_key'

Common patterns:

# Docker secret
access_key:
file: /run/secrets/api_key
# Kubernetes secret volume
access_key:
file: /etc/llmsoup/secrets/api-key

Executes a shell command and captures its stdout as the secret value. Useful for integration with secret managers like AWS Secrets Manager or 1Password CLI.

access_key:
command: "aws secretsmanager get-secret-value --secret-id openai-key --query SecretString --output text"

Behavior: Runs the command via the system shell, captures stdout, and trims whitespace. The command must exit with code 0.

Error (when disabled): Secret resolution failed: command-based secrets are disabled. Set LLMSOUP_ALLOW_COMMAND_SECRETS=1 to enable.

Error (when command fails): Secret resolution failed: command exited with non-zero status

access_key:
vault: "secret/data/openai/api-key"

Error: Secret resolution failed: vault secret resolution is not yet implemented

When multiple methods are specified in the same secret reference, they are resolved in order: envfilevaultcommand. The first successful resolution wins.

# Tries env first, falls back to file
access_key:
env: OPENAI_API_KEY
file: /run/secrets/openai_api_key

When a request exceeds the model’s effective context window (see Context headroom), llmsoup applies a context overflow strategy to fit the conversation within the limit. The strategy is set globally in defaults.

defaults:
context_overflow: truncate_middle
context_headroom_ratio: 0.1

llmsoup estimates request tokens with a ~4 characters/token heuristic, which under-counts code-dense content. To compensate, a fraction of each model’s declared context_window is reserved as headroom: the effective window used for all fit checks and truncation budgets is context_window * (1 - context_headroom_ratio).

The fit check covers the whole request: messages, tools, and the output budget (max_completion_tokens if set, else max_tokens, else 4096).

context_headroom_ratio defaults to 0.1 and must be between 0.0 and 0.5 — out-of-range values fail config validation at load time with a file:line:column error. Set it to 0.0 to disable the headroom and check against the raw window.

Keeps the leading system and developer messages (the “protected prefix”) and the most recent messages. Drops messages from the middle of the conversation.

Before (6 messages, context exceeded):

[system] You are a helpful assistant.
[user] What is Rust? ← DROPPED
[assistant] Rust is a systems language... ← DROPPED
[user] How about Go? ← DROPPED
[assistant] Go is a compiled language... ← kept
[user] Compare their async models. ← kept

After:

[system] You are a helpful assistant.
[assistant] Go is a compiled language...
[user] Compare their async models.

Best for: General-purpose use. Preserves the system/developer instructions and the most recent turns, which are usually the most relevant.

Drops the oldest non-system messages first, keeping the most recent conversation.

Before (6 messages, context exceeded):

[system] You are a helpful assistant.
[user] What is Rust? ← DROPPED
[assistant] Rust is a systems language... ← DROPPED
[user] How about Go? ← DROPPED
[assistant] Go is a compiled language... ← kept
[user] Compare their async models. ← kept

After:

[system] You are a helpful assistant.
[assistant] Go is a compiled language...
[user] Compare their async models.

Best for: Long-running conversations where only the recent context matters. Similar to a sliding window over the chat history.

Returns an error immediately without sending the request. No messages are dropped.

Behavior: Returns HTTP 400 in OpenAI format. The message cites the effective window — the declared window minus headroom — so clients that resize their prompt use a limit that is actually enforced:

{
"error": {
"message": "Request of ~130000 tokens exceeds model 'gpt-4o' effective context window of 115200 tokens (128000 declared minus headroom). Reduce the message length or use a model with a larger context window.",
"type": "invalid_request_error",
"code": "context_length_exceeded"
}
}

Best for: Applications that need to know when context is exceeded so they can handle it themselves (e.g., summarize the conversation before retrying).

When a matched decision uses a confidence or ratings algorithm to select among multiple model refs, llmsoup estimates the request’s tokens once (messages + tools + output budget) and skips refs whose effective window cannot hold the estimate — provided at least one ref on the decision fits. Oversized prompts route to a larger-window ref instead of losing messages to truncation.

If no ref fits, the algorithm’s normal ranking stands and the configured overflow strategy truncates the request as a last resort. Skipped refs are logged with the model name, estimated tokens, and effective window. Models without metadata.context_window are assumed to fit. See Algorithms.

The decision’s fallback chain is ordered the same way: refs that fit the request are listed before refs that don’t, so a transient failure falls back to another fitting model instead of one that would truncate.

Token estimates are heuristic, so an oversized request can still reach a model and be rejected upstream. When an upstream call returns a 400 whose body matches a context-length signature — a message containing “maximum context length” (case-insensitive) or error code context_length_exceeded — llmsoup truncates the request per the configured overflow strategy and retries the same model exactly once.

The retry is skipped — the error surfaces immediately — when a second attempt cannot succeed:

  • Truncation cannot actually shrink the request, e.g. content the token estimator cannot see, such as image parts.
  • The output budget (max_completion_tokens/max_tokens) alone exceeds the effective window — dropping messages cannot fix that.
  • Dropping messages would leave only system/developer messages; llmsoup never silently answers a gutted conversation.
  • The first attempt already consumed half or more of the per-ref time budget (defaults.request_timeout_ms, default 30000 ms), so a retry would be killed by the executor timeout.

If the retry also fails with a context-length error, execution advances to the decision’s remaining refs (fallback). Other 4xx errors remain terminal for that ref and are not retried. This applies to both buffered and streaming (SSE) upstream calls. Every retry decision is counted in llmsoup_context_length_retries_total.


Model metadata fields influence routing decisions beyond simple name matching.

models:
- name: gpt-5.2
metadata:
context_window: 128000
parameter_count: 520000
latency_seconds: 1.2
FieldTypeUsed by
context_windowintegerContext overflow strategies, fit-aware model selection, prompt validation.
parameter_countintegerConfidence algorithm’s escalation_order: size — models are escalated in order of increasing parameter count.
latency_secondsfloatInitial latency estimates before real TPOT data is collected.

When using the confidence algorithm with escalation_order: size, models in the model_refs list are sorted by parameter_count (ascending). If the first model’s response falls below the confidence threshold, the request escalates to the next larger model:

action:
algorithm:
type: confidence
confidence:
threshold: 0.8
escalation_order: size
model_refs:
- model: gpt-5-mini # parameter_count: 120000 → tried first
- model: gpt-5.2 # parameter_count: 520000 → escalation target

Both the confidence and ratings algorithms support an on_error field that controls behavior when a model call fails during algorithm evaluation.

algorithm:
type: confidence
confidence:
threshold: 0.8
on_error: skip
ValueDefaultBehavior
skipYesSkip the failed model and try the next one in the list.
failReturn an error immediately without trying other models.

When on_error is skip (the default) and a model returns an error or times out, the algorithm moves to the next model in the evaluation order instead of failing the entire request. When set to fail, the first model error aborts the algorithm and returns the error to the caller.

The confidence_method field controls how the confidence algorithm calculates confidence from model responses:

MethodDescription
marginUses the difference between the top token probability and the second-highest. Larger margins indicate higher confidence.
avg_logprobUses the average log probability across all output tokens. Higher average means more confident.
hybridCombines both methods using configurable weights.
algorithm:
type: confidence
confidence:
threshold: 0.75
confidence_method: hybrid
hybrid_weights:
logprob_weight: 0.6
margin_weight: 0.4

Both algorithm types support a per-rule cost_quality_tradeoff that overrides the global defaults.cost_quality_tradeoff:

rules:
- name: budget-route
action:
algorithm:
type: confidence
confidence:
threshold: 0.7
cost_quality_tradeoff: 0.8 # strongly prefer cheaper models
- name: quality-route
action:
algorithm:
type: ratings
ratings:
policy: highest
cost_quality_tradeoff: 0.1 # strongly prefer quality

The value ranges from 0.0 (pure quality) to 1.0 (pure cost savings). This requires cost_aware_routing: true in defaults and pricing configured on models.