Chaos Testing
aimock provides probabilistic failure injection to test how your application handles unreliable LLM APIs. Four probabilistic failure modes and one deterministic delay can be configured at the server, fixture, or per-request level.
Failure Modes
| Mode | Action | Description |
|---|---|---|
drop |
HTTP 500 |
Returns a 500 error with
{"error":{"message":"Chaos: request
dropped","type":"server_error","code":"chaos_drop"}}
|
malformed |
Broken JSON |
Returns HTTP 200 with invalid JSON body:
{malformed json: <<<chaos>>>
|
rateLimit |
HTTP 429 |
Returns a 429 with
{"error":{"message":"Chaos: rate limit
exceeded","type":"rate_limit_error","code":"chaos_ratelimit"}}
plus Retry-After: 1 and the static OpenAI-style
x-ratelimit-* headers every aimock 429 carries
|
disconnect |
Connection destroyed | Destroys the TCP connection immediately with no response |
The four modes are rolled from rates. A fifth field, latencyMs, is not a mode
and is not rolled: when set it always fires, delaying the request by that many
milliseconds before the four rates are drawn, so the delay applies to a request
whether or not a fault then fires.
Precedence
Chaos configuration is resolved with a three-level precedence hierarchy. Higher levels override lower ones:
- Per-request headers (highest) — override everything
- Fixture-level config — overrides server defaults
- Server-level defaults (lowest)
Each level overrides the one below field by field: a fixture that sets only
dropRate leaves the server's latencyMs in place, and a request
header that sets only x-aimock-chaos-disconnect leaves the rest of the
fixture's rates alone.
Within a single level, modes are rolled in order — drop, malformed, rateLimit,
disconnect — as four separate draws, and the first draw that hits wins and
short-circuits the rest. Each rate is therefore the chance that its draw hits
given that no earlier mode fired, not the share of requests that end in that
fault; only dropRate, first in the chain, is unconditional. So
{ "dropRate": 0.5, "rateLimitRate": 0.5 } produces about 50% drops and about
25% rate limits (0.5 × 0.5), not 50% of each — to exercise two faults at their
nominal frequency, configure them on separate requests.
One exception: the per-test scope. The server level is scoped by
X-Test-Id — a runtime override installed with
POST /__aimock/chaos applies only to traffic carrying the same tag (see the
Control API). Picking that scope is a selection, not a
merge: a scoped override replaces the server-wide baseline wholesale
rather than layering on it. A server started with --chaos-latency 500 whose
test installs { "dropRate": 1 } sees drops with no latency on that
test's traffic — restate latencyMs in the override to keep it.
GET /__aimock/chaos with the same tag always reports the config actually in
effect for the scope, and the server warns when an install drops a field that was in
effect.
Quick Start
import { LLMock } from "@copilotkit/aimock";
const mock = new LLMock();
mock.onMessage("hello", { content: "Hi!" });
// 50% of all requests will be dropped with a 500
mock.setChaos({ dropRate: 0.5 });
await mock.start();
// Later, remove chaos
mock.clearChaos();
Programmatic API
// Set server-level chaos (returns `this` for chaining)
mock.setChaos({
dropRate: 0.1, // 10% drop rate
malformedRate: 0.05, // 5% malformed rate
rateLimitRate: 0.05, // 5% rate-limit (429) rate
disconnectRate: 0.02, // 2% disconnect rate
latencyMs: 250, // always delay 250 ms first
});
// Remove all server-level chaos
mock.clearChaos();
Fixture-Level Chaos
Attach a chaos config to individual fixtures so only specific responses
experience failures:
{
"fixtures": [
{
"match": { "userMessage": "unstable" },
"response": { "content": "This might fail!" },
"chaos": {
"dropRate": 0.3,
"malformedRate": 0.2,
"disconnectRate": 0.1
}
},
{
"match": { "userMessage": "stable" },
"response": { "content": "This always works." }
}
]
}
Per-Request Headers
Override chaos on individual requests using HTTP headers — one header per field,
each carrying the same value the corresponding chaos field takes:
| Header | Controls | Accepted values |
|---|---|---|
x-aimock-chaos-drop |
dropRate — 500 with an error body |
0–1 |
x-aimock-chaos-malformed |
malformedRate — truncated/invalid JSON |
0–1 |
x-aimock-chaos-disconnect |
disconnectRate — socket destroyed mid-flight |
0–1 |
x-aimock-chaos-ratelimit |
rateLimitRate — 429 with Retry-After |
0–1 |
x-aimock-chaos-latency |
latencyMs — delay before the response |
0–30000, whole ms |
Header values are read as plain decimal text and are never clamped: a
value that is out of range, not a plain decimal (1e3, 0x1f,
Infinity), or fractional where a whole number of milliseconds is required is
rejected with a warning and simply does not apply — the level below it (fixture,
then server default) still does. The same rule governs values set in a fixture, in the
server defaults, on the --chaos-* CLI flags and through
POST /__aimock/chaos.
// Force 100% disconnect on this specific request
await fetch(`${mock.url}/v1/chat/completions`, {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-aimock-chaos-disconnect": "1.0",
},
body: JSON.stringify({ model: "gpt-4", messages: [{ role: "user", content: "hello" }] }),
});
CLI Flags
Set server-level chaos from the command line. The --chaos-* flags belong to
the llmock binary (the fixtures-and-flags CLI), which is also what the Docker
image runs.
$ npx -p @copilotkit/aimock llmock --fixtures ./fixtures \
--chaos-drop 0.1 \
--chaos-malformed 0.05 \
--chaos-ratelimit 0.05 \
--chaos-disconnect 0.02 \
--chaos-latency 250
$ docker run -d -p 4010:4010 \
-v ./fixtures:/fixtures \
ghcr.io/copilotkit/aimock \
-f /fixtures -h 0.0.0.0 \
--chaos-drop 0.1 \
--chaos-malformed 0.05 \
--chaos-ratelimit 0.05 \
--chaos-disconnect 0.02 \
--chaos-latency 250
The aimock binary takes no chaos flags — its options are
--config, --port, --host and
-h, --help. With aimock --config, set the server-level chaos
under llm.chaos in the config file (the same five fields, the same bounds),
or install it at runtime with POST /__aimock/chaos (see the
Control API).
Proxy Mode
When aimock is configured as a record/replay proxy (--record), chaos applies
to proxied requests too — so a staging environment pointed at real upstream APIs
still sees the failure modes your tests expect. Chaos is rolled once per request,
after fixture matching, with the same headers > fixture > server
precedence.
| Mode | When upstream is contacted | What the client sees |
|---|---|---|
drop |
Never — upstream not contacted | HTTP 500 chaos body; upstream is not called |
rateLimit |
Never — upstream not contacted | HTTP 429 chaos body; upstream is not called |
disconnect |
Never — upstream not contacted | Connection destroyed; upstream is not called |
malformed |
Called — post-response | Request proxies normally; the upstream response is captured, then the body is replaced with invalid JSON before relay. The recorded fixture (if recording) keeps the real upstream response — chaos is a live-traffic decoration, not a fixture mutation. |
SSE bypass. If upstream returns
Content-Type: text/event-stream, aimock streams chunks to the client
progressively. By the time malformed would fire, the bytes are already on the
wire — the chaos action cannot be applied. This bypass is observable via the
aimock_chaos_bypassed_total counter (see Prometheus Metrics below) and a
warning in the server log, so a configured chaos rate doesn't silently drop to 0% on SSE
traffic. Streaming mutation is planned for a future phase.
Journal Tracking
When chaos triggers, the journal entry includes a chaosAction field recording
which failure mode was applied:
{
"method": "POST",
"path": "/v1/chat/completions",
"response": {
"status": 500,
"source": "fixture",
"fixture": { "...": "elided for brevity" },
"chaosAction": "drop"
}
}
The chaosAction values are "drop", "malformed",
"rateLimit" and "disconnect". The recorded status codes are 500
for drop, 200 for malformed, 429 for rateLimit, and 0 for disconnect (connection
destroyed). A latency delay alone leaves no chaosAction: the entry is the
ordinary one for whatever the request then received.
Prometheus Metrics
When metrics are enabled (--metrics), each chaos trigger increments the
aimock_chaos_triggered_total counter, tagged with action and
source. source="fixture" means a fixture matched (or would have,
before chaos intervened); source="proxy" means the request was on the proxy
dispatch path.
# TYPE aimock_chaos_triggered_total counter
aimock_chaos_triggered_total{action="drop",source="fixture"} 3
aimock_chaos_triggered_total{action="malformed",source="fixture"} 1
aimock_chaos_triggered_total{action="disconnect",source="proxy"} 2
When a chaos action is rolled but can't be applied — today, only
malformed on an SSE proxy response — the bypass is recorded in a
separate counter so operators can distinguish "chaos didn't roll" from "chaos rolled but
was bypassed":
# TYPE aimock_chaos_bypassed_total counter
aimock_chaos_bypassed_total{action="malformed",source="proxy",reason="sse_streamed"} 4