Chaos Testing

aimock provides probabilistic failure injection to test how your application handles unreliable LLM APIs. Four probabilistic failure modes and one deterministic delay can be configured at the server, fixture, or per-request level.

Failure Modes

Mode Action Description
drop HTTP 500 Returns a 500 error with {"error":{"message":"Chaos: request dropped","type":"server_error","code":"chaos_drop"}}
malformed Broken JSON Returns HTTP 200 with invalid JSON body: {malformed json: <<<chaos>>>
rateLimit HTTP 429 Returns a 429 with {"error":{"message":"Chaos: rate limit exceeded","type":"rate_limit_error","code":"chaos_ratelimit"}} plus Retry-After: 1 and the static OpenAI-style x-ratelimit-* headers every aimock 429 carries
disconnect Connection destroyed Destroys the TCP connection immediately with no response

The four modes are rolled from rates. A fifth field, latencyMs, is not a mode and is not rolled: when set it always fires, delaying the request by that many milliseconds before the four rates are drawn, so the delay applies to a request whether or not a fault then fires.

Precedence

Chaos configuration is resolved with a three-level precedence hierarchy. Higher levels override lower ones:

  1. Per-request headers (highest) — override everything
  2. Fixture-level config — overrides server defaults
  3. Server-level defaults (lowest)

Each level overrides the one below field by field: a fixture that sets only dropRate leaves the server's latencyMs in place, and a request header that sets only x-aimock-chaos-disconnect leaves the rest of the fixture's rates alone.

Within a single level, modes are rolled in order — drop, malformed, rateLimit, disconnect — as four separate draws, and the first draw that hits wins and short-circuits the rest. Each rate is therefore the chance that its draw hits given that no earlier mode fired, not the share of requests that end in that fault; only dropRate, first in the chain, is unconditional. So { "dropRate": 0.5, "rateLimitRate": 0.5 } produces about 50% drops and about 25% rate limits (0.5 × 0.5), not 50% of each — to exercise two faults at their nominal frequency, configure them on separate requests.

One exception: the per-test scope. The server level is scoped by X-Test-Id — a runtime override installed with POST /__aimock/chaos applies only to traffic carrying the same tag (see the Control API). Picking that scope is a selection, not a merge: a scoped override replaces the server-wide baseline wholesale rather than layering on it. A server started with --chaos-latency 500 whose test installs { "dropRate": 1 } sees drops with no latency on that test's traffic — restate latencyMs in the override to keep it. GET /__aimock/chaos with the same tag always reports the config actually in effect for the scope, and the server warns when an install drops a field that was in effect.

Quick Start

chaos-quick-start.ts ts
import { LLMock } from "@copilotkit/aimock";

const mock = new LLMock();
mock.onMessage("hello", { content: "Hi!" });

// 50% of all requests will be dropped with a 500
mock.setChaos({ dropRate: 0.5 });

await mock.start();

// Later, remove chaos
mock.clearChaos();

Programmatic API

Programmatic chaos control ts
// Set server-level chaos (returns `this` for chaining)
mock.setChaos({
  dropRate: 0.1,        // 10% drop rate
  malformedRate: 0.05,  // 5% malformed rate
  rateLimitRate: 0.05,  // 5% rate-limit (429) rate
  disconnectRate: 0.02, // 2% disconnect rate
  latencyMs: 250,       // always delay 250 ms first
});

// Remove all server-level chaos
mock.clearChaos();

Fixture-Level Chaos

Attach a chaos config to individual fixtures so only specific responses experience failures:

chaos-fixture.json json
{
  "fixtures": [
    {
      "match": { "userMessage": "unstable" },
      "response": { "content": "This might fail!" },
      "chaos": {
        "dropRate": 0.3,
        "malformedRate": 0.2,
        "disconnectRate": 0.1
      }
    },
    {
      "match": { "userMessage": "stable" },
      "response": { "content": "This always works." }
    }
  ]
}

Per-Request Headers

Override chaos on individual requests using HTTP headers — one header per field, each carrying the same value the corresponding chaos field takes:

Header Controls Accepted values
x-aimock-chaos-drop dropRate — 500 with an error body 0–1
x-aimock-chaos-malformed malformedRate — truncated/invalid JSON 0–1
x-aimock-chaos-disconnect disconnectRate — socket destroyed mid-flight 0–1
x-aimock-chaos-ratelimit rateLimitRate — 429 with Retry-After 0–1
x-aimock-chaos-latency latencyMs — delay before the response 0–30000, whole ms

Header values are read as plain decimal text and are never clamped: a value that is out of range, not a plain decimal (1e3, 0x1f, Infinity), or fractional where a whole number of milliseconds is required is rejected with a warning and simply does not apply — the level below it (fixture, then server default) still does. The same rule governs values set in a fixture, in the server defaults, on the --chaos-* CLI flags and through POST /__aimock/chaos.

Per-request chaos via headers ts
// Force 100% disconnect on this specific request
await fetch(`${mock.url}/v1/chat/completions`, {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "x-aimock-chaos-disconnect": "1.0",
  },
  body: JSON.stringify({ model: "gpt-4", messages: [{ role: "user", content: "hello" }] }),
});

CLI Flags

Set server-level chaos from the command line. The --chaos-* flags belong to the llmock binary (the fixtures-and-flags CLI), which is also what the Docker image runs.

llmock chaos flags shell
$ npx -p @copilotkit/aimock llmock --fixtures ./fixtures \
  --chaos-drop 0.1 \
  --chaos-malformed 0.05 \
  --chaos-ratelimit 0.05 \
  --chaos-disconnect 0.02 \
  --chaos-latency 250
CLI chaos flags shell
$ docker run -d -p 4010:4010 \
  -v ./fixtures:/fixtures \
  ghcr.io/copilotkit/aimock \
  -f /fixtures -h 0.0.0.0 \
  --chaos-drop 0.1 \
  --chaos-malformed 0.05 \
  --chaos-ratelimit 0.05 \
  --chaos-disconnect 0.02 \
  --chaos-latency 250

The aimock binary takes no chaos flags — its options are --config, --port, --host and -h, --help. With aimock --config, set the server-level chaos under llm.chaos in the config file (the same five fields, the same bounds), or install it at runtime with POST /__aimock/chaos (see the Control API).

Proxy Mode

When aimock is configured as a record/replay proxy (--record), chaos applies to proxied requests too — so a staging environment pointed at real upstream APIs still sees the failure modes your tests expect. Chaos is rolled once per request, after fixture matching, with the same headers > fixture > server precedence.

Mode When upstream is contacted What the client sees
drop Never — upstream not contacted HTTP 500 chaos body; upstream is not called
rateLimit Never — upstream not contacted HTTP 429 chaos body; upstream is not called
disconnect Never — upstream not contacted Connection destroyed; upstream is not called
malformed Called — post-response Request proxies normally; the upstream response is captured, then the body is replaced with invalid JSON before relay. The recorded fixture (if recording) keeps the real upstream response — chaos is a live-traffic decoration, not a fixture mutation.

SSE bypass. If upstream returns Content-Type: text/event-stream, aimock streams chunks to the client progressively. By the time malformed would fire, the bytes are already on the wire — the chaos action cannot be applied. This bypass is observable via the aimock_chaos_bypassed_total counter (see Prometheus Metrics below) and a warning in the server log, so a configured chaos rate doesn't silently drop to 0% on SSE traffic. Streaming mutation is planned for a future phase.

Journal Tracking

When chaos triggers, the journal entry includes a chaosAction field recording which failure mode was applied:

Journal entry with chaos json
{
  "method": "POST",
  "path": "/v1/chat/completions",
  "response": {
    "status": 500,
    "source": "fixture",
    "fixture": { "...": "elided for brevity" },
    "chaosAction": "drop"
  }
}

The chaosAction values are "drop", "malformed", "rateLimit" and "disconnect". The recorded status codes are 500 for drop, 200 for malformed, 429 for rateLimit, and 0 for disconnect (connection destroyed). A latency delay alone leaves no chaosAction: the entry is the ordinary one for whatever the request then received.

Prometheus Metrics

When metrics are enabled (--metrics), each chaos trigger increments the aimock_chaos_triggered_total counter, tagged with action and source. source="fixture" means a fixture matched (or would have, before chaos intervened); source="proxy" means the request was on the proxy dispatch path.

Metrics output text
# TYPE aimock_chaos_triggered_total counter
aimock_chaos_triggered_total{action="drop",source="fixture"} 3
aimock_chaos_triggered_total{action="malformed",source="fixture"} 1
aimock_chaos_triggered_total{action="disconnect",source="proxy"} 2

When a chaos action is rolled but can't be applied — today, only malformed on an SSE proxy response — the bypass is recorded in a separate counter so operators can distinguish "chaos didn't roll" from "chaos rolled but was bypassed":

Bypass counter text
# TYPE aimock_chaos_bypassed_total counter
aimock_chaos_bypassed_total{action="malformed",source="proxy",reason="sse_streamed"} 4