Record & Replay
VCR-style record-and-replay support. When a request doesn't match any fixture, aimock proxies it to the real upstream provider, records the response as a fixture on disk and in memory, then replays it on subsequent identical requests.
How It Works
- Client sends a request to aimock
- aimock attempts fixture matching as usual
- On miss: the request is forwarded to the configured upstream provider
- The upstream response is relayed back to the client immediately
- The response is collapsed (if streaming) and saved as a fixture to disk and memory
- Subsequent identical requests match the newly recorded fixture
Proxy-Only Mode
Use --proxy-only instead of --record when you want unmatched
requests to always reach the real provider — no fixture files are written to disk
and no responses are cached in memory. Matched fixtures still work normally.
This is ideal for demos and live environments where you have canned
fixtures for repeatable demo scenarios you want to show off, but also want regular
interactions to work normally by proxying to the real provider. Without
--proxy-only, the first real API call would get recorded and cached, and
subsequent identical requests would get the stale recorded response instead of hitting the
live provider.
$ npx -p @copilotkit/aimock llmock -f ./fixtures \
--proxy-only \
--provider-openai https://api.openai.com
$ docker run -d -p 4010:4010 \
-v $(pwd)/fixtures:/fixtures \
ghcr.io/copilotkit/aimock \
-f /fixtures -h 0.0.0.0 \
--proxy-only \
--provider-openai https://api.openai.com
| Mode | Unmatched request | Writes to disk | Caches in memory |
|---|---|---|---|
--record |
Proxy → save → replay next time | Yes | Yes |
--proxy-only |
Proxy → relay → proxy again next time | No | No |
Quick Start
$ npx -p @copilotkit/aimock llmock -f ./fixtures \
--record \
--provider-openai https://api.openai.com \
--provider-anthropic https://api.anthropic.com
$ docker run -d -p 4010:4010 \
-v $(pwd)/fixtures:/fixtures \
ghcr.io/copilotkit/aimock \
-f /fixtures -h 0.0.0.0 \
--record \
--provider-openai https://api.openai.com \
--provider-anthropic https://api.anthropic.com
Record & Replay CLI Flags
| Flag | Description |
|---|---|
-f, --fixtures <path> |
Path to fixtures directory or file (default: ./fixtures) |
-p, --port <port> |
Port to listen on (default: 4010) |
-h, --host <host> |
Host to bind to (default: 127.0.0.1) |
--record |
Enable record mode (proxy, save, and cache on miss) |
--proxy-only |
Proxy mode (forward on miss, no saving or caching) |
--strict |
Strict mode: return 503 (not 404) on unmatched requests |
-w, --watch |
Watch fixture path for changes and reload |
--log-level <level> |
Log verbosity: silent, warn, info,
debug (default: info)
|
--validate-on-load |
Validate fixture schemas at startup |
--journal-max <n> |
Max request entries retained in memory (default: 1000 for both
serve and createServer() since v1.14.2; 0 =
unbounded; direct new Journal() instantiation still defaults to
unbounded for back-compat)
|
--fixture-counts-max <n> |
Max unique testIds retained in the fixture match-count map (default:
500; 0 = unbounded)
|
--provider-openai <url> |
Upstream URL for OpenAI |
--provider-anthropic <url> |
Upstream URL for Anthropic |
--provider-gemini <url> |
Upstream URL for Gemini |
--provider-vertexai <url> |
Upstream URL for Vertex AI |
--provider-bedrock <url> |
Upstream URL for Bedrock |
--provider-azure <url> |
Upstream URL for Azure OpenAI |
--provider-ollama <url> |
Upstream URL for Ollama |
--provider-cohere <url> |
Upstream URL for Cohere |
--provider-openrouter <url> |
Upstream URL for OpenRouter — the chat/router and video lifecycle record proxies |
--upstream-timeout-ms <ms> |
Connection idle timeout (ms) on the upstream request socket before the response body
begins (default: 30000). Increase for upstreams with slow initial
responses (reasoning models, queue-backed providers). Nuance: the OpenRouter video
surface's small-JSON lifecycle fetches (submit, status poll, models listing) apply
this value as a total deadline rather than a socket-idle timeout —
indistinguishable for envelope-sized bodies; the byte-bearing content fetches gate
only the response headers on it.
|
--body-timeout-ms <ms> |
Inter-chunk idle timeout (ms) on the upstream response body — fires if no
bytes arrive for this duration after the response has started streaming (default:
30000). Reasoning models under concurrent load can leave 30s+ gaps
between chunks; increase to e.g. 180000 in those setups.
|
--agui-record |
Enable AG-UI recording (proxy unmatched AG-UI requests) |
--agui-proxy-only |
AG-UI proxy mode (forward on miss, no saving or caching) |
--agui-upstream <url> |
Upstream AG-UI agent URL (used with --agui-record /
--agui-proxy-only)
|
--chaos-drop <rate> |
Probability (0–1) of dropping requests with 500
|
--chaos-malformed <rate> |
Probability (0–1) of returning malformed JSON |
--chaos-disconnect <rate> |
Probability (0–1) of destroying the connection
mid-stream
|
Programmatic API
import { LLMock } from "@copilotkit/aimock";
const mock = new LLMock();
await mock.start();
try {
// Enable recording — unmatched requests are proxied AND saved as fixtures
mock.enableRecording({
providers: {
openai: "https://api.openai.com",
anthropic: "https://api.anthropic.com",
},
fixturePath: "./fixtures/recorded",
});
// Make requests — unmatched ones are proxied and recorded
// ...
// Disable recording — recorded fixtures persist on disk
mock.disableRecording();
} finally {
// Always release the port, even if a test above threw
await mock.stop();
}
To proxy unmatched requests without writing fixtures to disk, set
proxyOnly: true and omit fixturePath. Despite the method name,
enableRecording({ proxyOnly: true }) does not write fixtures
— omit or set proxyOnly: false to actually record. Remember to
mock.stop() when done:
try {
mock.enableRecording({
providers: {
openai: "https://api.openai.com",
anthropic: "https://api.anthropic.com",
},
proxyOnly: true,
});
// ...make requests; unmatched ones proxy through without being saved
} finally {
await mock.stop();
}
GPT-Live Recording
GPT-Live uses the primary WebSocket endpoint /v1/live/sessions. Unlike HTTP
streams, Live recordings preserve both directions, event order, PCM audio, timing, and
identifier references. They do not collapse into a text response. Both client and managed
delegation use this endpoint.
Connect to ws://127.0.0.1:4010/v1/live/sessions for a default local server.
Send session.start first. Use the configuration and command sequence in the
WebSocket guide. The supported audio format is mono PCM16LE at
24,000 Hz. aimock replays supplied audio. It does not synthesize speech.
import { LLMock } from "@copilotkit/aimock";
const providerKey = process.env.OPENAI_API_KEY;
if (!providerKey) throw new Error("Set OPENAI_API_KEY before recording");
const mock = new LLMock({
live: { maxDurationMs: 120_000 },
});
await mock.start();
try {
mock.enableRecording({
providers: { openai: "https://api.openai.com" },
providerKeys: { openai: providerKey },
fixturePath: "./fixtures/live-recorded",
upstreamTimeoutMs: 30_000,
bodyTimeoutMs: 30_000,
});
const liveUrl = new URL("/v1/live/sessions", mock.url);
liveUrl.protocol = "ws:";
// Connect your Live client to liveUrl.href and complete one session.
// Await session.closed and successful fixture export before stopping.
} finally {
await mock.stop();
}
This fragment configures recording. Add your awaited client scenario inside the
try block. Only a complete, successful session.closed capture
becomes a fixture. A disconnect, timeout, failed export, or resource limit does not
produce a partial replay fixture.
The upstream base accepts https://api.openai.com or
https://api.openai.com/v1. It cannot contain credentials, a query, or a
fragment. aimock uses an upstream WebSocket connection and forwards only the selected
Authorization header. POST SDP, WebRTC, sideband connections, session forks, and recording
downloads are outside this support.
Local and Provider Credentials
auth.apiKeys controls access to aimock. It does not grant provider access. A
matching local key in an Authorization Bearer or Key header is removed before the upstream
request. Configure a separate record.providerKeys.openai value for provider
access.
After local-key removal, a caller credential takes precedence unless it starts with
sk-aimock-. If the caller credential is absent or has that dummy prefix,
aimock selects the configured provider key. AIMOCK_DUMMY_KEY_MARKER changes
the prefix. Without a configured provider key, the remaining caller header stays
unchanged.
The CLI reads AIMOCK_PROVIDER_OPENAI_KEY for the provider key. Programmatic
code supplies providerKeys.openai explicitly, as in this example.
OPENAI_API_KEY is read by the example, not automatically by
new LLMock(). Offline replay needs no provider key or recording
configuration.
Retained Content and Export Safety
Recording explicitly retains input and output audio, transcripts, prompts, tool arguments, tool results, and other semantic event content. These can contain personal or confidential information. Use audio and text that you are permitted to retain and share.
Before export, aimock removes known credential metadata and replaces opaque identifiers with stable fixture identifiers. Replay creates fresh server identifiers for each connection while preserving their references. Sanitization preserves semantic content and audio bytes.
Add known secret strings to live.secretValues. The recorder also checks the
selected upstream credential and configured local keys. If a secret occurs in semantic
content or cannot be removed safely, export fails with unsafe-export. There
is no unsafe-export override. A sanitized fixture is not a general personal-data scrub.
Recording journal entries retain status, failure category, and model metadata rather than the raw capture or caller headers. Replay diagnostics can contain fixture content. Configured literal-secret filtering does not identify arbitrary secrets or personal data. Review retained fixtures and diagnostics before sharing them.
With proxyOnly: true, unmatched Live sessions still reach the provider.
aimock does not save or cache a replay fixture. Temporary bounded capture state exists
during the session and is cleared at completion. Provider charges still apply.
Live Limits and Deadlines
Set these values in new LLMock({ live: { ... } }). All numeric Live limits
must be positive safe integers. Zero does not disable a limit. One MiB equals 1,048,576
bytes. Limits apply to each session unless the table states otherwise.
| Live option | Default | Scope |
|---|---|---|
maxSessions |
16 | Concurrent Live sessions per server, including pending startup. |
maxMessageBytes |
1 MiB | Each complete JSON message in either direction. |
maxBufferedBytes |
2 MiB | Each transport receive buffer, including fragmented input. Also bounds queued startup commands. |
maxWriteBytes |
1 MiB | Pending encoded writes on each connection, in either relay direction. |
maxDecodedAudioBytes |
16 MiB | Cumulative decoded PCM bytes, separately for client input and server output. |
maxDurationMs |
120,000 ms | Absolute session lifetime, including startup. Also bounds transcript timestamps. |
idleTimeoutMs |
30,000 ms | Replay inactivity and startup when recording is disabled. |
mismatchTimeoutMs |
30,000 ms | Replay wait for the next required client command or PCM prefix. |
Recording uses the following enableRecording() configuration instead of
replay idle and mismatch deadlines. The absolute live.maxDurationMs limit
remains active. Increase that limit too when you need a longer recording.
| Record option | Default | Live behavior |
|---|---|---|
upstreamTimeoutMs |
30,000 ms | Startup deadline and upstream connection/upgrade deadline. |
bodyTimeoutMs |
30,000 ms | Upstream byte-idle deadline after the upgrade. Partial frame bytes count as activity. |
maxProxyBufferBytes |
64 MiB | Combined serialized event bytes from both directions. Hard ceiling: 256 MiB. Pending startup uses the smaller of this value and maxBufferedBytes. |
maxProxyBufferFrames |
5,000,000 | Combined event count from both directions. Also bounds queued startup commands. |
Live recording closes the session when a capture limit is exceeded. It does not continue an unlimited relay after capture stops. Export and fixture loading also enforce a 64 MiB JSON limit, depth 32, and at most 5,000,000 array entries. A larger capture allowance does not remove those validation limits.
Restart and Replay Offline
After a successful export, start a fresh server with the saved directory. Do not enable
recording on this server. This also works with either authored example file:
fixtures/openai-live-client.json or
fixtures/openai-live-managed.json. Load one example mode per server because
both match the same model and endpoint.
import { LLMock } from "@copilotkit/aimock";
const mock = new LLMock();
mock.loadFixtureDir("./fixtures/live-recorded");
await mock.start();
try {
const liveUrl = new URL("/v1/live/sessions", mock.url);
liveUrl.protocol = "ws:";
// Await your client scenario against liveUrl.href here.
} finally {
await mock.stop();
}
Replay expects the recorded configuration and client commands. Use identifiers from
emitted events instead of copied fixture identifiers. In managed mode, send every pending
tool result before the explicit response.create continuation. Equivalent PCM
samples can use different chunk boundaries. Command order and required audio prefixes
still control when server events become eligible.
Saved recordings use recorded timing. liveTiming: "immediate" removes timed
waits but preserves client barriers. A provider access refusal or unavailable runtime does
not establish recording coverage. Offline authored examples establish replay behavior, not
live-provider compatibility.
Test Cleanup
Always close client sockets and await mock.stop() in a
finally block. To reuse a server, call mock.closeLiveSessions()
between tests. Pass a test ID to close only sessions assigned to that test.
The @copilotkit/aimock/vitest and
@copilotkit/aimock/jest helpers close their server's Live sessions before and
after each test. They reset match counts before each test.
resetMatchCounts() alone does not close sessions. A full
mock.reset() closes active sessions and clears fixtures.
Stream Collapsing
When the upstream provider returns a streaming response, aimock collapses it into a non-streaming fixture. Six streaming formats are supported:
| Format | Provider | Content-Type |
|---|---|---|
| OpenAI SSE | OpenAI, Azure, OpenRouter | text/event-stream |
| Anthropic SSE | Anthropic | text/event-stream |
| Gemini SSE | Gemini, Vertex AI | text/event-stream |
| Cohere SSE | Cohere | text/event-stream |
| Ollama NDJSON | Ollama | application/x-ndjson |
| Bedrock EventStream | AWS Bedrock | application/vnd.amazon.eventstream |
The collapse extracts text content and tool calls from streaming chunks and produces a
fixture response such as { content }, { toolCalls }, or
{ content, toolCalls }. For OpenAI-compatible streams it also captures the
provider-reported token usage from the final usage frame and attaches it to
whichever of those shapes results — usage is recorded independently of
whether the response carries content, tool calls, or both. Genuinely tool-first or
interleaved streams additionally gain an ordered
blocks array. See
Recording Block Order and
Recording Token Usage & Cost below.
Recording Block Order
When a recorded stream is genuinely tool-first or interleaved — a tool-call
delta arrives before the first content delta, or content arrives after a tool-call delta
— the collapser preserves that arrival order as a
blocks array on the fixture. This
works across OpenAI, Anthropic, Gemini, Ollama, Cohere, and Bedrock. Ordinary
text-then-tools streams are saved in the legacy { content, toolCalls } shape
with no blocks key, so existing recordings round-trip byte-identically.
Gemini Interactions is the exception: its record-side collapser
normalizes tool-call arguments only and does not reorder blocks on capture. Ordering is
still honored on replay from a hand-authored blocks fixture; it is simply not
reconstructed automatically from a recording. See the
per-provider observability matrix for how
faithfully block order is reconstructable on each provider's wire.
Recording Token Usage & Cost
OpenAI-compatible recordings (including OpenRouter) keep
the provider's reported usage on the fixture. For a non-streaming call that
is the envelope's usage object; for a stream it is the final
chat.completion.chunk — the one with an empty choices array —
that OpenAI emits when the request sets
stream_options: { include_usage: true } (OpenRouter emits it either way).
{
"match": { "userMessage": "summarize this" },
"response": {
"content": "…",
"usage": {
"prompt_tokens": 1234,
"completion_tokens": 567,
"total_tokens": 1801,
"cost": 0.0042
}
}
}
On replay those counts are served verbatim instead of aimock's
ceil(length / 4) estimate. The usage.cost shown above is
OpenRouter-shaped output: it is re-emitted only when the fixture is replayed through
aimock's OpenRouter endpoint
(/api/v1/chat/completions), so a test can assert a wallet or ledger deduction
against the amount the provider actually charged. On a plain OpenAI (/v1/…)
replay the cost key is still recorded and validated on the fixture but is
not served back (token counts round-trip while cost stays inert), so
copy this example into an OpenRouter fixture if you need the cost to appear. Extra
provider fields (cost_details, prompt_tokens_details,
completion_tokens_details, native_tokens_*, …) round-trip too.
Record with usage enabled to get cost. If the recorded request did not
ask for usage — and the provider therefore never sent a usage frame — the fixture is
written without a usage key and replay falls back to estimated token counts,
exactly as before. You can always hand-author response.usage on a fixture
instead.
Header Forwarding
Live forwards only the selected Authorization header, as described in GPT-Live Recording.
When proxying HTTP requests to upstream providers, aimock forwards the original request's
headers except for a small, fixed strip list: hop-by-hop headers (per RFC 2616 §13.5.1),
headers the upstream HTTP client must set from the target URL or body, and aimock's own
mock-internal control headers. Everything else passes through — custom gateway headers,
openai-organization, anthropic-version, and any other
provider-specific auth or routing header all reach the upstream as-is.
The following headers are stripped before proxying:
-
Hop-by-hop (RFC 2616 §13.5.1):
connection,keep-alive,transfer-encoding,te,trailer,upgrade,proxy-authorization,proxy-authenticate -
Set by the HTTP client from the target URL / body:
host,content-length -
Not relevant for LLM APIs (avoid leaking or mismatched encoding):
cookie,accept-encoding -
Mock-internal control headers (meaningless — and potentially confusing or leaky — on
a real provider's wire):
x-test-id,x-aimock-strict,x-aimock-context, and thex-aimock-chaos-*prefix family
Auth headers are never saved in recorded fixtures. The fixture contains
the match criteria (derived from the last user message) and the response content — plus,
when applicable, a metadata block (drift-detection hashes, below) and a
recordedTimings block (streaming frame timings) — never the request headers.
aimock-Owned Upstream Keys
By default, aimock forwards the caller's auth header to the upstream provider unchanged. If your tests can only send a dummy placeholder key — for example, an SDK that refuses to start without a non-empty API key — aimock can inject its own configured upstream key on a fixture-miss passthrough so the recorded/proxied call actually authenticates. This is opt-in and backward-compatible: with no key configured the feature is fully inert and the caller's header is forwarded as-is.
Set one or more of these env vars to aimock's real upstream keys. Each is independent — configure only the providers you record against:
| Env var | Provider | Injected header |
|---|---|---|
AIMOCK_PROVIDER_OPENAI_KEY |
OpenAI | Authorization: Bearer <key> |
AIMOCK_PROVIDER_OPENROUTER_KEY |
OpenRouter | Authorization: Bearer <key> |
AIMOCK_PROVIDER_COHERE_KEY |
Cohere | Authorization: Bearer <key> |
AIMOCK_PROVIDER_GROK_KEY |
Grok (xAI) | Authorization: Bearer <key> |
AIMOCK_PROVIDER_OLLAMA_KEY |
Ollama (Cloud / bearer-gated) | Authorization: Bearer <key> |
AIMOCK_PROVIDER_ANTHROPIC_KEY |
Anthropic | x-api-key: <key> |
AIMOCK_PROVIDER_GEMINI_KEY |
Gemini (and Gemini Interactions) | x-goog-api-key: <key> |
AIMOCK_PROVIDER_VEO_KEY |
Veo | x-goog-api-key: <key> |
AIMOCK_PROVIDER_AZURE_KEY |
Azure OpenAI | api-key: <key> |
AIMOCK_PROVIDER_ELEVENLABS_KEY |
ElevenLabs | xi-api-key: <key> |
AIMOCK_PROVIDER_FAL_KEY |
fal.ai | Authorization: Key <key> |
gemini-interactions reuses AIMOCK_PROVIDER_GEMINI_KEY (same
upstream API as Gemini). An empty-string value is treated as unset.
Injection fires only when both of these hold:
-
The request is a fixture-miss passthrough (record mode or
--proxy-only). -
The caller sent no credential for that provider,
or sent a dummy credential prefixed with
sk-aimock-. A caller credential that does not start with the marker is treated as a real key and forwarded verbatim — the caller always overrides aimock.
The dummy marker prefix is overridable via AIMOCK_DUMMY_KEY_MARKER for setups
that mint placeholder keys under a different prefix. Gemini is injected as an
x-goog-api-key header only (no query-param ?key= rewrite).
Signed and exchanged credentials are never rewritten. AWS Bedrock (SigV4), Vertex AI, and Azure AD carry OAuth/signed auth rather than a simple bearer or api-key header, so aimock always forwards their credentials unchanged regardless of these env vars.
Strict Mode
When --strict is enabled, an unmatched HTTP request returns
503 Service Unavailable before any proxy attempt — even
when an upstream is configured for that provider, nothing is forwarded and nothing is
recorded. Strict mode wins over record mode. This is useful for CI environments where you
want to catch unexpected API calls instead of silently recording new fixtures.
Strict mode also prevents Live fixture misses from reaching the provider. After the WebSocket upgrade, Live reports an aimock error and closes the session instead of sending an HTTP 503 response.
Strict mode does not change how MCP fake calls are answered. To fail on a tool that a
scenario does not declare, use "undeclaredTools": "deny" in its
mcpFakes block (see Closed world). An
mcpFakes block that aimock cannot honor fails the load in every mode, strict
or not, and with or without --validate-on-load.
The CLI check for a start with no fixtures counts mcpFakes blocks as loaded.
A start with only fakes does not abort under --strict or
--validate-on-load. It prints
Warning: No LLM fixtures loaded; LLM requests will return 404 and serves the
fakes.
Fixture Auto-Generation
Recorded fixtures are saved to disk with timestamped filenames:
// fixtures/recorded/openai-<YYYY-MM-DD>T<HH-MM-SS>-<ms>Z-<8-hex>.json (e.g. openai-2026-05-18T10-30-00-000Z-a1b2c3d4.json)
{
"fixtures": [
{
"match": {
"userMessage": "What is the weather?",
"model": "gpt-4o",
"turnIndex": 0,
"hasToolResult": false
},
"metadata": { "systemHash": "a7f3c291" },
"response": { "content": "I don't have real-time weather data..." }
}
]
}
Match criteria are derived from the original request: the last user message becomes
userMessage (or, for embedding requests, the input becomes
inputText), the normalized model becomes model, and for chat
requests the recorder also writes the multi-turn disambiguators
turnIndex (the number of assistant messages already in the conversation) and
hasToolResult (whether the current turn — the messages after the last
user message — contains a tool result). If no match criteria can be derived (e.g.,
empty messages), the fixture is saved to disk with a warning but not registered in memory.
Model-Aware Recording
When recording fixtures, aimock automatically includes the model name in match criteria. This prevents collisions when your app makes multiple LLM calls with the same user message but different models (e.g., Opus for chat + Haiku for title generation).
Model names are normalized by stripping date/version suffixes so fixtures survive provider version bumps:
| Request Model | Recorded As |
|---|---|
claude-opus-4-20250514 |
claude-opus-4 |
gpt-4o-2024-08-06 |
gpt-4o |
claude-3-5-sonnet-20241022 |
claude-3-5-sonnet |
llama3.1 |
llama3.1 (no date suffix — unchanged) |
Matching uses prefix comparison, so model: "claude-opus-4" in a fixture
matches requests for claude-opus-4-20250514,
claude-opus-4-20250915, or any future version.
To record the full model version instead (disabling normalization), set
recordFullModelVersion to true in the recording config:
{
"llm": {
"record": {
"providers": { "openai": "https://api.openai.com" },
"recordFullModelVersion": true
}
}
}
Or programmatically:
mock.enableRecording({
providers: { openai: "https://api.openai.com" },
fixturePath: "./fixtures/recorded",
recordFullModelVersion: true,
});
Context-Aware Recording
When a request carries an X-AIMock-Context header, the recorder automatically
captures the context value in match.context. On replay, fixtures with
context only match requests carrying that exact header value — fixtures
without context remain shared across all callers.
Directory routing
Without snapshot-style recording (X-Test-Id), recorded fixtures for a given
context are written to a <fixturePath>/<context>/ subdirectory:
fixtures/recorded/
openai-2026-05-18T10-30-00-000Z-a1b2c3d4.json # no context (shared)
langgraph-python/
openai-2026-05-18T10-30-01-000Z-e5f6a7b8.json # context = langgraph-python
crewai/
openai-2026-05-18T10-30-02-000Z-c9d0e1f2.json # context = crewai
When X-Test-Id is also present, snapshot-style paths take precedence and the
context is captured only in match.context within the fixture file, not in the
directory structure.
Sending X-AIMock-Context
// Set as a default header on your LLM client
const client = new OpenAI({
baseURL: "http://localhost:4010/v1",
apiKey: "mock",
defaultHeaders: { "X-AIMock-Context": "langgraph-python" },
});
Upstream Timeouts
By default, aimock aborts a proxied request if the upstream socket is idle for 30 seconds
(before the response body) or if no bytes arrive for 30 seconds during streaming.
Reasoning models under concurrent load can exceed these limits during the thinking phase.
Override with upstreamTimeoutMs and bodyTimeoutMs:
mock.enableRecording({
providers: { openai: "https://api.openai.com" },
fixturePath: "./fixtures/recorded",
bodyTimeoutMs: 180_000,
});
Or via CLI: --body-timeout-ms 180000. Values must be positive finite numbers;
zero, negative, NaN, and Infinity are rejected (CLI exits
non-zero; programmatic API falls back to the 30s default).
Drift Detection Metadata
Recorded fixtures include a metadata block with hashes of the system prompt
and tool definitions at recording time. These are informational only — not used for
matching — and help you detect when your prompts or tools have changed since the
fixture was recorded.
{
"match": { "userMessage": "hello", "model": "claude-opus-4" },
"metadata": { "systemHash": "a7f3c291", "toolsHash": "e4b12d08" },
"response": { "content": "Hi there!" }
}
When you re-record a fixture and the hashes differ from the previous version, it signals that your application’s prompts or tool definitions have evolved. This is useful for auditing fixture freshness — if the hashes don’t match, the recorded response may no longer reflect what the real provider would return for the current prompt.
Snapshot-Style Recording
When the X-Test-Id header is present on a request, aimock uses
snapshot-style recording instead of the default timestamp-based
filenames. Fixtures are organized by test, producing stable file paths that work well with
version control and PR diffs.
Directory structure
The test ID is slugified into a directory name, and each provider gets its own file within that directory:
fixtures/recorded/
agent-chat--handles-tool-call/
openai.json # All OpenAI fixtures for this test
anthropic.json # All Anthropic fixtures for this test
simple-test/
openai.json
The slugify rules: Common test file prefixes (.spec.ts,
.test.tsx, .e2e.js, etc.) are automatically stripped from the
test ID before slugifying, so my-app.spec.ts › greeting becomes
greeting. Then Playwright's › separator becomes
--, non-word characters become -, runs of 3+ dashes collapse to
--, and the result is lowercased. For example,
"agent chat › handles tool call" becomes
agent-chat--handles-tool-call.
Merge behavior on re-run
When you re-run a test, the new fixture is appended to the existing
<provider>.json file rather than overwriting it. This preserves
multi-turn conversations in a single file. If the existing file is corrupted (invalid
JSON, or a non-array fixtures field), it is replaced with a logged warning.
Sending X-Test-Id from test frameworks
// Playwright exposes testInfo.titlePath which joins suite + test titles
import { test } from "@playwright/test";
test("handles tool call", async ({ page }, testInfo) => {
// titlePath = ["agent chat", "handles tool call"]
const testId = testInfo.titlePath.join(" › ");
// Set on your OpenAI/Anthropic client config as a default header:
// headers: { "X-Test-Id": testId }
});
import { describe, it } from "vitest";
describe("agent chat", () => {
it("handles tool call", async () => {
// Pass X-Test-Id on each LLM request:
const resp = await fetch("http://localhost:4010/v1/chat/completions", {
headers: { "X-Test-Id": "agent chat › handles tool call" },
// ...body
});
});
});
Fallback behavior
When no X-Test-Id header is present (or the value is
__default__), recording falls back to the standard timestamp-based filename:
<provider>-<timestamp>-<8-hex>.json (the suffix is the
first 8 hex characters of a random UUID).
Fixture Lifecycle
-
On disk: Fixtures persist in the configured
fixturePathdirectory (default:./fixtures/recorded) - In memory: Recorded fixtures are immediately available for matching subsequent requests in the same session
- After restart: Load the recorded fixture directory to replay previous recordings
Local Development Workflow
Record once against real APIs, then replay from fixtures for fast, offline development.
# First run: record real API responses
$ npx -p @copilotkit/aimock llmock --record --provider-openai https://api.openai.com -f ./fixtures
# Subsequent runs: replay from recorded fixtures
$ npx -p @copilotkit/aimock llmock -f ./fixtures
# First run: record real API responses
$ docker run -d -p 4010:4010 \
-v $(pwd)/fixtures:/fixtures \
ghcr.io/copilotkit/aimock \
--record --provider-openai https://api.openai.com -f /fixtures -h 0.0.0.0
# Subsequent runs: replay from recorded fixtures
$ docker run -d -p 4010:4010 \
-v $(pwd)/fixtures:/fixtures \
ghcr.io/copilotkit/aimock \
-f /fixtures -h 0.0.0.0
CI Pipeline Workflow
Use the Docker image in CI with --strict mode to ensure every request matches
a recorded fixture. No API keys needed, no flaky network calls.
- name: Start aimock
run: |
docker run -d --rm --name aimock \
-v $(pwd)/fixtures:/fixtures \
-p 4010:4010 \
ghcr.io/copilotkit/aimock \
--strict -f /fixtures -h 0.0.0.0
- name: Run tests
env:
OPENAI_BASE_URL: http://localhost:4010/v1
run: pnpm test
- name: Stop aimock
if: always()
run: docker rm -f aimock
Request Transform
Prompts often contain dynamic data — timestamps, UUIDs, session IDs — that
changes between runs. This causes fixture mismatches on replay because the recorded key no
longer matches the live request. The requestTransform option normalizes
requests before both matching and recording, stripping out the volatile parts.
import { LLMock } from "@copilotkit/aimock";
const mock = new LLMock({
requestTransform: (req) => ({
...req,
messages: req.messages.map((m) => ({
...m,
content:
typeof m.content === "string"
? m.content.replace(/\d{4}-\d{2}-\d{2}T[\d:.+Z-]+/g, "").trim()
: m.content,
})),
}),
});
// Fixture uses the cleaned key (no timestamp)
mock.onMessage("tell me the weather", { content: "Sunny" });
// Request with a timestamp still matches after transform
await mock.start();
When requestTransform is set, string matching for userMessage,
systemMessage, and inputText switches from substring
(includes) to exact equality (===). This prevents shortened keys
from accidentally matching unrelated prompts. Without a transform, the existing
includes behavior is preserved for backward compatibility.
The transform is applied in both directions: recording saves the transformed match key (no timestamps in the fixture file), and matching transforms the incoming request before comparison. This means recorded fixtures and live requests always use the same normalized key.
Building Fixture Sets
A practical workflow for building and maintaining fixture sets:
- Run with
--recordagainst real APIs during development - Review recorded fixtures in
fixtures/recorded/ - Move and rename to organized fixture directories
- Switch to
--strictmode in CI - Re-record when upstream APIs change (drift detection catches this)
Recording Multi-Turn Conversations
The recorder is stateless across turns. Every incoming request is treated
as an independent unit: aimock derives fixture match criteria from a
single request at a time, and it doesn’t remember or hash prior turns of
the same test session. For chat requests the derived key combines the
last user message (match.userMessage), the normalized
model, and two shape disambiguators read off the request itself:
turnIndex (how many assistant messages the conversation already contains) and
hasToolResult (whether the current turn — the messages after the last
user message — contains a tool result). Embedding requests key on the input text
instead. There is no history-based fingerprinting of the full request body — but
because a growing conversation carries a growing turnIndex, successive turns
that happen to end with the same user message still record as distinct fixtures.
// src/recorder.ts (simplified; real implementation guards against a null last-user-message)
function buildFixtureMatch(request) {
if (request.embeddingInput) {
return { inputText: request.embeddingInput };
}
// Chat — key on the LAST user message, the normalized model, and the
// multi-turn disambiguators derived from the request shape
const lastUser = getLastMessageByRole(request.messages, "user");
const match = {
userMessage: getTextContent(lastUser.content),
model: normalizeModelName(request.model),
turnIndex: request.messages.filter((m) => m.role === "assistant").length,
// Scoped to the current turn (messages after the last user message), so it
// shares the exact predicate the matcher uses — see router.ts.
hasToolResult: currentTurnHasToolResult(request.messages),
};
// Capture context from X-AIMock-Context header if present
if (request._context) match.context = request._context;
return match;
}
What still collides: byte-identical repeats
Because turnIndex and hasToolResult come from the request shape,
shadowing only happens when the same request shape repeats — e.g. a test
that replays the identical conversation twice (retries, re-runs, or a loop that re-sends
the same turn). Those produce fixture entries with identical match keys; on replay, the
router picks the first fixture that matches (first-wins by file load
order) and the later entries are shadowed. For such true repeats, add
sequenceIndex (0,
1, …) post-record to differentiate them by call order.
Recommended workflow
-
Run the test once under
--recordand let aimock capture one fixture per turn — the recordedturnIndex/hasToolResultkeys keep ordinary multi-turn flows (including tool rounds) apart automatically. -
Review the recorded fixtures. For turns whose purpose is answering a tool call, you can
additionally key on
toolCallId— the canonical tool-round idiom when you want the match tied to a specific call rather than a turn position. -
For genuine byte-identical repeats of the same turn, add
sequenceIndexto each fixture. -
Move the hand-edited fixtures into your organized
fixture directory and switch to
--strictfor replay in CI.
Cross-Language Testing
The Docker image serves any language that speaks HTTP. Point your client at the mock server's URL instead of the real API.
# Docker image serves all languages
docker run -d -p 4010:4010 -v $(pwd)/fixtures:/fixtures ghcr.io/copilotkit/aimock -f /fixtures -h 0.0.0.0
# Python
import openai
client = openai.OpenAI(base_url="http://localhost:4010/v1", api_key="mock")
# Go — github.com/sashabaranov/go-openai
config := openai.DefaultConfig("mock")
config.BaseURL = "http://localhost:4010/v1"
client := openai.NewClientWithConfig(config)
# Rust — async-openai
let config = OpenAIConfig::new()
.with_api_base("http://localhost:4010/v1")
.with_api_key("mock");
let client = Client::with_config(config);