WebSocket APIs
aimock implements four WebSocket APIs with zero dependencies — real RFC 6455 framing built from scratch. Responses fixtures can drive HTTP and WebSocket transports; GPT-Live uses dedicated transcripts.
Endpoints
| Path | API | Protocol |
|---|---|---|
| /v1/live/sessions | OpenAI GPT-Live | WebSocket JSON messages, including base64 PCM |
| /v1/responses | OpenAI Responses API | WebSocket JSON messages |
| /v1/realtime | OpenAI Realtime API | WebSocket JSON messages |
| /ws/google.ai.generativelanguage.* | Gemini Live | WebSocket JSON messages |
OpenAI GPT-Live
GPT-Live uses a primary WebSocket upgrade: GET /v1/live/sessions. Connect to
the mock server and send session.start as the first JSON message. The fixture
must match endpoint: "openai-live" and model: "gpt-live-1".
Replay needs no provider key or network connection to OpenAI. It reproduces the saved conversation, including audio, delegation, usage, and session closure. It does not synthesize speech or run a backend model.
Run an offline client example
Use Node.js 22 or later for this example's built-in WebSocket. From an aimock
checkout, install dependencies with pnpm install --frozen-lockfile and build
with pnpm build. Save the following as live-client.mts and run
pnpm exec tsx live-client.mts.
The included fixtures/openai-live-client.json is an authored example, not a
provider recording. Its six PCM16 samples per direction are synthetic, not speech. Timings
and usage are illustrative.
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import { LLMock, type LiveTranscript } from "@copilotkit/aimock";
const file = "fixtures/openai-live-client.json";
const document: { fixtures: { response: { live: LiveTranscript } }[] } =
JSON.parse(readFileSync(file, "utf8"));
const transcript = document.fixtures[0].response.live;
const mock = new LLMock().loadFixtureFile(file);
await mock.start();
const ws = new WebSocket(`${mock.url.replace(/^http/, "ws")}/v1/live/sessions`);
const audio: Buffer[] = [];
const send = (event: object) => ws.send(JSON.stringify(event));
try {
await new Promise<void>((resolve, reject) => {
const timer = setTimeout(() => reject(new Error("Live replay timed out")), 5000);
const finish = (error?: Error) => {
clearTimeout(timer);
if (error) reject(error);
else resolve();
};
let closed = false;
ws.onopen = () => send(transcript.entries[0].event); // session.start
ws.onerror = () => finish(new Error("WebSocket failed"));
ws.onclose = () => {
if (!closed) finish(new Error("Missing session.closed"));
};
ws.onmessage = ({ data }) => {
try {
const event = JSON.parse(String(data));
if (event.type === "session.started") {
send({ type: "session.input_audio.append", audio: "AAABAP//AgD+/wAA" });
} else if (event.type === "session.delegation.created") {
const delegation_id = event.delegation.id; // Fresh ID from this connection
send({
type: "session.thinking.append", event_id: "thinking-1", delegation_id,
content: "The fictional red and blue sensor checks are still pending.",
});
send({
type: "session.commentary.append", event_id: "commentary-1", delegation_id,
content: "The fictional red sensor reads seven.",
});
} else if (event.type === "session.output_audio.delta") {
audio.push(Buffer.from(event.delta, "base64"));
} else if (event.type === "session.usage.updated") {
send({ type: "session.close", event_id: "close-1" });
} else if (event.type === "session.closed") {
assert.deepEqual(Buffer.concat(audio), Buffer.from("AAACAP7/AQD//wAA", "base64"));
closed = true;
finish();
} else if (event.type === "aimock.error" || event.type === "error") {
finish(new Error(JSON.stringify(event)));
}
} catch (error) {
finish(error instanceof Error ? error : new Error(String(error)));
}
};
});
console.log("Live client replay passed");
} finally {
ws.close();
await mock.stop();
}
The example loads one fixture file into one server. The managed example matches the same endpoint and model; do not load both examples together unless you add distinct matching rules.
// Alternative registration: use the transcript directly.
const mock = new LLMock().onLive({ model: "gpt-live-1" }, transcript);
onLive takes a LiveTranscript, not a
{ live: transcript } wrapper. The builder uses recorded timing by default.
The supplied files select liveTiming: "immediate". Use
mock.on({ endpoint: "openai-live", model: "gpt-live-1" }, { live: transcript,
liveTiming: "immediate" })
to select immediate timing explicitly.
Client and managed delegation
In client mode, session.delegation.created asks your application to handle
backend work. Send session.thinking.append and
session.commentary.append using the emitted delegation ID. Recorded null and
repeated commentary values remain part of the transcript.
For managed mode, load fixtures/openai-live-managed.json instead and use its
startup configuration. The transcript calls this mode managed; the wire
configuration uses delegation.type: "responses". Backend events arrive inside
response.event. Read complete function calls from nested
response.output_item.done events. Use each nested item's
call_id when returning its output.
The managed fixture requests read_red and read_blue. After both
complete calls arrive, return both outputs and then explicitly request continuation:
// Send each output using the call_id received in response.event.
send({
type: "response.item.create", event_id: "red-result",
item: {
type: "function_call_output", call_id: redCallId,
output: JSON.stringify({ sensor: "read_red", value: 7 }),
},
});
send({
type: "response.item.create", event_id: "blue-result",
item: {
type: "function_call_output", call_id: blueCallId,
output: JSON.stringify({ sensor: "read_blue", value: 9 }),
},
});
// Continue only after all pending tool outputs have been sent.
send({ type: "response.create", event_id: "continue-1" });
The managed application must handle these nested events; the client-mode handler above does not. Use IDs received on the current connection. Replay creates fresh server IDs and preserves their references.
Audio, timing, and lifecycle
- Audio is mono, 24,000 Hz, signed 16-bit little-endian PCM. JSON messages carry base64 audio; binary frames are not supported.
- Replay compares decoded PCM bytes, not base64 chunk boundaries. Rechunking consecutive audio is valid; moving audio across control commands is not.
- Server events wait for their recorded client command and audio-byte barriers. Immediate timing removes delays but preserves those barriers.
- Recorded timing preserves event delays. Replay does not infer speech, new tool calls, or an unrecorded interruption.
- Model, delegation mode, and audio format are fixed for a session. Only observed backend instruction updates are supported in managed mode.
-
A successful transcript ends with
session.closed. A transport close alone does not establish a complete recording.
Always close the client and call await mock.stop() in cleanup. Use
mock.closeLiveSessions(testId) to close sessions for an explicit
X-Test-Id, or omit the ID to close all Live sessions on that server. The
Vitest and Jest helpers close their owned Live sessions at test boundaries.
resetMatchCounts() only resets counters; it does not close sessions.
For provider authentication, recording, retained audio, sanitization, and finite session limits, see Record & Replay. For the transcript format, see Fixtures.
OpenAI Responses (WebSocket)
const instance = await createServer([
{ match: { userMessage: "hello" }, response: { content: "Hi there!" } }
]);
const ws = await connectWebSocket(instance.url, "/v1/responses");
// Send a response.create message
ws.send(JSON.stringify({
type: "response.create",
model: "gpt-4",
input: [{ role: "user", content: "hello" }],
}));
const messages = await ws.waitForMessages(9);
const events = messages.map(m => JSON.parse(m));
const types = events.map(e => e.type);
expect(types[0]).toBe("response.created");
expect(types).toContain("response.output_text.delta");
expect(types).toContain("response.completed");
OpenAI Realtime
The Realtime API uses a conversational protocol with session management. aimock implements
the
GA (General Availability) protocol natively — event names like
response.output_text.delta, conversation.item.added, and nested
audio session config are the defaults. The Beta protocol (OpenAI-Beta: realtime=v1) has been removed upstream and is mocked as removed — see
Beta shape: mocked as removed.
Supported Models
| Model | Session Types | Notes |
|---|---|---|
| gpt-realtime | conversation | Base alias — resolves to latest GA model |
| gpt-realtime-2 | conversation | Default model — GA successor to gpt-4o-realtime-preview |
| gpt-realtime-1.5 | conversation | Previous generation GA model |
| gpt-realtime-mini | conversation | Smaller, faster GA model |
| gpt-4o-transcribe | transcription, translation | Speech transcription and translation |
| gpt-4o-mini-transcribe | transcription, translation | Smaller transcription and translation model |
| whisper-1 | transcription | Legacy Whisper transcription model |
Session Types
- conversation (default) — Standard conversational interaction with text and audio modalities
-
transcription — Audio-to-text transcription (requires
gpt-4o-transcribe,gpt-4o-mini-transcribe, orwhisper-1) -
translation — Real-time speech translation (requires
gpt-4o-transcribeorgpt-4o-mini-transcribe)
GA Protocol Features
-
GA event names —
response.output_text.delta(wasresponse.text.delta),conversation.item.added(wasconversation.item.created), etc. -
Nested audio config — Session config uses
session.audio.voiceinstead of flatsession.voice -
Image input —
input_imagecontent parts inconversation.item.create -
Commentary phase —
phasefield onresponse.output_item.added/doneevents (final_answerorcommentary) -
conversation.item.done— New event emitted after each completed response item -
response.cancel— Client message to cancel in-flight responses
Beta shape: mocked as removed
OpenAI-Beta: realtime=v1 gets the sunset rejection the real API
returns today, which is what is left to test: that your code handles the removal. This
follows aimock’s
deprecated & removed APIs policy — a removed
surface is mocked as removed, never as a success.
The upgrade still succeeds (HTTP/1.1 101); the rejection arrives on the
socket. aimock replays exactly what a live Beta handshake against
wss://api.openai.com/v1/realtime returns: one error event, no
session.created, then a close frame.
{
"type": "error",
"event_id": "event_...",
"error": {
"type": "invalid_request_error",
"code": "beta_api_shape_disabled",
"message": "The Realtime Beta API is no longer supported. Please use /v1/realtime for the GA API.",
"param": null,
"event_id": null
}
}
// then: WebSocket CLOSE code 4000
// reason "invalid_request_error.beta_api_shape_disabled"
Tests that used to drive the Beta shape should move to the GA shape (drop the
OpenAI-Beta header) — that is the migration target, and the only
Realtime surface OpenAI still serves. The Beta shape is also excluded from aimock’s
drift pipeline, because there is no live Beta endpoint left to compare against.
const ws = await connectWebSocket(instance.url, "/v1/realtime?model=gpt-realtime-2");
// Server sends session.created on connect
const [sessionMsg] = await ws.waitForMessages(1);
const session = JSON.parse(sessionMsg);
expect(session.type).toBe("session.created");
expect(session.session.type).toBe("conversation");
expect(session.session.audio).toBeDefined();
// Configure session with nested audio config
ws.send(JSON.stringify({
type: "session.update",
session: {
modalities: ["text"],
audio: { voice: "alloy" }
}
}));
// Add a user message (supports input_text + input_image content)
ws.send(JSON.stringify({
type: "conversation.item.create",
item: {
type: "message",
role: "user",
content: [{ type: "input_text", text: "hello" }]
}
}));
// Request a response
ws.send(JSON.stringify({ type: "response.create" }));
// GA events: output_text instead of text, item.added instead of item.created
const msgs = await ws.waitForMessages(10);
const events = msgs.map(m => JSON.parse(m));
expect(events.some(e => e.type === "response.output_text.delta")).toBe(true);
expect(events.some(e => e.type === "conversation.item.added")).toBe(true);
expect(events.some(e => e.type === "conversation.item.done")).toBe(true);
Gemini Live
Bidirectional streaming for Google Gemini Live API.
const ws = await connectWebSocket(
instance.url,
"/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent"
);
// Send setup message
ws.send(JSON.stringify({
setup: { model: "models/gemini-2.0-flash-live" }
}));
// Send client content
ws.send(JSON.stringify({
clientContent: {
turns: [{ role: "user", parts: [{ text: "hello" }] }],
turnComplete: true,
}
}));
Implementation Details
- Built on raw RFC 6455 WebSocket framing — zero external dependencies
- JSON text messages; GPT-Live carries base64 PCM audio. No binary frames.
- Endpoint-specific fixture matching; GPT-Live requires a Live transcript
- All WebSocket connections are logged in the journal
Gemini Live text support is unverified — no text-capable Gemini Live model existed at time of implementation. The WebSocket framing and protocol messages follow the published API spec.
Provider WebSocket Support
Not all LLM providers offer WebSocket APIs. Here's the current landscape:
| Provider | WebSocket API | aimock Status |
|---|---|---|
| OpenAI GPT-Live | wss://api.openai.com/v1/live/sessions | Primary WebSocket replay and recording |
| OpenAI Realtime | wss://api.openai.com/v1/realtime | Supported ✓ |
| OpenAI Responses | wss://api.openai.com/v1/responses | Supported ✓ |
| Gemini Live | wss://...BidiGenerateContent | Implemented, awaiting text model |
| Anthropic Claude | None | N/A |
| Azure OpenAI | Uses OpenAI Realtime | Covered by OpenAI |
| Mistral / Groq / Cohere | None | N/A |
| AWS Bedrock | EventStream (not WebSocket) | N/A |
aimock includes drift canary tests that automatically detect when providers add new WebSocket capabilities. When a canary fires, it signals that aimock should be updated to support the new API.