Model Misbehavior
Use semantic faults to test invalid tool calls, empty answers, refusals, and clean length stops. The HTTP response and stream framing remain valid.
Chaos Testing changes the transport: errors, malformed bodies, delays, or disconnects. Misbehavior changes the model output inside a valid response.
Supported wires
All nine faults support streaming and non-streaming OpenAI Chat Completions. Azure OpenAI
and OpenRouter use this same openai-chat wire.
Support below describes this Unreleased build. A provider filter does not enable an unavailable renderer. The fault catalog below uses OpenAI Chat output as its example.
| Wire | Fault support | Restrictions |
|---|---|---|
openai-chat |
All nine faults; streaming and non-streaming | Refusal category is unsupported. Includes Azure and OpenRouter. |
openai-responses |
All nine faults; HTTP object/SSE and WebSocket responses | Refusal category is unsupported. |
openai-realtime |
All except refusal and reasoning-only |
GA WebSocket only. The retired beta API remains unsupported. |
anthropic, bedrock-invoke |
All except content-filter |
Invalid JSON and schema not-object require streaming. Non-streaming
tool input must remain an object.
|
bedrock-converse |
All except refusal |
Schema not-object requires streaming. Invalid arguments produce the
malformed_tool_use stop reason in non-streaming mode.
|
gemini generateContent, including the Vertex wrapper |
All except refusal |
Schema not-object is unsupported. Invalid arguments produce the
MALFORMED_FUNCTION_CALL finish reason. Duplicate IDs require an
ID-emitting tool-call mode.
|
gemini-interactions |
All except refusal and content-filter |
Invalid JSON and schema not-object require streaming. |
cohere |
All except refusal, content-filter, and
reasoning-only
|
Streaming and non-streaming, including schema not-object. |
ollama chat |
Invalid arguments, schema violations, unknown tools, empty responses, and reasoning-only responses |
Schema not-object is unsupported. Invalid arguments produce a
native-format error.
|
gemini-live |
Invalid arguments, schema violations, unknown tools, duplicate IDs, and empty responses | Schema not-object is unsupported. WebSocket only. |
Gemini and the Vertex wrapper share the gemini fault renderer. This local
support does not establish native Vertex fidelity. Bedrock Converse reasoning-only output
remains modeled/native-unverified, and recurring native drift checks remain pending.
Ollama /api/generate does not render semantic faults.
Realtime length-stop closure remains modeled/native-unverified, as described below. Recurring native drift checks for this contract remain pending.
Cohere streaming length stops follow a captured tool-call lifecycle; object-mode length stops remain modeled/native-unverified, as described below. Ollama native checks remain pending for invalid-argument error wording and reasoning-only output; the modeled exceptions do not cover those Ollama cases.
Unsupported fixture or request-header faults fail explicitly. Server and runtime faults
skip unsupported wires and record unsupported-on-wire. See
Errors and applicability.
Fault catalog
Common fields are rate, times, tool, and
providers. Each fault accepts only its own parameters.
| Fault | Parameters | OpenAI Chat output |
|---|---|---|
tool-args-invalid-json |
style: truncated (default), trailing-comma,
or single-quotes
|
Invalid JSON in the selected tool argument string. The envelope remains valid;
finish reason is tool_calls.
|
tool-args-schema-violation |
violation: missing-required (default),
wrong-type, extra-property, enum-mismatch, or
not-object; optional property
|
Valid JSON that violates a provable constraint in the request tool schema. Unresolved schema constructs do not count as proof. |
tool-unknown-name |
Optional replacement name |
A tool name absent from the request tool declarations. The default derives an unused name from the selected tool. |
tool-call-id-duplicate |
None | Two calls share an ID. With one call, aimock adds a second copy. |
stop-length-mid-tool |
at: fraction strictly between 0 and 1; default 0.5 |
A proper argument prefix, then finish_reason: "length". Calls after the
target disappear. Streams end with normal framing and [DONE].
|
empty-response |
None | No visible text or tool calls; finish reason is stop. |
refusal |
message defaults to "I can't help with that.";
category is unsupported on this wire
|
A refusal field, no visible answer or tools, and finish reason stop.
Streams carry delta.refusal.
|
content-filter |
None |
No visible answer or tools; finish reason is content_filter. Azure
filter-result metadata is not synthesized.
|
reasoning-only |
Optional reasoning; otherwise fixture reasoning or a fallback |
No visible answer or tools; finish reason is length. Native OpenAI
exposes no reasoning text. OpenRouter exposes reasoning output.
|
stop-length-mid-tool models a clean model length stop. Use
stream truncation to test an abrupt transport cut.
Inspect the runtime catalog and support table with
GET /__aimock/misbehavior/catalog.
Provider output differences
Responses length stops use status: "incomplete" with
incomplete_details.reason: "max_output_tokens". Reasoning-only responses
contain reasoning output and no visible answer or tool call.
Anthropic and Bedrock InvokeModel length stops use stop_reason: "max_tokens".
A non-streaming cut tool call has input: {}. Streaming retains the cut
argument prefix and omits the cut block's content_block_stop before the
message terminal.
Interactions length stops use status: "incomplete". Non-streaming output
omits the cut tool call. Streaming retains the cut arguments and closes the tool step with
step.stop. That streaming closure is modeled, as described below.
Duplicate-ID faults require a wire that emits tool-call IDs. For a wire that preserves authored IDs, the selected call must have a nonempty ID. Missing IDs can make the fault inapplicable even when the wire supports it.
Native captures and modeled contracts
A local SDK test proves how that client reads aimock's response. It does not prove that a live model produces the same fault. Captured responses establish only the observed request, model, and mode.
The eight earlier contracts below permit modeled behavior where native attempts did not produce the required target. A later Gemini capture confirms one mode, as noted below. Other listed targets remain modeled/native-unverified. The support table above controls availability. Four separately approved Vertex cells are described after these earlier contracts.
| Contract | Modeled output |
|---|---|
Realtime stop-length-mid-tool |
Cut argument deltas, then response.function_call_arguments.done with
that same prefix, before an incomplete response.done with
max_output_tokens. Native done-event occurrence is unverified.
|
Anthropic reasoning-only, object and streaming |
Thinking only, no text or tools, and max_tokens. An explicitly empty
reasoning string still produces a thinking block. Signature and block closure remain
present; placeholder signatures are not live-provider signatures.
|
Gemini Developer API stop-length-mid-tool, object and streaming |
Retains preceding text, omits the function call, and ends with
MAX_TOKENS. A malformed-function response does not establish this
length-stop contract.
|
Gemini Developer API reasoning-only, object and streaming |
A thought: true part, MAX_TOKENS, and no answer or tools.
A subsequent non-streaming native capture on 2026-10-08 matches this target.
Streaming still contains answer text and remains native-unverified. This
mode-specific observation does not establish Vertex fidelity.
|
Interactions stop-length-mid-tool, streaming only |
Cut arguments_delta, then step.stop and an incomplete
interaction.completed, with clean terminal framing. Non-streaming
length stops and both reasoning-only modes have separate native captures.
|
Bedrock Converse reasoning-only, object and streaming |
reasoningContent, max_tokens, and no answer or tools.
Native captures contained reasoning and answers; the no-answer outcome remains
unverified.
|
Bedrock InvokeModel reasoning-only, object and streaming |
Thinking only inside the Claude response/AWS envelope, with
max_tokens and no answer or tools. Native streaming reached that
terminal alongside text. The object capture ended with end_turn and
text; its no-answer outcome and target terminal remain modeled.
|
Cohere stop-length-mid-tool, object only |
Cut tool arguments with MAX_TOKENS. Native object captures did not
contain a cut tool; this object contract remains modeled/native-unverified.
Streaming has a separate observed contract and is not part of this exception.
|
For these exact cases, a complete valid native response that does not trigger the fault is
recorded as NOT_TRIGGERED. It is not a successful shape comparison. A
triggered target must match the contract. Access failures, transport errors, and malformed
or truncated captures cannot count as NOT_TRIGGERED.
These earlier exceptions do not cover Responses, Cohere streaming, or Ollama. They do not themselves extend to Vertex, whose separate approval is described below. They do not remove local SDK checks or recurring drift requirements. Gemini safety-filter attempts also did not capture the modeled safety ratings; those values remain unobserved.
The Cohere streaming capture on command-a-03-2025 contains an empty tool
argument opener, cut argument deltas, one tool-call-end, and a
TOOL_CALL message end followed by exactly one data: [DONE]. The
arguments remain incomplete JSON. The comparator requires this lifecycle and raw/SDK
agreement; it does not also accept MAX_TOKENS for streaming cuts. Only a
complete, valid object nontrigger may be NOT_TRIGGERED, and it does not
suppress the required stream comparison. Both recurring modes remain required; malformed
captures and genuine access or transport errors fail the comparison.
Each Cohere K5 invocation permits two requests, 128 requested output tokens per mode (256 total), a 45-second deadline and a 256 KiB response limit per request, with no SDK retries. The existing collector can repeat up to three times: up to six K5 requests and 768 requested output tokens, apart from existing ordinary checks and model discovery. Local and retained-capture tests passing, or native tests skipping without credentials, does not establish a successful hosted daily run.
The opt-in AWS comparison covers eight cases: Bedrock InvokeModel and Converse, each in
object and streaming modes, for K5 (stop-length-mid-tool) and K9
(reasoning-only). K5 compares the captured native length-stop contract;
aimock's exact truncated argument prefix is deterministic and is not a prediction of
native output bytes. K9 remains modeled: a complete, valid native response with answer
text or tools is NOT_TRIGGERED, not evidence of the reasoning-only target.
Without the opt-in, local SDK and captured-response checks run and native cases are
skipped.
AIMOCK_AWS_MISBEHAVIOR_LIVE=1 AWS_PROFILE=copilotkit-admin AWS_REGION=us-west-2 AIMOCK_AWS_MODEL_ID=us.anthropic.claude-sonnet-4-5-20250929-v1:0 PATH=/opt/homebrew/opt/node@20/bin:$PATH pnpm exec vitest run --config vitest.config.drift.ts src/__tests__/drift/bedrock-misbehavior-live.drift.ts --maxWorkers=1 --no-file-parallelism --retry=0 --bail=1
This command makes paid native requests for the exact profile, region, and model shown, after local prerequisites pass. It permits at most eight native requests without retries. It does not establish whole-AWS coverage or daily operation. The daily AWS execution destination remains unresolved; recurring native drift evidence is still required.
Vertex modeled contracts and recurring observations
Vertex has four separately approved modeled-contract cells:
stop-length-mid-tool and reasoning-only, each in object and
streaming modes. Their target shapes remain native-unverified. Length stops use
MAX_TOKENS without a function call; preceding ordinary text is allowed.
Reasoning-only output has nonempty thought: true text,
MAX_TOKENS, and no visible answer or tool call. Local Google SDK, value, and
mutation checks remain required.
Four bounded observations on Vertex gemini-2.5-flash did not produce these
target shapes. K5 object returned the undeclared
google:python_interpreter tool rather than the declared
archive tool. Its enclosing response was SDK-decodable, but this was not
normal declared-tool success. K5 streaming ended with
MALFORMED_FUNCTION_CALL. Both K9 modes contained reasoning and a visible
answer. These captured requests are examples, not guarantees of provider behavior or proof
that the target is impossible.
The comparator reports TARGET_COMPARED separately from explicitly approved
NOT_TRIGGERED outcomes. The interpreter exception matches only the retained
object response's exact tool name, program content, argument shape, and absence of extra
output. Changes to that semantic signature fail pending review. A
MALFORMED_FUNCTION_CALL non-trigger requires a valid enclosing response and a
diagnostic; malformed JSON still fails. Transport/access failures and contradictory
targets also fail. A non-trigger does not prove native fidelity and cannot weaken local
aimock assertions.
Live configuration requires AIMOCK_VERTEX_LIVE=1,
AIMOCK_VERTEX_PROJECT=llmock-drift-testing,
AIMOCK_VERTEX_LOCATION=us-central1,
AIMOCK_VERTEX_MODEL=gemini-2.5-flash, and a nonempty
AIMOCK_VERTEX_ACCESS_TOKEN. Missing enabled configuration fails. Disabled
local checks do not establish live acceptance. The existing daily main schedule and main
manual workflow use work-owned GitHub federation and a short-lived token, without a static
service-account key or generated ADC file.
Each invocation observes four cases in order: K5 object, K5 streaming, K9 object, and K9 streaming. Output caps are 16, 16, 64, and 64 tokens respectively (160 total). Each request has a 45-second deadline and a 256 KiB response limit, with no per-case retries, redirects, or fallback. Approved non-targets allow the next case to run. Genuine failures stop the invocation and leave the remaining cases recorded as unattempted. Existing outer retries can reach 12 calls and 480 requested output tokens; these caps are not a price guarantee.
Local comparator and workflow checks do not prove a successful hosted authentication
exchange or daily execution. Those remain pending until observed. Manual hosted execution
is separate from scheduled recurrence. The selected Vertex
gemini-2.5-flash model retires on October 20, 2026. Replacement remains a
maintenance follow-up, not a new acceptance gate or a migration requirement for other
providers.
Configuration
{
"seed": 42,
"faults": [
{
"fault": "tool-args-invalid-json",
"style": "truncated",
"rate": 0.5,
"times": 1,
"tool": "weather",
"providers": ["openai-chat"]
}
]
}
rateis between 0 and 1; the default is 1.-
timesis a positive integer that limits successful firings per test ID, configuration source, and fault entry. -
toolselects the first call with that name. Tool faults otherwise select the first call. -
providersfilters wire IDs, such asopenai-chat. An empty list excludes every wire. -
seedaccepts an integer or"random". Omission uses the deterministic default seed.
The shorthand "tool-args-invalid-json" means
{ "faults": [{ "fault": "tool-args-invalid-json" }] }. An empty
faults list explicitly disables faults at that configuration level.
Set the server baseline with new LLMock({ misbehavior: config }) or
mock.setMisbehavior(config). The config file uses
llm.misbehavior. The setter preserves named test overrides;
mock.clearMisbehavior() clears the baseline and all runtime overrides.
A fixture's misbehavior field accepts the same config or shorthand. Older aimock versions can ignore this unknown fixture field. Do not assume that a successful load proves fault support.
Scopes and precedence
- Request header:
X-AIMock-Misbehavior - Fixture
misbehavior - Runtime override for
X-Test-Id - Server baseline
The first present config wins as a whole. Fields and fault lists do not merge. An explicit empty list prevents fallback to lower levels.
Faults are evaluated in list order. At most one applies to a response. Later candidates see the original response, not another fault's rewritten output.
The header accepts one fault and semicolon-separated scalar parameters. It does not accept
times or providers.
X-AIMock-Misbehavior: tool-args-invalid-json;style=trailing-comma;rate=1
Realtime, Responses WS, and Gemini Live upgrades reject this header with HTTP 400. Use a
runtime override with the session's X-Test-Id; unsupported WebSocket wires
still skip runtime faults.
Fail once, then retry
Use times: 1 to exercise an application retry. SDK transport retries do not
parse tool arguments, so the example retries after JSON.parse fails.
import { LLMock } from "@copilotkit/aimock";
import OpenAI from "openai";
const mock = new LLMock({ port: 0 });
mock.onMessage("weather", {
toolCalls: [{ name: "weather", arguments: '{"city":"Paris"}' }],
});
await mock.start();
try {
const headers = { "X-Test-Id": "weather-retry" };
const installed = await fetch(`${mock.url}/__aimock/misbehavior`, {
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
body: JSON.stringify({
faults: [{ fault: "tool-args-invalid-json", times: 1 }],
}),
});
if (!installed.ok) throw new Error(await installed.text());
const client = new OpenAI({
apiKey: "local", baseURL: `${mock.url}/v1`, maxRetries: 0,
defaultHeaders: headers,
});
for (let attempt = 0; attempt < 2; attempt++) {
const result = await client.chat.completions.create({
model: "gpt-4o", messages: [{ role: "user", content: "weather" }],
});
const args = result.choices[0].message.tool_calls![0].function.arguments;
try {
console.log(JSON.parse(args)); // The second attempt prints { city: "Paris" }.
break;
} catch (error) {
if (attempt === 1) throw error;
}
}
} finally {
await mock.stop();
}
GET /__aimock/misbehavior returns the config in effect for the selected test.
DELETE removes its runtime override and restores the baseline.
POST {} is a route-only shorthand for an explicit empty list. Fixture and
programmatic configs require { "faults": [] }.
Determinism and ordering
A draw depends on the seed, test ID, configuration source, fault entry, and evaluation ordinal. A failed rate draw advances that ordinal; an applied fault also spends its firing budget.
Provider exclusions, inapplicable faults, unsupported wires, and exhausted budgets do not spend a firing. They do not advance the draw ordinal.
seed: "random" chooses and logs a process seed. Reuse the numeric seed with
the same fixture source and request order to replay the draws.
Within a test ID, request order matters. Concurrent requests and async response factories can change the order that reaches fault planning. A seed does not make that scheduling deterministic.
Use a separate test ID for each test. Preserve fixture order and source paths when replaying a run. A full reset clears fault counters; a journal-only reset preserves them. Replacing a config alone does not reset them.
Counter retention shares the journal's fixtureCountsMaxTestIds limit. FIFO
eviction removes the oldest test ID's match and fault counters; later traffic for that ID
starts again.
Observed SDK behavior
The local conformance tests use the official OpenAI Node SDK 4.104.0 with retries disabled. These observations describe that client against aimock, not every SDK version.
| Client path | Observed result |
|---|---|
chat.completions.create, stream or non-stream |
Invalid tool argument strings reach the caller. A separate
JSON.parse throws. Length, refusal, and content-filter fields remain
visible to the caller.
|
Strict-tool beta.chat.completions.parse or
.stream().finalChatCompletion()
|
A length stop throws LengthFinishReasonError. A content-filter stop
throws ContentFilterFinishReasonError.
|
Schema violations and unknown tool names still require application checks on the basic
create path. A valid HTTP response does not prove valid tool arguments.
Responses WebSocket checks use OpenAI Node SDK 6.49.0 alongside the existing Chat SDK. SDK helpers can transform or reject incomplete output; raw events and aggregated results are separate observations. The support table describes aimock's wire output, not a promise that every SDK helper returns a complete object.
Python integration checks use OpenAI 3.26.1 and Anthropic 1.12.1 against the local CLI. They exercise the existing fixture and journal APIs, streaming, and strict-parser failures. See the aimock-pytest local test instructions. These checks require a build with this Unreleased feature and no live provider key.
Inspect the prepared output
Read entry.response.misbehavior from the
request journal. It records the source, wire,
fault, evaluations, and skip or error reasons.
const entry = mock.journal.getAll().at(-1);
const fault = entry?.response.misbehavior;
console.log(fault?.applied, fault?.fault, fault?.servedToolCalls);
servedToolCalls contains the prepared names, argument strings, and emitted
IDs. It can be empty for faults that clear tool calls.
applied: true means that aimock prepared the fault output. It does not prove
delivery to the client. A disconnect can occur before the prepared call reaches the
network.
The original fixture remains unchanged. Terminal chaos and proxy responses bypass semantic
faults, with chaos-fired or proxied journal reasons.
Record/replay keeps the upstream response as the recording truth.
Metrics use aimock_misbehavior_total, labeled by fault,
wire, and outcome. Empty lists and provider-independent bypasses
do not invent a fault metric.
Errors and applicability
Unknown keys, invalid values, static inapplicability, and decidably unsupported wires fail
during fixture loading. aimock validate reports the logical source, fixture
index, and offending path.
At request time, an explicit fixture or header fault that cannot apply returns HTTP 501
with aimock_misbehavior_not_applicable or
aimock_misbehavior_unsupported. Malformed headers return HTTP 400 with
aimock_misbehavior_invalid.
Runtime and baseline configs skip inapplicable or unsupported faults so that ordinary text turns can continue. A provider exclusion also skips the entry.
One-character tool arguments cannot produce a nonempty proper prefix. Empty arguments or schema constraints that cannot prove a requested violation can also make a fault inapplicable.
Authored invalid JSON remains a validation error. Use
misbehavior: tool-args-invalid-json for intentional invalid arguments. See
fixture validation guidance.
Strict tools and realism
Use invalid JSON and schema violations to test non-strict tools and open models. Invalid JSON and schema violations are not realistic production outputs for OpenAI or Anthropic strict structured tools.
Aimock's strict routing mode concerns fixture matches and proxy fallback. It does not validate a tool's JSON schema or disable an authored semantic fault.