Model Misbehavior

Use semantic faults to test invalid tool calls, empty answers, refusals, and clean length stops. The HTTP response and stream framing remain valid.

Chaos Testing changes the transport: errors, malformed bodies, delays, or disconnects. Misbehavior changes the model output inside a valid response.

Supported wires

All nine faults support streaming and non-streaming OpenAI Chat Completions. Azure OpenAI and OpenRouter use this same openai-chat wire.

Support below describes this Unreleased build. A provider filter does not enable an unavailable renderer. The fault catalog below uses OpenAI Chat output as its example.

Wire Fault support Restrictions
openai-chat All nine faults; streaming and non-streaming Refusal category is unsupported. Includes Azure and OpenRouter.
openai-responses All nine faults; HTTP object/SSE and WebSocket responses Refusal category is unsupported.
openai-realtime All except refusal and reasoning-only GA WebSocket only. The retired beta API remains unsupported.
anthropic, bedrock-invoke All except content-filter Invalid JSON and schema not-object require streaming. Non-streaming tool input must remain an object.
bedrock-converse All except refusal Schema not-object requires streaming. Invalid arguments produce the malformed_tool_use stop reason in non-streaming mode.
gemini generateContent, including the Vertex wrapper All except refusal Schema not-object is unsupported. Invalid arguments produce the MALFORMED_FUNCTION_CALL finish reason. Duplicate IDs require an ID-emitting tool-call mode.
gemini-interactions All except refusal and content-filter Invalid JSON and schema not-object require streaming.
cohere All except refusal, content-filter, and reasoning-only Streaming and non-streaming, including schema not-object.
ollama chat Invalid arguments, schema violations, unknown tools, empty responses, and reasoning-only responses Schema not-object is unsupported. Invalid arguments produce a native-format error.
gemini-live Invalid arguments, schema violations, unknown tools, duplicate IDs, and empty responses Schema not-object is unsupported. WebSocket only.

Gemini and the Vertex wrapper share the gemini fault renderer. This local support does not establish native Vertex fidelity. Bedrock Converse reasoning-only output remains modeled/native-unverified, and recurring native drift checks remain pending. Ollama /api/generate does not render semantic faults.

Realtime length-stop closure remains modeled/native-unverified, as described below. Recurring native drift checks for this contract remain pending.

Cohere streaming length stops follow a captured tool-call lifecycle; object-mode length stops remain modeled/native-unverified, as described below. Ollama native checks remain pending for invalid-argument error wording and reasoning-only output; the modeled exceptions do not cover those Ollama cases.

Unsupported fixture or request-header faults fail explicitly. Server and runtime faults skip unsupported wires and record unsupported-on-wire. See Errors and applicability.

Fault catalog

Common fields are rate, times, tool, and providers. Each fault accepts only its own parameters.

Fault Parameters OpenAI Chat output
tool-args-invalid-json style: truncated (default), trailing-comma, or single-quotes Invalid JSON in the selected tool argument string. The envelope remains valid; finish reason is tool_calls.
tool-args-schema-violation violation: missing-required (default), wrong-type, extra-property, enum-mismatch, or not-object; optional property Valid JSON that violates a provable constraint in the request tool schema. Unresolved schema constructs do not count as proof.
tool-unknown-name Optional replacement name A tool name absent from the request tool declarations. The default derives an unused name from the selected tool.
tool-call-id-duplicate None Two calls share an ID. With one call, aimock adds a second copy.
stop-length-mid-tool at: fraction strictly between 0 and 1; default 0.5 A proper argument prefix, then finish_reason: "length". Calls after the target disappear. Streams end with normal framing and [DONE].
empty-response None No visible text or tool calls; finish reason is stop.
refusal message defaults to "I can't help with that."; category is unsupported on this wire A refusal field, no visible answer or tools, and finish reason stop. Streams carry delta.refusal.
content-filter None No visible answer or tools; finish reason is content_filter. Azure filter-result metadata is not synthesized.
reasoning-only Optional reasoning; otherwise fixture reasoning or a fallback No visible answer or tools; finish reason is length. Native OpenAI exposes no reasoning text. OpenRouter exposes reasoning output.

stop-length-mid-tool models a clean model length stop. Use stream truncation to test an abrupt transport cut.

Inspect the runtime catalog and support table with GET /__aimock/misbehavior/catalog.

Provider output differences

Responses length stops use status: "incomplete" with incomplete_details.reason: "max_output_tokens". Reasoning-only responses contain reasoning output and no visible answer or tool call.

Anthropic and Bedrock InvokeModel length stops use stop_reason: "max_tokens". A non-streaming cut tool call has input: {}. Streaming retains the cut argument prefix and omits the cut block's content_block_stop before the message terminal.

Interactions length stops use status: "incomplete". Non-streaming output omits the cut tool call. Streaming retains the cut arguments and closes the tool step with step.stop. That streaming closure is modeled, as described below.

Duplicate-ID faults require a wire that emits tool-call IDs. For a wire that preserves authored IDs, the selected call must have a nonempty ID. Missing IDs can make the fault inapplicable even when the wire supports it.

Native captures and modeled contracts

A local SDK test proves how that client reads aimock's response. It does not prove that a live model produces the same fault. Captured responses establish only the observed request, model, and mode.

The eight earlier contracts below permit modeled behavior where native attempts did not produce the required target. A later Gemini capture confirms one mode, as noted below. Other listed targets remain modeled/native-unverified. The support table above controls availability. Four separately approved Vertex cells are described after these earlier contracts.

Contract Modeled output
Realtime stop-length-mid-tool Cut argument deltas, then response.function_call_arguments.done with that same prefix, before an incomplete response.done with max_output_tokens. Native done-event occurrence is unverified.
Anthropic reasoning-only, object and streaming Thinking only, no text or tools, and max_tokens. An explicitly empty reasoning string still produces a thinking block. Signature and block closure remain present; placeholder signatures are not live-provider signatures.
Gemini Developer API stop-length-mid-tool, object and streaming Retains preceding text, omits the function call, and ends with MAX_TOKENS. A malformed-function response does not establish this length-stop contract.
Gemini Developer API reasoning-only, object and streaming A thought: true part, MAX_TOKENS, and no answer or tools. A subsequent non-streaming native capture on 2026-10-08 matches this target. Streaming still contains answer text and remains native-unverified. This mode-specific observation does not establish Vertex fidelity.
Interactions stop-length-mid-tool, streaming only Cut arguments_delta, then step.stop and an incomplete interaction.completed, with clean terminal framing. Non-streaming length stops and both reasoning-only modes have separate native captures.
Bedrock Converse reasoning-only, object and streaming reasoningContent, max_tokens, and no answer or tools. Native captures contained reasoning and answers; the no-answer outcome remains unverified.
Bedrock InvokeModel reasoning-only, object and streaming Thinking only inside the Claude response/AWS envelope, with max_tokens and no answer or tools. Native streaming reached that terminal alongside text. The object capture ended with end_turn and text; its no-answer outcome and target terminal remain modeled.
Cohere stop-length-mid-tool, object only Cut tool arguments with MAX_TOKENS. Native object captures did not contain a cut tool; this object contract remains modeled/native-unverified. Streaming has a separate observed contract and is not part of this exception.

For these exact cases, a complete valid native response that does not trigger the fault is recorded as NOT_TRIGGERED. It is not a successful shape comparison. A triggered target must match the contract. Access failures, transport errors, and malformed or truncated captures cannot count as NOT_TRIGGERED.

These earlier exceptions do not cover Responses, Cohere streaming, or Ollama. They do not themselves extend to Vertex, whose separate approval is described below. They do not remove local SDK checks or recurring drift requirements. Gemini safety-filter attempts also did not capture the modeled safety ratings; those values remain unobserved.

The Cohere streaming capture on command-a-03-2025 contains an empty tool argument opener, cut argument deltas, one tool-call-end, and a TOOL_CALL message end followed by exactly one data: [DONE]. The arguments remain incomplete JSON. The comparator requires this lifecycle and raw/SDK agreement; it does not also accept MAX_TOKENS for streaming cuts. Only a complete, valid object nontrigger may be NOT_TRIGGERED, and it does not suppress the required stream comparison. Both recurring modes remain required; malformed captures and genuine access or transport errors fail the comparison.

Each Cohere K5 invocation permits two requests, 128 requested output tokens per mode (256 total), a 45-second deadline and a 256 KiB response limit per request, with no SDK retries. The existing collector can repeat up to three times: up to six K5 requests and 768 requested output tokens, apart from existing ordinary checks and model discovery. Local and retained-capture tests passing, or native tests skipping without credentials, does not establish a successful hosted daily run.

The opt-in AWS comparison covers eight cases: Bedrock InvokeModel and Converse, each in object and streaming modes, for K5 (stop-length-mid-tool) and K9 (reasoning-only). K5 compares the captured native length-stop contract; aimock's exact truncated argument prefix is deterministic and is not a prediction of native output bytes. K9 remains modeled: a complete, valid native response with answer text or tools is NOT_TRIGGERED, not evidence of the reasoning-only target. Without the opt-in, local SDK and captured-response checks run and native cases are skipped.

opt-in AWS comparison sh
AIMOCK_AWS_MISBEHAVIOR_LIVE=1 AWS_PROFILE=copilotkit-admin AWS_REGION=us-west-2 AIMOCK_AWS_MODEL_ID=us.anthropic.claude-sonnet-4-5-20250929-v1:0 PATH=/opt/homebrew/opt/node@20/bin:$PATH pnpm exec vitest run --config vitest.config.drift.ts src/__tests__/drift/bedrock-misbehavior-live.drift.ts --maxWorkers=1 --no-file-parallelism --retry=0 --bail=1

This command makes paid native requests for the exact profile, region, and model shown, after local prerequisites pass. It permits at most eight native requests without retries. It does not establish whole-AWS coverage or daily operation. The daily AWS execution destination remains unresolved; recurring native drift evidence is still required.

Vertex modeled contracts and recurring observations

Vertex has four separately approved modeled-contract cells: stop-length-mid-tool and reasoning-only, each in object and streaming modes. Their target shapes remain native-unverified. Length stops use MAX_TOKENS without a function call; preceding ordinary text is allowed. Reasoning-only output has nonempty thought: true text, MAX_TOKENS, and no visible answer or tool call. Local Google SDK, value, and mutation checks remain required.

Four bounded observations on Vertex gemini-2.5-flash did not produce these target shapes. K5 object returned the undeclared google:python_interpreter tool rather than the declared archive tool. Its enclosing response was SDK-decodable, but this was not normal declared-tool success. K5 streaming ended with MALFORMED_FUNCTION_CALL. Both K9 modes contained reasoning and a visible answer. These captured requests are examples, not guarantees of provider behavior or proof that the target is impossible.

The comparator reports TARGET_COMPARED separately from explicitly approved NOT_TRIGGERED outcomes. The interpreter exception matches only the retained object response's exact tool name, program content, argument shape, and absence of extra output. Changes to that semantic signature fail pending review. A MALFORMED_FUNCTION_CALL non-trigger requires a valid enclosing response and a diagnostic; malformed JSON still fails. Transport/access failures and contradictory targets also fail. A non-trigger does not prove native fidelity and cannot weaken local aimock assertions.

Live configuration requires AIMOCK_VERTEX_LIVE=1, AIMOCK_VERTEX_PROJECT=llmock-drift-testing, AIMOCK_VERTEX_LOCATION=us-central1, AIMOCK_VERTEX_MODEL=gemini-2.5-flash, and a nonempty AIMOCK_VERTEX_ACCESS_TOKEN. Missing enabled configuration fails. Disabled local checks do not establish live acceptance. The existing daily main schedule and main manual workflow use work-owned GitHub federation and a short-lived token, without a static service-account key or generated ADC file.

Each invocation observes four cases in order: K5 object, K5 streaming, K9 object, and K9 streaming. Output caps are 16, 16, 64, and 64 tokens respectively (160 total). Each request has a 45-second deadline and a 256 KiB response limit, with no per-case retries, redirects, or fallback. Approved non-targets allow the next case to run. Genuine failures stop the invocation and leave the remaining cases recorded as unattempted. Existing outer retries can reach 12 calls and 480 requested output tokens; these caps are not a price guarantee.

Local comparator and workflow checks do not prove a successful hosted authentication exchange or daily execution. Those remain pending until observed. Manual hosted execution is separate from scheduled recurrence. The selected Vertex gemini-2.5-flash model retires on October 20, 2026. Replacement remains a maintenance follow-up, not a new acceptance gate or a migration requirement for other providers.

Configuration

misbehavior config json
{
  "seed": 42,
  "faults": [
    {
      "fault": "tool-args-invalid-json",
      "style": "truncated",
      "rate": 0.5,
      "times": 1,
      "tool": "weather",
      "providers": ["openai-chat"]
    }
  ]
}

The shorthand "tool-args-invalid-json" means { "faults": [{ "fault": "tool-args-invalid-json" }] }. An empty faults list explicitly disables faults at that configuration level.

Set the server baseline with new LLMock({ misbehavior: config }) or mock.setMisbehavior(config). The config file uses llm.misbehavior. The setter preserves named test overrides; mock.clearMisbehavior() clears the baseline and all runtime overrides.

A fixture's misbehavior field accepts the same config or shorthand. Older aimock versions can ignore this unknown fixture field. Do not assume that a successful load proves fault support.

Scopes and precedence

  1. Request header: X-AIMock-Misbehavior
  2. Fixture misbehavior
  3. Runtime override for X-Test-Id
  4. Server baseline

The first present config wins as a whole. Fields and fault lists do not merge. An explicit empty list prevents fallback to lower levels.

Faults are evaluated in list order. At most one applies to a response. Later candidates see the original response, not another fault's rewritten output.

The header accepts one fault and semicolon-separated scalar parameters. It does not accept times or providers.

HTTP request header text
X-AIMock-Misbehavior: tool-args-invalid-json;style=trailing-comma;rate=1

Realtime, Responses WS, and Gemini Live upgrades reject this header with HTTP 400. Use a runtime override with the session's X-Test-Id; unsupported WebSocket wires still skip runtime faults.

Fail once, then retry

Use times: 1 to exercise an application retry. SDK transport retries do not parse tool arguments, so the example retries after JSON.parse fails.

retry.ts ts
import { LLMock } from "@copilotkit/aimock";
import OpenAI from "openai";

const mock = new LLMock({ port: 0 });
mock.onMessage("weather", {
  toolCalls: [{ name: "weather", arguments: '{"city":"Paris"}' }],
});
await mock.start();
try {
  const headers = { "X-Test-Id": "weather-retry" };
  const installed = await fetch(`${mock.url}/__aimock/misbehavior`, {
    method: "POST",
    headers: { ...headers, "Content-Type": "application/json" },
    body: JSON.stringify({
      faults: [{ fault: "tool-args-invalid-json", times: 1 }],
    }),
  });
  if (!installed.ok) throw new Error(await installed.text());
  const client = new OpenAI({
    apiKey: "local", baseURL: `${mock.url}/v1`, maxRetries: 0,
    defaultHeaders: headers,
  });
  for (let attempt = 0; attempt < 2; attempt++) {
    const result = await client.chat.completions.create({
      model: "gpt-4o", messages: [{ role: "user", content: "weather" }],
    });
    const args = result.choices[0].message.tool_calls![0].function.arguments;
    try {
      console.log(JSON.parse(args)); // The second attempt prints { city: "Paris" }.
      break;
    } catch (error) {
      if (attempt === 1) throw error;
    }
  }
} finally {
  await mock.stop();
}

GET /__aimock/misbehavior returns the config in effect for the selected test. DELETE removes its runtime override and restores the baseline.

POST {} is a route-only shorthand for an explicit empty list. Fixture and programmatic configs require { "faults": [] }.

Determinism and ordering

A draw depends on the seed, test ID, configuration source, fault entry, and evaluation ordinal. A failed rate draw advances that ordinal; an applied fault also spends its firing budget.

Provider exclusions, inapplicable faults, unsupported wires, and exhausted budgets do not spend a firing. They do not advance the draw ordinal.

seed: "random" chooses and logs a process seed. Reuse the numeric seed with the same fixture source and request order to replay the draws.

Within a test ID, request order matters. Concurrent requests and async response factories can change the order that reaches fault planning. A seed does not make that scheduling deterministic.

Use a separate test ID for each test. Preserve fixture order and source paths when replaying a run. A full reset clears fault counters; a journal-only reset preserves them. Replacing a config alone does not reset them.

Counter retention shares the journal's fixtureCountsMaxTestIds limit. FIFO eviction removes the oldest test ID's match and fault counters; later traffic for that ID starts again.

Observed SDK behavior

The local conformance tests use the official OpenAI Node SDK 4.104.0 with retries disabled. These observations describe that client against aimock, not every SDK version.

Client path Observed result
chat.completions.create, stream or non-stream Invalid tool argument strings reach the caller. A separate JSON.parse throws. Length, refusal, and content-filter fields remain visible to the caller.
Strict-tool beta.chat.completions.parse or .stream().finalChatCompletion() A length stop throws LengthFinishReasonError. A content-filter stop throws ContentFilterFinishReasonError.

Schema violations and unknown tool names still require application checks on the basic create path. A valid HTTP response does not prove valid tool arguments.

Responses WebSocket checks use OpenAI Node SDK 6.49.0 alongside the existing Chat SDK. SDK helpers can transform or reject incomplete output; raw events and aggregated results are separate observations. The support table describes aimock's wire output, not a promise that every SDK helper returns a complete object.

Python integration checks use OpenAI 3.26.1 and Anthropic 1.12.1 against the local CLI. They exercise the existing fixture and journal APIs, streaming, and strict-parser failures. See the aimock-pytest local test instructions. These checks require a build with this Unreleased feature and no live provider key.

Inspect the prepared output

Read entry.response.misbehavior from the request journal. It records the source, wire, fault, evaluations, and skip or error reasons.

journal inspection ts
const entry = mock.journal.getAll().at(-1);
const fault = entry?.response.misbehavior;
console.log(fault?.applied, fault?.fault, fault?.servedToolCalls);

servedToolCalls contains the prepared names, argument strings, and emitted IDs. It can be empty for faults that clear tool calls.

applied: true means that aimock prepared the fault output. It does not prove delivery to the client. A disconnect can occur before the prepared call reaches the network.

The original fixture remains unchanged. Terminal chaos and proxy responses bypass semantic faults, with chaos-fired or proxied journal reasons. Record/replay keeps the upstream response as the recording truth.

Metrics use aimock_misbehavior_total, labeled by fault, wire, and outcome. Empty lists and provider-independent bypasses do not invent a fault metric.

Errors and applicability

Unknown keys, invalid values, static inapplicability, and decidably unsupported wires fail during fixture loading. aimock validate reports the logical source, fixture index, and offending path.

At request time, an explicit fixture or header fault that cannot apply returns HTTP 501 with aimock_misbehavior_not_applicable or aimock_misbehavior_unsupported. Malformed headers return HTTP 400 with aimock_misbehavior_invalid.

Runtime and baseline configs skip inapplicable or unsupported faults so that ordinary text turns can continue. A provider exclusion also skips the entry.

One-character tool arguments cannot produce a nonempty proper prefix. Empty arguments or schema constraints that cannot prove a requested violation can also make a fault inapplicable.

Authored invalid JSON remains a validation error. Use misbehavior: tool-args-invalid-json for intentional invalid arguments. See fixture validation guidance.

Strict tools and realism

Use invalid JSON and schema violations to test non-strict tools and open models. Invalid JSON and schema violations are not realistic production outputs for OpenAI or Anthropic strict structured tools.

Aimock's strict routing mode concerns fixture matches and proxy fallback. It does not validate a tool's JSON schema or disable an authored semantic fault.