Skip to content

Runs and streaming

View as Markdown

A run is one agent turn: the input goes in, the model and tools work for up to max_tool_rounds rounds, and the output comes out. Every ask, every answered session message and every tool-output submission produces a run.

queued ──► in_progress ──► completed | failed | cancelled | superseded
▲ │
│ ▼
requires_action (waiting for your client tool outputs)
  • status moves from queued to in_progress, may alternate with requires_action, and ends in exactly one terminal status.
  • outcome is set only when the run ends: answered, handed_off, failed, cancelled, superseded or client_tool_calls (the last one only on the OpenAI-compatible endpoint). requires_action is a status, never an outcome.
  • Never silent. A provider error is retried, then tried on your fallback_models; if all of them fail, the customer gets your fallback message and the conversation is handed to your team. failed is used only when no handoff path exists or fallback.handoff_on_failure is off.
  • Configuration problems — no model key, a model your plan doesn’t include, AI paused — are not provider failures: they are never retried or handed off, and API calls get 409 agent_not_ready.
Field Description
id, object run_…, "run"
agent {id, version} — the exact version that answered
session_id, end_user_id null for one-shot runs without an end user
mode, channel session, one_shot or playground; api, widget, playground or openai
status, outcome See the lifecycle above
input_message_ids The user messages this run answered (several, after an interrupt)
output, output_text The run’s assistant messages, and their text joined ("" when there is none)
handoff null or {id, reason_type, summary, status}
required_action While paused: {type: "submit_tool_outputs", tool_calls, expires_at}
error null or {code, message}
superseded_by The newer run that replaced this one
usage {input_tokens, cached_input_tokens, output_tokens, model_calls, units, weight}
config_hash, snapshot_hash Which settings, and which frozen prompt and tool list, produced this answer
warnings Non-fatal notes
created_at, started_at, completed_at Unix seconds

GET /v1/runs/{run} returns a run, GET /v1/runs lists them, and GET /v1/runs/{run}/steps returns the trace: every model call and tool call with arguments, results, latency and tokens (redacted where your settings say so).

Without streaming, ask, messages and submit_tool_outputs wait for the run:

  • wait_seconds sets how long (default 60, maximum 110).
  • 200 when the run has finished, is waiting for client tool outputs, or no run was needed (human mode, or a repeated client_message_id).
  • 202 only when the run is still queued or in_progress — because you sent background: true or the wait ran out. The Location header points to /v1/runs/{id}.

Runs keep going when your connection drops: a client disconnect never cancels a run. Only POST /v1/runs/{run}/cancel (or an interrupt from a newer message) stops one. Cancelling a finished run returns 409 run_already_completed.

Add "stream": true to ask, POST /v1/sessions, POST /v1/sessions/{session}/messages or submit_tool_outputs, and the response becomes a text/event-stream. You can also attach to a stream at any time:

  • Run stream — GET /v1/runs/{run}/events: one run’s events. It ends with exactly one of run.completed, run.failed, run.cancelled, run.superseded or run.requires_action, and then the server closes it.
  • Session stream — GET /v1/sessions/{session}/events: everything that happens in the session, run after run, including your team’s replies. It never closes on its own.
id: 4185
event: message.completed
data: {"message":{"id":"msg_01k6rz8e0h2k4n6q8s0v2x4z6b","role":"assistant","content":[{"type":"text","text":"Yes, within 7 days, as long as it's unopened and in its original packaging."}]}}
: ping
  • Each event has event: and data: (JSON). Durable events also have an id: — a number that only grows, shared by run and session streams.
  • message.delta events have no id:: they are delivered live and never stored or replayed.
  • A comment line (: ping) is sent every 15 seconds to keep proxies from closing the connection.

The Streaming in JavaScript guide has a complete parser.

Event Data Notes
session.created {session, created} First event of POST /v1/sessions with stream: true.
run.created {run} The run exists.
run.in_progress {run} The agent started working on it.
message.delta {message_id, delta} A piece of reply text. Best-effort preview.
message.completed {message, replaced?, discarded?, truncated?} The final text of one assistant message. Authoritative.
tool_call.created {call_id, name, arguments_preview} The model called a tool.
tool_call.completed {call_id, ok, summary} The tool returned.
handoff.requested {handoff} A handoff was recorded.
run.requires_action {run} Paused: run.required_action lists the client tool calls. Ends a run stream.
run.completed {run} Terminal.
run.failed {run} Terminal; run.error says why.
run.cancelled {run} Terminal.
run.superseded {run, superseded_by} Terminal: a newer message took over (interrupt).
message.created {message} Session streams: new user messages and your team’s public replies.
session.updated {id, status, mode} For example mode changed to human or back to agent.
session.closed {session} The session was closed.
handoff.assigned, handoff.resolved, handoff.expired, handoff.unclaimed {handoff} Handoff lifecycle, on session streams.
ticket.created, conversation.flagged Raised by the built-in tools.
note.created {message} Internal notes. Only on streams opened by your team, never to end users.
error {code} A stream fault, not a run failure: token_expired, stream_timeout or internal_error.
  • Deltas are previews. message.delta can arrive late, be dropped on a reconnect, or belong to an attempt that was abandoned.
  • message.completed is the truth for a message’s text. Replace whatever you built from deltas with it:
    • discarded: true — that attempt was abandoned (for example a retry on another model) and the client should drop it;
    • replaced: true — the text was replaced by your verbatim escalation reply;
    • truncated: true — the model ran out of reply tokens; you get the partial text.
  • run.* terminal events are the truth for how the turn ended.
  • Ignore what you don’t know. New event types and fields are added over time without a version change.

With the agent setting conversation.stream_mode: "auto" (the default), text is buffered one round at a time whenever an escalation reply is configured, so a streamed sentence can never contradict the reply your policy substitutes.

Reconnect with the last id you received, in the Last-Event-ID header or the after query parameter. You get every durable event after it, then the live stream:

نافذة الطرفية
curl -N https://api.k-agent.kerneltics.com/v1/runs/run_01k6rz7d9f1h3k5n7q9s1v3x5z/events \
-H "Authorization: Bearer $KAGENT_API_KEY" \
-H "Last-Event-ID: 4185"
  • A run stream with no cursor starts from the beginning of the run.
  • A session stream with no cursor starts live; ?after=0 replays everything still retained (events are kept for 7 days).
  • An idempotent retry of a streaming request re-attaches to the original run’s stream from the start.
  • An error before the stream starts (bad key, session_busy, validation) is a normal JSON error response with its HTTP status.
  • After the stream starts, a failed run always ends with run.failed. The error event is used only for stream faults; the client should reconnect with Last-Event-ID (and, for token_expired, a fresh token).
  • When no run is needed — the session is in human mode, or the message is a duplicate — the stream sends message.created, then session.updated, and closes.
  • Publishable keys and client tokens (browsers) receive a reduced set: message.*, run.created|completed|failed|cancelled|superseded, handoff.requested, session.updated with {id, status, mode}, and tool_call.created with {call_id, name}.