# Use the OpenAI SDK

> Point the official OpenAI SDKs for Python and JavaScript at K-Agent — the agent slug is the model — with sessions, streaming, end users and the kagent extension.

K-Agent speaks the OpenAI Chat Completions format at `/openai/v1`. Keep the SDK and code you already have, change two settings, and every request is answered by your agent — with its knowledge, tools, guardrails and handoffs.

| Setting | Value |
|---|---|
| Base URL | `https://api.k-agent.kerneltics.com/openai/v1` |
| API key | Your K-Agent **secret key** (`kt_sk_…`) |
| `model` | Your agent's slug (`store-assistant`) or ID (`agt_…`) |

## A first request

**Python**

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.k-agent.kerneltics.com/openai/v1",
    api_key=os.environ["KAGENT_API_KEY"],
)

completion = client.chat.completions.create(
    model="store-assistant",
    messages=[{"role": "user", "content": "How much is delivery?"}],
)
print(completion.choices[0].message.content)
print(completion.model_extra["kagent"])  # {"run_id": "run_…", "session": None, "handoff": None, "units": 0.25}
```

**JavaScript**

```js
import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://api.k-agent.kerneltics.com/openai/v1',
  apiKey: process.env.KAGENT_API_KEY,
});

const completion = await client.chat.completions.create({
  model: 'store-assistant',
  messages: [{ role: 'user', content: 'How much is delivery?' }],
});
console.log(completion.choices[0].message.content);
console.log(completion.kagent); // { run_id: 'run_…', session: null, handoff: null, units: 0.25 }
```

`client.models.list()` returns your agents, with their slugs as model IDs.

## Stateless or in a session

**Without a session** (the default), each request is independent, like a one-shot [`ask`](/docs/en/guides/one-shot/):

- the **last user message** is the input; earlier messages are passed to the agent as caller-supplied, unverified history;
- leading `system` or `developer` messages return `400 system_message_not_allowed` — the agent's own instructions apply. If the agent allows the `instructions_append` override, they are used as extra instructions for that request instead;
- it is billed like an `ask`.

**With a session**, K-Agent keeps the history and you send only the new message. Name the session with a top-level `session` field (`extra_body` in Python) or the `x-kagent-session` header. It is get-or-create, like `POST /v1/sessions`, and accepts a `sess_` ID or your own `external_id`:

**Python**

```python
completion = client.chat.completions.create(
    model="store-assistant",
    messages=[{"role": "user", "content": "Where is my order?"}],
    user="cus_1042",                      # your end user's ID: verified, since this is your server
    extra_body={"session": "order-8812"},  # get-or-create the session
)
```

**JavaScript**

```js
const completion = await client.chat.completions.create(
  {
    model: 'store-assistant',
    messages: [{ role: 'user', content: 'Where is my order?' }],
    user: 'cus_1042', // your end user's ID: verified, since this is your server
  },
  { headers: { 'x-kagent-session': 'order-8812' } }, // get-or-create the session
);
```

- The stored history is authoritative. K-Agent reads only the messages after the last `assistant` message in your request; if it ignored earlier ones, the response's `kagent.warnings` contains `"client_history_ignored"`.
- Using the same session with a different agent returns `409 session_agent_mismatch`.
- While your team has the conversation (human mode), the response is `200` with empty `content`, `finish_reason: "stop"` and `kagent.mode: "human"`.

## Streaming

`stream: true` returns the usual chunk stream. One K-Agent detail: just before `data: [DONE]` comes **a final chunk with an empty `choices` array** that carries the `kagent` object. Guard for it:

**Python**

```python
stream = client.chat.completions.create(
    model="store-assistant",
    messages=[{"role": "user", "content": "Can I change the delivery address?"}],
    extra_body={"session": "order-8812"},
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
    else:
        kagent = chunk.model_extra["kagent"]  # run_id, session, handoff, units
```

**JavaScript**

```js
const stream = await client.chat.completions.create(
  { model: 'store-assistant', messages: [{ role: 'user', content: 'Can I change the delivery address?' }], stream: true },
  { headers: { 'x-kagent-session': 'order-8812' } },
);
for await (const chunk of stream) {
  if (chunk.choices.length) process.stdout.write(chunk.choices[0].delta.content ?? '');
  else console.log(chunk.kagent); // run_id, session, handoff, units
}
```

When an escalation reply is configured, text is streamed one round at a time, so what you show never contradicts the reply your policy substitutes.

## The `kagent` extension

Every response carries a `kagent` object next to the standard fields (in TypeScript, read it as `(completion as any).kagent`):

| Field | Description |
|---|---|
| `run_id` | The K-Agent run. Fetch its trace with `GET /v1/runs/{run}/steps`. |
| `session` | The session, when you named one. |
| `handoff` | `null`, or the handoff the agent made (`{id, reason_type, summary, status}`). |
| `units` | AI conversation units this request used. |
| `mode`, `warnings` | Present when relevant, e.g. `"human"`, `["client_history_ignored"]`. |

`usage` adds up every model call the agent made for the reply.

## Tools

- The agent's own tools (built-in and HTTP) run inside K-Agent; you don't see them as `tool_calls`.
- To use **client tools**, declare them on the agent, then pass `tools` with the same names in a **stateless** request. When the model calls one, you get `finish_reason: "tool_calls"` as usual; send the result back as a `tool` message in your next request, which is a new stateless run.
- A tool name the agent doesn't declare returns `400 tool_not_declared`, and `tools` together with a session returns `400 tools_not_allowed_with_session`. In sessions, use the [native client tools flow](/docs/en/guides/client-tools/).

## Parameters

| Parameter | Behavior |
|---|---|
| `model` | Agent slug or ID. |
| `messages` | Up to 100. See stateless versus session above. |
| `user` | Maps to the end user's `external_id`, verified. |
| `stream` | Supported. |
| `temperature` | Used only if the agent allows the `temperature` override. |
| `max_tokens`, `max_completion_tokens` | Capped at the agent's `max_reply_tokens`. |
| `n` greater than 1, `logprobs`, `response_format` with `json_schema` | `400 unsupported_parameter`. |
| Anything else | Ignored. |

Errors use the OpenAI shape — `{"error": {"message", "type", "param", "code"}}` — with K-Agent's stable [error codes](/docs/en/reference/errors/). Only secret keys work here, and no CORS headers are sent: call it from your server.
