# Usage and billing

> The one billable unit — the AI conversation — how it is metered, how model weights apply, plans and quotas, and the usage API.

K-Agent bills one unit everywhere: the **AI conversation**. The dashboard, the API and the pricing page all use the same definition.

## The unit: AI conversation

> **Definition:** one end user (or one session when anonymous) with one agent, in a fixed 24-hour window starting at the agent's first reply, up to 50 agent replies. Reply 51, or any reply after 24 h, opens a new conversation.

- A **reply** is a run that ends with the outcome `answered`. Runs that end `handed_off`, `failed`, `cancelled` or `superseded` are never billed, never count as replies and never open a window.
- The window is per **end user**, not per session: a verified customer who opens three sessions with the same agent in one day uses one window. Without an end user, the session is the subject.
- A window also closes after **300,000 input tokens**; the next reply opens a new one.
- Human conversations are never billed. Handoffs and your team's replies cost nothing, and seats are not the meter.

## The meter

| Event | Units |
|---|---|
| First agent reply in a window | 0.25 × w |
| Second reply | + 0.75 × w (the window now totals 1.0 × w) |
| Replies 3–50 | 0 |
| One-shot `ask` | 0.25 × w × the number of started 50,000-token blocks of input (at least 1) |
| Runs that hand off, fail, are cancelled or are superseded | 0 |
| Runs that fail at the provider | 0 (the provider cost is still recorded) |
| The dashboard's test panel | 0 (the provider cost is still recorded) |

**w** is the model weight. It is fixed when the window opens: the weight of the model that produced the first reply. One-shot input counts every model call of the run, cached tokens included. Stateless calls to the OpenAI-compatible endpoint are billed like `ask`.

### Model weights

| Tier | Models | Weight |
|---|---|---|
| Standard | deepseek-flash, gpt-6-luna, Gemini 3.5 Flash-Lite | 1× |
| Advanced | Gemini 3.8 Flash, Claude Haiku 4.5, deepseek-v4-pro | 3× |
| Premium | Claude Sonnet 5.5 | 6× |
| Your own key (BYOK) | Any model, billed by your provider | 0.2× (platform fee only) |

The table is published and reviewed every quarter. `GET /v1/models` returns each model's tier, weight and whether your plan allows it.

### Worked examples

| What happened | Units |
|---|---|
| A customer asks one question on a Standard model | 0.25 |
| The same customer sends 8 messages that day | 1.0 |
| They come back 30 hours after the first reply and ask again | 1.0 + 0.25 |
| 60 replies within 24 hours | 1.0 (replies 1–50) + 1.0 (replies 51–52 open a new window; 53–60 are free) = 2.0 |
| A three-reply conversation on Claude Haiku 4.5 (Advanced) | 1.0 × 3 = 3.0 |
| A three-reply conversation on your own OpenAI key | 1.0 × 0.2 = 0.2 |
| A one-shot ask with 120,000 input tokens on a Standard model | 0.25 × 3 = 0.75 |
| The agent hands the first question to your team | 0 |

## Plans

Prices in Saudi riyals, excluding VAT.

| Plan | Price / month | Included AI conversations (Standard models) | Overage | Seats |
|---|---|---|---|---|
| Free | 0 | 100 (hard cap) | — | 2 |
| Starter | 299 | 500 | SAR 1.00, invoiced monthly | 5 |
| Business | 899 | 2,000 | SAR 1.00, invoiced monthly | 15 |
| Scale | 2,499 | 7,500 | SAR 1.00, invoiced monthly | 40 |
| Enterprise | Custom | Custom | Custom | Custom |

Advanced counts 3, Premium 6, your own key 0.2 per conversation.

- **Models per plan:** Free includes Standard models and your own keys; Starter adds Advanced; Business and above include every model.
- **Plan changes** are made by the Kerneltics team in v1 — contact us to upgrade. Online payment and ZATCA e-invoicing arrive in v1.1.

## Quotas

- The quota is **per organization**, per calendar month in Asia/Riyadh time.
- It is checked when a window **opens**. A window that is already open always finishes.
- **Free stops at 100 units.** Sessions then reply with a short notice and hand the conversation to your team — customers never see a billing error — and one-shot calls return `429 quota_exceeded` with `x-should-retry: false`.
- **Paid plans never stop** in v1. Units beyond the included amount are reported as `overage_units` and invoiced monthly. A contract can set a hard cap instead.
- Alerts fire at **75% and 100%**: a banner in the dashboard and the `usage.threshold_reached` webhook with `{percent}`.
- If the meter itself has a problem, it fails open: your agent keeps answering.

### Daily cost caps

To protect you from runaway spend, each organization also has a daily cap on model cost (Free $1, Starter $10, Business $30, Scale $100 by default; adjustable on request). Past the cap, sessions hand off with a notice and `ask` returns `429 cost_cap_exceeded` until the next day.

Test-panel runs on platform models are limited to 200 per project per day (`429 playground_limit_reached`); with your own provider key they are unlimited.

## Usage API

```bash
curl https://api.k-agent.kerneltics.com/v1/usage/summary \
  -H "Authorization: Bearer $KAGENT_API_KEY"
```

```json
{
  "object": "usage_summary",
  "plan": "starter",
  "period": { "start": 1790802000, "end": 1793480400, "timezone": "Asia/Riyadh" },
  "included": 500,
  "used": 212.5,
  "remaining": 287.5,
  "overage_units": 0,
  "hard_cap": false
}
```

`GET /v1/usage?group_by=day|agent|model|channel&from=…&to=…` breaks units, tokens and cost down over time. Both need the `usage:read` scope. Every run also reports its own `usage.units` and `usage.weight`.
