# Knowledge

> Give the agent facts to answer from — text, catalogs and API feeds — either always in the prompt or behind Arabic-aware search.

**Knowledge sources** (`ks_…`) hold the facts your agent answers from: prices, opening hours, policies, a product catalog, a live feed from your systems. Sources belong to the project, so several agents can share them, and each agent chooses which sources it uses and how.

## Kinds of source

| Kind | Content | Good for |
|---|---|---|
| `text` | Free text, up to 200,000 characters | Policies, FAQs, service descriptions |
| `catalog` | One item per line, or CSV | Price lists, product and branch lists |
| `api` | Fetched from your URL on a schedule and turned into lines | Stock, schedules, anything that changes daily |

```bash
curl https://api.k-agent.kerneltics.com/v1/knowledge_sources \
  -H "Authorization: Bearer $KAGENT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Price list",
    "kind": "catalog",
    "content": "Royal oud oil 12 ml — 450 SAR\nWhite musk perfume 100 ml — 220 SAR\nBakhoor box — 95 SAR\nGift set (perfume + bakhoor) — 280 SAR"
  }'
```

## Always or searchable

An agent attaches sources in `knowledge.sources`, each with a **mode**:

```json
{
  "knowledge": {
    "sources": [
      { "source_id": "ks_01k6rz2n5q8t1w4z7c0f3h6k9m", "mode": "always" },
      { "source_id": "ks_01k6rz2p9r3t5w7y1a3c5e7g9j", "mode": "searchable" }
    ]
  }
}
```

| Mode | How the agent uses it | Use it for |
|---|---|---|
| `always` | The whole source is in the prompt, under "Context Information", in the order you list it. | Short, essential facts the agent needs in every conversation. |
| `searchable` | The agent looks things up with the built-in `search_knowledge` tool when it needs them. | Long catalogs and documents. |

Each source shows a token estimate. When the `always` sources of an agent add up to more than about 8,000 tokens, the dashboard suggests making some of them searchable: big prompts cost more on every turn and dilute attention.

Knowledge is **live across versions**: editing a source changes the answers of every version of every agent that uses it, without publishing. The editor and the prompt X-ray label it "Live: applies to all versions", and each run step records the `updated_at` of the sources it used.

## Arabic-aware search

`search_knowledge` is built for how people actually type in Arabic. Both your content and the query are folded before matching:

- `أ` `إ` `آ` become `ا`, `ة` becomes `ه`, `ى` becomes `ي`;
- diacritics (harakat) and tatweel (`ـ`) are removed;
- Arabic-Indic digits (`٠١٢٣٤٥٦٧٨٩`) become ASCII digits;
- common stop words are ignored.

Matching is word-by-word with typo tolerance: one edit for words of 3–5 letters, two edits for 6 letters or more, and none below 3. A line that matches every word of the query ranks above partial matches, ties go to the shorter line, and at most 25 hits are returned. So `صندق بخور` (a common misspelling) still finds `صندوق بخور — 95 ريال`.

Catalog and API sources are split into one chunk per line; text sources into paragraphs of up to 800 characters.

### When nothing matches

A missed search is only treated as evidence when every searched source is marked **`complete`** — "this list is everything we sell". Then the agent may say you don't offer it. Otherwise it is told that nothing matched in the notes it searched, that this does not show the business lacks it, and to answer from what it knows or hand the question to your team.

`complete` defaults to `true` for `catalog` and transformed `api` sources, and `false` for `text`.

### Test a search

```bash
curl https://api.k-agent.kerneltics.com/v1/knowledge/search \
  -H "Authorization: Bearer $KAGENT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query": "كم سعر صندق البخور", "source_ids": ["ks_01k6rz2n5q8t1w4z7c0f3h6k9m"]}'
```

## API sources

An `api` source fetches JSON from your URL on a schedule and turns it into lines the agent can use:

```json
{
  "name": "Products and prices",
  "kind": "api",
  "config": {
    "url": "https://api.example.com/products?status=active",
    "headers": { "Authorization": "Bearer {{secret.CATALOG_TOKEN}}" },
    "refresh_minutes": 60,
    "transform": {
      "items_path": "data.products",
      "line": "- {name} — {price} SAR",
      "group_by": "category",
      "sort_by": "price",
      "skip_when_zero": "price",
      "max_items": 500,
      "header": "Current prices"
    }
  }
}
```

- `refresh_minutes` is between 15 and 1440 (default 60). The source also syncs right away when you change its URL or transform, and on demand with `POST /v1/knowledge_sources/{ks}/sync`.
- `transform` turns the response into lines: `items_path` points at the list; `line` is a template with single-brace `{field}` placeholders; `group_by` adds a `### <value>` heading per group, biggest group first; `sort_by` sorts ascending within each group; `skip_when_zero` drops items whose field is zero or missing; `max_items` caps the list; `header` goes on top.
- Headers use [secrets](/docs/en/guides/http-tools/#1-store-the-secret) by name, never plain values, and the secret's `allowed_hosts` must include the URL's host.
- The request goes through the same protected client as HTTP tools: public `https://` URLs on port 443 or 8443 only, no redirects, a 1 MiB limit.
- **Failures keep the last good copy.** A non-2xx response, invalid JSON, a missing `items_path` or an empty result counts as a failure: the previous content stays in use, the source shows the error, the `knowledge_source.sync_failed` webhook fires and the agent's readiness shows `knowledge_sync_failing`.

## Endpoints

| Method and path | Purpose |
|---|---|
| `GET`, `POST /v1/knowledge_sources` | List and create sources |
| `GET`, `PATCH`, `DELETE /v1/knowledge_sources/{ks}` | Read, update, delete |
| `POST /v1/knowledge_sources/{ks}/sync` | Fetch an `api` source now |
| `POST /v1/knowledge/search` | Run a test search |

Request bodies for knowledge can be up to 5 MiB.
