← Back to the blog
The index against the corpus: 2,509 tokens versus 49,056, or 5.1%

The prompt carries the map, not the territory: 49,000 tokens of docs in 2,500

When you hand documentation to an assistant, the instinct is to paste all of it. The instinct is wrong, and the reason isn't context size — it's the cache. Kobi ships an index worth 5% of the corpus and fetches a whole page only when the question needs it. Here is what it cost, latency bill included.

When you hand documentation to an assistant, the instinct is to paste all of it into the system prompt. It’s all there, the model doesn’t have to go looking, and there’s no retrieval layer that can get it wrong.

Our documentation is 43 pages. Measured: 196,224 characters, roughly 49,056 tokens. That fits comfortably inside any current model’s context window. So the instinct looks right.

It isn’t. And the reason has nothing to do with context size.

What sits underneath, in five lines:

  • The corpus is 49,056 tokens. The index that ships in the prompt is 2,509. That’s 5.1%.
  • What you pay for isn’t occupying context: it’s writing the cache.
  • The headings in the index are the part that makes the model pick correctly on the first try.
  • A fetched page stays in the conversation: follow-ups on the same topic cost nothing extra.
  • The real price is 4.6 seconds, and a visitor has already told us so.

A stable prompt is an asset, and padding it ruins it

Kobi’s system prompt carries a cache breakpoint with a one-hour lifetime. That means the prefix — instructions, tone, product facts, the index — is written once and subsequent conversations read it instead of resending it.

The price difference is not small. A cache read bills at 0.1x the input rate. A cache write bills at 1.25x. Twelve and a half times more expensive than reading it.

// tools goes BEFORE system: the tool definition lands in the same cached
// prefix, which is why it costs nothing per turn.
client.messages.stream({
  model: MODEL,
  tools: [READ_DOC],
  system: [{
    type: "text",
    text: system,                                   // instructions + product + the index
    cache_control: { type: "ephemeral", ttl: "1h" },
  }],
  messages,
});

That cache_control is the whole economics of the design in one line. Whatever sits above it is written once and read many times. Whatever you put inside it, you carry on every write.

That’s the problem with pasting the whole corpus. It isn’t that 49,000 tokens won’t fit: it’s that every time the cache goes cold and the prefix has to be rewritten, you pay for 49,000 tokens at 1.25x to carry text the overwhelming majority of conversations will never look at. Somebody asks about pricing, and their conversation pays for the RBAC page, the OpenShift page and the Falco page.

A stable prompt is an asset. Every token you add to it for the rare case is paid for in the common one.

The map, not the territory

What ships in the prompt is an index: one line per page with its slug, its title, a trimmed description and its second-level headings.

Full corpusIndex
Characters196,22410,036
Approximate tokens49,0562,509
Share100%5.1%

Twenty times smaller. And with that, the model knows what is documented and where, even without knowing exactly what it says.

Three real entries from the index. Each is a single line in the prompt; they’re folded here to be readable:

quickstart | Quick Start | Get KubeBolt running in under 2 minutes.
  | Fastest: Helm chart · Get the admin password · The first-login Setup
    Wizard · Docker Compose · Prefer the Cloud? · Next

api-tokens | API tokens | kbs_ service tokens and kbk_ API keys — how to
  create them, what scopes they get, and why a service token i...
  | Two kinds · Creating one · Scopes · Cluster targeting · The
    public-edge restriction · Revoking

troubleshooting | Troubleshooting | Symptom, cause, fix — for the failures
  that actually happen: empty dashboards, the limited-access banner, 5...
  | The agent is connected but the dashboards are empty · Amber banner:
    "Limited access…" · 503 "cluster not connected" · The admin password
    is lost · Kobi's panel is missing · Still stuck

Descriptions are trimmed to 110 characters and headings capped at nine per page. A visitor asking why their dashboard is empty has the answer located before a single page is fetched.

When a question needs a page’s contents — an install command, a Helm value, an environment variable, a specific symptom — it calls a tool that hands back the whole page. At most two pages per answer, and only slugs that appear in the index: the tool’s schema enumerates them, so the model cannot invent a page that doesn’t exist.

const READ_DOC = {
  name: "read_doc",
  input_schema: {
    type: "object" as const,
    properties: {
      slugs: {
        type: "array",
        items: { type: "string", enum: DOC_SLUGS },  // all 43, and only those
        minItems: 1,
        maxItems: 2,                                 // two pages per answer
        description: "Doc slugs from the index, without the /docs/ prefix.",
      },
    },
    required: ["slugs"],
  },
};

The enum is the important part, and it isn’t a typing convenience: it’s what makes a hallucinated path impossible. The model cannot ask for a page that doesn’t exist, because the schema won’t let it.

And here’s the part that makes the design hold up: a fetched page stays in that conversation’s message history. If the visitor asks a follow-up on the same topic, the text is already in front of it. Retrieval is paid once per conversation, not once per question.

The headings are the part that earns its keep

An index with only each page’s title and description also fits, and it’s shorter. But it produces worse behavior, and it’s worth understanding why.

Our index includes each page’s second-level headings: 220 of them across the 43 pages. They are 40% of its size, and they do two things.

The first is that many questions get answered without fetching anything. Somebody asks where token rotation is documented, and the index already carries that section by name: the correct answer is to link the page, not to read it. Without the headings, the model would have to fetch the page just to find out whether it’s in there.

The second is that when it does fetch, it hits on the first try. An index of generic titles forces a choice between similar candidates, and when in doubt the model fetches two pages to be safe. That doubles the cost of retrieval and adds irrelevant text to the context.

Put differently: the headings are 40% of the index, and they eliminate a share of the calls and nearly all of the misses.

The bill: 4.6 seconds and a thumbs-down

This is where this article parts ways with most write-ups about an optimization. The design has a price, it’s measurable, and we measured it.

Fetching a page requires a second round trip to the model. The first decides what to fetch, the second answers with the text in front of it. In production, across our own turns:

First tokenFull answer
No retrieval1,213 ms4,129 ms
With retrieval2,628 ms8,767 ms

It doubles. And that isn’t a theoretical figure: one of the first visitors rated an answer thumbs-down with the reason too_slow. The metrics table told us, without storing a single word of their question, that the turn in question had fetched documentation.

It’s a trade, not a clean win. We take it because 100% of conversations would pay the cache write for the corpus, while only 40% of turns end up fetching a page at all. But if those seven or eight seconds become a problem, the lever is lowering the per-page character cap, not removing the round trip.

Two things that are true today by a very thin margin

When we say this design leaves nothing out, it’s worth saying how far that claim reaches. There are two caveats, and both are close to stopping being true.

The per-page cap is 12,000 characters. It exists so a runaway page can’t blow the context or the bill. Today it never triggers: the longest page is 11,861 characters. A margin of 139 characters, or 1.2%. One more paragraph on that page and we start cutting content with nothing to warn us.

The index caps headings at nine per page. Of the 43 pages, exactly one exceeds that today: the API reference. Its sections past the ninth don’t appear on the map, so a question about one of them depends on the model fetching the page by its title rather than by the section.

Neither is a defect right now. Both are true by a thin margin, and neither complains when it stops being true. It’s exactly the third bucket we wrote about in auditing a site against its own code: not an error, but a claim that’s correct today with nobody watching it.

What we did about it was the honest minimum: write them down here, and leave both numbers measurable with one command.

What’s left

The obvious next step would be retrieving by section instead of by page: if the index already knows the headings, fetching only the relevant section would cut both the added text and the second of latency it costs to process it.

We haven’t done it, and the reason is that we don’t yet have data justifying it. With the median page at 4,013 characters, fetching the whole page is cheap and gives the model the surrounding context, which is sometimes exactly what was needed. When the metrics say we’re pulling large pages to answer small questions, that will be the moment.

In the meantime, the design fits in one sentence: the prompt carries the map, and the model asks for the territory when it needs it. And the part that shows up on no diagram is that we know what it costs, because we measured it.