# Context and compression

What the agent actually reads on every turn, what it costs, and the two different ways to make the brief smaller.

Source: https://docs.omazy.ai/how-to/agent/context/

import Figure from '../../../../components/Figure.astro'

Every reply your agent writes is produced from a bundle of text assembled fresh
for that turn. The brief is one part of that bundle. Knowing what the other
parts are explains most of the behaviour people find surprising, including why a
long brief is expensive and a long conversation is not.

## What the agent reads on a turn

Four layers are assembled, in this order, plus anything retrieved for the
question being asked.

| Layer | What it is | Changes |
|---|---|---|
| **1. System** | Your published brief, plus the model and channel | Only when you publish |
| **2. Long memory** | What is known about this customer across conversations | Slowly |
| **3. Session summary** | A condensed account of earlier in this same conversation | As the chat goes on |
| **4. Short memory** | The most recent turns, verbatim | Every turn |

<Figure
  label="The four context layers assembled for one turn"
  caption="Layers 2 to 4 are managed for you and vary per conversation. Layer 1 is your brief: identical on every turn, and the only one you author."
>
<svg viewBox="0 0 720 300" xmlns="http://www.w3.org/2000/svg">
  <text x="0" y="14" class="d-eyebrow">ASSEMBLED FRESH, EVERY TURN</text>

  <rect x="0" y="30" width="330" height="46" rx="8" class="d-box-accent d-pulse" />
  <text x="16" y="52" class="d-label">1 · System</text>
  <text x="16" y="68" class="d-sub">Your published brief. Model and channel.</text>
  <text x="248" y="52" class="d-accent-text">yours</text>
  <text x="248" y="68" class="d-sub">fixed cost</text>

  <rect x="0" y="86" width="330" height="46" rx="8" class="d-box" />
  <text x="16" y="108" class="d-label">2 · Long memory</text>
  <text x="16" y="124" class="d-sub">What is known about this customer.</text>

  <rect x="0" y="142" width="330" height="46" rx="8" class="d-box" />
  <text x="16" y="164" class="d-label">3 · Session summary</text>
  <text x="16" y="180" class="d-sub">Earlier in this same conversation.</text>

  <rect x="0" y="198" width="330" height="46" rx="8" class="d-box" />
  <text x="16" y="220" class="d-label">4 · Short memory</text>
  <text x="16" y="236" class="d-sub">The most recent turns, verbatim.</text>

  <rect x="0" y="254" width="158" height="38" rx="8" class="d-box" />
  <text x="16" y="272" class="d-label">Retrieved</text>
  <text x="16" y="286" class="d-sub">when relevant</text>

  <rect x="172" y="254" width="158" height="38" rx="8" class="d-box" />
  <text x="188" y="272" class="d-label">Tools</text>
  <text x="188" y="286" class="d-sub">what it may call</text>

  <path d="M338 160 H 396" class="d-arrow" />
  <polygon points="396,155 408,160 396,165" class="d-arrow-head" />

  <rect x="416" y="112" width="300" height="96" rx="12" class="d-box" />
  <text x="436" y="146" class="d-label">One reply</text>
  <text x="436" y="166" class="d-sub">The model reads all of the above,</text>
  <text x="436" y="182" class="d-sub">then writes a single answer.</text>
</svg>
</Figure>

On top of those, two more things arrive when they are relevant:

- **Retrieved knowledge.** Passages pulled from your knowledge base and from any
  files the customer shared into this conversation, merged by relevance and
  de-duplicated.
- **Tools.** Descriptions of what the agent is allowed to call.

Layers 2 and 3 exist so a long conversation does not become an expensive one.
Rather than resending an hour of chat, the platform keeps a summary of the older
part and only the recent turns in full. Roughly the last twenty messages carry
verbatim; anything purged for data protection is excluded and never rebuilt.

## The one layer you author is the one that never varies

This is the thing worth internalising.

Layers 2, 3 and 4 change per customer and per conversation, and the platform
manages their size for you. **Layer 1 is the brief, it is identical on every
turn, and it is entirely yours.**

So the brief is a fixed cost paid on every reply, to every customer, for as long
as the agent runs. A hundred wasted words in a Context prompt is not a hundred
wasted words. It is a hundred words times every conversation you will ever have.

That is the whole argument for keeping the brief terse, and it is why the editor
puts a number on it.

## Reading the numbers

The Prompts editor shows the running cost in the composed panel.

| Reading | Means |
|---|---|
| **Context cost** | What every turn will carry once you publish |
| **Currently live** | What the agent is paying right now |
| **After you publish** | What it will pay if you publish the working copy |

The number quoted is the *compacted* size, not the raw size, because compacting
is what publish actually writes. Quoting the raw figure would have you budgeting
against a bill the agent never receives.

If the two figures differ, you have unpublished changes.

## Two ways to make it smaller, and they are not the same

People treat these as one feature. They are opposites, and the difference is
about who is allowed to change your words.

### Compact: automatic, lossless, always on

Whitespace only. Extra blank lines and stray indentation are trimmed when you
publish. Nothing is reworded, nothing is dropped, and no model is involved.

You do not turn this on. It already happened. When the editor says spacing is
already trimmed, or that a prompt is as short as it safely gets, that is
compacting telling you there is no free space left.

### Shorten with AI: manual, lossy, proposal only

This is the lever that can actually shrink prose, and it is deliberately
awkward to use, because a reworded rule is a changed rule.

The rules it operates under are worth stating plainly:

- **It never writes.** You get a proposal and a diff. Accepting is a separate,
  explicit action, and accepting is a normal edit, so it creates a version.
- **It never runs on publish**, and it never runs during a customer conversation.
  Nothing is being quietly rewritten behind you.
- **Rules are excluded.** A Rule prompt gets whitespace tidying only, never
  rewording. This is the single most dangerous place for a model to be helpful,
  so it is not allowed to be.

### The dropped-invariant check

When a shortened version comes back, it is scanned against the original for
things that must not vanish: **numbers, prices, URLs, and code blocks**. Anything
present before and missing after is flagged for you.

This check is deterministic, needs no model, and exists because of one specific
failure: a price silently disappearing from a paragraph that otherwise reads
fine. A human skimming a diff will catch a changed sentence and miss a missing
number every time.

Flags do not block you. They tell you where to look before you accept.

## A practical order of operations

When the brief is too big, work in this order. It goes cheapest and safest
first.

1. **Disable prompts you are not sure about.** Free, reversible, and it answers
   the real question, which is whether a prompt is earning its tokens.
2. **Move facts to [knowledge](/how-to/agent/knowledge/).** A fact in knowledge is
   retrieved only when relevant. The same fact in a Context prompt is paid for on
   every turn, including all the turns where nobody asked.
3. **Cut samples to the ones that teach.** Three well-chosen examples beat twelve
   near-duplicates, and near-duplicates are the most common form of brief bloat.
4. **Then shorten with AI**, and read the diff.

Most briefs that feel too big have a Context prompt doing the job of a knowledge
base. Step 2 usually finds it.

## When answers get cut off mid-sentence

Worth knowing because it looks like a context problem and is not.

If replies stop partway through a sentence, the output limit is too low rather
than the context being too big. An unset maximum is not an unlimited one; leaving
it unset means inheriting a default that is smaller than most people expect. See
[Providers and models](/reference/llm-gateway/providers/).
