Tutorials

How to Reduce AI Agent Context and Tool-Call Overhead

A real AI-agent workflow audit shows how task routing, compact state checks, shared validation scripts, and smaller handoffs can reduce context and tool-call overhead.

Approximately 11 min read

Long-context models make it tempting to solve every agent problem by adding more context.

More instructions. More repository history. More tool definitions. More handoff notes. More shell output.

That works for a while.

Then the agent starts spending a surprising amount of its trajectory rediscovering state, rereading instructions, waiting for commands, and carrying old information that is no longer relevant to the current task.

I recently audited one real multi-step coding-agent workflow after noticing exactly that pattern.

The workflow itself was ordinary: a private repository, generated output, validation gates, browser checks, and a production deployment. The interesting part was not the application. It was the agent overhead around it.

After restructuring the workflow, the always-loaded repository instruction payload for a common publishing task fell from about 15.1 KB to 7.8 KB, a reduction of roughly 48%.

That is not a claim that total model token usage fell by 48%.

It is a narrower measurement: the repository guidance that had to be loaded for the task became about half as large.

The bigger lesson was that prompt length was only part of the problem. Repeated tool choreography was at least as important.

The agent loop has more overhead than the visible prompt

A user may type one short request, but a coding agent can receive much more than that.

OpenAI’s description of the Codex agent loop breaks the model input into instructions, tools, and user or task inputs. Tool definitions and repository instructions are part of what the model must work with, not free metadata outside the context. OpenAI also documents that Codex can automatically load AGENTS.md files into the conversation.

That means an agent session can accumulate several kinds of overhead:

persistent repository instructions
task-specific instructions
chat history
tool schemas
tool results
shell output
retrieved files
handoff notes
generated patches
validation logs

If you want a deeper explanation of why this matters at the model level, see What Does Context Length Mean in an LLM?.

The important point for agent engineering is simpler:

Context is not only the code you want the model to reason about. It is everything you force the model to carry while doing the work.

A larger context window raises the ceiling. It does not make irrelevant context free.

What I measured

I looked at two different sources of overhead.

The first was repository instruction load: how much guidance the agent was told to read before starting a common task.

The second was tool-call churn: how often the workflow repeated predictable terminal and verification operations.

For privacy, I am intentionally not publishing repository names, local paths, device identifiers, deployment URLs, authentication configuration, or raw logs. None of those details are necessary to reproduce the engineering pattern.

The instruction audit looked like this:

Workflow state Guidance loaded for a common task
Before ~15.1 KB
After ~7.8 KB
Reduction ~48%

Again, this is not total tokens. It is the size of the repository guidance that the workflow required the agent to load.

The tool audit was even more revealing.

Across one eight-day window, the tool bridge recorded 4,731 calls. Three terminal-process operations dominated the trace:

start process        1,570
interact with process 1,340
read process output  1,148

Together, those three operations represented the large majority of recorded calls.

Not all of them were waste. Long-running commands genuinely need process management.

But the trace showed a repeated pattern:

start shell command
wait
read output
run next validation command
wait
read output
run HTTP check
wait
read output

Much of that sequence was deterministic.

That means it belonged in scripts, not in the model’s reasoning loop.

Mistake 1: turning AGENTS.md into an encyclopedia

A repository instruction file usually starts small.

Then every incident adds another paragraph.

A deployment bug adds deployment rules. A broken image adds image rules. A privacy incident adds security guidance. A benchmark adds benchmarking notes. A future agent inherits all of it.

Eventually the file becomes a miniature operations manual that is injected into tasks whether those rules matter or not.

This creates three problems.

Irrelevant rules compete with relevant rules

If a task is editing one article, the model probably does not need the full deployment history, every security incident, every image-generation rule, and every benchmark convention.

The important constraints become harder to distinguish from background material.

Old state looks authoritative

Historical notes are especially dangerous because they often look like current instructions.

An agent may read:

current branch is ...
current build is ...
temporary workaround is ...

days or weeks after those statements stopped being true.

Stale state can be worse than missing state.

Humans stop maintaining giant instruction files

Once a guidance file becomes painful to review, people append instead of refactor.

The file becomes a log rather than an interface.

OpenAI described a similar lesson in its own agent-first engineering work: use AGENTS.md as a map to repository knowledge rather than a giant instruction manual.

That matched what I saw in practice.

Fix 1: make AGENTS.md a router

The first change was to reduce the root instruction file to durable rules and task routing.

Conceptually:

AGENTS.md
|
|-- always-applicable safety rules
|-- source-of-truth rules
|-- validation requirement
|
|-- publishing task  -> docs/workflows/publish.md
|-- image task       -> docs/workflows/images.md
|-- release check    -> docs/workflows/release.md
|-- security change  -> SECURITY.md

The root file no longer tries to teach the agent the entire repository.

It answers two questions:

  1. What rules are always true?
  2. Where should the agent look for this task?

This is a useful distinction.

An agent needs navigation context on every task.

It does not need every destination’s contents on every task.

Fix 2: load task-specific policy only when it becomes relevant

After turning the root file into a router, I moved operational detail into small workflow documents.

For example, a content-publishing task could load:

AGENTS.md
editorial policy
daily-publish workflow

An image task could add:

hero-image workflow

A security change could load:

SECURITY.md

This sounds obvious, but the effect is important.

The task no longer pays the context cost of every unrelated workflow.

The result was the measured change from approximately 15.1 KB of always-required repository guidance to about 7.8 KB for the common path I audited.

The exact numbers will vary by repository.

The architecture is the reusable part.

Fix 3: generate current state instead of writing current state

My first instinct was to create a CURRENT_STATE.md.

I decided against it.

A hand-maintained current-state file creates another synchronization problem. It can become stale the moment someone pushes a commit, switches branches, generates output, or leaves uncommitted changes.

Instead, I added a small preflight command that computes state when needed.

A generic version might report:

{
  "branch": "feature/example",
  "head": "abc123...",
  "upstream": "def456...",
  "ahead": 2,
  "behind": 0,
  "dirty": [],
  "build": "current-build-marker"
}

The principle is:

Dynamic state should usually be measured, not documented.

This reduces handoff size and removes a class of stale-context bugs.

A good handoff can now say:

Task: publish one article
Branch: feature/example
Changed: one source file
Validation: passed
Deployment: pending
Blocker: none

It does not need to retell the repository’s history.

Fix 4: collapse deterministic tool choreography into scripts

This was the most practical change.

The original workflow expected the agent to remember and run several validation commands individually.

Conceptually:

generate
generator --check
content-check
javascript-syntax-check
access-tests

That is fine for a human terminal session.

For an agent, each separate step can become:

model decision
tool call
tool result
model decision
tool call
tool result
...

If the order is fixed and the failure policy is fixed, the model adds little value by deciding the sequence every time.

So I replaced the repeated sequence with one wrapper:

python tools/validate.py

The wrapper runs the same gates in the same order and stops on failure.

The same idea applies to deployment verification.

Instead of repeatedly constructing several HTTP commands, a single verifier can check:

expected build marker
public page status
authentication boundary
required redirect behavior

The model still interprets failures.

It just does not need to reinvent the plumbing that produces them.

The right boundary: script mechanics, let the agent handle judgment

It would be a mistake to script everything.

The useful split is:

Scripts should handle deterministic mechanics.

Examples:

The agent should handle judgment.

Examples:

This division reduces tool-call overhead without turning the agent into a glorified shell alias.

A compact agent workflow

After the refactor, the common path became roughly:

preflight
   |
inspect task-specific source
   |
edit
   |
validate
   |
visual review if required
   |
merge
   |
verify production

Each phase has a clear boundary.

The model does not need a giant handoff explaining everything that happened before it.

The repository itself carries durable policy.

Scripts produce current state.

Version control carries the code history.

The agent context contains what is relevant now.

Why fewer tool calls can matter even when the model is fast

Tool-call overhead is not only about model tokens.

Every call can involve some combination of:

A 200-millisecond command can become a much larger end-to-end interaction when wrapped in an agent loop.

The effect is especially visible with shell processes that require a second call just to retrieve output.

This is why a one-command validator can improve the workflow even if the underlying validation commands take exactly the same amount of CPU time.

The optimization target is not necessarily:

make tests execute faster

It may be:

make the agent negotiate the test sequence fewer times

Those are different problems.

Tool availability should not mean tool visibility everywhere

There is a related lesson about tools and plugins.

An agent platform may support dozens of integrations.

That does not mean every task benefits from exposing every integration to the model.

OpenAI’s description of the Codex loop explicitly treats tools as part of the model-facing interface. Large tool surfaces can therefore add both cognitive and orchestration cost.

The best design is usually capability routing:

task needs Git
-> expose Git

task needs browser verification
-> expose browser verification

task does not need calendar, email, CRM, analytics, or deployment controls
-> keep them out of the active path

This is also a security advantage.

Reducing unnecessary capability reduces unnecessary trust. I discussed that side of the problem separately in Why Enterprises Restrict Open-Source AI Tools.

What I would measure next

The current audit establishes a useful baseline, but it is not a controlled benchmark of model efficiency.

A stronger experiment would capture, per task:

input tokens
output tokens
cached input tokens
tool-call count
tool-result bytes
wall-clock time
model turns
failed/retried calls
files read
bytes read
validation commands executed

Then run matched tasks under two harnesses:

A: monolithic instructions + manual command choreography
B: routed instructions + scripted deterministic operations

Keep the model, task, repository revision, and tool permissions fixed.

That would let us separate several effects that are currently mixed together:

Until that controlled comparison exists, claims about total token or time savings should stay conservative.

Practical rules that survived the audit

If I were cleaning up another coding-agent repository tomorrow, I would start with these rules:

  1. Use the root instruction file as a map, not a wiki.
  2. Keep durable policy separate from current state.
  3. Generate current state from Git and build artifacts.
  4. Move task-specific procedures behind explicit routes.
  5. Wrap deterministic multi-command sequences in one script.
  6. Keep visual, editorial, architectural, and security judgment with the agent.
  7. Measure tool calls as well as tokens.
  8. Do not report instruction savings as total agent savings.
  9. Keep handoffs short enough that a new session can actually use them.
  10. Treat old context as a liability unless the current task needs it.

The deeper lesson

Agent optimization is often framed as a model-selection problem:

Should I use a bigger model?
Should I use more reasoning?
Should I increase context?

Those questions matter.

But a capable model can still be wrapped in a wasteful harness.

If every task begins by loading a small encyclopedia, discovering stale state, and manually replaying the same command sequence, a larger context window mostly gives the system more room to carry the inefficiency.

The better question is:

What information and actions genuinely require model judgment on this task?

Everything else is a candidate for routing, measurement, or scripting.

That is the part of agent engineering I was underestimating.

The model did not need to get smaller.

The harness did.

Sources and further reading

Continue reading