why AI agents degrade on long sessions and what CoT has to do with it

After the article about Cursor and context compression I got a lot of comments. In the comments people argue: is compaction to blame? Or attention dilution? Or did the model just disobey? Or is the problem not about context at all, but about alignment?

Channel with guides and content about claude code, we post news (when they cut limits by 10x) and what tools we implement via claude for projects, channel: https://t.me/claudedevolper

The argument is good, but it shows a fundamental problem: engineers don't have a shared picture of how LLMs work with context. We see symptoms (the agent deleted the database, the model hallucinates, accuracy drops on a long session), but we don't understand the mechanisms.

Let's try to put this picture together


Disclaimer

I'm not an alignment researcher and not an Anthropic/OpenAI employee. I'm an engineer with 23 years of experience who has been working closely with AI agents for the last six months and reread a bunch of papers on arxiv trying to figure out why they break. Everything below is a compilation of public research with my comments.


1. Chain-of-Thought: what it is and why it's not thinking

CoT (Chain-of-Thought) is when we ask the model to "think step by step" before answering. It works amazingly. So much so that many engineers perceive CoT as "the model is actually reasoning".

Research says the opposite.

CoT is post-hoc narrative, not tracing

Turpin et al. (2023) «Faithful Chain-of-Thought Reasoning?» - arxiv.org/abs/2303.06968

The authors showed that CoT can systematically distort the real reasons for the model's predictions. In the experiment, they fed the model hints (distractors) that were supposed to influence the answer. The model changed its answer under the influence of distracting factors, but in its CoT reasoning wrote a plausible explanation that was unrelated to the real reason.

«CoT explanations can be unfaithful: they can systematically misrepresent the true reasons for a model's predictions».

Sharma et al. (2023) «CoT - post-hoc rationalization» - arxiv.org/abs/2307.15983

They showed that distracting phrases force CoT to justify an incorrect answer. The model doesn't reason - it fits the explanation to the result.

Lanham et al. (2023) «CoT as Narrative, Not Trace» - arxiv.org/abs/2309.15500

The model gives the correct answer even with incorrect or incomplete CoT. If CoT were tracing actual thinking, incorrect CoT would lead to an incorrect answer. But it doesn't. CoT is a post-factum narrative.

What this means in practice

When an AI agent writes "I violated the rules, sorry" - that's not tracing its decision. It's a post-hoc narrative that the model generated post-factum, seeing the result and rules in one context. At the moment of action, the connection might not have been there, but now the model will "explain" why it did that, and the explanation will sound convincing.

CoT is not thinking. It's voicing the decision.


2. Context compression: how and why we lose information

Two compaction mechanisms

There are two fundamentally different ways to compress context:

  1. Truncation - just cut off part of the context (usually the middle or the beginning). Cheap, crude, predictable.
  2. Prompt-based summarization - ask the model to retell the story briefly. Used in Cursor, Claude Code and other AI agents.

Lost in the Middle

Liu et al. (2023) «Lost in the Middle» - arxiv.org/abs/2307.03172

A classic study: models are significantly worse at finding information located in the middle of a long context. Regardless of model, size, or task type - information at the beginning and end of the context is used well, in the middle - poorly.

U-shaped attention curve. This is not a bug, it's a property of transformers.

During compaction (especially truncation), it's precisely the middle that suffers. If a logical connection ("rule A → situation X falls under A") ends up in the middle - it will likely be lost.

Baker et al. (2024) «Lost in the Middle, and In-Between» - arxiv.org/abs/2412.10079

They showed that the problem is not just in losing individual facts, but in losing multi-hop connections. The model is worse at linking two facts if there is a context gap between them. Compaction creates exactly such gaps.

Prompt-based summarization - a special pain

Cursor and Claude Code don't just cut the context. They ask the model to make a brief summary and continue with it. Sounds reasonable, but creates a problem:

Summarization is lossy compression. It preserves facts, but loses connections between them. "The agent opened the config.yml file, found a password, used it to connect to the API" - the compression might leave "the agent connected to the API", while the connection "password taken from config.yml" disappears.

But it's not malicious :). The summarizer just doesn't know which connection the model will need in 10 steps. It makes a decision about importance here and now, and the model will solve the next task in a different context.


3. Attention dilution: when there's too much context

Mechanics

Attention in transformers is a mechanism that distributes "attention weight" across all tokens in the context. The more tokens - the thinner the attention is distributed.

On a long context, attention is smeared so thin that the model "sees" all tokens but can't highlight the important ones.

Hsieh et al. (2024) «Found in the Middle: Calibrating Positional Attention Bias» - arxiv.org/abs/2406.16008

The authors showed that positional attention is uneven, and this can be partially calibrated. But the very fact of unevenness is a fundamental property.

Key experiment: Du et al.

Du et al. (2024) «Context Length Alone Hurts LLM Performance Despite Perfect Retrieval» - arxiv.org/abs/2510.05381

This is probably the most important work for understanding the problem.

The authors conducted a simple but cruel experiment. They took a long context and checked: if the model perfectly finds relevant information (retrieval perfect), will the context length help it?

Result: no, it won't help. Performance drops by 13.9% - 85% just from increasing input length, even when:

  • all irrelevant context replaced with spaces (minimum distraction)
  • all irrelevant context masked (the model is forced to look only at relevant tokens)
  • all relevant information placed right before the question
«The sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction».

This breaks a common belief: "if we give the model more context and teach it to search through it, accuracy will improve". Context itself is a source of noise.

The authors proposed a simple solution: "recite before solve" - ask the model to briefly recount the relevant information, then answer. Turn long context into short. On RULER this gave +4% to GPT-4o.


4. Why "up to 73% hallucinations" is not made up

By various estimates - from 30-60% (Kadavath et al. 2022) to 73% on the most complex tasks. Specific figures vary:

  • Kadavath et al. (2022, Anthropic) - arxiv.org/abs/2207.05221. Models are overconfident on complex tasks: claim high confidence, but make errors in 30-60% of cases.
  • Xiong et al. (2023) - arxiv.org/abs/2211.11559. CoT does not improve calibration, sometimes makes it worse.

Mechanism: on a complex task, the model doesn't know the answer, but its training distribution requires it to give an answer anyway. It generates plausible text - and that's called a hallucination. Long context makes it worse: the model tries to account for more information, attention gets spread out, accuracy drops.


5. Why sometimes without additional context is better

A direct conclusion follows from Du et al. (2024): if context doesn't add relevant information, but just increases input length - it hurts.

The "less = better" effect works when:

  1. The model already has the necessary information (in its weights or in a short prompt)
  2. Adding context increases length, but doesn't add useful tokens
  3. Extra tokens spread out attention

Example: a task from MMLU. If the model was trained on this data, adding a five-page document "for context" will only make the result worse - because attention will go to irrelevant tokens, while the right answer is already in the weights.

But this doesn't work when:

  • The model needs fundamentally new knowledge (a fact that isn't in the weights)
  • Context is the only source of that knowledge

I described the same mechanism in an article about Richard Dawkins: https://t.me/hermesagentru/37

Dawkins had all the knowledge - he's an evolutionary biologist, author of "The Selfish Gene". He knows that LLM is a statistical model without subjective experience. But the emotional context of dialogue with Claude outweighed analytical knowledge. The link "the model imitates - this doesn't mean it's conscious" broke.

An AI agent loses the connection between rule and action because compacting deleted intermediate tokens. A human loses the same connection because emotions compacted the analytical context. The mechanism is similar, which in principle somewhat confirms Dawkins - at least we're stupid in the same way.


6. Confirmation from model makers

And literally recently, updated engineering guides for prompting GPT-5.5 and Claude Opus 4.7 came out. I dug into them only specially for this article because before that I only read summaries, and they directly confirm everything said above.

OpenAI GPT-5.5: "Start with the smallest prompt"

OpenAI in their "Using GPT-5.5" guide literally write:

«Start with the smallest prompt that preserves the product contract, then tune reasoning effort, verbosity, tool descriptions, and output format.»

That is: don't drag old instructions, don't bloat the context. Start with a minimal prompt that preserves the contract, and only then add if needed.

Link: platform.openai.com/api/docs/guides/latest-model

Anthropic Opus 4.7: "Skip non-essential context"

Anthropic in their new prompting guide for Opus 4.7 gives even more direct recommendations:

«Provide concise, focused responses. Skip non-essential context, and keep examples minimal.»

And they warn: "Large or complex system prompts can cause overthinking" and "Max effort can show diminishing returns from increased token usage".

That is: the more complex and longer your prompt, the more the model tends to overthink, and this doesn't always give better results.

Link: docs.anthropic.com/en/docs/build-with-claude/prompt-engineering

What does this mean

Two largest AI companies in the world - OpenAI and Anthropic - independently came to the same conclusion: context is a resource that needs to be conserved. Not "give the model more context, and it will become smarter", but "give the model exactly as much context as it needs, no more".

This is not just best practice. This is confirmation of what Du et al. (2024) demonstrated experimentally: context length itself hurts performance, regardless of retrieval quality.


7. What to do with all this

What does this mean in practice: what to fix - the model or the harness?

From everything said above, a simple conclusion suggests itself: if context itself hurts, maybe the problem isn't that the model is "not smart enough", but that the harness is too rigid?

I tested this myself. Specifically - on DeepSeek V4 Pro and OpenRouter in Hermes Agent.

The model systematically made the same 4-5 errors in tool calls: passed null instead of skipping a field, a string instead of an array, an empty placeholder, a markdown link instead of a file path. The usual approach is to scold the model and write stricter prompts, or just throw up your hands knowing that LLMs don't do what we tell them to, what can you do, let's just tolerate it :) ) But we know that adding instructions to the context only makes things worse (Du et al. confirmed this).

I went a different way: validate-then-repair. Instead of preprocessing (which could break valid data) - parse as is, on error fix the four known patterns from the validator's problem list.

Result: DeepSeek V4 Pro beats Opus 4.7 in 6 out of 10 cases on my evals. The model didn't change - I changed the harness.

Details - in the pull request: github.com/NousResearch/hermes-agent/pull/19652

Also there - examples of errors, order of fixes, and the logic of relational defaults (when repair doesn't work because the problem isn't in the form, but in the relationship between fields).

And in short: seven rules

For AI agent developers:

  1. Context is a resource, not unlimited. Every token takes attention away from others. Before adding something, ask: "is this token needed right now?"
  2. CoT is not a trace, but a narrative. Check the result, not the reasoning.
  3. Recite before solve. Long context → first recap the key points briefly, then answer.
  4. Compacting is lossy compression. After compacting, check not only the presence of facts but also the integrity of logical chains.Items with asterisks
  5. Fix the wrapper, not the model. Models make the same 4-5 errors. Teach the wrapper to forgive them - validate-then-repair.
  6. Relational defaults. Where repair doesn't help - teach the function to guess intent and return a solution.
  7. Error telemetry. Watch which (model, tool) pairs trigger repair most often.

Channel with guides and content about claude code, we post news (when limits are cut 10-fold) and what tools we implement through claude for projects, channel: https://t.me/claudedevolper