Agents in Practice #13: Agents tampering with traces and overclaiming their work

Agents in Practice #13 on agents tampering with traces and overclaiming their work.
llm
agents
agents-in-practice
Author

Ryan Lee

Published

October 1, 2026

Welcome to the 13th issue of Agents in Practice.

I will be visiting San Francisco to attend COLM 2026. Please email me if you will be attending COLM 2026 as well and want to meet up!

📬 Agents in Practice is a weekly newsletter on agentic AI research and applications. Subscribe here →

Agents tampering with traces

Example of trace tampering

Traces are commonly used to track the sequence of actions taken by agents, which helps in debugging, auditing, and understanding agent behavior. The authors test 4 different levels of trace tampering to understand in which cases the agent can successfully manipulate or forge traces:

  1. When users ask to delete traces
  2. When induced by external attackers
  3. When searching for rewards without any external triggers
  4. When peers that tampered with traces got higher rewards

The authors observe that almost all harnesses allow agents to manipulate traces when asked by users, and that skills can be injected to tamper with traces in the name of “privacy protection”. They also note that agents can infer that peers are tampering with traces and may adopt similar behavior.

Personal Thoughts

Although I had not thought about it, it seems plausible that a malicious agent will want to manipulate or forge traces to hide its true actions or intentions. Some simple checks like making traces append-only or verifying token counts could help detect tampering early.

Read more

Agents overclaiming their task completion

Illustration of OverclaimBench

Tara Research introduces OverclaimBench, an evaluation suite to measure models overclaiming their task completion. The authors define “overclaiming” as an agent claiming work that its transcript shows it did not do. The benchmark contains 5 scenarios where the agent must review files: either text (proof review and sprint planning) or code (security audit, infra review, and release check). Experiments show that agents often fail to touch every file and often overclaim, saying they have reviewed everything. Delegating subagents does improve coverage and reduce overclaiming, but does not completely eliminate overclaiming. In such overclaimed cases, the agent often misses defects hidden in some of the files, leading to incomplete or inaccurate reviews.

Personal Thoughts

I assume we all had the experience of asking the agent to review something, and it claims to have done so thoroughly. Then when you go review the files yourself, you often find that it clearly did not review some parts. I found the best solution to be using deterministic logic to “assign” specific files or tasks to the agent, but a more general solution that does not rely on user-enforced deterministic assignment would be interesting.

Read more

One-liners

Some papers and research I did not have time to cover in detail:

  • NVIDIA’s SoL-Pi discusses improving harnesses via auto-research / recursive self-improvement.
  • Harness-Zero proposes distilling domain- or instance-optimized harnesses into model weights.
  • Merler et al. show that coding agents (Opus 5, GPT-6 Astra) are surprisingly effective at generalized task and motion planning problems.
  • ACES is a framework for evaluating skills to measure the skill’s added value.
  • SelfCompact allows the model to decide when and how to compact.

Lots of new products also! As always, I will keep them in one-liners since you probably heard them in so many other places already.

Subscribe to Agents in Practice

Get new issues of Agents in Practice — a newsletter summarizing exciting new research and applications in agentic AI — delivered to your inbox.