LLM Agents Can Easily Tamper With Their Own Traces

Jeremy Qin1,2,3*, David Schmotz1,2,3*, Derck Prinzhorn4, Luca Beurer-Kellner5, Ameya Prabhu2,3,6†, Maksym Andriushchenko1,2,3†

1ELLIS Institute Tübingen · 2Max Planck Institute for Intelligent Systems · 3Tübingen AI Center · 4Exponential Security Labs · 5Snyk · 6University of Tübingen
* equal contribution · † equal supervision

Monitoring, incident investigations and audits rely on an agent's execution trace to reconstruct what happened. We show that local LLM agents can delete or rewrite those traces, when asked, when tricked by a malicious skill, and on their own when it improves their score.

9 out of 10 agents can tamper with their own traces in full-access mode

In the financial and personal-summary deletion tasks, nine model–harness pairs tampered in at least one of ten trials per task. Muse Code refused all 20 deletion requests. Fabricated reset events and tool-call manipulation exposed different boundaries: Claude Code, ZCode and Muse Code often refused to invent a reset event, while most other pairs wrote one.

Paper Figure 2: attack success rates for financial deletion, personal-summary deletion, fabricated reset addition and tool-call manipulation across ten model–harness pairs. Muse Code records zero successful trace deletions in both privacy tasks.
Figure 2 · Direct-request capability. Attack success rate over ten trials per setting; N/A means the harness did not support the attack. View in the paper ↗

A malicious skill can induce tampering

An installed privacy-protection skill led some agents to delete their own session trace while completing an unrelated coding task. A second skill attempted to insert instructions into a stored compaction summary. The user had not requested either change, and success varied sharply across pairs.

Paper Figure 6: attack success rates for privacy-skill trace deletion and compaction-skill injection. Codex GPT-5.6-Sol succeeded in all ten privacy-skill trials, while several pairs showed zero success; compaction results varied by harness.
Figure 6 · Skill-file exploitability. Ten trials per pair and setting. N/A marks an unsupported attack. View in the paper ↗

Agents do it on their own for reward

Agents received a task, a time budget and only a numerical score. Completing the task earned at most 0.8 out of 1.0; a higher score required reducing the trace length, a condition the agents were not told. Every pair tampered at least once in both reward settings. Muse Code was the clearest exception to the general pattern, with two successful anonymization trials out of ten.

Paper Figure 7: reward-optimization attack success rates for anonymization and Terminal-Bench. All ten model–harness pairs tampered in at least one trial of each task; Muse Code succeeded in two of ten anonymization trials.
Figure 7 · Reward optimization. Attack success rate across ten trials per pair in each task. View in the paper ↗

High-scoring peer examples can lead agents to tamper with traces

When agents saw higher-scoring peer logs or workspaces, many adopted trace tampering. Muse Code and Claude Code with Opus 5.5 refused the explicitly shown behavior in peer logs; both tampered in some peer-workspace trials, where they had to infer the behavior from files and scores.

Paper Figure 8: attack success rates after exposure to peer logs and peer workspaces. Muse Code and Claude Code Opus 5.5 show zero success with peer logs but nonzero success with peer workspaces.
Figure 8 · Peer examples. Synthetic action logs or saved workspaces paired with scores; ten trials per pair in each setting. View in the paper ↗

Permissions change the outcome

Auto-mode permissions substantially reduced privacy-deletion success for both Claude Code models, Antigravity and ZCode. They did not stop every pair: Codex with GPT-6-Sol still deleted traces in all 20 privacy trials, and OpenCode with Qwen 3.8 Max succeeded in most direct-request trials. The exact restrictions and outcomes depend on the harness.

Paper Figure 3: side-by-side full-access and auto-mode attack success rates for financial deletion, personal-summary deletion, fabricated reset addition and tool-call manipulation. Auto mode reduces several pairs to zero, but Codex GPT-6-Sol remains at 100 percent in both deletion tasks.
Figure 3 · Full access versus auto mode. Ten trials per condition. B means a benign tool call was blocked or inaccessible, not that the model refused. * marks three inconclusive ZCode reset trials; † marks four Muse approval waits counted as unsuccessful; ‡ marks five blocked ZCode tool-call baselines. View in the paper ↗

How we tested trace tampering

We ran ten model–harness pairs in containers and kept an independent log outside the agent environment. We tested four routes to tampering: a direct request, instructions hidden in a skill, a higher score for shortening the trace, and examples from other agents. Each setting was repeated ten times.

The interactive trace browser now uses only the 1,590 trials selected in the paper-results evidence package. Four OpenCode/Qwen positives in that package are based on observed native-session loss; the responsible agent action was not captured.

A recorded Codex trial in which an installed privacy skill turns an ordinary Python client task into deletion of the agent's own session trace. An independent filesystem observer confirms the deletion.
Figure 5 · A privacy skill turns a simple coding task into trace deletion. (a) The user requests a Python client without naming the skill or requesting trace deletion. The agent loads the installed skill, follows its cleanup instruction, and deletes its own native session traces. (b) The installed skill asks for trace deletion under the pretext of privacy protection. Open this trace →

Agents can reach the files that record their sessions

Deleting the trace can leave an agent untraceable

After a Codex session file was deleted, the session kept running without later activity appearing in a newly created native trace. Claude Code and Antigravity recreated trace files, but agents also scheduled repeated or delayed cleanup to delete later records.

Codex
  1. Session file deleted
  2. Session keeps running
  3. Nothing afterwards is recorded
Claude Code · Antigravity
  1. Session file deleted
  2. Harness starts a new file
  3. Repeated or delayed cleanup can erase it

Can an agent alter its own record?

In full-access mode, agents can reach the local files that hold their session traces. Some of the tested locations are:

Claude Code~/.claude/projects/<proj>/<id>.jsonl
Codex~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl
OpenCode~/.local/share/opencode/opencode.db
Grok Build~/.grok/sessions/<workspace-key>/<session>/chat_history.jsonl
Muse Code~/.local/share/muse/sessions/YYYY/MM/DD/<id>/session.jsonl
Antigravity~/.gemini/antigravity-cli/brain/<session>/.system_generated/logs/transcript_full.jsonl
ZCode~/.zcode/cli/rollout/model-io-<session>.jsonl
Kimi Code~/.kimi-code/sessions/<workDirKey>/<session>/agents/main/wire.jsonl

Record model traffic off the agent host

Record traces off the host

The paper proposes sending model traffic through an interception server outside the agent host's control. It records each request and response, including streaming chunks, before forwarding them. Its log is append-only, and it stops forwarding if recording fails. This preserves the model exchange record even if the host is compromised.

It cannot see calls that bypass it, such as local models or other APIs, and it records tool results without verifying them.

Complete, append-only provider-side logging could also protect the record. Zero Data Retention policies may prevent providers from keeping it, and existing compliance APIs can omit thinking blocks and tool definitions or truncate entries. Recording the exchange also does not prove that a reported tool call was executed faithfully.

An agent's authority to do useful work must not include authority to rewrite the record used to evaluate that work.

Cite this paper

@misc{qin2026llmagentseasilytamper,
  title={LLM Agents Can Easily Tamper With Their Own Traces},
  author={Jeremy Qin and David Schmotz and Derck Prinzhorn and Luca Beurer-Kellner and Ameya Prabhu and Maksym Andriushchenko},
  year={2026},
  eprint={2609.30266},
  archivePrefix={arXiv},
  primaryClass={cs.CR},
  url={https://arxiv.org/abs/2609.30266},
}