- SANS ISC's Jim Clausing released two Python scripts that rebuild opencode and Hermes AI agent activity from the files each tool leaves on disk.
- Each tool stores records differently, and Clausing says neither script proves every action was recorded. Transcripts may also contain secrets.
- Inventory approved AI agents, find where each stores records, and test whether responders can collect them from a disk image.
When an AI coding agent changes something on a company laptop, someone may later need to know what it was asked, what it did and what it saw. The answer sits in files. Each tool stores those files in its own place and its own format. That is the problem a SANS instructor set out to ease this week.
Jim Clausing of the SANS Internet Storm Center published two Python scripts for reviewing the records left by two agents, opencode and Hermes. He wrote that his SANS course FOR577 has been updated to teach how to investigate AI use during incident response. Its fifth day now covers eight popular coding assistants and agents. Clausing names Claude Code, Codex and Cursor among them.
The idea: an agent's memory is not an audit log
Executives often assume that if an agent did something, a clean record exists. The SANS post points the other way. The records are by-products of how each tool works. They are scattered, formatted differently, and cover only what that tool chose to save.
The lesson here is that agent accountability is an evidence-collection problem before it is a policy problem. A rule that says "review what agents do" means little if no one knows where the record lives or what it leaves out. Clausing makes the same point in his closing line: the useful question is which source supports a finding, and what context sits outside the export.
How the two scripts work
The opencode script reads the tool's SQLite database, a single-file store. It rebuilds a session as a readable transcript, in Markdown by default, or as JSON or JSONL for analysis. A stored turn can hold text, the assistant's reasoning, tool calls, errors, and token and cost figures. The script handles both of the tool's storage layouts. It prefers the newer consolidated one when a session has data there.
The Hermes script casts a wider net. It pulls from three places and writes one JSON object per line. A tag on every line names the kind of evidence it holds. The tags include session, message, model_usage, request_dump and log_entry.
The request dumps matter most. They hold the full payloads sent to and received from the model provider, including headers, tools and messages. That is a different view from the saved chat. It shows what the model actually received.
Protecting the evidence
Opening a database can change it. Before querying, the opencode script copies the database and its write-ahead sidecar files into a temporary folder. It then reads that copy in read-only mode and deletes the folder on exit. Clausing says this avoids altering evidence through partly staged writes.
The Hermes extractor opens its database read-only. Clausing notes that its README does not describe an equivalent snapshot step. Both scripts can point at a mounted disk image or a collected archive, so analysis does not have to happen on the suspect machine.
Limits that leaders should know
The tools are for review, not for re-running an agent's actions. Clausing also states plainly that neither proves every action was recorded.
Transcripts can hold secrets. The opencode script never opens the tool's auth.json file. That does not make its output safe, though. Clausing cautions that the tool inputs and outputs it captured could still expose secrets. A transcript should be handled as sensitive evidence.
That cap can be removed. The detail is a reminder that how an export is configured shapes what a reviewer sees. Clausing also wrote that the scripts were largely generated by the agents they examine, with differences in design that reflect it. These are one practitioner's tools, not a vetted product.
What to ask your team
Start with an inventory. Which AI coding agents run on company machines, and which are approved? For each, where does it store sessions, request payloads and logs?
Then test collection. Could your incident responders recover a past session from a disk image without installing the tool? How long does the data persist before it is overwritten?
Finally, set handling rules. Transcripts and request dumps may carry secrets, so decide who may read them and where they are stored.
An agent that acts for your staff is only as accountable as the records you can find. Locate them before you need them.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: SANS Internet Storm Center.





