Open Sourcing the AI Session Collector
Miguel Martinez
Git records what changed in a commit, but the work behind that commit now happens in an AI session that git never sees. The session holds its inputs, meaning the specs, tasks, skills and screenshots the agent worked from, plus the prompts and every model, tool and command the agent ran, along with what it cost.
Today we are open sourcing our AI session collector, chainloop trace, under Apache 2.0. It records every session from Claude Code, OpenCode or Cursor, strips the secrets, signs it, and keeps it tamper-evident in storage you own.
With a record of every session, you can:
- Verify. Guardrails check each session on push for approved models, allowed MCP servers, leaked secrets and dangerous commands. The findings go back to the agent, so its next session follows your rules.
- Reproduce. Each session vendors the specs, tasks, skills and screenshots the agent worked from as signed evidence, so they can’t disappear or be changed after the fact.
- Trace. Each session is the first signed link to its pull request, build and deployment, so any change in production traces back to the session that wrote it.
- Own. Sessions from every agent and every model land in one store on your infrastructure, or in Chainloop Platform.
The evidence store and the policy engine were already open source, so with the collector the whole stack for governing agentic coding is open source end to end. That same stack already runs in production at Fortune 500 companies.
What a session holds
A session is the full journal of one piece of agent work. For each one, the collector keeps:
- The conversation, every prompt and every response.
- The model and its usage, with token counts and the estimated cost.
- The tools, meaning every tool call, MCP server, skill and shell command.
- The code changes, as a snapshot of each file before and after every edit. That snapshot is how each changed line gets attributed to the agent or to a person.
- The agent configuration, the skills, rules and MCP servers the agent was set up with.
- What the session worked from, meaning the specs and plans, the tasks from Jira or Linear, the skills the agent loaded, and the screenshots you gave it.
An OpenCode session running Kimi K3 in Chainloop Platform. Every line is attributed to AI, 10 files changed in 1 hour 19 minutes for $13.45 in tokens, and the session links to the pull request it produced.
What you can do with it
See how AI-native your SDLC is
Your new code gets written in the session, which makes it the most direct measure of how far your team has moved to agentic development. You see how much of the code is AI-authored per repository and per pull request, plus which agents, models and MCP servers are in use, by whom, and at what cost. A leader needs that baseline before handing agents more of the work.
Seven days of AI coding in one organization on the Chainloop Platform dashboard, with sessions, cost, active developers, AI-assisted pull requests, AI-authored share, and usage by model and harness.
Set the rules agents work by
Policies run against the session evidence, so they see what the diff hides, such as the models used, the MCP servers and skills called, the commands run, and any secret that passed through. You can allow only approved models and MCP servers, flag dangerous commands or leaked secrets, set a token budget, or cap the share of AI-authored lines in a change.
On a violation, Chainloop blocks the push or reports it, whichever you choose. Either way the findings go back to the developer and to the agent, so the agent follows the rules in its next session. Once the system enforces the limits, teams give more people access instead of relying on a reviewer to catch problems by hand.
A Claude Code session in Chainloop Platform, scored and checked. Three of four policies pass, the no-secrets check fails, and the verification criterion notes that validation stopped at unit tests.
Reproduce how a change was made
Teams vendor their dependencies so a build still works when an upstream package disappears, because you can’t reproduce software without every input. The same holds for an agent session. Specs, tasks, skills and screenshots are inputs to your code, and a link to them is a weak record. The Linear task gets edited, the skill gets a new version, and the spec lived on a laptop that has since been wiped.
So the collector vendors those inputs into the session evidence. It gathers them automatically while the agent works and signs them into the same attestation as the transcript, stored by digest in storage you own. A vendored input can’t disappear, and any edit to it breaks the signature, so the spec in the evidence is the spec the agent read. The agent saves the text of each spec, task or plan as it works, and the collector attaches it, which means Chainloop needs no connectors and no credentials to fetch it. Months later, a reviewer or an auditor reads the code next to the exact instructions it came from.
The context of one Claude Code session in Chainloop Platform. Next to the transcript sit the task it worked on, the merged design spec for issue 3543, and the
br:eng-spec skill the agent loaded.
Prove what shipped and how
A recorded session becomes the first signed link in the same chain as the commit, the build and the release, tied to every line it changed. Policies such as pr-review-required add proof that a person reviewed and approved the change. When AI-written code ships, an auditor or a Cyber Resilience Act assessor can follow it from the spec and the prompt to production.
One Claude Code session in Chainloop Platform, followed from its pull request through the release pipelines and the IaC, secrets and vulnerability scans that ran on the commit, to the images it shipped and the findings on them.
Why keeping it is hard
The session runs on a laptop before CI, so CI telemetry never sees it. Teams mix Claude Code, OpenCode and Cursor, and some run agents in CI jobs or sandboxes, so any collector has to sit next to each of them.
Agents also read .env files, paste tokens into commands and print credentials in tool output, and all of it lands in the transcript. Ship those files to a shared bucket and you have built a credential leak, so the secrets have to come out before anything leaves the machine.
An auditor will only trust a session that nobody could edit after the fact. That means signing it at the source and keeping it in storage you own, where anyone with the public key can verify it.
And a platform team won’t install a tool that reads transcripts on a few hundred laptops without knowing what it captures, what it redacts and when anything leaves the machine. Infrastructure that holds your compliance evidence has to be open source, so you can run it and check it yourself.
What we open sourced
Chainloop already covered SDLC governance in open source with the CLI that collects evidence from your pipelines, the policy engine that checks it, and the Evidence Store that keeps it signed in storage you own. Regulated companies run it in production as the evidence and policy layer for their software supply chain.
The collector is the piece that connects that system to coding agents. It ships as the trace command in the Chainloop CLI and brings each AI session into the same evidence store as your builds, scans and releases.
Hooks capture the session while you work. On push, the CLI redacts it, runs your policies and signs it, and the control plane stores it in a content-addressable storage backend you own.
You run chainloop trace init once per repository and keep working as before. Every step is in the code, from the provider implementations to the secret redaction.
Nothing leaves the machine until the session is attested on git push. Only pushes that contain commits from an AI session build an attestation, so hand-written commits pass through untouched. If recording fails, say because you are offline, the push still goes through with a warning.
Claude Code, for example, tells you at startup that the session is being recorded and where the evidence will go.
Claude Code at startup after chainloop trace init. The hook names the Chainloop instance, organization and project the evidence goes to.
After the push, the session is stored as signed evidence that you can read back with the open source CLI.
chainloop wf run describe on one push shows a verified attestation holding the agent configuration, the coding session and the spec it worked from, each
with its policy results.
Beyond coding sessions: single-task agents
Some agents run one task and exit, like an agent in a CI job that triages an issue or reviews a pull request. chainloop trace run wraps that single invocation, records it and attests it when the agent exits. It needs no git repository and records the session even when nothing was committed, so it fits CI jobs, containers and sandboxed agents. It is the mode behind our Docker Sandboxes partnership, where the Chainloop kit records a sandboxed Claude Code session and posts it to the pull request it produced.
What Chainloop Platform adds
AI Coding Governance in Chainloop Platform adds the session views and the org-wide dashboard shown above, pull request correlation with a merge check and the AI Session Score, a curated policy library, agentic review policies, and lineage from the session to production. The Open Source vs. Platform page has the full split.
The collector, CLI, evidence store and policy engine are open source. The curated policy catalog, pull request check, dashboard and score come with Chainloop Platform.
Next steps
The AI Sessions quickstart takes you from installing the CLI to your first recorded session, and the trace guide covers policies, trace run and troubleshooting. To run the whole stack yourself, start from the Chainloop repository and the Evidence Store deployment guide.
Claude Code and OpenCode are supported today with full metrics, including token usage and cost. Cursor support is experimental and doesn’t capture token usage or cost yet. Codex, GitHub Copilot, Gemini, Windsurf, Amp, Junie and Droid are next.
If your agent is missing, open an issue or talk to us. Contributions are welcome, from new agent providers to policies, and if you like what we are building, star the repository.