Open Sourcing the AI Session Collector

Miguel Martinez
Our open source AI session collector records, redacts, signs, and stores every coding agent session on infrastructure you own.

Git records what changed in a commit, but the work behind that commit now happens in an AI session that git never sees. The session holds its inputs, meaning the specs, tasks, skills and screenshots the agent worked from, plus the prompts and every model, tool and command the agent ran, along with what it cost.

Today we are open sourcing our AI session collector, chainloop trace, under Apache 2.0. It records every session from Claude Code, OpenCode or Cursor, strips the secrets, signs it, and keeps it tamper-evident in storage you own.

With a record of every session, you can:

  • Verify. Guardrails check each session on push for approved models, allowed MCP servers, leaked secrets and dangerous commands. The findings go back to the agent, so its next session follows your rules.
  • Reproduce. Each session vendors the specs, tasks, skills and screenshots the agent worked from as signed evidence, so they can’t disappear or be changed after the fact.
  • Trace. Each session is the first signed link to its pull request, build and deployment, so any change in production traces back to the session that wrote it.
  • Own. Sessions from every agent and every model land in one store on your infrastructure, or in Chainloop Platform.

The evidence store and the policy engine were already open source, so with the collector the whole stack for governing agentic coding is open source end to end. That same stack already runs in production at Fortune 500 companies.

What a session holds

A session is the full journal of one piece of agent work. For each one, the collector keeps:

  • The conversation, every prompt and every response.
  • The model and its usage, with token counts and the estimated cost.
  • The tools, meaning every tool call, MCP server, skill and shell command.
  • The code changes, as a snapshot of each file before and after every edit. That snapshot is how each changed line gets attributed to the agent or to a person.
  • The agent configuration, the skills, rules and MCP servers the agent was set up with.
  • What the session worked from, meaning the specs and plans, the tasks from Jira or Linear, the skills the agent loaded, and the screenshots you gave it.
The overview tab of an OpenCode session in Chainloop running Kimi K3 through OpenRouter: 100 percent AI authorship, 1675 lines added and 19 removed across 10 files, a duration of 1 hour 19 minutes, 1,054,318 tokens costing 13.45 dollars, a change summary, and the list of changed files, each attributed to AI, linked to pull request 6737

An OpenCode session running Kimi K3 in Chainloop Platform. Every line is attributed to AI, 10 files changed in 1 hour 19 minutes for $13.45 in tokens, and the session links to the pull request it produced.

What you can do with it

See how AI-native your SDLC is

Your new code gets written in the session, which makes it the most direct measure of how far your team has moved to agentic development. You see how much of the code is AI-authored per repository and per pull request, plus which agents, models and MCP servers are in use, by whom, and at what cost. A leader needs that baseline before handing agents more of the work.

The AI Governance overview dashboard in Chainloop over seven days: total sessions and their cost, active developers, AI-assisted pull requests, the share of AI-authored code, the AI score trend across alignment, planning, scope, quality, verification and trust criteria, and breakdowns by contributor, model and coding harness

Seven days of AI coding in one organization on the Chainloop Platform dashboard, with sessions, cost, active developers, AI-assisted pull requests, AI-authored share, and usage by model and harness.

Set the rules agents work by

Policies run against the session evidence, so they see what the diff hides, such as the models used, the MCP servers and skills called, the commands run, and any secret that passed through. You can allow only approved models and MCP servers, flag dangerous commands or leaked secrets, set a token budget, or cap the share of AI-authored lines in a change.

On a violation, Chainloop blocks the push or reports it, whichever you choose. Either way the findings go back to the developer and to the agent, so the agent follows the rules in its next session. Once the system enforces the limits, teams give more people access instead of relying on a reviewer to catch problems by hand.

The score and policies tab of a Claude Code session in Chainloop: an AI session score of 90 percent with a per-criterion breakdown for scope discipline, user trust signal, alignment, solution quality, context and planning, and verification, followed by four policy results where allowed agents, no dangerous commands, and allowed MCP servers pass and no secrets fails

A Claude Code session in Chainloop Platform, scored and checked. Three of four policies pass, the no-secrets check fails, and the verification criterion notes that validation stopped at unit tests.

Reproduce how a change was made

Teams vendor their dependencies so a build still works when an upstream package disappears, because you can’t reproduce software without every input. The same holds for an agent session. Specs, tasks, skills and screenshots are inputs to your code, and a link to them is a weak record. The Linear task gets edited, the skill gets a new version, and the spec lived on a laptop that has since been wiped.

So the collector vendors those inputs into the session evidence. It gathers them automatically while the agent works and signs them into the same attestation as the transcript, stored by digest in storage you own. A vendored input can’t disappear, and any edit to it breaks the signature, so the spec in the evidence is the spec the agent read. The agent saves the text of each spec, task or plan as it works, and the collector attaches it, which means Chainloop needs no connectors and no credentials to fetch it. Months later, a reviewer or an auditor reads the code next to the exact instructions it came from.

The context tab of a Claude Code session in Chainloop for pull request 3562, feat(cas): cache blob existence checks for uploads. A sidebar lists the full transcript, the task issue 3543, the spec and plan Spec issue-3543: Existence cache in the Artifact CAS, and the skill br:eng-spec. The main panel shows the spec with its summary, problem, and goals and non-goals sections

The context of one Claude Code session in Chainloop Platform. Next to the transcript sit the task it worked on, the merged design spec for issue 3543, and the br:eng-spec skill the agent loaded.

Prove what shipped and how

A recorded session becomes the first signed link in the same chain as the commit, the build and the release, tied to every line it changed. Policies such as pr-review-required add proof that a person reviewed and approved the change. When AI-written code ships, an auditor or a Cyber Resilience Act assessor can follow it from the spec and the prompt to production.

The lineage tab of a Claude Code session in Chainloop, in four columns: the source, with the AI session, its policy and score results, the pull request and the released commit; the workflows and scans that ran on that commit, including release pipelines, IaC scanning, secrets detection, vulnerability scanning, a release gate, and pull request validation; the container images and SARIF evidence they produced and attested; and the deployments and the findings scanned on them, with critical and high CVEs and their VEX status

One Claude Code session in Chainloop Platform, followed from its pull request through the release pipelines and the IaC, secrets and vulnerability scans that ran on the commit, to the images it shipped and the findings on them.

Why keeping it is hard

The session runs on a laptop before CI, so CI telemetry never sees it. Teams mix Claude Code, OpenCode and Cursor, and some run agents in CI jobs or sandboxes, so any collector has to sit next to each of them.

Agents also read .env files, paste tokens into commands and print credentials in tool output, and all of it lands in the transcript. Ship those files to a shared bucket and you have built a credential leak, so the secrets have to come out before anything leaves the machine.

An auditor will only trust a session that nobody could edit after the fact. That means signing it at the source and keeping it in storage you own, where anyone with the public key can verify it.

And a platform team won’t install a tool that reads transcripts on a few hundred laptops without knowing what it captures, what it redacts and when anything leaves the machine. Infrastructure that holds your compliance evidence has to be open source, so you can run it and check it yourself.

What we open sourced

Chainloop already covered SDLC governance in open source with the CLI that collects evidence from your pipelines, the policy engine that checks it, and the Evidence Store that keeps it signed in storage you own. Regulated companies run it in production as the evidence and policy layer for their software supply chain.

The collector is the piece that connects that system to coding agents. It ships as the trace command in the Chainloop CLI and brings each AI session into the same evidence store as your builds, scans and releases.

How the AI session collector works in three steps. One, capture: hooks in the coding agent and in git record the prompts and full transcript, the specs, Jira tickets and links, the code changes line by line, and the models, tools and MCP calls, for Claude Code, OpenCode and Cursor. Two, redact, check, sign: on git push the Chainloop CLI redacts secrets from the transcript, runs your policies on the session, and signs it as an in-toto attestation with your keys, a KMS or your CA, before anything leaves your machine. Three, store: the Chainloop control plane, self-hosted or ours, stores it by digest in your own OCI, S3, GCS or Azure storage, tamper-evident and next to builds, scans and releases. All three steps are open source in github.com/chainloop-dev/chainloop.

Hooks capture the session while you work. On push, the CLI redacts it, runs your policies and signs it, and the control plane stores it in a content-addressable storage backend you own.

You run chainloop trace init once per repository and keep working as before. Every step is in the code, from the provider implementations to the secret redaction.

Nothing leaves the machine until the session is attested on git push. Only pushes that contain commits from an AI session build an attestation, so hand-written commits pass through untouched. If recording fails, say because you are offline, the push still goes through with a warning.

Claude Code, for example, tells you at startup that the session is being recorded and where the evidence will go.

Claude Code starting in the chainloop/platform repository with a SessionStart hook message: Chainloop Trace is recording this session, and evidence will be sent to app.chainloop.dev for the chainloop organization and the chainloop-platform project

Claude Code at startup after chainloop trace init. The hook names the Chainloop instance, organization and project the evidence goes to.

After the push, the session is stored as signed evidence that you can read back with the open source CLI.

Output of chainloop wf run describe for one AI coding session: a verified in-toto attestation with passing policies, holding three materials. The first is the Claude agent configuration with its skills and rules, the second is the OpenCode coding session on the Kimi K3 model, and the third is the spec the session worked from. Each material lists its digest, annotations, and policy results such as no secrets, no dangerous commands, and allowed MCP servers

chainloop wf run describe on one push shows a verified attestation holding the agent configuration, the coding session and the spec it worked from, each with its policy results.

Beyond coding sessions: single-task agents

Some agents run one task and exit, like an agent in a CI job that triages an issue or reviews a pull request. chainloop trace run wraps that single invocation, records it and attests it when the agent exits. It needs no git repository and records the session even when nothing was committed, so it fits CI jobs, containers and sandboxed agents. It is the mode behind our Docker Sandboxes partnership, where the Chainloop kit records a sandboxed Claude Code session and posts it to the pull request it produced.

What Chainloop Platform adds

AI Coding Governance in Chainloop Platform adds the session views and the org-wide dashboard shown above, pull request correlation with a merge check and the AI Session Score, a curated policy library, agentic review policies, and lineage from the session to production. The Open Source vs. Platform page has the full split.

Agents and the Chainloop CLI produce signed attestations that land in the trusted evidence store; the policy engine evaluates them against policies as code; results surface in the session view, pull request check, dashboard, MCP and API

The collector, CLI, evidence store and policy engine are open source. The curated policy catalog, pull request check, dashboard and score come with Chainloop Platform.

Next steps

The AI Sessions quickstart takes you from installing the CLI to your first recorded session, and the trace guide covers policies, trace run and troubleshooting. To run the whole stack yourself, start from the Chainloop repository and the Evidence Store deployment guide.

Claude Code and OpenCode are supported today with full metrics, including token usage and cost. Cursor support is experimental and doesn’t capture token usage or cost yet. Codex, GitHub Copilot, Gemini, Windsurf, Amp, Junie and Droid are next.

If your agent is missing, open an issue or talk to us. Contributions are welcome, from new agent providers to policies, and if you like what we are building, star the repository.

Continue Reading

; ---