RunLedger
Early stage · Building in public

Every agent run, on the record.

RunLedger is a black box recorder for AI agents. It turns each run into a shareable receipt: what changed, what ran, what it cost, and what looked risky, in plain language.

Sample receipt — illustrative data

RUN RECEIPT · #rl-0042

Claude Code · repo: payments-service

branch: fix/retry-logic · duration 4m 12s

  1. 1Read 6 files to understand retry logicHaiku3.1k tokens$0.004
  2. 2Edited src/payments/retry.ts: added exponential backoff (+38 −12)Sonnet12.4k tokens$0.061
  3. 3Ran npm test — 142 passedHaiku1.2k tokens$0.002
  4. 4Deleted tests/payments/retry.legacy.spec.ts▲ Flagged: test deleted
  5. 5Read .env.local● Flagged: secrets file

Risk score62 / 100 · Medium

  • Deleted a test file
  • Touched a secrets file (.env.local)
  • No actions outside the working folder ✓

Total 18.7k tokens · $0.071 · Files changed 3

The problem

Agents now ship code. Nobody reads the transcript.

01

Runs happen unsupervised

Agents edit files, run shell commands and call APIs while you're in another tab.

02

Logs aren't explanations

Raw transcripts are thousands of lines. Reviewers need the what and the why, fast.

03

Risk hides in the details

A deleted test, a peek at a secrets file, a write outside the project folder: easy to miss, costly to ship.

How it works

Record, explain, then decide.

01

During the run

Record

RunLedger captures every step of an agent run: files read and written, commands, API calls, model and tokens.

02

After the run

Explain & score

Each step is summarised in plain language, and the run gets a risk score with the reasons behind it.

03

Your call

Share, approve, roll back

Send the receipt to a teammate, apply permission rules, and roll the run back if something looks wrong.

Features

Everything a reviewer needs, in one receipt.

Shareable run receipts

One link per run: files changed, commands and API calls, explained in plain language.

Cost & tokens per step

See which model handled each step and what it cost.

Risk score

Flags deleted tests, touched secrets, and actions outside the working folder.

Permission rules

Define what an agent may and may not touch.

Rollback

Undo a run's changes in one step.

Built for Claude Code first

Then MCP-based tools and other agents.

Claude-first

Built Claude-first.

RunLedger uses the right Claude model for each job: Claude Haiku writes fast, low-cost summaries of every step, while Claude Sonnet and Claude Opus handle the harder work: evaluating risk and explaining diffs.

Haiku
Step summaries
Sonnet
Risk evaluation
Opus
Deep diff explanations

First integration: Claude Code Next: MCP other agents

Claude, Claude Code, Haiku, Sonnet and Opus are products of Anthropic. RunLedger is an independent project, not affiliated with or endorsed by Anthropic.

Where we are

RunLedger is an early-stage, bootstrapped project founded in October 2026. The product is in active development; the receipt shown on this page is an illustrative mockup, not live data.