Open-source agent tracing

Agent tracing, without the platform.

A TypeScript library for LangGraph and LangChain. Traces land in your files or AWS account, and a local viewer opens them.

Get started
The m4trix trace viewer showing a LangGraph run tree with a selected tool call, its metadata, input and output payloads.

Fits the stack you already run

  • LangGraph
  • LangChain
  • Model Context Protocol
  • Claude
  • Cursor
  • Docker
  • Kubernetes

Let Claude Code debug the run

The viewer serves an MCP endpoint on localhost. Claude Code reads the run tree, searches payloads and points at the call that broke.

terminalbash
claude mcp add --transport http \  m4trix-traces http://127.0.0.1:4319/mcp
  • list_traces
  • find_runs
  • search_payloads
  • get_trace
  • get_run
  • get_payload
  • get_conversation
  • analyze_trace
  • compare
  • annotate
  • load_trace_payloads
claude~/support-agent
> why did the last support-agent run fail?
⏺ analyze_trace (MCP)(traceId: "latest")
⎿ 2 error runs, 1 root cause, 1 unfinished run
⏺ get_trace (MCP)(traceId: "aaaa1111")
⎿ 6 runs shown, 1 hidden
⏺ search_payloads (MCP)(query: "cache miss")
⎿ 1 distinct value, 3 occurrences in cache_get
⏺ get_conversation (MCP)(runId: "a-llm")
⎿ 2 messages, 1 tool call to lookup_order
⏺ It failed in lookup_order. Outline of run aaaa1111:
 
support-agent 5.00s ✗ chain
├─ cache_get ×3 5ms ✓ "cache miss" ◀ same key, 3 calls
├─ ChatOpenAI 1.20s ✓ 580 tok
├─ lookup_order 3.00s ✗ TimeoutError ◀ root cause
└─ notify_user - … never ended ◀ unfinished
 
Context of ChatOpenAI (500 in / 80 out)
┌───────────────────────────────────────┐
│ system ██░░░░░░░░░░░░░░░░░░ 62 tok │
│ user █░░░░░░░░░░░░░░░░░░░ 14 tok │
│ tools █████████████████░░░ 424 tok │ ◀ 3 identical results
├───────────────────────────────────────┤
│ output ███░░░░░░░░░░░░░░░░░ 80 tok │
└───────────────────────────────────────┘
 
Fixes
1. lookup_order hits its 3s timeout. Retry with backoff.
2. cache_get repeats the same key. Cache the miss once.
3. notify_user is never awaited on the error path.
>

One callback. Every step recorded.

The tracer speaks the LangChain callback API. Pass it to invoke and every chain, model and tool call lands in ./.traces.

graph.invoke(input, { callbacks: [tracer] })
chainintake
chainplan_work
tooldocumentation_search
toolrepository_search
llmdraft_answer
chainfinal_review

.traces/traces/f19bca0d/runs.ndjson

{"name":"intake","type":"chain","status":"success"}

{"name":"plan_work","type":"chain","status":"success"}

{"name":"documentation_search","type":"tool","status":"success"}

{"name":"repository_search","type":"tool","status":"success"}

{"name":"draft_answer","type":"llm","status":"success"}

{"name":"final_review","type":"chain","status":"success"}

agent.tstypescript
import {  FsPayloadStoreAdapter,  FsStructureStoreAdapter,  TraceStore,  Tracer,  toLangGraph,} from '@m4trix/tracing';
const traceStore = TraceStore.of({  structureStoreAdapter: new FsStructureStoreAdapter({ path: './.traces' }),  payloadStoreAdapter: new FsPayloadStoreAdapter({ path: './.traces' }),});
const tracer = Tracer.from(traceStore).adapt(toLangGraph);
await graph.invoke(input, { callbacks: [tracer] });await tracer.flush();
installbash
pnpm add @m4trix/tracing
open the viewerbash
npx @m4trix/trace-viewer \  --adapter fs --path ./.traces

Open a run, see what happened

Small rows for lists and filters, full payloads when you open a run, and notes that stay with the trace.

Payloads that read like conversations

Profiles turn raw JSON into messages, tool calls and tables. A model drafts the mapping from samples of your traces, with your own key, straight from the browser.

The New AI profile dialog in the trace viewer, sampling the current trace and sending it to the Claude API with a key held in browser memory.

Plain files on disk

No sign-up, no API key, no upload queue. Grep it, commit a fixture, or mount it in Docker.

.traces/traces/
└─ f19bca0d/
   ├─ trace.json
   ├─ runs.ndjson
   └─ payloads/

Split storage, one store

Structure rows stay small, so lists and filters are fast. Prompts and completions are blobs, fetched by ref when you open a run. The adapters that wrote them serve both back.

Structure rows (files or DynamoDB)

chainplan_workinputRef
tooldocumentation_searchinputRef
llmdraft_answerinputRef

Payload blobs (files or S3)

plan_work/input.json{"task":"Explain how the trace…
documentation_search/input.json{"query":"trace viewer setup…
draft_answer/input.json{"messages":[{"role":"user",…

Review in place

Annotate traces and single runs after the fact. Notes live next to the structure rows, not in another tool.

annotation: { review: 'approved' }

Start on a laptop. Ship to AWS when you need to.

Same Tracer, same viewer at every stage. Only the adapters change.

Write locally

Your appTracer + Fs adapters
/tracestrace.json, runs.ndjson

Filesystem adapters write to a folder. On a laptop that is the whole setup.

Ship with a sidecar

Sidecarm4trix-tracing-sidecar
payloads firststructure after

A companion container uploads payloads before structure, so the app needs no AWS credentials.

Read from AWS

S3payloads
DynamoDBstructure
Viewer--adapter aws-stack

The same viewer and MCP server read straight from your account.

Six primitives. No platform.

Swap any adapter without touching the tracer or the viewer.

Drop-in LangGraph and LangChain callbacks

Tracer.from(traceStore) implements the callback surface LangGraph expects. Pass tracer.adapt(toLangGraph) to callbacks. Every chain, LLM, tool, and retriever span lands in your store without rewriting agent code.
  • Handles chain, LLM, chat model, tool, and retriever events
  • Batches pending runs on flush for efficient writes
  • Typed LangGraph adapter via toLangGraph
agent.tstypescript
import { Tracer, toLangGraph } from '@m4trix/tracing';
const tracer = Tracer.from(traceStore);const lgTracer = tracer.adapt(toLangGraph);
await graph.invoke(input, { callbacks: [lgTracer] });await lgTracer.flush();

Questions

Does anything leave my machine?

Not by default. The filesystem adapters write to a folder you choose. Data only moves if you configure the S3 and DynamoDB adapters or run the sidecar. AI profiles call the model provider directly from your browser, with your key.

Do I have to use LangGraph?

No. The Tracer implements the LangChain callback methods without importing LangChain, so anything that emits those callbacks works. LangGraph gets a typed adapter through toLangGraph.

How is this different from LangSmith or Langfuse?

Those are platforms: a hosted service, or a database and web app you operate. m4trix tracing is a library. Traces are files or rows in your own AWS account, and the viewer is a CLI you start when you need it.

Is it ready for production?

The sidecar pattern is built for it: the app writes locally and a companion container ships to S3 and DynamoDB. Sampling, PII redaction, viewer auth and retention policies are not included yet.

What does it cost?

Nothing. It is MIT licensed with no seats or usage tiers. You pay only for the storage you choose to use.

Part of the m4trix toolkit

Each package works on its own. Together they share one TypeScript model.

Trace your next run

One callback, one folder, one command to open it.