New

Reasoning traces are live

See how your
model thinks.

Lumen shows the steps behind every AI answer, so your team can debug faster, prove quality, and ship with confidence.

Free up to 100k traced tokens · No card needed

Free up to 100k traced tokens · No card needed

Lumen Traces dashboard with trace volume, latency, eval pass rate and a list of recent traces

Drops into the stack you already run

  • Any LLM API

  • Open-weight models

  • Vector databases

  • Agent frameworks

  • Your own evals

  • SSO & audit logs

  • Any LLM API

  • Open-weight models

  • Vector databases

  • Agent frameworks

  • Your own evals

  • SSO & audit logs

Why we built
Lumen

AI should make things clearer, not louder. Lumen is the observability layer for teams shipping AI into real products. Every answer comes with the steps behind it: what the model planned, what it retrieved, which tools it called and why. When something breaks, you see where, fix it once and prove it stays fixed.

Trusted by teams shipping AI in production

OpenAI
Meta
MoneyGram
Canva
Anthropic
Vercel
DoorDash
Coinbase

Features

Fewer black boxes.
More answers you can trust.

Trace every step

Open any answer and see the plan, every tool call and the reasoning on one timeline.

Reasoning trace with plan, retrieve, tool, reason and answer steps

Trace every step

Open any answer and see the plan, every tool call and the reasoning on one timeline.

Reasoning trace with plan, retrieve, tool, reason and answer steps

Catch the spike

Monitors watch latency, cost and quality, and flag a regression before users feel it.

P95 latency chart with a spike flagged by a monitor

Catch the spike

Monitors watch latency, cost and quality, and flag a regression before users feel it.

P95 latency chart with a spike flagged by a monitor

Jump anywhere with ⌘K

Traces, evals, guardrails and the playground, one keystroke away. Search, open, replay.

Command bar searching for refund across traces, evals, guardrails and playground

Jump anywhere with ⌘K

Traces, evals, guardrails and the playground, one keystroke away. Search, open, replay.

Command bar searching for refund across traces, evals, guardrails and playground

How it works

From first trace to trusted in an afternoon.

Step 1

Connect

Wrap your existing model client. Two lines, any provider, no changes to your prompts.

import lumen
client = lumen.wrap(client)

Step 2

Watch

Every request becomes a trace you can open, replay and share with your team.

lumen.trace(“refund-check”)
→ 14 steps · 1.24s

Step 3

Improve

Turn good and bad traces into evals, then gate each deploy on the result.

lumen eval run golden-set
✓ 96.3% passed

Platform

Better insights.
Fewer tools.

Works with the stack you already run

Lumen plugs into any model provider, open-weight model, vector store or agent framework. Traces, evals and guardrails share one timeline, so nobody stitches dashboards together again.

Lumen trace detail for a refund eligibility check

Traces

Every request,
every step.

See every call your AI makes, with timing, cost and status. Filter by model or outcome, then open any trace in one click.

See every call your AI makes, with timing, cost and status. Filter by model or outcome, then open any trace in one click.

Lumen Traces view listing recent requests with model, steps, latency and status

Trace detail

Answers you can
trace to the source.

See what the model read, which tools it called and why it decided. Share one link and the whole team sees the same thing.

See what the model read, which tools it called and why it decided. Share one link and the whole team sees the same thing.

Lumen trace detail showing each reasoning step, the model reasoning and a waterfall timeline

Playground

Replay it. Fix it.
Ship it.

Rerun any trace with a new model, prompt or setting. When the answer is right, save it as an eval so it stays right.

Rerun any trace with a new model, prompt or setting. When the answer is right, save it as an eval so it stays right.

Lumen Playground replaying a customer question with run settings

Testimonials

Teams that
stopped guessing.

“We used to screenshot logs into Slack to argue about why the agent refunded someone. Now we paste a trace link and the argument is over in a minute.”

Priya Raman

Staff Engineer, Northwind

“Two lines of code and we had traces in production the same afternoon. No observability tool has ever installed that quietly for us.”

Tomás Weber

Founding Engineer, Quill

“Lumen caught a latency regression in our retrieval step before a single customer noticed. The deploy gate paid for itself in the first week.”

Marcus Hale

Head of Platform, Fieldnote

“The command bar is how I move through everything. Open a trace, replay it in the playground, ship the fix.”

Dev Patel

Senior Engineer, Orbit Labs

“Our compliance team signs off on AI features without a two-week review now. Every answer links to its sources, and that settles the conversation.”

Elena Ortiz

VP Product, Ledgerline

“Evals used to live in a notebook nobody trusted. Now they run on every pull request, and product can read the results.”

Hannah Cho

ML Lead, Brightside

“We used to screenshot logs into Slack to argue about why the agent refunded someone. Now we paste a trace link and the argument is over in a minute.”

Priya Raman

Staff Engineer, Northwind

“Two lines of code and we had traces in production the same afternoon. No observability tool has ever installed that quietly for us.”

Tomás Weber

Founding Engineer, Quill

“Lumen caught a latency regression in our retrieval step before a single customer noticed. The deploy gate paid for itself in the first week.”

Marcus Hale

Head of Platform, Fieldnote

“The command bar is how I move through everything. Open a trace, replay it in the playground, ship the fix.”

Dev Patel

Senior Engineer, Orbit Labs

“Our compliance team signs off on AI features without a two-week review now. Every answer links to its sources, and that settles the conversation.”

Elena Ortiz

VP Product, Ledgerline

“Evals used to live in a notebook nobody trusted. Now they run on every pull request, and product can read the results.”

Hannah Cho

ML Lead, Brightside

Pricing

Start free.
Scale when it matters.

$0

free forever

Starter · for side projects and first traces

100k traced tokens a month

7-day trace retention

1 project

Evals on up to 50 cases

Community support

$49

per seat / month

Team · for AI running in production

10M traced tokens a month

90-day trace retention

Unlimited projects

Evals and deploy gates

Guardrails and PII redaction

Replay in the playground

SSO and role-based access

Slack and email alerts

Email support, 1-day response

FAQ

Questions,
answered.

Short answers to what teams ask before they start tracing. Can’t find yours? An engineer will reply within one working day.

Common questions

Which models and providers does Lumen support?

How long does setup take?

Where is our trace data stored?

Can Lumen redact personal data?

How is pricing calculated?

Can we self-host Lumen?

Your first trace
is two lines away.

Your first trace
is two lines away.

Wrap your model client, send a request and open the trace. Free up to 100k traced tokens a month, no card needed.

Create a free website with Framer, the website builder loved by startups, designers and agencies.