Files
redsen-lean-harness/docs/adr/0006-local-ndjson-telemetry.md
T
mozempkandCopilot 383129f571 feat: scaffold redsen-lean-harness v0.1.0
Recovered from crashed session (Node OOM). Repo contains full P0-P6
scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006,
lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10
.github/agents, 12 CLI skills, instructions, context7 mcp.json, and
unit/e2e test suite.

Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs
use --input-tokens/--output-tokens, so telemetry totals were always 0.
Now accepts both forms.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-09 22:44:15 +02:00

2.5 KiB

ADR 0006 — Local NDJSON telemetry, no server, no budget cap

  • Status: Accepted
  • Date: 2026-09-09

Context

Requirement: the harness must be fast, observable and benchmarkable. Scope was explicitly set to telemetry only — per-run token counts, latency, tool-call traces, pass/fail — not an external benchmark suite.

Separately, the operator chose no budget cap on parallel fan-out, combined with fully dynamic re-planning. That combination is the single largest cost risk in the system.

Decision

  1. Append-only NDJSON at .agents/runs/<runId>/events.ndjson. One JSON object per line, written with O_APPEND so concurrent lanes cannot interleave partial lines.
  2. OpenTelemetry GenAI-inspired attribute names (gen_ai.usage.input_tokens, gen_ai.request.model, …) without an OTEL dependency. The data is portable to a real tracing backend later; today it costs nothing.
  3. Tolerant reader. Malformed lines are skipped, never thrown on. A crashed agent must not destroy the log.
  4. Live board. board.md is regenerated on every event and shows lanes in flight, ralph iterations, elapsed time, tokens, and — prominently — burn rate in tokens/min. Board rendering is wrapped in try/catch so it can never break a telemetry write.
  5. lh report renders the markdown summary; lh report --journal renders the committed, self-documenting per-run journal.

On the absent budget cap

No hard cap is enforced, per the operator's explicit choice. The mitigation is visibility, not prevention: live burn rate in board.md, a post-run cost breakdown in lh report, and an advisory concurrency.maxWriteLanes ceiling that warns and only blocks under --strict.

Consequences

  • Zero infrastructure. Telemetry works offline and in CI with no setup.
  • lh run event is on the hot path (dozens of calls per run) and must stay well under 100 ms.
  • Cost overrun is possible by design. Accepted and documented; visibility is the control.
  • Because the schema is OTEL-shaped, exporting to Langfuse/OpenLLMetry later is a transform, not a re-instrumentation.

Alternatives considered

  • OpenTelemetry SDK — rejected: heavy dependency and a collector to run, for data we currently only read locally.
  • SQLite event store — rejected: native dependency (see ADR 0005), and NDJSON is already append-safe and greppable.
  • Hard credit cap — rejected by the operator. Revisit if real runs overrun.