Recovered from crashed session (Node OOM). Repo contains full P0-P6 scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006, lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10 .github/agents, 12 CLI skills, instructions, context7 mcp.json, and unit/e2e test suite. Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs use --input-tokens/--output-tokens, so telemetry totals were always 0. Now accepts both forms. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2.5 KiB
2.5 KiB
ADR 0006 — Local NDJSON telemetry, no server, no budget cap
- Status: Accepted
- Date: 2026-09-09
Context
Requirement: the harness must be fast, observable and benchmarkable. Scope was explicitly set to telemetry only — per-run token counts, latency, tool-call traces, pass/fail — not an external benchmark suite.
Separately, the operator chose no budget cap on parallel fan-out, combined with fully dynamic re-planning. That combination is the single largest cost risk in the system.
Decision
- Append-only NDJSON at
.agents/runs/<runId>/events.ndjson. One JSON object per line, written withO_APPENDso concurrent lanes cannot interleave partial lines. - OpenTelemetry GenAI-inspired attribute names (
gen_ai.usage.input_tokens,gen_ai.request.model, …) without an OTEL dependency. The data is portable to a real tracing backend later; today it costs nothing. - Tolerant reader. Malformed lines are skipped, never thrown on. A crashed agent must not destroy the log.
- Live board.
board.mdis regenerated on every event and shows lanes in flight, ralph iterations, elapsed time, tokens, and — prominently — burn rate in tokens/min. Board rendering is wrapped in try/catch so it can never break a telemetry write. lh reportrenders the markdown summary;lh report --journalrenders the committed, self-documenting per-run journal.
On the absent budget cap
No hard cap is enforced, per the operator's explicit choice. The mitigation is visibility, not
prevention: live burn rate in board.md, a post-run cost breakdown in lh report, and an
advisory concurrency.maxWriteLanes ceiling that warns and only blocks under --strict.
Consequences
- Zero infrastructure. Telemetry works offline and in CI with no setup.
lh run eventis on the hot path (dozens of calls per run) and must stay well under 100 ms.- Cost overrun is possible by design. Accepted and documented; visibility is the control.
- Because the schema is OTEL-shaped, exporting to Langfuse/OpenLLMetry later is a transform, not a re-instrumentation.
Alternatives considered
- OpenTelemetry SDK — rejected: heavy dependency and a collector to run, for data we currently only read locally.
- SQLite event store — rejected: native dependency (see ADR 0005), and NDJSON is already append-safe and greppable.
- Hard credit cap — rejected by the operator. Revisit if real runs overrun.