feat: scaffold redsen-lean-harness v0.1.0
Recovered from crashed session (Node OOM). Repo contains full P0-P6 scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006, lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10 .github/agents, 12 CLI skills, instructions, context7 mcp.json, and unit/e2e test suite. Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs use --input-tokens/--output-tokens, so telemetry totals were always 0. Now accepts both forms. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
@@ -0,0 +1,50 @@
|
||||
# ADR 0006 — Local NDJSON telemetry, no server, no budget cap
|
||||
|
||||
- **Status**: Accepted
|
||||
- **Date**: 2026-09-09
|
||||
|
||||
## Context
|
||||
|
||||
Requirement: the harness must be *fast, observable and benchmarkable*. Scope was explicitly set to
|
||||
**telemetry only** — per-run token counts, latency, tool-call traces, pass/fail — not an external
|
||||
benchmark suite.
|
||||
|
||||
Separately, the operator chose **no budget cap** on parallel fan-out, combined with **fully
|
||||
dynamic** re-planning. That combination is the single largest cost risk in the system.
|
||||
|
||||
## Decision
|
||||
|
||||
1. **Append-only NDJSON** at `.agents/runs/<runId>/events.ndjson`. One JSON object per line,
|
||||
written with `O_APPEND` so concurrent lanes cannot interleave partial lines.
|
||||
2. **OpenTelemetry GenAI-inspired attribute names** (`gen_ai.usage.input_tokens`,
|
||||
`gen_ai.request.model`, …) **without an OTEL dependency**. The data is portable to a real
|
||||
tracing backend later; today it costs nothing.
|
||||
3. **Tolerant reader.** Malformed lines are skipped, never thrown on. A crashed agent must not
|
||||
destroy the log.
|
||||
4. **Live board.** `board.md` is regenerated on every event and shows lanes in flight, ralph
|
||||
iterations, elapsed time, tokens, and — prominently — **burn rate in tokens/min**. Board
|
||||
rendering is wrapped in try/catch so it can never break a telemetry write.
|
||||
5. **`lh report`** renders the markdown summary; **`lh report --journal`** renders the committed,
|
||||
self-documenting per-run journal.
|
||||
|
||||
### On the absent budget cap
|
||||
|
||||
No hard cap is enforced, per the operator's explicit choice. The mitigation is **visibility, not
|
||||
prevention**: live burn rate in `board.md`, a post-run cost breakdown in `lh report`, and an
|
||||
advisory `concurrency.maxWriteLanes` ceiling that warns and only blocks under `--strict`.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Zero infrastructure. Telemetry works offline and in CI with no setup.
|
||||
- `lh run event` is on the hot path (dozens of calls per run) and must stay well under 100 ms.
|
||||
- Cost overrun is possible by design. Accepted and documented; visibility is the control.
|
||||
- Because the schema is OTEL-shaped, exporting to Langfuse/OpenLLMetry later is a transform, not a
|
||||
re-instrumentation.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **OpenTelemetry SDK** — rejected: heavy dependency and a collector to run, for data we currently
|
||||
only read locally.
|
||||
- **SQLite event store** — rejected: native dependency (see ADR 0005), and NDJSON is already
|
||||
append-safe and greppable.
|
||||
- **Hard credit cap** — rejected by the operator. Revisit if real runs overrun.
|
||||
Reference in New Issue
Block a user