feat: scaffold redsen-lean-harness v0.1.0

Recovered from crashed session (Node OOM). Repo contains full P0-P6
scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006,
lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10
.github/agents, 12 CLI skills, instructions, context7 mcp.json, and
unit/e2e test suite.

Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs
use --input-tokens/--output-tokens, so telemetry totals were always 0.
Now accepts both forms.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
2026-09-09 22:44:15 +02:00
co-authored by Copilot
commit 383129f571
79 changed files with 7855 additions and 0 deletions
@@ -0,0 +1,42 @@
# ADR 0001 — Markdown owns behaviour, Node owns determinism
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
The harness must be *simple*, *lean on context*, *distributable as a plugin*, and run on two
hosts (GitHub Copilot CLI, VS Code Copilot) that both already contain a capable agent runtime.
There is a strong temptation to build an orchestrator process that drives the model directly.
That path produces a second agent runtime we would have to maintain, and it cannot be shipped
as a plugin because plugins run *inside* the host.
## Decision
Split the system in three, with a hard boundary:
| Layer | Owns | Artifact |
| --- | --- | --- |
| Markdown | Behaviour — agents, skills, instructions | `.github/` |
| Node CLI (`lh`) | Determinism — index, gates, lanes, telemetry, memory | `src/` |
| Host | Execution — subagents, fleet, tools, models | Copilot CLI / VS Code |
`lh` never calls a model. Agents never do arithmetic, parsing, git plumbing, or bookkeeping
by hand. **We do not reimplement an agent runtime.**
## Consequences
- The same `.github/` tree works in both hosts; only the scheduler differs (see ADR 0003).
- `lh` is trivially testable — it is pure I/O with no model in the loop.
- Anything the host cannot do, we cannot do. Accepted: host capability detection (`lh host`)
makes the limitation explicit and degrades rather than failing.
- Behaviour changes are markdown diffs, reviewable in a PR without running anything.
## Alternatives considered
- **Standalone orchestrator binary** — rejected: not plugin-distributable, duplicates the host.
- **Everything in markdown, no CLI** — rejected: indexing, worktrees, structural analysis and
token accounting are not things an LLM should do by hand; they are slow, expensive and unreliable.
- **Everything in Node, markdown as prompts only** — rejected: opaque, unreviewable, and it
breaks the self-documenting requirement.
+45
View File
@@ -0,0 +1,45 @@
# ADR 0002 — web-tree-sitter (WASM) for the code index, with a regex fallback
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
The harness must work on brownfield repos in *any* language, be OS-agnostic, and install with
no build step. It needs a token-budgeted repo map so agents stop reading whole files.
Candidates evaluated:
| Option | Verdict |
| --- | --- |
| `tree-sitter` (native bindings) | Requires a native compile toolchain — fails "easy install" |
| **`web-tree-sitter` (WASM)** | Pure npm, no compiler, runs anywhere Node runs |
| `universal-ctags` | External binary the user must install; 200+ languages but not npm-installable |
| `ast-grep` | Prebuilt binary via npm, good, but heavier and rule-oriented rather than map-oriented |
| SCIP / LSIF | Per-language indexers — far too much install surface |
| Zoekt | Requires Docker |
| Host `/lsp` | Excellent fidelity but CLI-only and not available in every host/language |
## Decision
Use **`web-tree-sitter`**, the sole runtime dependency.
Grammars are **not vendored** and **not fetched at import time**. They are cached lazily under
`.agents/.cache/wasm/` and only downloaded when the user passes `--fetch`. When a grammar is
absent, the indexer **silently degrades to a per-language regex extractor** and reports the
degradation in `lh index --stats`.
The repo map is ranked with a hand-rolled PageRank over the reference graph and truncated to
`config.index.budget` tokens, degrading full signatures → names → file-level counts.
## Consequences
- `lh index` works **fully offline** on first run, at lower fidelity. This is a hard requirement.
- No compiler, no Docker, no external binary, no service. `npm install` is the whole setup.
- Fidelity varies by language and by whether a grammar has been fetched. `--stats` makes this visible.
- We own a small glob matcher and a small PageRank implementation rather than taking dependencies.
## Alternatives considered
Rejected `--fetch`-by-default: it would make the first run fail on an air-gapped machine and
would surprise users with network traffic. Opt-in is the safer default.
+49
View File
@@ -0,0 +1,49 @@
# ADR 0003 — git worktree per write-lane; pluggable isolation backend
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
The harness drives **parallel dynamic workflows on its own**: the conductor re-plans and
re-fans-out at every checkpoint. Multiple builder agents therefore write code concurrently.
Concurrent writes to a single checkout corrupt work — two agents editing the same file, or one
agent's partial state being read by another, produces failures that are extremely hard to
diagnose and that waste far more tokens than they save.
## Decision
1. **Read-only lanes share the main checkout.** Recon and review never write, so they are safe
to fan out N-wide with no isolation.
2. **Every write lane gets its own `git worktree` + branch** (`lh/<runId>/<laneId>`).
3. **File-scope leases.** Every lane declares `scope` globs up front. `lh lane create` performs a
glob-intersection check against all live write lanes and **rejects overlapping scopes**. The
check is deliberately conservative: when intersection is ambiguous, it rejects.
4. **Sequential integration.** A dedicated integrator agent merges lane branches one at a time and
runs one full verify. On conflict, `lh lane merge` does **not** auto-resolve — it marks the lane
blocked and returns the conflicted paths for escalation.
5. **The isolation backend is pluggable** — `worktree` (default), `inplace`, and a stubbed
`devcontainer`, selected by `config.isolation.backend`.
Scope compliance is also a Ralph exit criterion: `lh lane status` reports files changed outside
the declared scope, and a lane with out-of-scope changes cannot pass.
## Consequences
- Parallel writes are safe by construction rather than by convention.
- Lanes are cheap (a branch and a worktree), so the fully dynamic re-planning model can spawn and
drop them freely.
- Merge conflicts surface as an explicit, escalatable state instead of silent corruption.
- Worktrees require git ≥ 2.5 and a real repository. `lh doctor` and `lh host` check for this and
degrade to `inplace` + sequential execution when unavailable.
- Moving to per-lane devcontainers in v2 is a backend swap, not a rewrite.
## Alternatives considered
- **Shared checkout with file locks** — rejected: locks are advisory, agents forget them, and a
crashed agent leaves stale locks.
- **Devcontainer per lane now** — deferred to v2. Strongest isolation but a heavy dependency and a
slow inner loop; the pluggable backend keeps the door open.
- **Read-only parallelism only** — rejected: it caps the speedup at exactly the phase that is
already cheap.
@@ -0,0 +1,45 @@
# ADR 0004 — Do not vendor `pi-context-mode`; adopt techniques only
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
Two prior-art projects were investigated for the memory and context layer:
| Project | Repo | License | Storage |
| --- | --- | --- | --- |
| `pi-hermes-memory` | `chandra447/pi-hermes-memory` | unconfirmed | Markdown + SQLite FTS5 |
| `pi-context-mode` | `FluidLogicLabs/pi-context-mode` | **Elastic License 2.0 (ELv2)** | SQLite FTS5 + event log |
`pi-context-mode` is genuinely good at what we need — auto-compaction near token limits,
anchors and checkpoints, tree-structured session pruning, pre-compaction hooks.
But **ELv2 is not an open-source licence.** It forbids providing the software to third parties
as a managed service and forbids circumventing licence-key functionality. Bundling ELv2 code
into an MIT-licensed plugin that we publish to a marketplace creates a licence conflict and
would misrepresent the terms downstream consumers are bound by.
## Decision
1. **Do not vendor, bundle, fork, or take a dependency on `pi-context-mode`.**
2. **Adopt its techniques**, which are ideas and not protected by the licence: token budgeting
with graceful degradation, compaction of the oldest entries when approaching a budget,
checkpoint anchors, and progressive disclosure instead of bulk loading.
3. Optionally **document it as a user-installed MCP server** for users who want it. The user
installs it themselves under their own licence terms; we ship no ELv2 code.
4. From `pi-hermes-memory`, adopt the *ideas* — categorised memory shards (failures, corrections,
insights, conventions, quirks), secret scanning on write, auto-consolidation, and
memory-*policy* prompting rather than loading everything — but **take no dependency** (see ADR 0005).
## Consequences
- The plugin stays cleanly MIT with one runtime dependency (`web-tree-sitter`).
- We implement compaction ourselves in `src/lib/memory.mjs`. It is less sophisticated than
`pi-context-mode`'s, which is an accepted trade for licence cleanliness and zero dependencies.
- Attribution: both projects are credited as inspiration in the README and here.
## Follow-up
Both projects were identified via an automated research pass. **Re-verify the repositories and
their licences directly before citing them in any shipped documentation.**
+49
View File
@@ -0,0 +1,49 @@
# ADR 0005 — File-only memory with a single always-loaded index
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
The harness must be *lean on token and context size* and use a *lean, OS-agnostic, easily
configurable memory system*. It must also be **self-documenting** and work on greenfield and
brownfield repos.
The dominant failure mode of agent memory is loading all of it every turn. A 30 KB memory file
costs ~8 000 tokens on every single request, which is exactly the cost the harness exists to avoid.
## Decision
**Plain markdown files. No database, no vector store, no embedding model, no server.**
```
.agents/memory/
INDEX.md # one line per shard + counts + token estimate — the ONLY always-loaded file
seed.md failures.md corrections.md insights.md conventions.md quirks.md
```
1. **Progressive disclosure.** Only `INDEX.md` enters context automatically. Shards are pulled on
demand via `lh memory get <shard>` or scored retrieval with `lh memory get --query`.
2. **Categorised shards**, so retrieval is targeted rather than semantic-guessy.
3. **Budgeted compaction.** When the estimated total exceeds
`memory.tokenBudget × memory.compactAtPercent`, `lh memory compact` deterministically merges the
oldest entries of the largest shards. The newest entries are never touched.
4. **Secret scanning is mandatory on write.** `putEntry` refuses to persist an entry containing
anything matching a credential pattern. Previews are redacted.
5. **Committed by default.** Durable memory, the architecture doc, conventions, specs and run
journals are committed so they are reviewable in PRs and shared with the team. Volatile
artifacts (`.cache/`, `events.ndjson`, `board.md`) are gitignored. The split is configurable.
Deliberately rejected: SQLite (a native dependency), embeddings (a model dependency and
non-determinism), and any MCP memory server that needs a running service.
## Consequences
- Zero install cost, zero runtime cost, works offline, identical on Linux/macOS/Windows.
- Memory is human-readable and diffable — this *is* the self-documenting requirement, not a
separate feature.
- Retrieval is lexical, not semantic. Accepted: shards are small and categorised, and the agent
knows which category it wants. Semantic search can be added later behind the same
`getMemory()` interface without changing the file format.
- Compaction is lossy by design. Mitigated by never compacting recent entries and by keeping the
full history in git.
+50
View File
@@ -0,0 +1,50 @@
# ADR 0006 — Local NDJSON telemetry, no server, no budget cap
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
Requirement: the harness must be *fast, observable and benchmarkable*. Scope was explicitly set to
**telemetry only** — per-run token counts, latency, tool-call traces, pass/fail — not an external
benchmark suite.
Separately, the operator chose **no budget cap** on parallel fan-out, combined with **fully
dynamic** re-planning. That combination is the single largest cost risk in the system.
## Decision
1. **Append-only NDJSON** at `.agents/runs/<runId>/events.ndjson`. One JSON object per line,
written with `O_APPEND` so concurrent lanes cannot interleave partial lines.
2. **OpenTelemetry GenAI-inspired attribute names** (`gen_ai.usage.input_tokens`,
`gen_ai.request.model`, …) **without an OTEL dependency**. The data is portable to a real
tracing backend later; today it costs nothing.
3. **Tolerant reader.** Malformed lines are skipped, never thrown on. A crashed agent must not
destroy the log.
4. **Live board.** `board.md` is regenerated on every event and shows lanes in flight, ralph
iterations, elapsed time, tokens, and — prominently — **burn rate in tokens/min**. Board
rendering is wrapped in try/catch so it can never break a telemetry write.
5. **`lh report`** renders the markdown summary; **`lh report --journal`** renders the committed,
self-documenting per-run journal.
### On the absent budget cap
No hard cap is enforced, per the operator's explicit choice. The mitigation is **visibility, not
prevention**: live burn rate in `board.md`, a post-run cost breakdown in `lh report`, and an
advisory `concurrency.maxWriteLanes` ceiling that warns and only blocks under `--strict`.
## Consequences
- Zero infrastructure. Telemetry works offline and in CI with no setup.
- `lh run event` is on the hot path (dozens of calls per run) and must stay well under 100 ms.
- Cost overrun is possible by design. Accepted and documented; visibility is the control.
- Because the schema is OTEL-shaped, exporting to Langfuse/OpenLLMetry later is a transform, not a
re-instrumentation.
## Alternatives considered
- **OpenTelemetry SDK** — rejected: heavy dependency and a collector to run, for data we currently
only read locally.
- **SQLite event store** — rejected: native dependency (see ADR 0005), and NDJSON is already
append-safe and greppable.
- **Hard credit cap** — rejected by the operator. Revisit if real runs overrun.