\`lh doctor\`'s context7 check is a required, env-only check by design (ADR: never inferred, never persisted). Tests that shell out to \`lh\` or \`scripts/onboard.mjs\` inherited whatever CONTEXT7_API_KEY the developer's shell happened to export, so \`doctor passes once the repo is initialised\` and the onboard-script e2e test only ever passed on machines with a real key set — never verified in a clean environment until this CI run (no secret configured, correctly). Fixed by injecting an obviously-fake fixture key (TEST_ENV in helpers.mjs, exported and reused by onboard.test.mjs) into every subprocess these tests spawn, so behaviour no longer depends on the ambient shell. Verified locally with \`env -u CONTEXT7_API_KEY\` to reproduce the CI environment exactly: 58/58 pass either way now. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
redsen-lean-harness
An imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot.
It imposes a spec-driven pipeline, runs the implement→verify stage as a Ralph loop, drives parallel dynamic workflows on its own using git worktree lanes, and writes down everything it did. Mostly markdown. One runtime dependency.
Why it exists
Coding agents fail in predictable ways: they guess at requirements, load far too much context, lose what they learned between sessions, work sequentially when the work is parallel, and leave no trace of why anything was done.
This harness fixes each of those with a specific mechanism, not with prompt-engineering hope.
| Failure | Mechanism |
|---|---|
| Guessing at requirements | interrogator asks questions with recommended answers; the pipeline blocks until answered |
| Context bloat | Only memory/INDEX.md is ever auto-loaded; everything else is pulled on demand |
| Amnesia between runs | Categorised memory shards, committed to the repo |
| Sequential work | Dynamic DAG + git worktree lanes with file-scope leases |
| No audit trail | ADRs, living specs, and a per-run journal — written by a cheap agent, continuously |
| Unverifiable "done" | Five explicit Ralph exit criteria, all machine-checkable |
Design
Three layers, one hard boundary — we never reimplement an agent runtime (ADR 0001).
| Layer | Owns | Where |
|---|---|---|
| Markdown | Behaviour — agents, skills, instructions | .github/ |
Node CLI (lh) |
Determinism — index, gates, lanes, telemetry, memory | src/ |
| Host | Execution — subagents, fleet, tools, models | Copilot CLI / VS Code |
lh never calls a model. Agents never do git plumbing, parsing, or token arithmetic by hand.
Install
GitHub Copilot CLI
# direct from the repo
copilot plugin install redsentech/lean-harness
# or via the marketplace
copilot plugin marketplace add redsentech/lean-harness
copilot plugin install redsen-lean-harness@redsen
VS Code Copilot
Clone the repo and point VS Code's Copilot customization settings at it, or copy the .github/
tree into your project. .github/agents/, .github/instructions/ and .github/mcp.json are
read by both hosts.
.github/skills/is Copilot CLI only. VS Code gets equivalent behaviour through.github/agents/and.github/prompts/.
The lh CLI
npm install -g @redsen/lean-harness # or: npx @redsen/lean-harness <command>
lh doctor # verify the environment
Context7
Mandatory before any external-library work. The key is read from the environment only — the harness never writes secrets to disk.
export CONTEXT7_API_KEY="<your key>" # add to ~/.bashrc or ~/.zshrc
Quick start
cd your-repo
lh init # first-run wizard; writes .agents/harness.config.json
lh doctor # confirm everything is wired
Then, in Copilot CLI or VS Code Copilot:
@conductor build a rate limiter for the public API
For a small brownfield change, skip the ceremony:
/fast-track fix the off-by-one in pagination
To learn an unfamiliar codebase first:
/onboard
Onboarding a new project
Before the plugin is published to a marketplace (or if you want to try it against a local repo
first), scripts/onboard.mjs does the whole bootstrap in one shot: copies the behaviour layer
(.github/agents, skills, instructions, prompts, mcp.json, copilot-instructions.md,
AGENTS.md) into a target repo, npm links the lh CLI so it's on PATH there, runs
lh init and lh doctor, and prints next steps.
# from inside this repo
node scripts/onboard.mjs /path/to/target-repo --dry-run # preview, writes nothing
node scripts/onboard.mjs /path/to/target-repo --yes # do it
| Flag | Effect |
|---|---|
--yes |
Non-interactive: auto git init if needed, run lh init without asking |
--force |
Overwrite target files that already exist and differ (default: skip + warn) |
--dry-run |
Print the plan, write and run nothing |
--no-npm-link |
Skip npm link; print the absolute node .../src/cli.mjs invocation instead |
Safe to re-run: files that are already identical are left alone, and re-running lh init on an
already-initialized repo warns instead of failing (rerun with lh init --force to reset).
npm install -g @redsen/lean-harness (once published) replaces steps 3–5 of the script with a
normal global install.
Pipeline
design → interrogator asks, user answers → decisions.md [USER GATE]
plan → architect writes spec + acceptance criteria
splitter emits plan.dag.json (lanes + file scopes)
build → lh host picks the strategy
read lanes → shared checkout, N-wide scout fan-out
write lanes → lh lane create → git worktree + branch
builder ⇄ verifier (RALPH loop)
checkpoint → conductor RE-PLANS, may spawn/kill/re-scope lanes
failure → isolate → bounded retry → re-plan around it
integrate → integrator merges lanes SEQUENTIALLY, then one full verify
document → scribe writes journal, ADRs, spec, conventions (throughout)
Ralph exit criteria
A lane is done only when all of these hold. Otherwise it iterates; at max iterations it escalates to you.
- Every declared verify command exits 0
- Every acceptance criterion is individually checked off
lh graphstructural gate passesreviewerapproves- No files were changed outside the lane's declared scope
Agents
| Agent | Tier | Invokable | Responsibility |
|---|---|---|---|
conductor |
strong | yes | Entry point. Owns the pipeline and the dynamic DAG. |
interrogator |
strong | yes | Design-phase Q&A with recommended answers. |
scout |
cheap | – | Read-only recon, fanned out N-wide. |
architect |
strong | yes | Spec, acceptance criteria, architecture doc, ADRs. |
splitter |
mid | – | Decomposes the spec into lanes with file-scope globs. |
builder |
mid | – | Implements one lane inside its worktree. |
verifier |
cheap | – | Runs verify commands + the structural gate. |
reviewer |
strong | yes | Acceptance-criteria and scope gate. |
integrator |
strong | – | Sequential merge, conflict escalation, full verify. |
scribe |
cheap | – | Journal, ADRs, living spec, conventions. |
Tiers map to models in harness.config.json and are fully overridable. The default map is
cheap → claude-haiku-4.5, mid → claude-sonnet-5, strong → claude-opus-5. Expensive
exploration is deliberately pushed onto the cheap tier; only summaries return to the main context.
lh commands
| Command | Purpose |
|---|---|
lh init |
First-run wizard → .agents/harness.config.json |
lh index [--budget N] [--focus g] [--fetch] |
Token-budgeted tree-sitter repo map |
lh graph |
Structural gate, terse output by default. Exit 1 on violations |
lh lane create|list|status|merge|drop |
Worktree lanes + file-scope leases |
lh run start|event|end |
NDJSON telemetry + live board |
lh memory get|put|compact|scan|list |
Memory shards, compaction, secret scanning |
lh host |
Detect host capabilities, print the orchestration strategy |
lh report [runId] [--journal] |
Markdown telemetry report / run journal |
lh doctor |
Environment checks |
Full flag reference
Every flag each subcommand reads (click to expand)
lh init [--yes] [--force] [--json]—--yesskips prompts with defaults;--forcere-initializes an existing.agents/harness.config.json.lh index [--budget N] [--focus <glob>] [--fetch] [--force] [--stats] [--json]—--fetchlazily downloads/caches missing tree-sitter grammars;--forcerebuilds the cache instead of reusing it;--statsprints token counts instead of the map.lh graph [--severity <level>] [--fix-manifest] [--json]— terse brief is the default output;--fix-manifestrewritescodegraph.manifest.json-style baselines.lh lane create --id <id> [--title t] [--kind read\|write] [--scope glob...] [--depends-on id] [--base ref] [--json]lh lane list [--status s] [--run-id id] [--json]lh lane status --id <id> [--json]lh lane merge --id <id> [--abort] [--strategy s] [--json]lh lane drop --id <id> [--force] [--json]lh run start --objective "..." [--spec slug] [--host h] [--strategy s] [--attrs '{"k":"v"}'] [--json]lh run event --run <id> --type <t> [--name n] [--status ok\|fail] [--agent a] [--lane id] [--model m] [--operation op] [--input-tokens N] [--output-tokens N] [--duration-ms N] [--exit-code N] [--attrs '{"k":"v"}'] [--json]lh run end --run <id> [--status ok\|fail] [--summary "..."] [--json]lh run list [--json]/lh run show [runId] [--json]lh memory list [--shard s] [--json]lh memory get --shard s [--query q] [--limit N] [--json]lh memory put --shard s --title t [--body "..." | piped via stdin] [--tags a,b] [--allow-secrets] [--json]— writes are refused if a secret pattern matches unless--allow-secretsis set.lh memory compact [--force] [--json]lh memory scan— secret scan only, no writelh host [--strategy s] [--json]lh report [runId] [--run-id id] [--json] [--board] [--journal] [--spec slug] [--write]lh doctor [--json]
Repository footprint
The harness writes into one configurable directory:
.agents/
harness.config.json
architecture.md # living, ADR log inside [committed]
conventions.md # living [committed]
memory/INDEX.md # the ONLY always-loaded file [committed]
memory/{seed,failures,corrections,insights,conventions,quirks}.md
specs/<slug>/{questionnaire,decisions,spec}.md · plan.dag.json
runs/<id>/journal.md [committed]
runs/<id>/{board.md,events.ndjson} [gitignored]
.cache/ # repomap, symbols, wasm, lanes [gitignored]
The committed/gitignored split is chosen during lh init.
Observability
lh run event appends OTEL-GenAI-shaped NDJSON. No collector, no server, works offline.
board.mdregenerates on every event: lanes in flight, ralph iterations, elapsed, tokens, and live burn rate.lh reportgives the post-run breakdown by phase, agent, model and lane.journal.mdis the committed, human-readable record of what happened and why.
There is no hard budget cap — a deliberate choice. The control is visibility: live burn rate while running, full cost breakdown after. See ADR 0006.
Token discipline
- Only
memory/INDEX.mdis always loaded; shards are pulled on demand. lh index --budgetcaps the repo map and degrades signatures → names → counts.- Cheap-tier subagents absorb exploration; only their summaries re-enter the main context.
lh graphandlh reportemit briefs, never dumps.- Skill bodies stay short; procedures live in
templates/, loaded only when used. lh memory compactmerges the oldest entries when the budget is approached.
Architecture decisions
| ADR | Decision |
|---|---|
| 0001 | Markdown owns behaviour, Node owns determinism |
| 0002 | web-tree-sitter WASM index with an offline regex fallback |
| 0003 | git worktree per write-lane; pluggable isolation backend |
| 0004 | No vendoring of Elastic-licensed code |
| 0005 | File-only memory with a single always-loaded index |
| 0006 | Local NDJSON telemetry, no server, no budget cap |
Development
npm install
npm run validate # manifests + every markdown frontmatter block
npm test
node src/cli.mjs --help
node scripts/onboard.mjs <target-dir> --dry-run # preview bootstrapping another repo
npm run validate is the distribution gate: it checks plugin.json, marketplace.json,
agent/skill/instruction frontmatter, delegation targets, and scans .github/mcp.json for
committed secrets. npm test covers config, memory, lane scope safety, glob matching, secret
scanning, telemetry, and the onboarding script end to end (tests/onboard.test.mjs).
Prior art
Ideas taken (no code, no dependencies) from pi-hermes-memory (categorised memory shards,
secret scanning, consolidation) and pi-context-mode (token budgeting, compaction, checkpoint
anchors — Elastic Licensed, deliberately not vendored, see ADR 0004). The structural gate is
modelled on codegraph; the Ralph loop and run-state telemetry on rapid-prototyping-agent;
the packaging on redsen-copilot-agents.
Roadmap
v1 — spec pipeline, Ralph loop, memory, index + structural gate, worktree lanes, integrator, telemetry.
v2 — devcontainer isolation backend with a generator/manager/updater ecosystem
(lh lane is already backend-pluggable), semantic memory retrieval behind the existing
getMemory() interface, and OTEL export.
License
MIT © Redsen