- add OPERATIONS table to conductor.agent.md (init, doctor, onboard, index, memory, telemetry, design, plan, build, verify, integrate, fast-track) so the conductor agent runs any named operation directly instead of only the full pipeline - thin every skill file to a one-line pointer into conductor's OPERATIONS table, removing duplicated procedure text (contributor rule: no duplicated behavior in prompts/skills) - document the operations table in README.md and AGENTS.md Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
redsen-lean-harness
An imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot.
It imposes a spec-driven pipeline, runs the implement→verify stage as a Ralph loop, drives parallel dynamic workflows on its own using git worktree lanes, and writes down everything it did. Mostly markdown. One runtime dependency.
Why it exists
Coding agents fail in predictable ways: they guess at requirements, load far too much context, lose what they learned between sessions, work sequentially when the work is parallel, and leave no trace of why anything was done.
This harness fixes each of those with a specific mechanism, not with prompt-engineering hope.
| Failure | Mechanism |
|---|---|
| Guessing at requirements | conductor asks questions itself with recommended answers; the pipeline blocks until answered |
| Context bloat | Only memory/INDEX.md is ever auto-loaded; everything else is pulled on demand |
| Amnesia between runs | Categorised memory shards, committed to the repo |
| Sequential work | Dynamic DAG + git worktree lanes with file-scope leases |
| No audit trail | ADRs, living specs, and a per-run journal — written by a cheap agent, continuously |
| Unverifiable "done" | Five explicit Ralph exit criteria, all machine-checkable |
Design
Three layers, one hard boundary — we never reimplement an agent runtime (ADR 0001).
| Layer | Owns | Where |
|---|---|---|
| Markdown | Behaviour — agents, skills, instructions | .github/ |
Node CLI (lh) |
Determinism — index, gates, lanes, telemetry, memory | src/ |
| Host | Execution — subagents, fleet, tools, models | Copilot CLI / VS Code |
lh never calls a model. Agents never do git plumbing, parsing, or token arithmetic by hand.
Install
GitHub Copilot CLI
# preferred: via the marketplace this repo publishes (verified end-to-end, no warning)
copilot plugin marketplace add redsentech/lean-harness
copilot plugin install redsen-lean-harness@redsen
# also works, but Copilot CLI prints a deprecation warning for direct-source installs
copilot plugin install redsentech/lean-harness
Either form installs only the behaviour layer (agents, skills, instructions, MCP config),
host-wide across every repo you open with that Copilot CLI account. You still need the lh
CLI on PATH separately — see below.
VS Code Copilot
Clone the repo and point VS Code's Copilot customization settings at it, or copy the .github/
tree into your project. .github/agents/, .github/instructions/ and .github/mcp.json are
read by both hosts.
.github/skills/is Copilot CLI only. VS Code gets equivalent behaviour through.github/agents/and.github/prompts/.
The lh CLI
The behaviour layer above only works if the agents can actually call lh — it's what runs
lh init/lh doctor/lh graph/etc. under the hood. Published to GitHub Packages
(npm.pkg.github.com), a private, org-scoped registry — not the public npm registry, so it
needs one extra step. scripts/setup-npm-registry.mjs (zero dependencies) does it for you:
git clone git@github.com:redsentech/lean-harness.git && cd lean-harness && npm install
node scripts/setup-npm-registry.mjs # finds a token (flag > env > `gh auth token` > browser-
# assisted classic-PAT creation + paste), verifies it,
# writes ~/.npmrc, confirms npm can reach the registry —
# never prints the token in full
npm install -g @redsentech/lean-harness # or: npx @redsentech/lean-harness <command>
lh doctor # verify the environment
Generating a GitHub token
If no token is found via --token, NPM_REGISTRY_TOKEN/GITHUB_TOKEN/GH_TOKEN, or
gh auth token, the script walks you through creating one — no manual scope-hunting required:
- It opens
github.com/settings/tokens/newin your default browser, pre-filled with theread:packagesscope and a memorable description (falls back to printing the URL if it can't launch a browser — e.g. over SSH, in a container, or in CI). - On that page, set:
- Note — anything memorable, e.g.
npm-registry - Expiration — your choice (90 days is a reasonable default)
- Scopes —
read:packagesis pre-checked (required to install); also checkwrite:packagesif you need to publish
- Note — anything memorable, e.g.
- Click Generate token, copy it (starts with
ghp_) — GitHub only shows it once. - Paste it back into the terminal prompt (input is hidden).
Only classic PATs work reliably with GitHub Packages — fine-grained tokens don't yet support package scopes, so the script always links to the classic token page.
Pass --no-open to skip the automatic browser launch and just print the URL (useful in
headless/SSH sessions where there's no display to open a browser on).
Equivalent by hand, if you already have a token with read:packages:
echo "@redsentech:registry=https://npm.pkg.github.com" >> ~/.npmrc
echo "//npm.pkg.github.com/:_authToken=<your token>" >> ~/.npmrc
Undo either with node scripts/setup-npm-registry.mjs --unset.
No published version yet? Clone and link instead:
git clone git@github.com:redsentech/lean-harness.git && cd lean-harness && npm install
npm link # puts `lh` on PATH globally
lh doctor # verify the environment
scripts/onboard.mjs (see Onboarding a new project) does the
clone-free equivalent of npm link, scoped to one target repo, in a single command.
Context7
Mandatory before any external-library work. The key is read from the environment only — the harness never writes secrets to disk.
export CONTEXT7_API_KEY="<your key>" # add to ~/.bashrc or ~/.zshrc
Quick start
Full walkthrough — install paths compared, troubleshooting, and exactly when you (vs. the agents) need to run
lhyourself: docs/QUICKSTART.md
cd your-repo
lh init # optional: first-run wizard; writes .agents/harness.config.json
lh doctor # optional: confirm everything is wired
conductor runs lh doctor itself and self-heals a missing config with lh init --yes
before doing anything else, so the two commands above are an optional manual pre-flight, not
a required step.
conductor is the single entry point — it asks every design question itself, directly, in
its own turn. There is no separate design agent to invoke first. Start with:
Use the conductor agent to build a rate limiter for the public API
If .agents/specs/<slug>/decisions.md is missing or incomplete, conductor asks its own
numbered questions right there in its response (each with a recommended answer, a Why:, and
a freeform Other: option), then stops and waits for your real reply. It never guesses,
infers, or delegates the question to a subagent — subagent calls in Copilot CLI are stateless
(they run to completion and return one final result; they cannot pause mid-task for a live
human reply), so any question that needs a real answer is asked directly, never through the
agent tool.
How to invoke a custom agent (per the official docs):
@in Copilot CLI only mentions files, never agents. There are three ways:
/agent→ pickconductorfrom the list, then press Enter to confirm the selection before typing your build prompt (selecting and prompting are two separate steps).- Name it in your prompt —
Use the conductor agent to ...— Copilot infers which agent you mean.copilot --agent=NAME -p "..."— force a specific agent non-interactively.Agent name note: inside this repo (or any repo where the harness lives natively in
.github/agents), the bare nameconductorresolves. Once installed as a plugin into another project, Copilot CLI namespaces agents only — the resolvable name becomesredsen-lean-harness:conductor,redsen-lean-harness:architect, etc. Skills are never namespaced —/fast-track,/design,/build,/verify,/onboard, etc. work as plain slash commands regardless of install method.If
/agentdoesn't list it, or invoking it reports "not found": the plugin is stale. Runcopilot plugin update redsen-lean-harness(orcopilot plugin marketplace update redsen && copilot plugin update redsen-lean-harnessif that alone doesn't refresh it), then start a brand-newcopilotsession — an already-running session keeps the plugin snapshot it loaded at startup and won't pick up the update until restarted.
Every operation this harness performs is reachable through conductor — select it once, then
name what you want in plain language:
Say to conductor |
It runs |
|---|---|
onboard |
env check → repo index → memory index → graph → summary |
doctor |
environment/config/host capability check, no mutation |
init |
lh init --yes + baseline confirmation |
index |
token-budgeted repo map |
memory |
shard list/get/put/scan/compact |
telemetry |
run start/event/end + report |
design |
gated Q&A, one question per turn |
plan |
spec + lane DAG |
build |
worktree lanes, Ralph loop |
verify |
commands + lh graph + acceptance/scope check, no edits |
integrate |
sequential lane merge + final verify |
fast-track |
small brownfield change, single lane, same Ralph/journal discipline |
| (anything else) | the full pipeline, start to finish |
For a small brownfield change, skip the ceremony:
Use the conductor agent: fast-track — fix the off-by-one in pagination
To learn an unfamiliar codebase first:
Use the conductor agent to onboard me on this repo
.github/skills/ still ships one thin skill (/fast-track, /onboard, etc.) per operation for
convenience — each is a one-line pointer into conductor.agent.md § OPERATIONS, not a separate
implementation, so behaviour never drifts between the two entry points.
Onboarding a new project
Before the plugin is published to a marketplace (or if you want to try it against a local repo
first), scripts/onboard.mjs does the whole bootstrap in one shot: copies the behaviour layer
(.github/agents, skills, instructions, prompts, mcp.json, copilot-instructions.md,
AGENTS.md) into a target repo, npm links the lh CLI so it's on PATH there, runs
lh init and lh doctor, and prints next steps.
# from inside this repo
node scripts/onboard.mjs /path/to/target-repo --dry-run # preview, writes nothing
node scripts/onboard.mjs /path/to/target-repo --yes # do it
| Flag | Effect |
|---|---|
--yes |
Non-interactive: auto git init if needed, run lh init without asking |
--force |
Overwrite target files that already exist and differ (default: skip + warn) |
--dry-run |
Print the plan, write and run nothing |
--no-npm-link |
Skip npm link; print the absolute node .../src/cli.mjs invocation instead |
Safe to re-run: files that are already identical are left alone, and re-running lh init on an
already-initialized repo warns instead of failing (rerun with lh init --force to reset).
npm install -g @redsentech/lean-harness (once your .npmrc points @redsentech at GitHub
Packages — see The lh CLI) replaces steps 3–5 of the script with a normal
global install.
Pipeline
design → conductor asks (inline), user answers → decisions.md [USER GATE]
plan → architect writes spec + acceptance criteria
splitter emits plan.dag.json (lanes + file scopes)
build → lh host picks the strategy
read lanes → shared checkout, N-wide scout fan-out
write lanes → lh lane create → git worktree + branch
builder ⇄ verifier (RALPH loop)
checkpoint → conductor RE-PLANS, may spawn/kill/re-scope lanes
failure → isolate → bounded retry → re-plan around it
integrate → integrator merges lanes SEQUENTIALLY, then one full verify
document → scribe writes journal, ADRs, spec, conventions (throughout)
Ralph exit criteria
A lane is done only when all of these hold. Otherwise it iterates; at max iterations it escalates to you.
- Every declared verify command exits 0
- Every acceptance criterion is individually checked off
lh graphstructural gate passesreviewerapproves- No files were changed outside the lane's declared scope
Agents
| Agent | Tier | Invokable | Responsibility |
|---|---|---|---|
conductor |
strong | yes | Single entry point. Asks design questions itself, then owns the pipeline and the dynamic DAG. |
scout |
cheap | – | Read-only recon, fanned out N-wide. |
architect |
strong | – | Spec, acceptance criteria, architecture doc, ADRs. |
splitter |
mid | – | Decomposes the spec into lanes with file-scope globs. |
builder |
mid | – | Implements one lane inside its worktree. |
verifier |
cheap | – | Runs verify commands + the structural gate. |
reviewer |
strong | – | Acceptance-criteria and scope gate. |
integrator |
strong | – | Sequential merge, conflict escalation, full verify. |
scribe |
cheap | – | Journal, ADRs, living spec, conventions. |
Tiers map to models in harness.config.json and are fully overridable. The default map is
cheap → claude-haiku-4.5, mid → claude-sonnet-5, strong → claude-opus-5. Expensive
exploration is deliberately pushed onto the cheap tier; only summaries return to the main context.
lh commands
| Command | Purpose |
|---|---|
lh init |
First-run wizard → .agents/harness.config.json |
lh index [--budget N] [--focus g] [--fetch] |
Token-budgeted tree-sitter repo map |
lh graph |
Structural gate, terse output by default. Exit 1 on violations |
lh lane create|list|status|merge|drop |
Worktree lanes + file-scope leases |
lh run start|event|end |
NDJSON telemetry + live board |
lh memory get|put|compact|scan|list |
Memory shards, compaction, secret scanning |
lh host |
Detect host capabilities, print the orchestration strategy |
lh report [runId] [--journal] |
Markdown telemetry report / run journal |
lh doctor |
Environment checks |
Full flag reference
Every flag each subcommand reads (click to expand)
lh init [--yes] [--force] [--json]—--yesskips prompts with defaults;--forcere-initializes an existing.agents/harness.config.json.lh index [--budget N] [--focus <glob>] [--fetch] [--force] [--stats] [--json]—--fetchlazily downloads/caches missing tree-sitter grammars;--forcerebuilds the cache instead of reusing it;--statsprints token counts instead of the map.lh graph [--severity <level>] [--fix-manifest] [--json]— terse brief is the default output;--fix-manifestrewritescodegraph.manifest.json-style baselines.lh lane create --id <id> [--title t] [--kind read\|write] [--scope glob...] [--depends-on id] [--base ref] [--json]lh lane list [--status s] [--run-id id] [--json]lh lane status --id <id> [--json]lh lane merge --id <id> [--abort] [--strategy s] [--json]lh lane drop --id <id> [--force] [--json]lh run start --objective "..." [--spec slug] [--host h] [--strategy s] [--attrs '{"k":"v"}'] [--json]lh run event --run <id> --type <t> [--name n] [--status ok\|fail] [--agent a] [--lane id] [--model m] [--operation op] [--input-tokens N] [--output-tokens N] [--duration-ms N] [--exit-code N] [--attrs '{"k":"v"}'] [--json]lh run end --run <id> [--status ok\|fail] [--summary "..."] [--json]lh run list [--json]/lh run show [runId] [--json]lh memory list [--shard s] [--json]lh memory get --shard s [--query q] [--limit N] [--json]lh memory put --shard s --title t [--body "..." | piped via stdin] [--tags a,b] [--allow-secrets] [--json]— writes are refused if a secret pattern matches unless--allow-secretsis set.lh memory compact [--force] [--json]lh memory scan— secret scan only, no writelh host [--strategy s] [--json]lh report [runId] [--run-id id] [--json] [--board] [--journal] [--spec slug] [--write]lh doctor [--json]
Repository footprint
The harness writes into one configurable directory:
.agents/
harness.config.json
architecture.md # living, ADR log inside [committed]
conventions.md # living [committed]
memory/INDEX.md # the ONLY always-loaded file [committed]
memory/{seed,failures,corrections,insights,conventions,quirks}.md
specs/<slug>/{questionnaire,decisions,spec}.md · plan.dag.json
runs/<id>/journal.md [committed]
runs/<id>/{board.md,events.ndjson} [gitignored]
.cache/ # repomap, symbols, wasm, lanes [gitignored]
The committed/gitignored split is chosen during lh init.
Observability
lh run event appends OTEL-GenAI-shaped NDJSON. No collector, no server, works offline.
board.mdregenerates on every event: lanes in flight, ralph iterations, elapsed, tokens, and live burn rate.lh reportgives the post-run breakdown by phase, agent, model and lane.journal.mdis the committed, human-readable record of what happened and why.
There is no hard budget cap — a deliberate choice. The control is visibility: live burn rate while running, full cost breakdown after. See ADR 0006.
Token discipline
- Only
memory/INDEX.mdis always loaded; shards are pulled on demand. lh index --budgetcaps the repo map and degrades signatures → names → counts.- Cheap-tier subagents absorb exploration; only their summaries re-enter the main context.
lh graphandlh reportemit briefs, never dumps.- Skill bodies stay short; procedures live in
templates/, loaded only when used. lh memory compactmerges the oldest entries when the budget is approached.
Architecture decisions
| ADR | Decision |
|---|---|
| 0001 | Markdown owns behaviour, Node owns determinism |
| 0002 | web-tree-sitter WASM index with an offline regex fallback |
| 0003 | git worktree per write-lane; pluggable isolation backend |
| 0004 | No vendoring of Elastic-licensed code |
| 0005 | File-only memory with a single always-loaded index |
| 0006 | Local NDJSON telemetry, no server, no budget cap |
Development
npm install
npm run validate # manifests + every markdown frontmatter block
npm test
node src/cli.mjs --help
node scripts/onboard.mjs <target-dir> --dry-run # preview bootstrapping another repo
npm run validate is the distribution gate: it checks plugin.json, marketplace.json,
agent/skill/instruction frontmatter, delegation targets, and scans .github/mcp.json for
committed secrets. npm test covers config, memory, lane scope safety, glob matching, secret
scanning, telemetry, and the onboarding script end to end (tests/onboard.test.mjs).
Prior art
Ideas taken (no code, no dependencies) from pi-hermes-memory (categorised memory shards,
secret scanning, consolidation) and pi-context-mode (token budgeting, compaction, checkpoint
anchors — Elastic Licensed, deliberately not vendored, see ADR 0004). The structural gate is
modelled on codegraph; the Ralph loop and run-state telemetry on rapid-prototyping-agent;
the packaging on redsen-copilot-agents.
Roadmap
v1 — spec pipeline, Ralph loop, memory, index + structural gate, worktree lanes, integrator, telemetry.
v2 — devcontainer isolation backend with a generator/manager/updater ecosystem
(lh lane is already backend-pluggable), semantic memory retrieval behind the existing
getMemory() interface, and OTEL export.
License
MIT © Redsen