mozempkandCopilot 3cbd929901
CI / test (20) (push) Successful in 13s
CI / test (22) (push) Successful in 13s
feat(conductor): consolidate all operations behind single entry point
- add OPERATIONS table to conductor.agent.md (init, doctor, onboard,
  index, memory, telemetry, design, plan, build, verify, integrate,
  fast-track) so the conductor agent runs any named operation directly
  instead of only the full pipeline
- thin every skill file to a one-line pointer into conductor's
  OPERATIONS table, removing duplicated procedure text (contributor
  rule: no duplicated behavior in prompts/skills)
- document the operations table in README.md and AGENTS.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-10 03:20:37 +02:00

redsen-lean-harness

An imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot.

It imposes a spec-driven pipeline, runs the implement→verify stage as a Ralph loop, drives parallel dynamic workflows on its own using git worktree lanes, and writes down everything it did. Mostly markdown. One runtime dependency.


Why it exists

Coding agents fail in predictable ways: they guess at requirements, load far too much context, lose what they learned between sessions, work sequentially when the work is parallel, and leave no trace of why anything was done.

This harness fixes each of those with a specific mechanism, not with prompt-engineering hope.

Failure Mechanism
Guessing at requirements conductor asks questions itself with recommended answers; the pipeline blocks until answered
Context bloat Only memory/INDEX.md is ever auto-loaded; everything else is pulled on demand
Amnesia between runs Categorised memory shards, committed to the repo
Sequential work Dynamic DAG + git worktree lanes with file-scope leases
No audit trail ADRs, living specs, and a per-run journal — written by a cheap agent, continuously
Unverifiable "done" Five explicit Ralph exit criteria, all machine-checkable

Design

Three layers, one hard boundary — we never reimplement an agent runtime (ADR 0001).

Layer Owns Where
Markdown Behaviour — agents, skills, instructions .github/
Node CLI (lh) Determinism — index, gates, lanes, telemetry, memory src/
Host Execution — subagents, fleet, tools, models Copilot CLI / VS Code

lh never calls a model. Agents never do git plumbing, parsing, or token arithmetic by hand.

Install

GitHub Copilot CLI

# preferred: via the marketplace this repo publishes (verified end-to-end, no warning)
copilot plugin marketplace add redsentech/lean-harness
copilot plugin install redsen-lean-harness@redsen

# also works, but Copilot CLI prints a deprecation warning for direct-source installs
copilot plugin install redsentech/lean-harness

Either form installs only the behaviour layer (agents, skills, instructions, MCP config), host-wide across every repo you open with that Copilot CLI account. You still need the lh CLI on PATH separately — see below.

VS Code Copilot

Clone the repo and point VS Code's Copilot customization settings at it, or copy the .github/ tree into your project. .github/agents/, .github/instructions/ and .github/mcp.json are read by both hosts.

.github/skills/ is Copilot CLI only. VS Code gets equivalent behaviour through .github/agents/ and .github/prompts/.

The lh CLI

The behaviour layer above only works if the agents can actually call lh — it's what runs lh init/lh doctor/lh graph/etc. under the hood. Published to GitHub Packages (npm.pkg.github.com), a private, org-scoped registry — not the public npm registry, so it needs one extra step. scripts/setup-npm-registry.mjs (zero dependencies) does it for you:

git clone git@github.com:redsentech/lean-harness.git && cd lean-harness && npm install
node scripts/setup-npm-registry.mjs   # finds a token (flag > env > `gh auth token` > browser-
                                       # assisted classic-PAT creation + paste), verifies it,
                                       # writes ~/.npmrc, confirms npm can reach the registry —
                                       # never prints the token in full
npm install -g @redsentech/lean-harness   # or: npx @redsentech/lean-harness <command>
lh doctor                                 # verify the environment

Generating a GitHub token

If no token is found via --token, NPM_REGISTRY_TOKEN/GITHUB_TOKEN/GH_TOKEN, or gh auth token, the script walks you through creating one — no manual scope-hunting required:

  1. It opens github.com/settings/tokens/new in your default browser, pre-filled with the read:packages scope and a memorable description (falls back to printing the URL if it can't launch a browser — e.g. over SSH, in a container, or in CI).
  2. On that page, set:
    • Note — anything memorable, e.g. npm-registry
    • Expiration — your choice (90 days is a reasonable default)
    • Scopes — read:packages is pre-checked (required to install); also check write:packages if you need to publish
  3. Click Generate token, copy it (starts with ghp_) — GitHub only shows it once.
  4. Paste it back into the terminal prompt (input is hidden).

Only classic PATs work reliably with GitHub Packages — fine-grained tokens don't yet support package scopes, so the script always links to the classic token page.

Pass --no-open to skip the automatic browser launch and just print the URL (useful in headless/SSH sessions where there's no display to open a browser on).

Equivalent by hand, if you already have a token with read:packages:

echo "@redsentech:registry=https://npm.pkg.github.com" >> ~/.npmrc
echo "//npm.pkg.github.com/:_authToken=<your token>" >> ~/.npmrc

Undo either with node scripts/setup-npm-registry.mjs --unset.

No published version yet? Clone and link instead:

git clone git@github.com:redsentech/lean-harness.git && cd lean-harness && npm install
npm link                              # puts `lh` on PATH globally
lh doctor                             # verify the environment

scripts/onboard.mjs (see Onboarding a new project) does the clone-free equivalent of npm link, scoped to one target repo, in a single command.

Context7

Mandatory before any external-library work. The key is read from the environment only — the harness never writes secrets to disk.

export CONTEXT7_API_KEY="<your key>"   # add to ~/.bashrc or ~/.zshrc

Quick start

Full walkthrough — install paths compared, troubleshooting, and exactly when you (vs. the agents) need to run lh yourself: docs/QUICKSTART.md

cd your-repo
lh init            # optional: first-run wizard; writes .agents/harness.config.json
lh doctor          # optional: confirm everything is wired

conductor runs lh doctor itself and self-heals a missing config with lh init --yes before doing anything else, so the two commands above are an optional manual pre-flight, not a required step.

conductor is the single entry point — it asks every design question itself, directly, in its own turn. There is no separate design agent to invoke first. Start with:

Use the conductor agent to build a rate limiter for the public API

If .agents/specs/<slug>/decisions.md is missing or incomplete, conductor asks its own numbered questions right there in its response (each with a recommended answer, a Why:, and a freeform Other: option), then stops and waits for your real reply. It never guesses, infers, or delegates the question to a subagent — subagent calls in Copilot CLI are stateless (they run to completion and return one final result; they cannot pause mid-task for a live human reply), so any question that needs a real answer is asked directly, never through the agent tool.

How to invoke a custom agent (per the official docs): @ in Copilot CLI only mentions files, never agents. There are three ways:

  1. /agent → pick conductor from the list, then press Enter to confirm the selection before typing your build prompt (selecting and prompting are two separate steps).
  2. Name it in your prompt — Use the conductor agent to ... — Copilot infers which agent you mean.
  3. copilot --agent=NAME -p "..." — force a specific agent non-interactively.

Agent name note: inside this repo (or any repo where the harness lives natively in .github/agents), the bare name conductor resolves. Once installed as a plugin into another project, Copilot CLI namespaces agents only — the resolvable name becomes redsen-lean-harness:conductor, redsen-lean-harness:architect, etc. Skills are never namespaced — /fast-track, /design, /build, /verify, /onboard, etc. work as plain slash commands regardless of install method.

If /agent doesn't list it, or invoking it reports "not found": the plugin is stale. Run copilot plugin update redsen-lean-harness (or copilot plugin marketplace update redsen && copilot plugin update redsen-lean-harness if that alone doesn't refresh it), then start a brand-new copilot session — an already-running session keeps the plugin snapshot it loaded at startup and won't pick up the update until restarted.

Every operation this harness performs is reachable through conductor — select it once, then name what you want in plain language:

Say to conductor It runs
onboard env check → repo index → memory index → graph → summary
doctor environment/config/host capability check, no mutation
init lh init --yes + baseline confirmation
index token-budgeted repo map
memory shard list/get/put/scan/compact
telemetry run start/event/end + report
design gated Q&A, one question per turn
plan spec + lane DAG
build worktree lanes, Ralph loop
verify commands + lh graph + acceptance/scope check, no edits
integrate sequential lane merge + final verify
fast-track small brownfield change, single lane, same Ralph/journal discipline
(anything else) the full pipeline, start to finish

For a small brownfield change, skip the ceremony:

Use the conductor agent: fast-track — fix the off-by-one in pagination

To learn an unfamiliar codebase first:

Use the conductor agent to onboard me on this repo

.github/skills/ still ships one thin skill (/fast-track, /onboard, etc.) per operation for convenience — each is a one-line pointer into conductor.agent.md § OPERATIONS, not a separate implementation, so behaviour never drifts between the two entry points.

Onboarding a new project

Before the plugin is published to a marketplace (or if you want to try it against a local repo first), scripts/onboard.mjs does the whole bootstrap in one shot: copies the behaviour layer (.github/agents, skills, instructions, prompts, mcp.json, copilot-instructions.md, AGENTS.md) into a target repo, npm links the lh CLI so it's on PATH there, runs lh init and lh doctor, and prints next steps.

# from inside this repo
node scripts/onboard.mjs /path/to/target-repo --dry-run   # preview, writes nothing
node scripts/onboard.mjs /path/to/target-repo --yes        # do it
Flag Effect
--yes Non-interactive: auto git init if needed, run lh init without asking
--force Overwrite target files that already exist and differ (default: skip + warn)
--dry-run Print the plan, write and run nothing
--no-npm-link Skip npm link; print the absolute node .../src/cli.mjs invocation instead

Safe to re-run: files that are already identical are left alone, and re-running lh init on an already-initialized repo warns instead of failing (rerun with lh init --force to reset). npm install -g @redsentech/lean-harness (once your .npmrc points @redsentech at GitHub Packages — see The lh CLI) replaces steps 3–5 of the script with a normal global install.

Pipeline

design    → conductor asks (inline), user answers     → decisions.md      [USER GATE]
plan      → architect writes spec + acceptance criteria
            splitter emits plan.dag.json (lanes + file scopes)
build     → lh host picks the strategy
              read lanes  → shared checkout, N-wide scout fan-out
              write lanes → lh lane create → git worktree + branch
                            builder ⇄ verifier   (RALPH loop)
            checkpoint  → conductor RE-PLANS, may spawn/kill/re-scope lanes
            failure     → isolate → bounded retry → re-plan around it
integrate → integrator merges lanes SEQUENTIALLY, then one full verify
document  → scribe writes journal, ADRs, spec, conventions   (throughout)

Ralph exit criteria

A lane is done only when all of these hold. Otherwise it iterates; at max iterations it escalates to you.

  1. Every declared verify command exits 0
  2. Every acceptance criterion is individually checked off
  3. lh graph structural gate passes
  4. reviewer approves
  5. No files were changed outside the lane's declared scope

Agents

Agent Tier Invokable Responsibility
conductor strong yes Single entry point. Asks design questions itself, then owns the pipeline and the dynamic DAG.
scout cheap – Read-only recon, fanned out N-wide.
architect strong – Spec, acceptance criteria, architecture doc, ADRs.
splitter mid – Decomposes the spec into lanes with file-scope globs.
builder mid – Implements one lane inside its worktree.
verifier cheap – Runs verify commands + the structural gate.
reviewer strong – Acceptance-criteria and scope gate.
integrator strong – Sequential merge, conflict escalation, full verify.
scribe cheap – Journal, ADRs, living spec, conventions.

Tiers map to models in harness.config.json and are fully overridable. The default map is cheap → claude-haiku-4.5, mid → claude-sonnet-5, strong → claude-opus-5. Expensive exploration is deliberately pushed onto the cheap tier; only summaries return to the main context.

lh commands

Command Purpose
lh init First-run wizard → .agents/harness.config.json
lh index [--budget N] [--focus g] [--fetch] Token-budgeted tree-sitter repo map
lh graph Structural gate, terse output by default. Exit 1 on violations
lh lane create|list|status|merge|drop Worktree lanes + file-scope leases
lh run start|event|end NDJSON telemetry + live board
lh memory get|put|compact|scan|list Memory shards, compaction, secret scanning
lh host Detect host capabilities, print the orchestration strategy
lh report [runId] [--journal] Markdown telemetry report / run journal
lh doctor Environment checks

Full flag reference

Every flag each subcommand reads (click to expand)
  • lh init [--yes] [--force] [--json] — --yes skips prompts with defaults; --force re-initializes an existing .agents/harness.config.json.
  • lh index [--budget N] [--focus <glob>] [--fetch] [--force] [--stats] [--json] — --fetch lazily downloads/caches missing tree-sitter grammars; --force rebuilds the cache instead of reusing it; --stats prints token counts instead of the map.
  • lh graph [--severity <level>] [--fix-manifest] [--json] — terse brief is the default output; --fix-manifest rewrites codegraph.manifest.json-style baselines.
  • lh lane create --id <id> [--title t] [--kind read\|write] [--scope glob...] [--depends-on id] [--base ref] [--json]
  • lh lane list [--status s] [--run-id id] [--json]
  • lh lane status --id <id> [--json]
  • lh lane merge --id <id> [--abort] [--strategy s] [--json]
  • lh lane drop --id <id> [--force] [--json]
  • lh run start --objective "..." [--spec slug] [--host h] [--strategy s] [--attrs '{"k":"v"}'] [--json]
  • lh run event --run <id> --type <t> [--name n] [--status ok\|fail] [--agent a] [--lane id] [--model m] [--operation op] [--input-tokens N] [--output-tokens N] [--duration-ms N] [--exit-code N] [--attrs '{"k":"v"}'] [--json]
  • lh run end --run <id> [--status ok\|fail] [--summary "..."] [--json]
  • lh run list [--json] / lh run show [runId] [--json]
  • lh memory list [--shard s] [--json]
  • lh memory get --shard s [--query q] [--limit N] [--json]
  • lh memory put --shard s --title t [--body "..." | piped via stdin] [--tags a,b] [--allow-secrets] [--json] — writes are refused if a secret pattern matches unless --allow-secrets is set.
  • lh memory compact [--force] [--json]
  • lh memory scan — secret scan only, no write
  • lh host [--strategy s] [--json]
  • lh report [runId] [--run-id id] [--json] [--board] [--journal] [--spec slug] [--write]
  • lh doctor [--json]

Repository footprint

The harness writes into one configurable directory:

.agents/
  harness.config.json
  architecture.md          # living, ADR log inside       [committed]
  conventions.md           # living                       [committed]
  memory/INDEX.md          # the ONLY always-loaded file  [committed]
  memory/{seed,failures,corrections,insights,conventions,quirks}.md
  specs/<slug>/{questionnaire,decisions,spec}.md · plan.dag.json
  runs/<id>/journal.md                                    [committed]
  runs/<id>/{board.md,events.ndjson}                      [gitignored]
  .cache/                  # repomap, symbols, wasm, lanes [gitignored]

The committed/gitignored split is chosen during lh init.

Observability

lh run event appends OTEL-GenAI-shaped NDJSON. No collector, no server, works offline.

  • board.md regenerates on every event: lanes in flight, ralph iterations, elapsed, tokens, and live burn rate.
  • lh report gives the post-run breakdown by phase, agent, model and lane.
  • journal.md is the committed, human-readable record of what happened and why.

There is no hard budget cap — a deliberate choice. The control is visibility: live burn rate while running, full cost breakdown after. See ADR 0006.

Token discipline

  1. Only memory/INDEX.md is always loaded; shards are pulled on demand.
  2. lh index --budget caps the repo map and degrades signatures → names → counts.
  3. Cheap-tier subagents absorb exploration; only their summaries re-enter the main context.
  4. lh graph and lh report emit briefs, never dumps.
  5. Skill bodies stay short; procedures live in templates/, loaded only when used.
  6. lh memory compact merges the oldest entries when the budget is approached.

Architecture decisions

ADR Decision
0001 Markdown owns behaviour, Node owns determinism
0002 web-tree-sitter WASM index with an offline regex fallback
0003 git worktree per write-lane; pluggable isolation backend
0004 No vendoring of Elastic-licensed code
0005 File-only memory with a single always-loaded index
0006 Local NDJSON telemetry, no server, no budget cap

Development

npm install
npm run validate     # manifests + every markdown frontmatter block
npm test
node src/cli.mjs --help
node scripts/onboard.mjs <target-dir> --dry-run   # preview bootstrapping another repo

npm run validate is the distribution gate: it checks plugin.json, marketplace.json, agent/skill/instruction frontmatter, delegation targets, and scans .github/mcp.json for committed secrets. npm test covers config, memory, lane scope safety, glob matching, secret scanning, telemetry, and the onboarding script end to end (tests/onboard.test.mjs).

Prior art

Ideas taken (no code, no dependencies) from pi-hermes-memory (categorised memory shards, secret scanning, consolidation) and pi-context-mode (token budgeting, compaction, checkpoint anchors — Elastic Licensed, deliberately not vendored, see ADR 0004). The structural gate is modelled on codegraph; the Ralph loop and run-state telemetry on rapid-prototyping-agent; the packaging on redsen-copilot-agents.

Roadmap

v1 — spec pipeline, Ralph loop, memory, index + structural gate, worktree lanes, integrator, telemetry.

v2 — devcontainer isolation backend with a generator/manager/updater ecosystem (lh lane is already backend-pluggable), semantic memory retrieval behind the existing getMemory() interface, and OTEL export.

License

MIT © Redsen

S
Description
redsen-lean-harness: imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot
Readme MIT
190 KiB
Languages
JavaScript 100%