mozempkandCopilot cd36cc0efc fix: telemetry/memory flag mismatches, doc drift; add CI; self-dogfood
Parallel fleet audit (2 background agents) + direct work:

- src/commands/memory.mjs: `memory put --allow-secrets` was read as
  `flags.allowSecrets` (camelCase) but the CLI parser only emits
  kebab-case keys, so the flag was always undefined/false. Fixed to
  `flags['allow-secrets']`.
- Docs: `lh graph --brief` was referenced 14x across 5 agent.md files,
  harness.instructions.md, 4 SKILL.md files, onboard.prompt.md, and
  README, but `lh graph` has no --brief flag (terse output is already
  the default, --json/--severity are the only flags). Corrected every
  reference to match actual CLI surface.
- .github/workflows/ci.yml: run npm install/validate/test on node 20+22.
- Self-dogfooded `lh init --yes` in this repo, producing .agents/
  (harness.config.json, architecture.md, conventions.md, memory
  shards). `lh doctor` now reports all required checks green.
- Deduped .gitignore harness block against init's auto-managed block.

Verified: 53/53 tests pass, validate 0 errors, doctor all-green.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-09 22:50:07 +02:00

redsen-lean-harness

An imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot.

It imposes a spec-driven pipeline, runs the implement→verify stage as a Ralph loop, drives parallel dynamic workflows on its own using git worktree lanes, and writes down everything it did. Mostly markdown. One runtime dependency.


Why it exists

Coding agents fail in predictable ways: they guess at requirements, load far too much context, lose what they learned between sessions, work sequentially when the work is parallel, and leave no trace of why anything was done.

This harness fixes each of those with a specific mechanism, not with prompt-engineering hope.

Failure Mechanism
Guessing at requirements interrogator asks questions with recommended answers; the pipeline blocks until answered
Context bloat Only memory/INDEX.md is ever auto-loaded; everything else is pulled on demand
Amnesia between runs Categorised memory shards, committed to the repo
Sequential work Dynamic DAG + git worktree lanes with file-scope leases
No audit trail ADRs, living specs, and a per-run journal — written by a cheap agent, continuously
Unverifiable "done" Five explicit Ralph exit criteria, all machine-checkable

Design

Three layers, one hard boundary — we never reimplement an agent runtime (ADR 0001).

Layer Owns Where
Markdown Behaviour — agents, skills, instructions .github/
Node CLI (lh) Determinism — index, gates, lanes, telemetry, memory src/
Host Execution — subagents, fleet, tools, models Copilot CLI / VS Code

lh never calls a model. Agents never do git plumbing, parsing, or token arithmetic by hand.

Install

GitHub Copilot CLI

# direct from the repo
copilot plugin install redsentech/lean-harness

# or via the marketplace
copilot plugin marketplace add redsentech/lean-harness
copilot plugin install redsen-lean-harness@redsen

VS Code Copilot

Clone the repo and point VS Code's Copilot customization settings at it, or copy the .github/ tree into your project. .github/agents/, .github/instructions/ and .github/mcp.json are read by both hosts.

.github/skills/ is Copilot CLI only. VS Code gets equivalent behaviour through .github/agents/ and .github/prompts/.

The lh CLI

npm install -g @redsen/lean-harness   # or: npx @redsen/lean-harness <command>
lh doctor                             # verify the environment

Context7

Mandatory before any external-library work. The key is read from the environment only — the harness never writes secrets to disk.

export CONTEXT7_API_KEY="<your key>"   # add to ~/.bashrc or ~/.zshrc

Quick start

cd your-repo
lh init            # first-run wizard; writes .agents/harness.config.json
lh doctor          # confirm everything is wired

Then, in Copilot CLI or VS Code Copilot:

@conductor build a rate limiter for the public API

For a small brownfield change, skip the ceremony:

/fast-track fix the off-by-one in pagination

To learn an unfamiliar codebase first:

/onboard

Pipeline

design    → interrogator asks, user answers          → decisions.md      [USER GATE]
plan      → architect writes spec + acceptance criteria
            splitter emits plan.dag.json (lanes + file scopes)
build     → lh host picks the strategy
              read lanes  → shared checkout, N-wide scout fan-out
              write lanes → lh lane create → git worktree + branch
                            builder ⇄ verifier   (RALPH loop)
            checkpoint  → conductor RE-PLANS, may spawn/kill/re-scope lanes
            failure     → isolate → bounded retry → re-plan around it
integrate → integrator merges lanes SEQUENTIALLY, then one full verify
document  → scribe writes journal, ADRs, spec, conventions   (throughout)

Ralph exit criteria

A lane is done only when all of these hold. Otherwise it iterates; at max iterations it escalates to you.

  1. Every declared verify command exits 0
  2. Every acceptance criterion is individually checked off
  3. lh graph structural gate passes
  4. reviewer approves
  5. No files were changed outside the lane's declared scope

Agents

Agent Tier Invokable Responsibility
conductor strong yes Entry point. Owns the pipeline and the dynamic DAG.
interrogator strong yes Design-phase Q&A with recommended answers.
scout cheap – Read-only recon, fanned out N-wide.
architect strong yes Spec, acceptance criteria, architecture doc, ADRs.
splitter mid – Decomposes the spec into lanes with file-scope globs.
builder mid – Implements one lane inside its worktree.
verifier cheap – Runs verify commands + the structural gate.
reviewer strong yes Acceptance-criteria and scope gate.
integrator strong – Sequential merge, conflict escalation, full verify.
scribe cheap – Journal, ADRs, living spec, conventions.

Tiers map to models in harness.config.json and are fully overridable. The default map is cheap → claude-haiku-4.5, mid → claude-sonnet-5, strong → claude-opus-5. Expensive exploration is deliberately pushed onto the cheap tier; only summaries return to the main context.

lh commands

Command Purpose
lh init First-run wizard → .agents/harness.config.json
lh index [--budget N] [--focus g] [--fetch] Token-budgeted tree-sitter repo map
lh graph Structural gate, terse output by default. Exit 1 on violations
lh lane create|list|status|merge|drop Worktree lanes + file-scope leases
lh run start|event|end NDJSON telemetry + live board
lh memory get|put|compact|scan|list Memory shards, compaction, secret scanning
lh host Detect host capabilities, print the orchestration strategy
lh report [runId] [--journal] Markdown telemetry report / run journal
lh doctor Environment checks

Repository footprint

The harness writes into one configurable directory:

.agents/
  harness.config.json
  architecture.md          # living, ADR log inside       [committed]
  conventions.md           # living                       [committed]
  memory/INDEX.md          # the ONLY always-loaded file  [committed]
  memory/{seed,failures,corrections,insights,conventions,quirks}.md
  specs/<slug>/{questionnaire,decisions,spec}.md · plan.dag.json
  runs/<id>/journal.md                                    [committed]
  runs/<id>/{board.md,events.ndjson}                      [gitignored]
  .cache/                  # repomap, symbols, wasm, lanes [gitignored]

The committed/gitignored split is chosen during lh init.

Observability

lh run event appends OTEL-GenAI-shaped NDJSON. No collector, no server, works offline.

  • board.md regenerates on every event: lanes in flight, ralph iterations, elapsed, tokens, and live burn rate.
  • lh report gives the post-run breakdown by phase, agent, model and lane.
  • journal.md is the committed, human-readable record of what happened and why.

There is no hard budget cap — a deliberate choice. The control is visibility: live burn rate while running, full cost breakdown after. See ADR 0006.

Token discipline

  1. Only memory/INDEX.md is always loaded; shards are pulled on demand.
  2. lh index --budget caps the repo map and degrades signatures → names → counts.
  3. Cheap-tier subagents absorb exploration; only their summaries re-enter the main context.
  4. lh graph and lh report emit briefs, never dumps.
  5. Skill bodies stay short; procedures live in templates/, loaded only when used.
  6. lh memory compact merges the oldest entries when the budget is approached.

Architecture decisions

ADR Decision
0001 Markdown owns behaviour, Node owns determinism
0002 web-tree-sitter WASM index with an offline regex fallback
0003 git worktree per write-lane; pluggable isolation backend
0004 No vendoring of Elastic-licensed code
0005 File-only memory with a single always-loaded index
0006 Local NDJSON telemetry, no server, no budget cap

Development

npm install
npm run validate     # manifests + every markdown frontmatter block
npm test
node src/cli.mjs --help

npm run validate is the distribution gate: it checks plugin.json, marketplace.json, agent/skill/instruction frontmatter, delegation targets, and scans .github/mcp.json for committed secrets.

Prior art

Ideas taken (no code, no dependencies) from pi-hermes-memory (categorised memory shards, secret scanning, consolidation) and pi-context-mode (token budgeting, compaction, checkpoint anchors — Elastic Licensed, deliberately not vendored, see ADR 0004). The structural gate is modelled on codegraph; the Ralph loop and run-state telemetry on rapid-prototyping-agent; the packaging on redsen-copilot-agents.

Roadmap

v1 — spec pipeline, Ralph loop, memory, index + structural gate, worktree lanes, integrator, telemetry.

v2 — devcontainer isolation backend with a generator/manager/updater ecosystem (lh lane is already backend-pluggable), semantic memory retrieval behind the existing getMemory() interface, and OTEL export.

License

MIT © Redsen

S
Description
redsen-lean-harness: imperative, token-lean, self-documenting agent harness for GitHub Copilot CLI and VS Code Copilot
Readme MIT
190 KiB
Languages
JavaScript 100%