Recovered from crashed session (Node OOM). Repo contains full P0-P6 scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006, lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10 .github/agents, 12 CLI skills, instructions, context7 mcp.json, and unit/e2e test suite. Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs use --input-tokens/--output-tokens, so telemetry totals were always 0. Now accepts both forms. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2.5 KiB
2.5 KiB
ADR 0005 — File-only memory with a single always-loaded index
- Status: Accepted
- Date: 2026-09-09
Context
The harness must be lean on token and context size and use a lean, OS-agnostic, easily configurable memory system. It must also be self-documenting and work on greenfield and brownfield repos.
The dominant failure mode of agent memory is loading all of it every turn. A 30 KB memory file costs ~8 000 tokens on every single request, which is exactly the cost the harness exists to avoid.
Decision
Plain markdown files. No database, no vector store, no embedding model, no server.
.agents/memory/
INDEX.md # one line per shard + counts + token estimate — the ONLY always-loaded file
seed.md failures.md corrections.md insights.md conventions.md quirks.md
- Progressive disclosure. Only
INDEX.mdenters context automatically. Shards are pulled on demand vialh memory get <shard>or scored retrieval withlh memory get --query. - Categorised shards, so retrieval is targeted rather than semantic-guessy.
- Budgeted compaction. When the estimated total exceeds
memory.tokenBudget × memory.compactAtPercent,lh memory compactdeterministically merges the oldest entries of the largest shards. The newest entries are never touched. - Secret scanning is mandatory on write.
putEntryrefuses to persist an entry containing anything matching a credential pattern. Previews are redacted. - Committed by default. Durable memory, the architecture doc, conventions, specs and run
journals are committed so they are reviewable in PRs and shared with the team. Volatile
artifacts (
.cache/,events.ndjson,board.md) are gitignored. The split is configurable.
Deliberately rejected: SQLite (a native dependency), embeddings (a model dependency and non-determinism), and any MCP memory server that needs a running service.
Consequences
- Zero install cost, zero runtime cost, works offline, identical on Linux/macOS/Windows.
- Memory is human-readable and diffable — this is the self-documenting requirement, not a separate feature.
- Retrieval is lexical, not semantic. Accepted: shards are small and categorised, and the agent
knows which category it wants. Semantic search can be added later behind the same
getMemory()interface without changing the file format. - Compaction is lossy by design. Mitigated by never compacting recent entries and by keeping the full history in git.