Files
redsen-lean-harness/docs/adr/0005-file-only-memory.md
T
mozempkandCopilot 383129f571 feat: scaffold redsen-lean-harness v0.1.0
Recovered from crashed session (Node OOM). Repo contains full P0-P6
scaffold: plugin.json/marketplace.json, AGENTS.md, ADRs 0001-0006,
lh CLI (init/index/graph/lane/run/memory/host/report/doctor), 10
.github/agents, 12 CLI skills, instructions, context7 mcp.json, and
unit/e2e test suite.

Fixed: run.mjs read --in-tokens/--out-tokens but tests and CLI docs
use --input-tokens/--output-tokens, so telemetry totals were always 0.
Now accepts both forms.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-09 22:44:15 +02:00

50 lines
2.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ADR 0005 — File-only memory with a single always-loaded index
- **Status**: Accepted
- **Date**: 2026-09-09
## Context
The harness must be *lean on token and context size* and use a *lean, OS-agnostic, easily
configurable memory system*. It must also be **self-documenting** and work on greenfield and
brownfield repos.
The dominant failure mode of agent memory is loading all of it every turn. A 30 KB memory file
costs ~8 000 tokens on every single request, which is exactly the cost the harness exists to avoid.
## Decision
**Plain markdown files. No database, no vector store, no embedding model, no server.**
```
.agents/memory/
INDEX.md # one line per shard + counts + token estimate — the ONLY always-loaded file
seed.md failures.md corrections.md insights.md conventions.md quirks.md
```
1. **Progressive disclosure.** Only `INDEX.md` enters context automatically. Shards are pulled on
demand via `lh memory get <shard>` or scored retrieval with `lh memory get --query`.
2. **Categorised shards**, so retrieval is targeted rather than semantic-guessy.
3. **Budgeted compaction.** When the estimated total exceeds
`memory.tokenBudget × memory.compactAtPercent`, `lh memory compact` deterministically merges the
oldest entries of the largest shards. The newest entries are never touched.
4. **Secret scanning is mandatory on write.** `putEntry` refuses to persist an entry containing
anything matching a credential pattern. Previews are redacted.
5. **Committed by default.** Durable memory, the architecture doc, conventions, specs and run
journals are committed so they are reviewable in PRs and shared with the team. Volatile
artifacts (`.cache/`, `events.ndjson`, `board.md`) are gitignored. The split is configurable.
Deliberately rejected: SQLite (a native dependency), embeddings (a model dependency and
non-determinism), and any MCP memory server that needs a running service.
## Consequences
- Zero install cost, zero runtime cost, works offline, identical on Linux/macOS/Windows.
- Memory is human-readable and diffable — this *is* the self-documenting requirement, not a
separate feature.
- Retrieval is lexical, not semantic. Accepted: shards are small and categorised, and the agent
knows which category it wants. Semantic search can be added later behind the same
`getMemory()` interface without changing the file format.
- Compaction is lossy by design. Mitigated by never compacting recent entries and by keeping the
full history in git.