Search papers, labs, and topics across Lattice.
This paper argues for integrating memory solutions into the agentic development lifecycle (ADLC) by leveraging Git's version control system, which provides inherent guarantees like ground truth and verification. The authors conducted an eight-corpus retrieval study to evaluate various ranking mechanisms, ultimately achieving a best configuration with a pooled mean reciprocal rank (MRR) of ~0.31, significantly outperforming traditional methods. They propose a novel routing system that combines git-anchored structural maps and confidence-gated episodes to enhance answer sufficiency, achieving a score of 0.83 on a production system while maintaining low token usage per question.
Git's version control could revolutionize memory management in coding agents, achieving 60x better retrieval performance than traditional methods.
Coding agents now produce a growing share of a team's code, while the reasoning behind each change -- the alternatives weighed, the constraints discovered, the approaches rejected -- is trapped in assistant transcripts that vanish with the session. Memory for this setting, the agentic development lifecycle (ADLC), is usually posed as one retrieval problem and built as machinery: tiered stores, memory graphs, compiled wikis, model-judged admission. We argue memory should instead be git-bound -- built into the repository's version control, inheriting the guarantees the machinery struggles to construct: ground truth from commits, freshness from rebuild, verification from the merge, containment from review. On this ledger we solve two problems separately, then combine them. Seed supply is closed as an eight-corpus retrieval study under a pre-registered ship discipline: five imported ranking mechanisms rejected, two kept, and a best configuration of ~0.31 pooled MRR -- ~60x the raw-transcript grep floor, ~15x an honest parsed-turn floor. Answer assembly is where ranking stops helping: single-shot retrieval scores only 0.07-0.20 answer-sufficiency on real developer questions, and ungated episode injection measurably degrades good answers. A router dispatches breadth to a git-anchored structural map, pointed lookups to confidence-gated episodes, and rationale to decision synthesis, which reconstructs why-arcs no single session contains (0.83 sufficiency on a young ~50k-LOC production system). Routed, the system answers at 382-980 tokens per question -- three orders of magnitude below the recorded history. Because ground truth is mined from commit-session links rather than annotated, every result is replicable on any user's own history at zero labeling cost. The remaining constraint is capture. Code, benchmark, and paper source: github.com/rekal-dev/rekal-cli.