Untitled
The distinction you’ve drawn (the decision is made by the deterministic engine; the LLM merely explains it) matches something I’ve seen in another field. I was testing the agent’s memory: the schema of a vehicle changes mid-run, and the agent must make a call appropriate to the new schema at the end. Results: - A 20-line ‘the latest definition wins’ rule: 40/40 - Providing the entire history to the model: 38/40 - Retrieval-based memory: Mem0 (open-source) 3/15, TF-IDF + recency 1/40 When presented with the correct information, the model used it without issue. What it couldn’t reliably do was decide which information was still valid. I’m curious about versioning on your end. When a rule changes, what happens to decisions made under the old rule? If someone asks six months later, “Why was this application rejected?”, does the explanation refer to the rule version at the time of the decision, or to the current one? Does the RAG side know that a policy passage has become invalid, or could it retrieve the old document to explain a new decision?
Comments
1 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
Great data point, and I think it's the same failure mode: retrieval is good at "what's relevent", bad at "what's still in force". Validity is a property you have to assign, not something similarity can recover. Honest anser on where I am: the rule side is the stronger half. Each rule can carry a citation (the exact policy sentence it encodes), and the full_audit response records every rule evalution, the facts it tested, and the chunks retrieved. So the verdict and the clause behind it are captured at decision time, independent of any later retrieval. What I don't hae yet is first-class versioning. Saving rules replaces the set, and a decision doesn't record which version of the rules it ran against. Each decision gets an ID and the audit response is self-contained, but today there's no endpoint to fetch it back later, so the caller has to store it. For "why was this rejected six months ago", the answer has to come from that stored, not a re-run. The RAG side has exactly the gap you describe. Chunks carry a source but no effective date or supersession, so nothing stops it from retrieving an outdated passage to narrate a new decision. Your "latest defination wins" result pushes me towards the obvious fix: make validity deterministic too, rather than hoping ranking sorts it out. The plan: 1. Immutable rule-set versions, with a hash stamped on every decision. 2. Documents with effective_from/supperseded_by, filtered by the decisions built from that record and the citations of that version, never from live re-retrieval. Curious whether your 38/40 with full history failed on the same cases each time, or randomly? That would say a lot about whethe it's a context-length problem or a "which one wins" problem.