fix(hermes): make prefetch recall query-aware - #181
Conversation
|
Caution Review failedThe pull request is closed. ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (17)
📝 WalkthroughWalkthroughHermes prefetch now uses ChangesRanked context-pack recall
Estimated code review effort: 4 (Complex) | ~45 minutes Suggested reviewers: Sequence Diagram(s)sequenceDiagram
participant Client
participant HermesProvider
participant RecallGate
participant BrainContextPack
Client->>HermesProvider: prefetch turn with query
HermesProvider->>RecallGate: check turn_id and telemetry_host
HermesProvider->>BrainContextPack: request ranked context pack
BrainContextPack-->>HermesProvider: return items, warnings, receipt_id
HermesProvider-->>Client: return recalled content and receipt metadata
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
`brain_context_pack` accepted a `query`, but read it as a case-insensitive SUBSTRING match of the whole string against `topic + principle`. Handing that argument a natural-language turn `filter-miss`ed every candidate and returned an empty pack, which is why a query-aware Hermes prefetch looked like it had to leave the lane for raw `brain_search` - and lose the containment guard, the curated preference pool, the tombstone and supersession-tip filters, owner-scope delivery, the enforced token budget, and the server-issued receipt that only exist here. The seam is inside the lane, at the `filter-miss` branch. `query_mode: "ranked"` takes the query out of the filter and into the sort: the already-collected candidates are ORDERED by structural token overlap - the same deterministic, stopword-free, language-agnostic kernel `matchRatedDecisions` ranks with, so it needs no model and no embedding and works on an install that has never indexed a vector - and none is excluded. Relevance sits between session focus and density in the comparator, under tier: a bound focus is a standing operator-established target and still dominates, density is a static content heuristic an explicit per-call query outranks, and a peripheral page never outranks a core one. The budget, not a lexical accident, then decides what is dropped, and every other property of the lane is inherited untouched. An omitted `query_mode` is byte-identical to every release before this one, pinned by a test. The mode requires a query: the schema states the pairing as `dependentRequired` and the handler refuses a mode with nothing to read rather than accepting a silent no-op. The prompt-prefix segment carries the mode only when the caller named one, so an unnamed mode keeps the prefix hash it had. Also fixed in passing: the tool discarded `report.warnings`. The core computes injection-time tension warnings and the owner-scope observation, `brain_pre_compress_pack` has forwarded them all along, and the one surface that puts vault text in front of a model every turn was the one that could not say a memory it injected is contested. Absent when empty, so a warning-free pack stays byte-identical. `plugins/hermes/_schemas.py` is re-vendored for the changed `brain_context_pack` schema, regenerated from a live `o2b mcp` `tools/list` rather than hand-edited. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd
Prefetch's recall goes back to one lane, `brain_context_pack`, and now carries the turn on it: `query` plus `query_mode: "ranked"`, so the curated candidates are ordered by relevance to the prompt instead of being injected query-blind. The diagnosis behind the previous shape was right - a query-blind pack on every gated turn is a real defect - but raw `brain_search` bought query-awareness by leaving the only path that runs `guardBrainContextSnippet`, enumerates the preference directory rather than post-filtering a ranker, drops tombstoned and superseded pages, honours owner-scope delivery, enforces a token budget with named skip reasons, and issues an auditable receipt. Ranked mode puts the query inside that path, so none of it is traded away. What follows from the lane being singular: - The sample id is the server-issued `receipt_id` again, never a locally hashed `hermes-search-<sha256>`. An id with no receipt behind it makes every outcome the agent posts resolve `unresolved / sample_absent` forever, which quietly undid what itechmeat#179 shipped one release earlier. The metadata marker returns to `[O2B context-pack metadata]`. - `_search_sample_id`, `_search_text`, and `_recall_body` are deleted with the `hashlib` import and the prefetch sequence counter that only existed to feed the fabricated id. `_recall_body` could not work as designed: search `content` is a 600-char ellipsized window, so the frontmatter it hunted for was usually not in the string, and when a body line happened to start `principle:` it returned that line alone and threw the note away. A regression test pins a plain note through. - The local character budget is gone. `_PREFETCH_MAX_TOKENS` is a TOKEN budget the server enforces against the bodies it emits; spending it in Python at four bytes per token over-injected by two to four times on Cyrillic and CJK vaults, and cut the joined output mid-string with no marker. Pinned for both scripts. - Degradations name themselves. The pack's own `warnings` - injection- time tension warnings, the owner-scope observation - are logged instead of discarded, and a gated turn that recalls nothing says so rather than passing for a healthy injection. - `handle_tool_call`'s correlation defaults are scoped to `operation == "post"`. `host` and `session_id` are also read filters on `list`/`summary`, so defaulting them there silently narrowed an agent's query to this host's rows. Kept from the previous shape: the `brain_recall_gate` enrichment (`telemetry_host`, `session_id`, `turn_id`), and preference-first ordering, which the pack lane provides structurally by walking the preference directory. The three itechmeat#179 regression pins now assert which lane ran: `FakeBrainBridge` answers `{}` for an unregistered tool, so a pin that only reads the output could keep passing while the provider recalled through something else entirely. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd
A behaviour change to the recall lane plus a new `brain_context_pack` argument, so MINOR. `package.json` is the source of truth and `scripts/sync-version.ts` propagated it to the seven mirrored manifests; `--check` is clean. The entry credits @Yori940619 and itechmeat#181 for the diagnosis - query-blind prefetch recall - and for the preference-first idea, and states plainly that the mechanism was redesigned into the pack lane so the guard, the budget, owner scope, and the receipts are inherited rather than bypassed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd
solaitken
left a comment
There was a problem hiding this comment.
Reviewed as maintainer: the rework moves query-awareness into the context-pack lane (ranked mode), keeps the server receipt as the sample id, and restores the anti-drift gate honestly. Full suite 11708/0, Python suite 146 OK.
|
Thank you for this one too - the diagnosis was exactly right: prefetch recalled the same query-blind pack no matter what the user asked. We moved the mechanism into the context-pack lane (ranked query mode, shipped in v1.53.0 with your commit as the base of the branch) so the guard, budget, owner scope, and receipt come along for free. Your preference-first idea and the recall-gate enrichment survived as-is. Released: https://github.com/itechmeat/open-second-brain/releases/tag/v1.53.0 |
Summary
Verification
test_memory_provider.pytests passruff --select E9,Fpassesgit diff --checkpassesFollow-up to #179.
Summary by CodeRabbit
New Features
Bug Fixes