Skip to content

fix(hermes): make prefetch recall query-aware - #181

Merged
solaitken merged 4 commits into
itechmeat:mainfrom
Yori940619:fix/hermes-prefetch-query-aware
Aug 27, 2026
Merged

solaitken merged 4 commits into
itechmeat:mainfrom
Yori940619:fix/hermes-prefetch-query-aware

Conversation

@Yori940619

@Yori940619 Yori940619 commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • use query-aware brain_search during Hermes prefetch
  • prioritize confirmed Brain preferences while retaining ordinary recall
  • carry a deterministic non-reversible sample id for explicit outcome recording
  • keep legacy context-pack fallback only when the search surface is unavailable

Verification

  • 105/105 test_memory_provider.py tests pass
  • ruff --select E9,F passes
  • TypeScript pre-push typecheck passes
  • git diff --check passes

Follow-up to #179.

Summary by CodeRabbit

  • New Features

    • Added ranked context-pack queries that prioritize relevant results while retaining available context.
    • Context-pack responses now include applicable recall warnings.
    • Added query-mode validation and clearer handling for incomplete or unsupported requests.
    • Upgraded to version 1.53.0.
  • Bug Fixes

    • Improved recall traceability with server-issued receipt identifiers.
    • Corrected outcome reporting so session and host details are applied only when appropriate.

@coderabbitai

coderabbitai Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 9b9a2398-52e7-42d5-afca-004097a33300

📥 Commits

Reviewing files that changed from the base of the PR and between e7b353a and 4d2fab3.

📒 Files selected for processing (17)
  • .claude-plugin/plugin.json
  • .codex-plugin/plugin.json
  • CHANGELOG.md
  • docs/mcp.md
  • openclaw.plugin.json
  • package.json
  • plugin.yaml
  • plugins/codex/.codex-plugin/plugin.json
  • plugins/hermes/_schemas.py
  • plugins/hermes/plugin.yaml
  • plugins/hermes/provider.py
  • pyproject.toml
  • src/core/brain/context-pack.ts
  • src/mcp/brain/pack-tools.ts
  • tests/core/brain/context-pack-ranked-query.test.ts
  • tests/mcp/context-pack-tool.test.ts
  • tests/python/test_memory_provider.py

📝 Walkthrough

Walkthrough

Hermes prefetch now uses brain_context_pack as its only recall lane. Ranked queries retain and order candidates by token overlap. The flow propagates turn context, warnings, server receipt IDs, and operation-specific outcome correlation defaults.

Changes

Ranked context-pack recall

Layer / File(s) Summary
Context-pack ranked query contract and ordering
src/core/brain/context-pack.ts, tests/core/brain/context-pack-ranked-query.test.ts
Adds substring and ranked query modes. Ranked mode orders retained candidates by token overlap while preserving tier ordering.
Context-pack tool validation and warnings
src/mcp/brain/pack-tools.ts, tests/mcp/context-pack-tool.test.ts
Validates query_mode, requires query when a mode is supplied, forwards the mode to packContext, and returns non-empty warnings.
Hermes pack-based prefetch integration
plugins/hermes/provider.py, tests/python/test_memory_provider.py
Replaces search-based prefetch with one context-pack call. It propagates turn_id, uses ranked mode, logs warnings, preserves note bodies, applies the server token budget, and uses receipt_id for metadata.
Release metadata and behavior documentation
plugins/hermes/_schemas.py, docs/mcp.md, CHANGELOG.md, package.json, plugin.yaml, *.plugin.json, pyproject.toml
Documents the new query mode and warnings behavior. Version metadata changes to 1.53.0.

Estimated code review effort: 4 (Complex) | ~45 minutes

Suggested reviewers: solaitken, itechmeat

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant HermesProvider
  participant RecallGate
  participant BrainContextPack
  Client->>HermesProvider: prefetch turn with query
  HermesProvider->>RecallGate: check turn_id and telemetry_host
  HermesProvider->>BrainContextPack: request ranked context pack
  BrainContextPack-->>HermesProvider: return items, warnings, receipt_id
  HermesProvider-->>Client: return recalled content and receipt metadata
Loading
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 14 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: updating Hermes prefetch recall to use query-aware search.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

itechmeat and others added 3 commits August 27, 2026 02:53
`brain_context_pack` accepted a `query`, but read it as a
case-insensitive SUBSTRING match of the whole string against
`topic + principle`. Handing that argument a natural-language turn
`filter-miss`ed every candidate and returned an empty pack, which is why
a query-aware Hermes prefetch looked like it had to leave the lane for
raw `brain_search` - and lose the containment guard, the curated
preference pool, the tombstone and supersession-tip filters, owner-scope
delivery, the enforced token budget, and the server-issued receipt that
only exist here.

The seam is inside the lane, at the `filter-miss` branch. `query_mode:
"ranked"` takes the query out of the filter and into the sort: the
already-collected candidates are ORDERED by structural token overlap -
the same deterministic, stopword-free, language-agnostic kernel
`matchRatedDecisions` ranks with, so it needs no model and no embedding
and works on an install that has never indexed a vector - and none is
excluded. Relevance sits between session focus and density in the
comparator, under tier: a bound focus is a standing operator-established
target and still dominates, density is a static content heuristic an
explicit per-call query outranks, and a peripheral page never outranks a
core one. The budget, not a lexical accident, then decides what is
dropped, and every other property of the lane is inherited untouched.

An omitted `query_mode` is byte-identical to every release before this
one, pinned by a test. The mode requires a query: the schema states the
pairing as `dependentRequired` and the handler refuses a mode with
nothing to read rather than accepting a silent no-op. The prompt-prefix
segment carries the mode only when the caller named one, so an unnamed
mode keeps the prefix hash it had.

Also fixed in passing: the tool discarded `report.warnings`. The core
computes injection-time tension warnings and the owner-scope
observation, `brain_pre_compress_pack` has forwarded them all along, and
the one surface that puts vault text in front of a model every turn was
the one that could not say a memory it injected is contested. Absent
when empty, so a warning-free pack stays byte-identical.

`plugins/hermes/_schemas.py` is re-vendored for the changed
`brain_context_pack` schema, regenerated from a live `o2b mcp`
`tools/list` rather than hand-edited.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd
Prefetch's recall goes back to one lane, `brain_context_pack`, and now
carries the turn on it: `query` plus `query_mode: "ranked"`, so the
curated candidates are ordered by relevance to the prompt instead of
being injected query-blind. The diagnosis behind the previous shape was
right - a query-blind pack on every gated turn is a real defect - but
raw `brain_search` bought query-awareness by leaving the only path that
runs `guardBrainContextSnippet`, enumerates the preference directory
rather than post-filtering a ranker, drops tombstoned and superseded
pages, honours owner-scope delivery, enforces a token budget with named
skip reasons, and issues an auditable receipt. Ranked mode puts the
query inside that path, so none of it is traded away.

What follows from the lane being singular:

- The sample id is the server-issued `receipt_id` again, never a locally
  hashed `hermes-search-<sha256>`. An id with no receipt behind it makes
  every outcome the agent posts resolve `unresolved / sample_absent`
  forever, which quietly undid what itechmeat#179 shipped one release earlier.
  The metadata marker returns to `[O2B context-pack metadata]`.
- `_search_sample_id`, `_search_text`, and `_recall_body` are deleted
  with the `hashlib` import and the prefetch sequence counter that only
  existed to feed the fabricated id. `_recall_body` could not work as
  designed: search `content` is a 600-char ellipsized window, so the
  frontmatter it hunted for was usually not in the string, and when a
  body line happened to start `principle:` it returned that line alone
  and threw the note away. A regression test pins a plain note through.
- The local character budget is gone. `_PREFETCH_MAX_TOKENS` is a TOKEN
  budget the server enforces against the bodies it emits; spending it in
  Python at four bytes per token over-injected by two to four times on
  Cyrillic and CJK vaults, and cut the joined output mid-string with no
  marker. Pinned for both scripts.
- Degradations name themselves. The pack's own `warnings` - injection-
  time tension warnings, the owner-scope observation - are logged
  instead of discarded, and a gated turn that recalls nothing says so
  rather than passing for a healthy injection.
- `handle_tool_call`'s correlation defaults are scoped to `operation ==
  "post"`. `host` and `session_id` are also read filters on
  `list`/`summary`, so defaulting them there silently narrowed an
  agent's query to this host's rows.

Kept from the previous shape: the `brain_recall_gate` enrichment
(`telemetry_host`, `session_id`, `turn_id`), and preference-first
ordering, which the pack lane provides structurally by walking the
preference directory. The three itechmeat#179 regression pins now assert which
lane ran: `FakeBrainBridge` answers `{}` for an unregistered tool, so a
pin that only reads the output could keep passing while the provider
recalled through something else entirely.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd
A behaviour change to the recall lane plus a new `brain_context_pack`
argument, so MINOR. `package.json` is the source of truth and
`scripts/sync-version.ts` propagated it to the seven mirrored manifests;
`--check` is clean.

The entry credits @Yori940619 and itechmeat#181 for the diagnosis - query-blind
prefetch recall - and for the preference-first idea, and states plainly
that the mechanism was redesigned into the pack lane so the guard, the
budget, owner scope, and the receipts are inherited rather than
bypassed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017dYpaSgWTK5L6o6AzcLcBd

@solaitken solaitken left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed as maintainer: the rework moves query-awareness into the context-pack lane (ranked mode), keeps the server receipt as the sample id, and restores the anti-drift gate honestly. Full suite 11708/0, Python suite 146 OK.

@solaitken
solaitken merged commit c2e13f9 into itechmeat:main Aug 27, 2026
1 of 2 checks passed
@solaitken

Copy link
Copy Markdown
Collaborator

Thank you for this one too - the diagnosis was exactly right: prefetch recalled the same query-blind pack no matter what the user asked. We moved the mechanism into the context-pack lane (ranked query mode, shipped in v1.53.0 with your commit as the base of the branch) so the guard, budget, owner scope, and receipt come along for free. Your preference-first idea and the recall-gate enrichment survived as-is. Released: https://github.com/itechmeat/open-second-brain/releases/tag/v1.53.0

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants