Skip to content

feat: add claude-vault-capture plugin - #16

Open
lhoupert wants to merge 23 commits into
mainfrom
feat/claude-vault-capture-plugin
Open

lhoupert wants to merge 23 commits into
mainfrom
feat/claude-vault-capture-plugin

Conversation

@lhoupert

@lhoupert lhoupert commented Oct 2, 2026 •

Copy link
Copy Markdown

Adds the claude-vault-capture plugin to the marketplace under skills/claude-vault-capture/. On SessionEnd it scrubs secrets from the transcript, has Claude Sonnet 5.5 curate one note (decision, runbook, gotcha or spec, or nothing for low-signal sessions) and writes it to <vault>/Inbox/auto/; it also ships the /vault-save skill. This supersedes #5 together with the review fixes stacked on it in alukach#1, so it can be reviewed and merged from this repo. 290 tests pass; ruff and shellcheck are clean.

For reviewers

Unlike the other skills here, this one runs by itself: each qualifying session costs a model call (~$0.11 median, or billed to a Claude Pro/Max plan). Sonnet 5.5 rejects thinking: disabled, so curation uses between_tools at effort low; in subscription mode that relies on two undocumented CLI env vars, CLAUDE_CODE_EXTRA_BODY and CLAUDE_CODE_DISABLE_THINKING, to re-check when bumping the claude-agent-sdk pin.

Author attestation

  • I am a human, these are my changes, and I have reviewed and understood every change and can explain why each is correct.

AI-assisted: the fixes came from multi-agent reviews, with each finding verified against the code; reviewed via the diff and the test suite.

🤖 Generated with Claude Code

lhoupert and others added 23 commits June 6, 2026 23:56
Repackage claude-vault-capture as a Claude Code marketplace plugin under
skills/claude-vault-capture/. It auto-captures Claude Code sessions into an
Obsidian vault on SessionEnd (scrub secrets → summarize with Sonnet → write a
curated note to Inbox/auto/) and ships the /vault-save skill for on-demand
exports to claude-docs/.

Adapted from https://github.com/lhoupert/claude-vault-capture (MIT, © Loïc
Houpert). Changes made for marketplace distribution:

- plugin.json with userConfig (vault path + sensitive API/OAuth tokens) replaces
  install.sh's interactive prompt and capture.env.
- hooks/hooks.json registers the SessionEnd hook via ${CLAUDE_PLUGIN_ROOT},
  replacing the manual settings.json edit.
- vault-save becomes skills/vault-save/SKILL.md; its auto-trigger lives in the
  skill description, so the old global CLAUDE.md injection is gone.
- curate.py carries PEP 723 inline deps and the hook launches it with `uv run`
  (subscription mode adds claude-agent-sdk via `uv run --with`), so there is no
  separate `uv sync` step. A pre-built .venv is still preferred when present, so
  standalone/dev use and the existing test suite are unchanged.
- Runtime state (dedup index, logs) moves to ${CLAUDE_PLUGIN_DATA} via
  CAPTURE_STATE_DIR / SCRUB_FAILURES_PATH, surviving plugin updates.
- LICENSE and a provenance note are preserved; the pytest suite travels along.

Register the plugin in marketplace.json and list it in the README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
tests/test_timeout.py imports anthropic/httpx at module level, but the PR
moved the runtime dep into PEP 723 metadata only, so 'uv sync && uv run
pytest' (the documented flow) failed at collection. Pinning the same
version as curate.py's inline block also unbreaks the standalone .venv
hook path, which ran curate.py without anthropic installed. uv.lock is
committed for reproducibility (dev deps are exact-pinned).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-mode tests

The uv-run branch expanded an empty WITH array under 'set -u', a fatal
'unbound variable' on bash < 4.4 — macOS ships 3.2, so every capture from
a marketplace install (no .venv, API-key mode) died before launching the
worker. Build the RUN array incrementally instead.

Fallback credential files (~/.claude_vault_token, ~/.claude_vault_oauth_token)
are now refused unless owner-only (mode *00), with a CAPTURE_TOKEN_FILE_PERMS
marker logged; README documents umask/chmod 600.

New tests/test_plugin_mode.py runs the hook copy under /bin/bash with a stub
uv and no .venv — the exact configuration the old suite never exercised (its
harness always built a .venv, so the uv branch was dead code in CI). Covers
CLAUDE_PLUGIN_OPTION_* mapping, CLAUDE_PLUGIN_DATA state dir, the
subscription --with arm, env-precedence, and token-file perms; 6 of 7 fail
against the unfixed script.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ject JSON path

render_frontmatter interpolated scalars unquoted, so the house-style
'Decision: …' titles produced frontmatter yaml.safe_load (and Obsidian)
reject; every string scalar is now json.dumps-quoted, which is valid,
inert YAML. Model-supplied tags/type were the one output-to-disk channel
that bypassed sanitization: tags are now scrubbed and slug-coerced
([a-z0-9-], max 40 chars, 10 tags), type is allowlisted. Valid-but-non-
object model JSON (quoted string, list, number) now routes down the
malformed_json path with usage attached instead of an AttributeError
that dropped token accounting.

conftest.parse_frontmatter now uses yaml.safe_load — the old partition-
on-colon parser is what masked the invalid-YAML bug; the suite now fails
exactly when Obsidian would. New round-trip and sanitizer tests cover
colon titles, hostile tags/type, and the non-object JSON cases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
env_var was ^-anchored, so 'export KEY=…' lines and assignments on a
message's first line (which curate.py prefixes with '[USER]: ') passed
through unredacted, and quoted values leaked everything past the first
space. The rule now matches after whitespace/quote/paren/colon, accepts
an export/set/env prefix, and consumes quoted values whole; lowercase
keys still never match. token_prefix's sk- charset stopped at the first
dash, redacting only the public 'sk-proj' prefix of modern OpenAI keys
while the key body stayed in clear; it now consumes _/- and covers
github_pat_ fine-grained PATs. A new aws_secret rule catches the
lowercase ~/.aws/credentials forms (aws_secret_access_key,
aws_session_token) the uppercase-only env_var rule missed.

The with-secrets fixture no longer stores a real-format PEM block:
gitleaks/push-protection match the BEGIN/END markers themselves, so
load_fixture() assembles them at load time instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…stale references

The repo README's manual-install/symlink instructions silently fail for
this entry (nested SKILL.md, hook only registers under /plugin install),
so the entry is now marked code-bearing/plugin-only with a caveat in the
manual-install section, and the intro acknowledges the marketplace hosts
both plain-language skills and code-bearing plugins. plugin.json's
repository now points at this marketplace (where the packaged code
lives); homepage keeps the upstream attribution link. curate.py comments
no longer reference install.sh flows or files that don't ship here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Brings the vendored copy up to developmentseed/claude-vault-capture@c03ccda
(the repo moved from lhoupert/; the old URL now redirects). Ports five
upstream commits that landed after the v0.2.0 the plugin was vendored from:
configurable CAPTURE_TIMEOUT_SECONDS, tool activity surfaced to the curator,
the W29/W30 no-capture fixes (JSON salvage from prose-wrapped replies,
subscription error_max_turns, unknown-usage propagation), and the
simplification refactor. Test count goes 189 -> 223.

Three merge conflicts resolved rather than taken wholesale:

- scrub_rules.py: upstream replaced the {sentinel, replace_value_only,
  replace_group} rule schema with a single 'replacement' regex template.
  The review fixes' rules were re-expressed in it — env_var keeps its widened
  boundary (export/set/env prefixes, first-line '[USER]: ' spans, quoted
  values) with the pre-value span captured as \g<pre>, and aws_secret gains a
  'replacement'. Without this the new engine's _compile_rules would have
  caught the missing key and SILENTLY SKIPPED the AWS rule.
- curate.py: upstream's JSON salvage and the review's non-dict handling are
  complementary; merged so a hard parse failure and valid-JSON-of-the-wrong-
  type both attempt salvage before raising malformed_json with usage attached.
  A literal null payload now retries like a raw "null" instead of crashing.
- session-end-capture.sh: upstream's timestamp hoist kept; its git-based
  deploy-drift guard scoped to standalone mode, since an installed plugin
  can't drift that way ( resolves to the marketplace clone) and it would
  spend forks on the close path to log an irrelevant SHA. Verified it still
  fires standalone and stays quiet under CLAUDE_PLUGIN_ROOT.

Version 0.2.0 -> 0.3.0 (upstream's own tag is still v0.2.0; these commits are
unreleased there). New knobs documented in the plugin README. Provenance URLs
follow the repo move; the MIT copyright (© Loïc Houpert) is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n mode

Two gaps a cross-check of the port surfaced:

The headline ported feature was unreachable. CAPTURE_TIMEOUT_SECONDS is the
fix for the upstream failure that motivated it (subscription-mode timeouts on
large transcripts) — and surfacing tool activity makes every transcript
bigger — but plugin config only reaches curate.py through userConfig, which
had no entry for it. Adds timeout_seconds and maps it in the hook. The
mapping validates: curate.py reads the var with a *string* default, so an
empty or non-numeric value would be a ValueError at import — a silent
no-capture — rather than a fallback to 30. Junk is logged and dropped.

Deploy identity is now logged in both modes instead of being suppressed for
plugin installs. Scoping the git guard to standalone (previous commit) threw
away the W29 lesson — you could no longer tell from hooks.log which version
produced a capture, in the mode where that is hardest to reconstruct. Plugin
installs now log the plugin version from plugin.json; git stays standalone-only
because "git -C" walks up and would describe whatever repo encloses the plugin,
reporting false STALE_DEPLOYs against an unrelated origin/main.

Also hoists the mkdir/NOW pair above all config handling so every logging path
has its destination, and documents the new knobs plus the tool-activity change
in the plugin README.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…urates it

Brings the vendored copy to developmentseed/claude-vault-capture@7577892, one
commit past the previous port. Upstream's root-cause fix for every
malformed_json in its archive: the prompt was an unterminated chat log, so the
model wrote the conversation's next turn instead of an artifact. Ports all four
layers — the _TRANSCRIPT_TAIL terminator appended by _invoke_model (so both
transports get it), _salvage_artifact's offset-scanning replacement for the old
outermost-brace window, the shared PATH_A_RESAMPLES budget that now retries an
unparseable reply once, and the held `unparsed` error so a trailing null can't
relabel a malformed reply in log.md. Suite 227 -> 237.

One conflict, in the same block the last port resolved. Upstream rewrote
_call_path_a's failure handling; the review's non-dict guard is still needed
because upstream does not have it — a reply that parses to "null", [1,2] or 42
never enters the except branch, so it reaches data.update() and dies as an
AttributeError, losing the usage accounting. Rather than keep it as a separate
raise-immediately path, the wrong-type class is now folded into upstream's
recovery route: salvage, then one resample, then malformed_json with both
attempts billed. That is upstream's own argument for retrying prose replies —
a wrong-shaped reply is non-deterministic too — and it keeps one failure path
instead of two. A literal parsed null still resamples as a null, matching
upstream. Verified directly: all three wrong-type shapes raise JSONDecodeError
with usage after 2 attempts, no AttributeError.

_salvage_object (added by the previous port) is deleted — upstream's
_salvage_artifact supersedes it and is strictly better: it scans every brace
offset instead of taking a first-to-last window that prose could widen, and it
requires the title/type/body contract, which closes a junk-write hole where an
unshaped dict would land in Inbox/ as an empty "untitled" note.

Version stays 0.3.0: that release is still unmerged, so this folds into it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ber gaps, unlogged losses

A 26-agent review with adversarial verification confirmed 21 findings against
this branch (0 refuted). The security-critical one and the defects that lose
data or corrupt notes are fixed here; each fix was reproduced before and
verified after. Suite 237 -> 255.

CRITICAL — secrets survived the scrubber via truncation. render_transcript caps
tool blocks (_BASH_CMD_CAP, _EDIT_DIFF_CAP, _ERROR_CAP, CAPTURE_SUCCESS_HEAD_CHARS)
and the caps ran BEFORE the scrub, so a cut secret matched no rule: private_key
needs its closing -----END----- and AIza…{35}/AKIA…{16} need their full body.
A PEM pasted into a failing command reached the model and the note in clear
(reproduced on [OUT], [TOOL] Write, and [ERROR]). Scrubbing now happens per block
before each cap, via _scrub_cap; render_transcript takes an optional dict to
collect those counts so `redactions:` still totals them.

Scrubber, two rules:
- env_var never matched a bare keyword name. A mandatory leading [A-Z] sat
  outside the alternation, so the keyword could never start at offset 0 —
  PASSWORD=, TOKEN=, SECRET=, PASS=, KEY_ID= all passed through in clear while
  the frontmatter reported 0 redactions. These are the commonest .env forms.
- the widened sk- alternative corrupted ordinary prose. Adding `-` to the body
  charset let it run across segments: "risk-averse-approach" became
  "ri<redacted:token_prefix>" in the archived note. Now word-boundary anchored,
  no `-` in the body, 16-char minimum. Both directions are now tested.

Unlogged session losses. Only the salvage path enforced the title/type/body
contract, so a directly-parsed reply with a non-string body (or title, or a
non-list tags/source_links) raised inside the write path — after the model call
was paid for and before append_log ran, leaving no log.md row at all, which is
the failure class the module comments call out as invisible to the no-capture
alarm. The contract is now shared by both paths via _is_artifact, and the
containers are coerced. Verified: all four shapes now log rather than crash.

Hook script:
- stat now follows symlinks (-L) and probes GNU -c before BSD -f. A symlinked
  token was refused with a chmod hint that could not fix it, and on Linux the
  BSD-first probe wrote a multi-line stat blob into hooks.log as the mode.
- the plugin.json version sed gets `|| true`: under set -euo pipefail a missing
  file made it abort the hook, skipping the capture over a cosmetic log line.
- an installed plugin now always uses `uv run`, never a .venv. Running `uv sync`
  inside an installed plugin — which the README's own Tests section invited —
  left a dev-only venv that shadowed the PEP 723 deps and silently broke
  subscription mode. README now says to run tests from a clone.
- a missing python3 is logged (CAPTURE_NO_PYTHON3) instead of exiting silently,
  and python3 is listed under Prerequisites.

Tests: the plugin-mode stub now writes atomically (rename) so the poll cannot
read a half-written record, and its timeout probe distinguishes unset from
set-and-empty — that assertion previously could not fail for the reason it
existed. One upstream assertion pinned a specific sentinel label on a value two
rules both claim; it now asserts that nothing leaks, which is the real contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…fig can migrate

Found while migrating a real standalone install: of the five settings a typical
capture.env carries, only three had a plugin equivalent. CAPTURE_EXCLUDED_COMMANDS
and CAPTURE_MAX_EST_TOKENS were readable by curate.py but unreachable through
plugin config, so switching to the plugin silently dropped them — the excluded
workflow commands (/daily-devlog, /weekly-recap) would start being captured into
the vault, and a raised ceiling (120000) would fall back to 50000, skipping the
long sessions it was raised for. Neither failure is visible at switch time.

Adds both as userConfig and maps them in the hook. The numeric validation added
for timeout_seconds is factored into _export_int_setting and now covers
max_est_tokens too: curate.py reads both with string defaults, so an empty or
non-numeric value is a ValueError at import — a silent no-capture — rather than a
fallback. The rejection marker is now CAPTURE_BAD_SETTING (was CAPTURE_BAD_TIMEOUT)
since it covers more than the timeout. excluded_commands needs no validation:
curate.py splits on "," and drops empties.

Verified by running the hook with all five values set as CLAUDE_PLUGIN_OPTION_*
and confirming each reaches the worker unchanged. Suite 255 -> 258.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Swaps MODEL_A to claude-sonnet-5, which covers both transports — the API-key
path and the subscription path both pass this one constant through.

Two changes the swap requires, neither of which the model ID alone gives you:

Thinking is now explicitly disabled on the API-key call. Sonnet 5 runs adaptive
thinking whenever `thinking` is omitted, where Sonnet 4.6 ran without it, and
max_tokens caps thinking and the reply together. Left unset, the curation call
would spend the artifact's budget on reasoning and truncate the JSON — arriving
as malformed_json and a lost capture rather than an error. Curation is a
one-shot extraction against a hard timeout, so thinking buys little here; the
comment records adaptive-at-low-effort as the tunable if capture quality ever
needs it.

MAX_TOKENS_A goes 2000 -> 3000. Sonnet 5's tokenizer emits roughly 30% more
tokens for the same text, so artifacts that previously fit would otherwise
truncate at the same ceiling.

Nothing else in the request had to change: the call passes no temperature/top_p/
top_k (rejected on Sonnet 5) and no assistant prefill (also rejected). List
pricing is unchanged at $3/$15 per MTok, so the cost estimator stays correct —
noted inline that introductory rates make it over-report through 2026-08-31.

The chars/4 token heuristic now runs ~30% low against the new tokenizer, so the
CAPTURE_MAX_EST_TOKENS ceiling admits somewhat larger transcripts than the
number implies. Left deliberately as-is and documented: tightening it would
silently start skipping sessions that are captured today.

New tests pin the request shape — thinking explicitly disabled, no removed
sampling params, and output headroom for the new tokenizer. Verified the first
fails when the thinking line is removed. Suite 258 -> 261.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A newcomer-focused documentation review (agents, each finding verified
against the vendored code) returned 33 confirmed issues. The ones that
actually broke a first install:

- The Install section said Claude Code prompts for a vault path "and,
  optionally, an API key. That's it." Credentials are not optional:
  curate.py exits before the model call with no note AND no log row, so
  the only trace is one line in hooks.log. Pro/Max users — who typically
  have no ANTHROPIC_API_KEY anywhere — were the most likely victims.
  The same "leave blank if you set it in your environment" advice was in
  the install prompt itself, which is worse, because Claude Code strips
  the environment it spawns hooks with.
- "Verify it's working" could not distinguish working from broken. Both
  its checks pass on a dead install: SESSION_END_RECEIVED is written
  before the worker starts, and an empty Inbox/auto/ is the correct
  result for a low-signal session. It now leads with log.md (the only
  file that answers the question), spells out how to resolve
  ${CLAUDE_PLUGIN_DATA} from a shell, and adds two tables mapping every
  skip_reason and every hooks.log marker to what to do.
- Cost and privacy were unstated. A plugin that makes an unattended paid
  call at every session close now says so up front, with the real figure
  (~$0.12 median, measured over 287 captures), and says plainly that the
  transcript — including tool output — is sent to the API and that regex
  scrubbing is not a guarantee.

Also: the curation prompt advertised a fifth artifact type,
devlog-snippet, that sanitize_type() silently rewrites to "decision", so
those notes were mislabeled in the vault — prompt aligned to the four
types the code, README and manifest all agree on. /vault-save now states
that it bypasses the scrubber (it writes directly; no Python in that
path). Install prompt reordered to ask for credentials before tuning
knobs, with every default and consequence stated per field. Added a
"Turning it off" section. Fixed the repo README's symlink fallback,
which produced a skill with an unexpanded ${user_config.vault_dir}, and
flagged there that this plugin is always-on unlike its neighbours. Stale
claude-sonnet-4-6 example updated after the Sonnet 5 move.

Suite: 261 passed, 1 skipped.
…pricing

A json.dump without ensure_ascii=False had rewritten marketplace.json,
plugin.json and mock-responses.json: literal characters became \u escapes
(including the unrelated devseed-poster entry) and inline arrays were
exploded one item per line. Formatting is restored from the PR base;
content is unchanged (verified by json.loads equality).

The ~$0.12 median was measured on Sonnet 4.6 rows priced at $3/$15. Sonnet
5 is $2/$10 (the introductory rate became standard; the planned September
increase was cancelled). Re-measured on 160 post-cutover Sonnet 5 captures:
~$0.13 median. Numeric config fields now say "positive whole number".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ent bloat

- cost estimate: $2/$10 per MTok (was Sonnet 4.6's $3/$15, a 1.5x
  over-report in log.md and note frontmatter); pinned by a test.
- scrub: vendor-prefixed sk- keys (sk-proj-, sk-svcacct-, sk-admin-) may
  contain '-' anywhere in the body; the old rule left everything after the
  first in-body dash in clear. Bare sk- keeps the dash-free body that
  protects hyphenated prose.
- hook: timeout_seconds / max_est_tokens reject 0 (every call would time
  out, or every session skip) as well as non-numeric values.
- render: MultiEdit falls through to the generic tool renderer; it carries
  an `edits` list, so the Edit branch rendered an empty diff.
- _call_path_a: same outcomes with one `malformed` variable instead of
  bad/unparsed/data juggling.
- Comments and docstrings trimmed to the why, dropping incident history,
  week numbers and session IDs; a dangling "SPEC §7" reference removed;
  test PEM markers assembled at runtime; ruff format applied.

Suite: 265 passed, 1 deselected (live).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ments

- make_slug: same output in 8 lines instead of 23.
- session-end-capture.sh: $NOW is always set before _read_secret_file
  runs, so drop its date fallback; three verbose comments cut to one line.
- tests: the identical _wait_for in test_plugin_mode and
  test_session_end_hook moves to conftest.wait_for.

Suite: 265 passed, 1 deselected (live).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Subscription mode (the production path):
- setting_sources=[]: the curation call no longer loads user settings, so
  this plugin's own SessionEnd hook stops firing on the curation session
  (the source of the phantom `threshold` rows in log.md). Thinking is
  disabled and output is capped at MAX_TOKENS_A, matching API mode.
- Return at the ResultMessage instead of waiting for CLI teardown inside
  the timeout; raise on AssistantMessage.error and on an error result with
  no reply, instead of parsing "Invalid API key" text as a reply.

Scrubber:
- basic_auth_url matches any scheme and an empty user (postgres://,
  redis://:pw@, mongodb+srv://); sk-or-v1-/sk-lf- keys; AWS JSON
  "SecretAccessKey"/"SessionToken".
- Generic tool inputs are scrubbed per string before json.dumps, whose
  escaped newlines hid KEY=value lines (MCP inputs, MultiEdit edits).
- Fewer prose false positives: `X == 3000` comparisons, "highs_and_lows"
  (GitHub tokens are 36+ chars), "Bearer token".

Pipeline:
- API stop_reason refusal/max_tokens -> skip `refusal`/`truncated` with
  usage, not resampled (was IndexError / two paid identical calls).
- Every exception from _call_path_a carries usage across attempts.
- Salvage searches the whole reply, not the first fenced block; a first
  line of `null` is a null; empty title/body is not an artifact; types
  are lowercased before the allowlist; a model `_null` key is ignored.
- Timeouts and errors are no longer indexed, so a resumed session retries.
- A failed note write, or any error in main(), still logs a log.md row.
- excluded_commands matches Claude Code's <command-name> markup.
- derive_project names the main repo from a linked worktree, even after
  the worktree is removed.
- Numeric settings: junk or empty values fall back to defaults.
- Dead code removed: HOOKS_LOG, _STATE_LOCK (flock suffices),
  LOG_REQUIRED_KEYS (moved into its test).

Hook: one python3 parse instead of three (a JSON-null path no longer
becomes "None"); `~` in vault_dir expanded, relative paths refused
(they resolved inside the session's project); empty int settings unset.

/vault-save: title/summary/project/tags written as quoted strings (bare
scalars broke YAML on ": "), project from --git-common-dir, empty slug
falls back to `untitled`. README: new skip reasons and markers; uninstall
deletes ${CLAUDE_PLUGIN_DATA} unless --keep-data; Templater warning.

Suite: 304 passed, 1 deselected (live). Subscription changes verified
against claude-agent-sdk 0.2.89's option handling and its bundled CLI's
flag parsing, not with a live call.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e test

- _invoke_via_subscription returns at the ResultMessage, which left the
  CLI to exit during asyncio.run's shutdown ("Loop ... that handles pid
  is closed" on every capture, i.e. hooks.log noise). The stream is now
  closed explicitly inside the loop, after the reply and outside the
  timeout, so a slow teardown still can't discard a finished reply.
- test_live_e2e crashed before any model call: the autouse temp_vault
  fixture already creates vault/Inbox/auto. mkdir(exist_ok=True).

Verified live (CAPTURE_LIVE_TESTS=1, real subscription call): passes, no
asyncio warning, no orphaned CLI, and no new SESSION_END_RECEIVED line or
log.md row from the installed plugin, so setting_sources=[] stops the
plugin's own hook firing on the curation session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rgrft2uNyq6DfUkcbE5xMD
A set `version` pins installs until it changes, so installs taken from
this branch at 0.3.0 would never pick up the fixes since 29 July.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rgrft2uNyq6DfUkcbE5xMD
Sonnet 5.5 rejects thinking {"type": "disabled"} with a 400; its
thinking-off setting is {"type": "between_tools"}, run at effort low.

- API mode sends between_tools plus output_config effort "low".
- Subscription mode: the CLI has no --thinking between_tools, so it gets
  CLAUDE_CODE_EXTRA_BODY (the thinking override) and
  CLAUDE_CODE_DISABLE_THINKING (drops the clear_thinking context edit
  the API rejects alongside between_tools), plus effort "low". Both env
  vars are undocumented CLI internals: re-check them on SDK pin bumps.
- claude-agent-sdk 0.2.89 -> 0.2.161. The old bundled CLI did not honour
  --thinking disabled either: 10 of 22 Sonnet 5 curation calls thought,
  3 hit the output cap.

A/B on 20 real transcripts, subscription mode, 0.3.1 vs this change:
20/20 notes vs 19/20; median 11 s vs 20 s (p90 15 s vs 41 s); median
cost $0.106 vs $0.126 per capture; blind judges preferred Sonnet 5.5
17-1 on usefulness and 9-0 on accuracy (8 distortions, all Sonnet 5).
Live check: the curation transcript has 0 thinking blocks.

Suite: 305 passed, 1 deselected (live, passed separately).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rgrft2uNyq6DfUkcbE5xMD
- Drop standalone mode: no capture.env, no .venv interpreter, no git
  STALE_DEPLOY check. The plugin is installed only via /plugin and the
  upstream repo is being retired, so that path had no users. The deploy
  line now always logs the plugin version. test_session_end_hook.py is
  removed; its happy-path, guard, unconfigured-vault and OAuth-token-file
  cases move to test_plugin_mode.py.
- umask 077 in the hook: hooks.log, log.md, the index and new notes hold
  transcript excerpts and were created world-readable (0644).
- API mode keeps the SDK's retries (max_retries=2): a single 529/5xx lost
  the capture for good. TIMEOUT_SECONDS now bounds each attempt.
- README: the threshold row describes what the gate actually counts. The
  documented "3 real turns / 1500 chars of your own text" would have
  skipped 17 of 116 sessions that produced notes (measured on the live
  log), so the docs change rather than the code.

Suite: 302 passed, 1 deselected (live).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rgrft2uNyq6DfUkcbE5xMD
…pend

- sanitize_summary had one caller: sanitize_title now takes max_len.
- make_filename was one f-string with one caller: inlined; the e2e test
  already checks the filename format.
- append_log and _append_index share _locked_append; the explicit
  flush()/LOCK_UN before close is dropped (close flushes, then releases
  the lock).
- _write_artifact passes render_frontmatter's keyword args through as
  **fm instead of re-declaring ten of them.

Drops the 13 tests that duplicated TestSanitizeTitle or tested
make_filename. Suite: 290 passed, 1 skipped (live).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@lhoupert
lhoupert requested a review from dannybauman October 2, 2026 09:51
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Claude finished @lhoupert's task in 36s —— View job


✅ No blocking issues — safe to merge.

I read hooks/session-end-capture.sh, hooks/curate.py, scrub.py, scrub_rules.py, hooks.json and plugin.json. I did not run the test suite, and I did not read the tests or SKILL.md.

Scrub-before-cap ordering, the model-output allowlists (type, tags, JSON-quoted frontmatter), the 600 check on token files, and the skip/index/log bookkeeping all look right. A few non-blocking points:

  • session-end-capture.sh:106-114, subscription mode may bill the API key. In subscription mode an ANTHROPIC_API_KEY that is already in the environment is passed through to the worker and the CLI. The CLI prefers that key over CLAUDE_CODE_OAUTH_TOKEN, so the call could be billed to the API instead of the plan. Adding unset ANTHROPIC_API_KEY in that branch makes the "bill to subscription" promise hold:
    if [[ "${CAPTURE_USE_SUBSCRIPTION:-}" == "1" ]]; then
        unset ANTHROPIC_API_KEY
  • curate.py:447-458, "hard wall" is not one. timeout=TIMEOUT_SECONDS applies per attempt. With max_retries=2 and one resample, the worst case is about 6 attempts, roughly 180 s at the default. Nothing is blocked, since the worker is backgrounded, but the timeout_seconds description ("Hard wall … on a single model call") undersells it.
  • curate.py:854 and :934, dedup race. The session is only recorded in the index after the model call. Two SessionEnd events for the same session within the call window (a quick resume or /clear) will both curate and both pay. The window is small, and you could write a "pending" index row before the call.
  • curate.py:892-900, redactions from model output are dropped. Redactions found while scrubbing the model's title, body and tags are discarded (_), so the frontmatter redactions: counts understate them. Worth knowing if you use those counts as a leak signal.
  • curate.py:678, source_links are written verbatim as markdown list items. A prompt-injected link can add arbitrary markdown or links to the note. The impact is low, but you could restrict it to http(s):// strings.

Simplify (ponytail)

  • curate.py *_a names, path_a, MODEL_A and PATH_A_RESAMPLES: the A/B experiment is retired. Rename to model, path and RESAMPLES, and keep the log schema keys only if readers need them.
  • session-end-capture.sh:39-59 and curate.py:54 _env_int: the same positive-int validation exists in both. Keep one. _env_int already falls back to the default, so the shell version can just export the raw value.
  • curate.py:1014 _log_best_effort and _log_error: these are two small wrappers around stderr plus the log. Inline them into one try/except in main.
  • scrub_rules.py is its own module for a plain list of dicts. Folding RULES into scrub.py removes one import and one file.
  • curate.py:108-117 and :132-150 _strip_fences plus _salvage_artifact: _salvage_artifact already finds the last artifact-shaped object, fenced or not. Dropping _strip_fences leaves one JSON-extraction path.
  • STATE_DIR and SCRUB_FAILURES_PATH fall back to eval/state. That fallback only matters for manual runs, and the plugin always sets CAPTURE_STATE_DIR.

💰 Estimated review cost: $0.27 · 0m35s · 6 turns

@dannybauman

Copy link
Copy Markdown
Member

thanks @lhoupert for bringing this over. I read through the hook, curate.py and the scrubber, a few things:

  • the scrubber catches the formats in its rules (private keys, known token prefixes, JWTs, export NAME=value, bearer tokens, passwords in URLs), but I tried a few common ones that aren't covered and they came through as is: a key in JSON like "OPENAI_KEY": "..." (the shape of .mcp.json and settings env blocks), password: ..., Authorization: Basic ..., curl -u user:pass, and Stripe, GitLab, HuggingFace and npm tokens. the model's note goes through the same rules, so if it repeats one of those, it could end up in the vault. could the rules cover more of these? or are there other ways or libraries to do this more comprehensively?
  • on the bot's billing point, the hooks docs say a hook inherits the parent environment apart from OTEL_* (https://code.claude.com/docs/en/hooks), the pinned agent SDK passes that environment on to the CLI, and the auth order puts Bedrock/Vertex/Foundry, ANTHROPIC_AUTH_TOKEN and ANTHROPIC_API_KEY all above CLAUDE_CODE_OAUTH_TOKEN (https://code.claude.com/docs/en/authentication). should the subscription branch unset all of those?
  • feature request: could there be a setting to skip some folders? I have a few repos and a restricted part of my vault I'd never want summarized
  • a session goes in the index after its first capture, and resuming keeps the same session id, so if I resume a session the next day, the new work is skipped as a duplicate. is that intended?
  • the 3 turn minimum counts tool results as user turns, so any session that runs a few tools passes it. and SessionEnd also fires for claude -p and Agent SDK runs, so scripted runs over the threshold get captured too. would it make sense to skip those, maybe with a matcher on the hook?

small one: CLAUDE_CODE_EXTRA_BODY and CLAUDE_CODE_DISABLE_THINKING are both in the env vars docs now (https://code.claude.com/docs/en/env-vars), so that line in the PR description could be updated

@lhoupert

lhoupert commented Oct 5, 2026

Copy link
Copy Markdown
Author

Thank you for your comments Danny! I will take the time to look into it later this week.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants