fix(runtime-api): keep file modes, allow PUT preflight, undo rejected provider switch, list all memory - #6771
Merged
Merged
Conversation
… switch, list all memory
Four verified Runtime API route defects:
- PUT /v1/workspace/files rewrote an existing file as 0600 (dropping the
executable bit and group/other read) and created new files 0600 with
0700 parents. The route now opens through WorkspaceFile::open_shared:
replacement keeps the target's permission bits (set-id bits are not
carried) and creation follows the umask. Private stores still use
WorkspaceFile::open and stay owner-only.
- The CORS layer did not advertise PUT, so browser preflight blocked the
six PUT routes (provider key, workspace file, session, operate plan,
thread goal, account model access). PUT is now allowed.
- POST /v1/providers/{id}/switch persisted the selection before reload
validation; a 400 "Config reload rejected" left the switch on disk for
the next restart or reload. The pre-switch document is captured under
the config lock and restored with the existing compare-and-swap writer
when load or reload fails; a newer concurrent save is never overwritten.
- GET /v1/memory with scope omitted or "all" listed only global notes.
It now includes this repository's workspace notes (list and search).
scope=workspace without an origin remote returns 400 for GET and
DELETE instead of 500, matching create.
Tests (targeted, macOS, shared target, 2 jobs):
- cargo test -p codewhale-tui --lib -- fleet::files cors_layer
switch_provider memory_ config_persistence commands::groups::config:
242 passed, 0 failed.
- Negative control with each source fix reverted and the new tests kept:
6 failed (shared_replacement_keeps_the_existing_permission_bits,
shared_creation_follows_the_umask_for_the_file_and_its_parents,
cors_layer_advertises_exact_supported_headers_and_never_an_extra,
switch_provider_rejected_by_reload_leaves_the_config_file_unchanged,
memory_all_scope_includes_workspace_notes,
memory_workspace_scope_without_a_remote_is_a_client_error), 67 passed.
- cargo fmt --all -- --check: pass. cargo clippy -p codewhale-tui --lib
--tests --all-features with CI allowances and -D warnings: clean.
- scripts/check-blocking-calls-budget.py, check-dead-code-budget.py,
check-reqwest-builders.py: pass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…erialized; memory scope follow-ups Review follow-ups on the runtime-api-routes fixes: - The undo token for a refused provider switch was read back with std::fs::read_to_string after the config lock was released, so a save landing in between became the token. The config crate now returns a ConfigDocumentUndo from mutate_config_document_undoable, holding the exact bytes the write left and the pre-write document, both taken under the lock. undo() restores only while the file still holds that write; when the file changed since, the newer save is kept and the error says so instead of advising a retry. - A switch on a missing config file left an empty config.toml behind on undo (which skips the first-run template). The undo now removes the file it created. - Save, reload and undo run as one tokio task that the request only awaits, so a client that disconnects mid-reload no longer skips the undo or leaves the engines on the new config while state.config keeps the old one. Switches are serialized by a runtime mutex so one refused switch cannot restore over, or under, another. The reply reports the route that switch applied. A process crash between save and reload remains a window. - The error text names the real config path, not "config.toml". - Memory: scope=all resolves this repository's identity only for GET, once per request (search reuses it via NativeMemoryStore::search_in_workspace), and logs when git is unavailable. DELETE scope=all no longer does an unused git lookup; its wider reach (every local scope) is documented. - Workspace file PUT: route test that 0755 and 0600 survive an edit and a setuid 04755 file comes back 0755; unit test that the shared replacement drops setgid (06750 -> 0750) and keeps 0600. - check-blocking-calls-budget: config_document.rs std_fs 2 -> 3. The new site is ConfigDocumentUndo::undo's file removal, which is synchronous by design and runs under spawn_blocking in switch_provider. Tests (local, shared target, build slot held): - cargo test -p codewhale-config --lib config_document: 9 passed, 0 failed - cargo test -p codewhale-runtime --lib memory: 1 passed, 0 failed - cargo test -p codewhale-tui --lib (fleet::files, runtime_api switch_provider/ memory/workspace_file/cors, config_persistence, commands::groups::config): 201 passed, 0 failed - rustfmt --check clean; clippy: no findings on changed lines; check-blocking-calls-budget, check-dead-code-budget, check-reqwest-builders pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Preserve the original cancellation-safe switch and CAS undo. Its existing owned task awaits the serialization mutex, then one spawn_blocking worker performs config persistence, load and rollback; async engine reload uses the existing Tokio handle. Capture test environment scope for that worker. Add a real current-thread loopback regression with an exclusive filesystem config lock, bounded negative-control release and async progress assertion. Focused governed proof, source-revert control and exact restoration plus npm/web and CI-policy Clippy are pending; no hosted or native provider acceptance claimed. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Strengthen the provider-switch regression after its first negative control incorrectly passed: the server helper owns a separate current-thread product runtime, so a client-side timer cannot prove server responsiveness. Probe the same server health endpoint while the real config file lock is held, preserving cleanup and CI-scaled budgets. Production worker fix is unchanged. Prior13/0 and npm774/0/check:web receipts remain dated evidence; the new affected control and final Clippy are pending. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Existing provider_switches serialization and disconnect-surviving task retain the same writer, exact undo token, reload and CAS semantics. One owned spawn_blocking worker performs Config persistence/load/undo with its joined environment gate and existing runtime reload. Frozen source b593596, tree 37f3dbf. Same-serving current-thread GET /health regression: 1 passed, 0 failed (0.23s). Original whole-function negative control: 0 passed, 1 failed (3.21s; serving worker delayed 2.863087875s). Exact production bytes restored; re-pass 1/0 (0.24s). CI-policy TUI all-targets/all-features locked clippy passed (2m12s). Earlier source 97f focused provider fixtures: 13/0. Its initial client-side health control also passed without the fix and is rejected evidence, not acceptance; all weak-control logs preserved. npm test 774/0 and check:web passed at 97f; only the Rust regression test changed since that gate, so Node/web proof is explicitly dated rather than claimed rerun at b593. Artifacts: ci-contributors/pr-6771.qualification.json and pr-6771.blocking-proof-v2.json. Exhaustive three-OS coverage belongs to fresh exact-head CI after this original-branch push. No release claim. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No-Issue: Repair verified Runtime API filesystem, provider-switch and memory route defects.
The Runtime API preserves ordinary workspace file permissions on replacement and the process umask on creation, advertises PUT in its CORS preflight, restores the exact prior config when a provider switch is rejected, and includes the repository workspace in memory scope=all and scoped search. Missing workspace identity produces a client error; private-store file creation remains owner-only.
Provider switches use the existing serialized, disconnect-surviving owner task. All real config file-lock, persistence, load and exact undo work runs in one blocking worker; the joined environment gate and existing runtime reload preserve the current writer and compare-and-swap semantics. A refused switch removes a config file it created, and an intervening newer write is never overwritten. Process crash between save and reload remains an existing durability window.
Current repair qualification (macOS, frozen b593596):
Original author historical route/config/memory/file-mode fixtures are preserved in commit history. Fresh exact-head Linux/macOS/Windows CI remains the exhaustive merge gate. No full local suite, native Windows, real-provider, deployment or release acceptance is claimed.