Skip to content

feat(agent): add experimental local document assistant - #249

Draft
chaxus wants to merge 183 commits into
mainfrom
feat/local-multilingual-assistant
Draft

chaxus wants to merge 183 commits into
mainfrom
feat/local-multilingual-assistant

Conversation

@chaxus

@chaxus chaxus commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Merge Verdict

[NEEDS_CHANGE] — Draft checkpoint for the experimental local document assistant. Do not merge as a completed AI release: factual-writing quality and repository lint remain unresolved.

3,176 files · +905,968 / -3,389 · agent-core, chat-ui, editor integration, local runtime assets, evaluation evidence

Summary

  • Add an editor-side conversational assistant with local WebGPU inference, CPU initialization fallback, model loading/cancellation, model selection and optional persistent conversation history.
  • Integrate bounded Word, spreadsheet and presentation operations, native history handling and text readback verification; support cached local-runtime workflows and hosting isolation policies.
  • Preserve model/browser evaluation evidence, including failed semantic-quality cases. Redact personal home paths from diagnostic receipts and resolve the compatibility probe dependency from the current project directory.

Risk Analysis

Risk Level Mitigation
Factual rewriting and summary reliability High No tested candidate establishes acceptable seven-language factual-writing quality. Keep experimental status; successful document application is not proof of semantic accuracy.
Offline/device coverage Medium Archived evidence is scoped to specific warmed desktop caches and operations; mobile, WebKit and broader PWA behavior remain incomplete.
Large review surface High Most added lines are evaluation receipts. Review was partitioned across core, UI and cross-cutting integration; it is not an exhaustive audit of all artifacts.

Design Decisions

  • Use local inference in the browser product entry, with initialization fallback rather than replaying an interrupted generation.
  • Bind bounded document operations to editor state and native history; preserve explicit source literals and refuse invalid plans.
  • The initial feature history was consolidated into a sanitized checkpoint on current main. Subsequent reviewed fixes and frozen experiments are separate commits; superseded personal-path receipts are absent from the PR commit ancestry.

Code Concerns

Status: personal-path exposure and tool cancellation fixed; custom URL storage remains a reviewer decision.

  1. ✅ Fixed [Must Fix] Personal home paths in diagnostic receipts and the compatibility probe.
    Redacted plain and URL-encoded paths in five JSON reports and made the probe use the current project directory. Tracked-file byte scans and evaluation JSON parsing passed. This does not claim removal from previously published commit caches.
    — Author reply: fixed in this checkpoint, including the PR history.
  2. ✅ Fixed [Needs Decision] General runtime dropped tool cancellation.
    Forward the run signal to cooperative in-flight tools, preserving single-argument calls without a signal. The delayed-write regression reproduced an unwanted write before the fix and now passes at iteration limits 1 and 8. Independent scoped review found no critical or important issues.
    — Author reply: fixed; this does not claim forcible cancellation of tools that ignore their signal or new native-editor cleanup certification.
  3. [Needs Decision] Custom model URLs persist in localStorage (lib/agent-plugin/ui/panel.ts:292, :666).
    Isolated browser probes confirm synthetic authenticated URLs persist in settings and SDK cache names/metadata. No real credentials were used. Decide session-only/no-cache behavior and a migration policy for existing authenticated entries before remediation. See the October 7 settings/cache evidence.

Verification

Check Result
PR preflight Passed for the initial PR checkpoint; subsequent clean follow-up commits are on the same branch
Personal-path scan Passed: no personal username/local user path found in additions across the current full PR diff; model binaries remain excluded
Evaluation JSON syntax Passed: 1,301 tracked evaluation files parsed on the current tree
TypeScript Passed after the latest generic loading-failure translation change
Unit tests Native readiness fix: full suite 4,533 passed / 141 files. Latest translation change: 78 related loading/presentation tests passed; no new full-suite claim for the subsequent translation-only change.
Production build Latest production build passed: core1791361242/vendorb6864850e7b3. Native allocation failure, exact generic Chinese copy and same-panel small-model recovery passed. Model-screen runtime was preserved until its run completed.
Repository lint Failed: latest local Oxlint reports 68 diagnostics, all in frozen evaluation scripts. Formatting/evidence-byte handling remains a decision; no repository-wide green-check claim.
Diff whitespace Failed: existing diagnostic text and patch payload whitespace; these artifacts were preserved
Browser/device/semantic acceptance Fresh native GPU 21-task screen, desktop WebKit readiness and narrow current-build warmed Chromium XLSX lifecycle were run; seven-language quality still fails and physical mobile/full matrix are missing.

Conclusion basis

  • Verified: source inspection, partitioned static review, Git topology/tree-equivalence checks, privacy scanning, JSON parsing, typecheck, unit tests and production build.
  • Verified undesirable persistence: synthetic authenticated URL settings and SDK cache entries. Remediation behavior remains a design decision.
  • Not covered: exhaustive review of all raw receipts and binary assets, physical mobile devices and the full offline/editor matrix.

October 7 continuation

  • Tool cancellation fix is verified by deterministic regressions, TypeScript, changed-file lint and a fresh full unit suite; independent review is scoped to that fix.

  • Gemma 4 completed three remaining native Word development summaries. All native writes/Undo/Redo matched raw outputs, but two summaries omitted or transformed explicitly requested current states. The model is not adopted. Historical first-case partial evidence retains its separate build/cleanup limitations. See docs/evaluations/2026-10-07-gemma4-summary-analysis.md.

  • The frozen two-source comparison completed all four native Word writes and exact Undo/Redo checks. The candidate leaves an explicitly required Chinese actor/action/condition relationship implicit and is not adopted. No production prompt/default/dependency changes. See docs/evaluations/2026-10-07-gemma4-consistent-analysis.md. Overall seven-language fidelity, broader operations and device/offline acceptance remain incomplete.

  • Current-build warmed desktop Chromium offline Word chain passed cached CPU inference, exact replacement, Undo/Redo, Save, independent DOCX XML inspection and a separate offline browser restart/reopen. Aborted resource requests are retained. This is browser offline emulation, not physical network isolation or the complete device/editor matrix. See docs/evaluations/2026-10-07-cpu-offline-word-analysis.md.

  • Current-build warmed Chromium offline PPT chain now verifies default CPU AI text-box insertion, exact native Undo/Redo, Save, independent slide XML inspection and a separate offline browser restart/reopen with an identical slide text tree. Layout and full device/editor coverage are not certified. See docs/evaluations/2026-10-07-offline-ppt-analysis.md.

  • A frozen Gemma sampling comparison on one known Chinese development source now retains the explicitly pending budget state with official recommended sampling, while product greedy sampling repeats the omission. All other request fields match, native writes/Undo/Redo pass; observed latency is 127–129 seconds. This is narrow positive evidence, not adoption or seven-language acceptance. See docs/evaluations/2026-10-07-gemma4-sampling-analysis.md.

  • Frozen recommended-sampling transfer is running 21 rewrite/summary/translation tasks over seven correlated previously unused source-language versions. This is an in-progress screen, not acceptance; translation targets cover English/Chinese only. The model decision index now reconciles October 7 results without changing historical counts. See docs/evaluations/2026-10-07-gemma4-sampling-seven-language-protocol.md.

  • First Chinese transfer rows: rewrite and English translation retain the examined roles/negatives/permission; summary adds unsupported causal linkage and omits the signer. This rejects broad acceptance of the configuration; the same frozen process continues remaining rows. No default promotion.

  • Formatting-only cleanup fixes 13 application/test/document files. Related 67 tests across six files, TypeScript, changed-file lint and formatting checks pass. Hash-bound diagnostic bytes and the served model-screen build are preserved. Repository-wide formatting/lint remains unresolved. No production rebuild was performed during the frozen inference screen.

  • Current evidence reconciliation preserves the October 5 historical source bindings and separately binds present source plus completed Word/PPT and model diagnostics. Its verifier passes with semantic/device acceptance explicitly false. The ongoing transfer remains excluded from completed-evidence bindings. See docs/evaluations/2026-10-07-current-state-reconciliation.md.

  • Transfer setup audit found three Korean rows with empty source/selection and zero inference, not model-quality failures. A native-only probe reproduces asynchronous font/input readiness, successful font requests and complete source after waiting/reselection. Corrected Korean driver is frozen and will run only after the original process finishes; original failures remain preserved. Partial validation now explicitly separates valid inference from empty-source setup failures and rejects tampered sampling/history receipts. See docs/evaluations/2026-10-07-gemma4-sampling-korean-corrected-protocol.md. Chinese/Japanese/German summary omissions remain genuine semantic failures.

Latest pushed evidence: the completed multilingual sampling attempt records 18 valid inference executions and three Korean empty-source setup failures. Its evidence verifier passes; this is not semantic acceptance. The original failures are retained, and model quality, device coverage, and repository formatting/CI gaps remain open. No personal local paths or usernames were found in the PR commit diffs; generic redacted log paths and WebAssembly virtual filesystem paths are retained.

Corrected Korean transfer is now archived separately: three valid native executions completed, with exact Undo/Redo and unchanged frozen assets. Rewrite and translation retain checked facts; summary omits Mei Tan and the technical team. Combined coverage is 21 valid task executions (24 attempts including three original setup failures), with 11 narrow passes, six failures and four uncertain manual reviews. Six of seven summaries fail. These are correlated versions of one source, not an independent accuracy estimate; no model adoption follows. Evidence and model-index verifiers pass.

Thinking-mode feasibility contrast completed with process exit 1: both actual loads used four CPU threads. Baseline reproduced the Korean summary actor omissions; thinking timed out at 240 seconds. The timeout catch path did not retain SDK request/intermediate output or native document snapshots, so no request-parity or semantic assertion is made for that row. The original failed receipt and explicit evidence limits are archived; no production model/settings change follows.

Current preserved-build Excel offline lifecycle passed: online warmup and process close, offline restart with service-worker homepage and cached default CPU model, exact B2 tool write, native Undo/Redo, actual editor Save, independent XLSX CRC and sheet1 B2 shared-string validation, then another offline browser restart/file-chooser reopen with the same native B2 value. Aborted resource requests are retained. Scope is one explicit operation in warmed desktop Chromium with offline emulation; no fresh-install, physical-disconnection, PWA/mobile/WebKit or generative-quality claim.

Actual backend date diagnostic ended with exit 1 and both contexts closed. Metal reproduced the prior date-22-to-unrelated-address error after correctly copying date-23. SwiftShader loaded the same model record but its first completion timed out at 300 seconds; no output comparison or Metal-causality claim is justified. Worker-level adapter/progress snapshots and prelaunch artifact hashes were not captured. Known-source date copying is diagnostic only; no model or writing-quality acceptance.

Frozen chat-state contrast completed all three sequences: date22 cold, after successful date23, and after successful date23 plus awaited resetChat(false). Each date22 result was the same unrelated address; date23 remained correct. Worker GPU vendor was apple, and frozen client/Worker/driver/request bytes matched prelaunch hashes. Prior-request state reuse is therefore not necessary for this known error, and reset does not repair it. No production reset or model promotion is added; the root cause and full quality/device acceptance remain unresolved.

Pinned compiled-library contrast completed: actual Apple Metal Worker runs loaded verified upstream base and sg32 WASM bytes. With identical requests/model configuration except model_lib, both copied date23 correctly and produced the same unrelated address for date22. Switching to the subgroup-enabled library is not a sufficient repair. Frozen served client/Worker bytes match prelaunch hashes; weight cache bytes were not independently rehashed. No production library or quality-acceptance change follows.

Frozen original-policy placement diagnostic

The Qwen3-4B-Instruct-2507 CPU development screen completed all 21 tasks with process exit 0, closed contexts/browser, unchanged frozen artifact hashes and full request/native-history verification. Six rows pass bounded manual checks and fifteen fail, including wrong-language, ISO-date, name-spelling, sentence-count and introduced grammar failures. These are development cases without native-speaker certification, not a population accuracy estimate or seven-language quality acceptance. Full raw receipts, hash-bound manual reviews, analysis and bindings are archived. The pre-frozen six-row known-source contrast is now running after that process closed: move the original writing preamble verbatim from user to system, preserving the original task JSON, model, schema and sampling. Its verifier rejects altered policy content. No production prompt, model or default changes are adopted from this diagnostic.

Native CPU readiness and compound lifecycle follow-up

The official pinned Qwen2.5 7B shards passed full byte/hash verification, but all four browser attempts failed native CPU buffer allocation before any completion request. These are setup failures, not semantic test results. A resolved SDK load could still show a loaded note with zero vocabulary/context. The provider now requires a native token-count/context preflight before reporting readiness and releases initialization failures through the existing cleanup path. The new regression fails before the fix and passes after; production build, TypeScript and all 141 test files / 4,533 tests pass. Real Chromium shows failed 7B initialization, zero completion calls and one runtime exit, while the ordinary .6B model still loads and chats.

A supported literal Excel read-range-then-set-cell sequence also passes warmed offline restart, exact whole-table Undo/Redo, native Save, independent OOXML CRC/shared-string inspection and separate offline process reopen. This does not certify arbitrary multi-step planning or physical network disconnection.

Signed/parameterized model-URL persistence remains open: both settings and SDK cache metadata can retain original URLs. A storage/cache policy is awaiting user choice. No seven-language model adoption or complete device/offline/privacy acceptance is claimed.

Fixed CPU memory ceiling and higher-precision diagnostic

The exact matched default WASM imports shared memory64 with a maximum of 65,536 pages (4 GiB); compatibility memory32 has the same maximum. The official 7B run requested a single 4,677,120,000-byte model buffer, exceeding the entire configured memory before context/other allocations. This identifies a current artifact ceiling, not insufficient physical host RAM. No larger runtime has been adopted.

A known-source Q4/Q6 Instruct-2507 contrast is frozen with original production messages/schema and matched sampling. The first Q6 download truncated at 841,791,698 bytes and failed full integrity validation; it was never loaded. Explicit HTTP range continuation is in progress after an exact single-byte 206 response preflight. Only a matching final whole-file hash permits browser loading. This diagnostic has no completion or quality-acceptance claim.

Latest diagnostic evidence

  • Verify the complete pinned Q6 model download and compare Q4/Q6 GGUF metadata; model binaries remain local and are excluded from Git. These checks do not establish multilingual acceptance.
  • Preserve isolated browser probes showing that synthetic authenticated model URLs persist in settings and SDK cache names/metadata. No real credentials were used. URL persistence remediation remains pending; the cache probe records a request-count assertion failure alongside the persistence evidence.
  • Scan additions across the full PR diff and the four newly pushed commits for personal usernames and local user paths: no matches found.

Latest completed regressions and GPU model diagnostic

  • Complete the eight-row Q4/Q6 known-source contrast: Q6 narrowly passes the German rewrite, but Chinese summary, Japanese-to-Korean actor/name preservation and Spanish source-language rewrite still fail. No candidate is adopted.
  • Verify native small-model readiness/count/chat in desktop Playwright WebKit; this does not cover physical Safari/mobile or semantic writing.
  • Reproduce the unsupported compound read-sort-sum request. Current sequence parsing and execution restrict earlier writing steps; broader multi-step failure/Stop behavior remains a design decision.
  • Verify two separate deterministic sort and sum commands, native Undo/Redo, Save, independent OOXML inspection and offline process reopen. Re-run with actual editor-a29E5U4S.js assertions after observing the earlier profile had an older active worker. Current build passes this narrow warmed Chromium XLSX flow.
  • Pinned Qwen2.5-7B GPU artifacts loaded on an actual Apple adapter. The four-known-source SDK diagnostic completed (three narrow passes, one sentence-form failure); the subsequent native 21-task screen still rejects global adoption. Binaries remain ignored locally.

Completed larger GPU screen and current recovery fix

Pinned Qwen2.5-7B loaded on an actual Apple GPU and completed all 21 original-product Word development tasks. Bounded manual review: 9 narrow passes, 11 failures, 1 uncertain. Missing actors/current status, changed recipients/actions/objects, wrong language/date formats and target-language issues prevent default promotion. Native writes and Undo/Redo pass separately; this is not unused-source transfer or native-speaker/full-device acceptance.

Seven generic loading-failure translations now provide accurate retry/model-switch guidance without implying a first-download network context. Related tests (78) and TypeScript pass. The current production browser probe confirms failed CPU allocation cleans up, then the same panel loads the small model and completes chat; a separate small-model control also passes.

Additional GPU diagnostics (042a14a)

  • Preserve the six-case policy-placement comparison: all six candidate outputs still fail their bounded checks. The run exited 1 because its original network-fetch gate did not observe the cached library; the failure is retained. A separate cache probe confirms the pinned library bytes, without certifying that run as successful.
  • Freeze the seven-source, three-variant rewrite diagnostic and its exact request bank. These commits contain the protocol and driver; no completed quality result is claimed.
  • The four added commits contain evaluation artifacts only. Added-content privacy scans against both the PR base and the previously pushed head found no personal usernames or user-directory paths. The PR remains a draft with the existing NEEDS_CHANGE verdict.

Completed fixed-example rewrite contrast (71866d1)

The original diagnostic exited 1 before generation because its Worker interception omitted production isolation headers. Its timeout is preserved. A separately bound correction retains the production response headers and completes all 21 SDK requests with actual pinned-library cache-read hashes and no Worker errors. Bounded manual review: product 4 narrow passes/2 failures/1 uncertain; minimal policy 3/4/0; fixed examples 5/2/0. Fixed examples produce Chinese from an English rewrite and omit the Japanese operating-team actor. No variant passes the seven known-source rewrites; no product default changes or full acceptance are claimed. A read-only cached-weight audit is in progress; its driver is committed, with no completed audit result claimed yet.

Late GPU failure feedback repair (448a9ad / ab22fba)

Actual device-loss testing found a stale loaded note after interrupted streaming. A failing idle-invalidation regression confirms that request finalization alone misses later Worker failure notifications. The provider now forwards its existing failure signal through an optional callback, and the panel updates readiness under controller-generation ownership. Full suite: 141 files/4538 tests pass; types, changed-file lint and production build pass. Current production native test destroys an actual GPUDevice after streaming starts: loaded status clears, Load model is shown, partial text is retained, input unlocks, document stays unchanged and explicit reload/new greeting succeeds. Original failed probes and pre-fix stale-status evidence remain preserved. This does not certify spontaneous loss, CPU replay monitoring or the full device matrix.

The read-only Qwen2.5-7B cache audit also completes: all 88 shard sizes/MD5s match the pinned manifest (4,284,263,424 bytes), process exit 0. This is post-inference current-cache integrity, not exact earlier consumption or semantic quality. Full PR added-content privacy scan passes; seven-language writing acceptance and broader device/compound-operation gates remain open.

Latest pushed evidence (2026-10-07, 509085d)

  • Current diff: 3,254 files, +918,772 / -3,389 lines.
  • Added-content privacy scan against origin/main passed for personal usernames and macOS/Windows user-home paths. Ignored local compiler environments and model caches are excluded.
  • Added native idle GPU-loss recovery, explicit CPU retry, and current CPU offline restart receipts with narrowly scoped verification.
  • The isolated compiler dependency inventory records dependency resolution only; custom model compilation and compatibility remain unverified.
  • This remains an experimental draft with the existing NEEDS_CHANGE verdict and outstanding multilingual quality and device acceptance gaps.

Latest model validation evidence (2026-10-07)

Locally compiled Qwen2.5 7B and 14B WebGPU libraries loaded on the observed Apple Metal adapter. The 14B structured request timed out with its default configuration; an explicit 2048-token context completed the same request in diagnostic controls. This does not change the product defaults.

The completed Chinese native writing triplet contains two failures (ISO date format changed in rewrite; delivery-team actor omitted in summary) and one narrowly passing translation sample. The broader language screen is still in progress and is not claimed as passing. Physical Windows/mobile coverage and full offline acceptance remain incomplete. This PR remains experimental and [NEEDS_CHANGE].

All added diff content was scanned before this push for the local username and macOS/Windows user-home paths; no matches were found. Local model binaries, toolchains, profiles, and raw scratch logs remain excluded.

Latest bounded-read repair and validation (2026-10-07)

Fixes WebKit cached model Blob failures by assembling worker buffers from sequential reads of at most 8 MiB. Exact requested bytes and failed-subread cleanup are regression-tested. Related Wllama tests: 10 files / 65 tests passed; TypeScript, changed-test lint/format and production build passed. Native artifact pairs remain unchanged and the client patch can be reversed to the original accepted hash.

On isolated Playwright 1.65.0-alpha-2026-10-07 / WebKit 27.2, the actual rebuilt product completed default CPU model loading, online-page closure, offline new-page navigation, cached-model restoration and exact WEBKIT_OFFLINE_OK reply with unchanged document. Offline spelling-script/update request failures remain recorded. This is one desktop page-close lifecycle, not physical Safari/mobile, process restart, save/reopen or complete offline acceptance. A corrected explicit tools-mode run now verifies offline insertion, Undo and Redo; the original chat-mode observer and correction are preserved. Project Playwright dependencies are unchanged.

The completed diagnostic Qwen2.5-14B 21-task screen gives 7 narrow passes, 12 failures and 2 uncertain outputs. A known-source policy-placement contrast does not repair actor/date/summary failures; no model/default is adopted. Authenticated model URL persistence and directory-query propagation gaps are documented and remain unresolved. The PR remains draft and [NEEDS_CHANGE]. Added content was scanned again for local username and user-home paths before push; no matches.

Offline saved-file reopen repair (2026-10-07)

A fresh WebKit offline save/reopen preserved native text but exposed an uncached font-menu sprite and page load errors. This PR now precaches the ten binary locale/density UI sprites (4,448,211 bytes) in the existing vendor-versioned cache. A regression failed before the repair and passed afterward. Full unit suite: 142 files / 4541 tests passed (asynchronous rejection-handling warnings were observed); TypeScript, changed-file lint/format and production build passed.

Fresh rebuilt-product WebKit 27.2 verified all ten cached sprites, default CPU restoration offline, exact insertion in tools mode, native Undo/Redo, DOCX save with ZIP/XML checks, and native reopening with exact text and no page errors. Original failure evidence is retained. This remains a same-context page lifecycle: process restart, physical Safari/mobile, broader operation coverage and seven-language writing fidelity are not certified. Service-worker update and spelling-script offline request failures remain recorded. No model/default or semantic acceptance was changed.

The complete added diff was scanned for personal username and macOS/Windows user-home paths before push. This PR remains draft and [NEEDS_CHANGE].

Cached WebKit browser restart and writing diagnostics (2026-10-07)

WebKit 27.2 now verifies closure/disconnection of the first browser, a new persistent-context launch, offline navigation served by the service worker, default CPU restoration, exact chat, native insertion/history, DOCX save and reopening with exact text and no page errors. A separate read-only probe found existing same-origin OPFS model bytes in a new configured persistent directory before model initialization; this proves cached restoration, not cold-cache isolation. Physical Safari/mobile, PWA installation and the complete offline/device matrix remain incomplete.

Four known-source fact-ledger writing contrasts yielded two narrow passes and two failures. The same narrow improvements also occur with the extra rendition prompt alone, so extraction adds no demonstrated benefit on those cases. Three protected-date rewrite contrasts restore ISO dates exactly, but Korean loses a permission actor present in both controls. No writing pipeline or model/default was adopted; seven-language factual writing remains unresolved. Raw outputs, mechanical request comparisons, limited manual reviews and failed evidence are preserved.

Full added-content scan before push: 950,683 lines, no local username or macOS/Windows user-home path matches. The draft remains [NEEDS_CHANGE].

Latest verification and evaluation (2026-10-07)

  • Added an editor CSP repair permitting embedded image byte reads via connect-src data:. Script and worker policies remain restrictive. Native browser controls verified image reads and blocked data scripts/workers. Both Excel frozen-pane E2E tests passed against the rebuilt product.
  • Validation: 142 unit-test files / 4,542 tests passed; TypeScript, changed-file lint/format checks, and production build passed. Full repository formatting/lint gates remain unresolved; the previous remote Pages/Docker failures have not yet been verified by a complete remote rerun.
  • A frozen seven-language, 42-request segmented-writing comparison completed. Bounded manual judgments: baseline 3 pass / 16 fail / 2 uncertain; candidate 5 pass / 9 fail / 7 uncertain. These are development-screen observations, not general accuracy or native-speaker certification. No candidate was adopted as the product default.
  • This PR remains draft with NEEDS_CHANGE pending reliable multilingual writing, remaining privacy decisions, device coverage, and complete CI acceptance.

Archive integrity and update observation repair (2026-10-07)

  • Stylistic checks now verify SHA-256 for 1,473 explicitly listed historical evaluation artifacts, plus JSON parsing and diagnostic script syntax, before excluding only those exact paths from formatting/lint. New evidence and product code retain normal gates. This preserves recorded bytes without claiming semantic acceptance. The maintained decision index remains formatted.
  • Local complete format, oxlint and TypeScript checks passed; 143 unit-test files / 4,547 tests passed (an existing asynchronous rejection warning remains). Integrity regression tests reject changed/missing bytes, invalid JSON/script syntax, duplicate entries and traversal.
  • The completed preceding remote run passed all five Pages-semantics shards. Ordinary E2E shards 1/2 failed when the expected automatic update reload destroyed a VERSION probe context. The observer now retries only that navigation error; its native Chromium test passed in 17.6 seconds. Product update behavior was not changed.
  • That run also exposed Docker presentation-theme tests calling a non-function currentThemeId; this remains under investigation. Latest commits require a fresh remote CI result. Draft NEEDS_CHANGE remains in effect, including multilingual writing and device/privacy acceptance gaps.

Theme readiness and local writing candidates (2026-10-07)

  • Theme observation now waits for the vendor theme method to become callable, then asserts the actual applied theme. All 10 local Chromium presentation-theme cases passed; TypeScript and changed-file lint/format checks passed. The Docker daemon is unavailable locally, so Docker validation remains remote.
  • The preceding remote run passed Lint and Validate with archive integrity checks. Theme-method readiness failures remained in Docker/Pages, and ordinary serial update tests exposed a separate persistent dev-controller result after navigation retry. A matching local Vite-build run reproduced it: new worker installed/waiting with e2e-next while dev still controls the page and the heal flag is set. Activation remains under investigation; no complete CI success is claimed.
  • An independent revision pass completed 21 14B SDK requests: bounded manual review 2 narrow passes / 19 failures. The two passing tasks already passed their baselines; no revision pipeline adoption.
  • Fixed SHA-256-verified Hy-MT2-1.8B Q4_K_M bytes loaded in the current browser CPU SDK and completed seven cyclic translations: bounded manual review 1 narrow pass / 4 failures / 2 uncertain. Date/name changes and terminology/grammar gaps remain. Model binaries were not committed; no product/default, editor-application, rewrite/summary or device/offline acceptance is claimed.
  • Archive integrity now pins 1,479 exact evidence files. Complete local formatting, oxlint and TypeScript gates passed with the new artifacts. The PR remains draft NEEDS_CHANGE.

Version-query cleanup and activation controls (2026-10-07)

  • Worker version queries now release both message ports and clear their timeout after a reply, timeout or sending failure. Three regression tests failed before the fix and passed afterward. The E2E version observer also releases its resources.
  • Validation: 71 relevant unit tests and the full 143-file / 4,550-test suite passed; TypeScript, changed-file lint/format and production build passed. The native silent-update test passed against the fully stamped production build (16.2 seconds).
  • The completed preceding remote run for 7fc6698 passed lint, ordinary E2E, Docker E2E and Pages-semantics E2E. This predates the cleanup; current commits require fresh CI.
  • Bare Vite-build update controls still sometimes stall with dev controlling the page while e2e-next waits. A diagnostic observed receipt of SKIP_WAITING with an unresolved skipWaiting Promise; event-lifetime instrumentation changed timing and succeeded. A stable-dev-cache-only control also failed. No causal activation repair is claimed, and unsaved-document protections were preserved.
  • Multilingual writing, complete device/offline coverage and pending privacy decisions remain incomplete. Draft NEEDS_CHANGE remains in effect.

Quoted cell text and desktop WebKit verification (2026-10-07)

Quoted English/Chinese single-cell assignments now constrain the model plan to the exact address, literal and text type; mismatches are rejected before execution. This repairs the observed leading-zero conversion. Unquoted numeric entry retains native parsing. Full unit suite: 143 files / 4,557 tests passed; type/lint, format and production build passed. Existing asynchronously handled rejection warnings remain.

Desktop WebKit 27.2 process restart with offline enabled before navigation now passes CPU chat plus exact XLSX/PPTX edits, Undo/Redo, native Save and offline reopen. Saved XML independently preserves 00123 and the PPT literal. These scoped checks do not establish physical/mobile Safari, arbitrary operations, full font/layout coverage or seven-language writing acceptance. Paired Hy-MT2 literal protection still exhibits semantic failures and is not adopted. Original failed receipts and the successful rerun remain hash-bound in the evaluation archive.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Deploying document with  Cloudflare Pages  Cloudflare Pages

Latest commit: 645ebe8
Status: ✅  Deploy successful!
Preview URL: https://bf067e4b.document-7hm.pages.dev
Branch Preview URL: https://feat-local-multilingual-assi.document-7hm.pages.dev

View logs

chaxus added 28 commits October 7, 2026 11:50
chaxus added 30 commits October 7, 2026 19:45

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant