Repository navigation
Conversation
Deploying document with
|
| Latest commit: |
645ebe8
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://bf067e4b.document-7hm.pages.dev |
| Branch Preview URL: | https://feat-local-multilingual-assi.document-7hm.pages.dev |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge Verdict
[NEEDS_CHANGE] — Draft checkpoint for the experimental local document assistant. Do not merge as a completed AI release: factual-writing quality and repository lint remain unresolved.
Summary
Risk Analysis
Design Decisions
Code Concerns
[Must Fix] Personal home paths in diagnostic receipts and the compatibility probe.Redacted plain and URL-encoded paths in five JSON reports and made the probe use the current project directory. Tracked-file byte scans and evaluation JSON parsing passed. This does not claim removal from previously published commit caches.
— Author reply: fixed in this checkpoint, including the PR history.
[Needs Decision] General runtime dropped tool cancellation.Forward the run signal to cooperative in-flight tools, preserving single-argument calls without a signal. The delayed-write regression reproduced an unwanted write before the fix and now passes at iteration limits 1 and 8. Independent scoped review found no critical or important issues.
— Author reply: fixed; this does not claim forcible cancellation of tools that ignore their signal or new native-editor cleanup certification.
lib/agent-plugin/ui/panel.ts:292,:666).Isolated browser probes confirm synthetic authenticated URLs persist in settings and SDK cache names/metadata. No real credentials were used. Decide session-only/no-cache behavior and a migration policy for existing authenticated entries before remediation. See the October 7 settings/cache evidence.
Verification
Conclusion basis
October 7 continuation
Tool cancellation fix is verified by deterministic regressions, TypeScript, changed-file lint and a fresh full unit suite; independent review is scoped to that fix.
Gemma 4 completed three remaining native Word development summaries. All native writes/Undo/Redo matched raw outputs, but two summaries omitted or transformed explicitly requested current states. The model is not adopted. Historical first-case partial evidence retains its separate build/cleanup limitations. See
docs/evaluations/2026-10-07-gemma4-summary-analysis.md.The frozen two-source comparison completed all four native Word writes and exact Undo/Redo checks. The candidate leaves an explicitly required Chinese actor/action/condition relationship implicit and is not adopted. No production prompt/default/dependency changes. See
docs/evaluations/2026-10-07-gemma4-consistent-analysis.md. Overall seven-language fidelity, broader operations and device/offline acceptance remain incomplete.Current-build warmed desktop Chromium offline Word chain passed cached CPU inference, exact replacement, Undo/Redo, Save, independent DOCX XML inspection and a separate offline browser restart/reopen. Aborted resource requests are retained. This is browser offline emulation, not physical network isolation or the complete device/editor matrix. See
docs/evaluations/2026-10-07-cpu-offline-word-analysis.md.Current-build warmed Chromium offline PPT chain now verifies default CPU AI text-box insertion, exact native Undo/Redo, Save, independent slide XML inspection and a separate offline browser restart/reopen with an identical slide text tree. Layout and full device/editor coverage are not certified. See
docs/evaluations/2026-10-07-offline-ppt-analysis.md.A frozen Gemma sampling comparison on one known Chinese development source now retains the explicitly pending budget state with official recommended sampling, while product greedy sampling repeats the omission. All other request fields match, native writes/Undo/Redo pass; observed latency is 127–129 seconds. This is narrow positive evidence, not adoption or seven-language acceptance. See
docs/evaluations/2026-10-07-gemma4-sampling-analysis.md.Frozen recommended-sampling transfer is running 21 rewrite/summary/translation tasks over seven correlated previously unused source-language versions. This is an in-progress screen, not acceptance; translation targets cover English/Chinese only. The model decision index now reconciles October 7 results without changing historical counts. See
docs/evaluations/2026-10-07-gemma4-sampling-seven-language-protocol.md.First Chinese transfer rows: rewrite and English translation retain the examined roles/negatives/permission; summary adds unsupported causal linkage and omits the signer. This rejects broad acceptance of the configuration; the same frozen process continues remaining rows. No default promotion.
Formatting-only cleanup fixes 13 application/test/document files. Related 67 tests across six files, TypeScript, changed-file lint and formatting checks pass. Hash-bound diagnostic bytes and the served model-screen build are preserved. Repository-wide formatting/lint remains unresolved. No production rebuild was performed during the frozen inference screen.
Current evidence reconciliation preserves the October 5 historical source bindings and separately binds present source plus completed Word/PPT and model diagnostics. Its verifier passes with semantic/device acceptance explicitly false. The ongoing transfer remains excluded from completed-evidence bindings. See
docs/evaluations/2026-10-07-current-state-reconciliation.md.Transfer setup audit found three Korean rows with empty source/selection and zero inference, not model-quality failures. A native-only probe reproduces asynchronous font/input readiness, successful font requests and complete source after waiting/reselection. Corrected Korean driver is frozen and will run only after the original process finishes; original failures remain preserved. Partial validation now explicitly separates valid inference from empty-source setup failures and rejects tampered sampling/history receipts. See
docs/evaluations/2026-10-07-gemma4-sampling-korean-corrected-protocol.md. Chinese/Japanese/German summary omissions remain genuine semantic failures.Latest pushed evidence: the completed multilingual sampling attempt records 18 valid inference executions and three Korean empty-source setup failures. Its evidence verifier passes; this is not semantic acceptance. The original failures are retained, and model quality, device coverage, and repository formatting/CI gaps remain open. No personal local paths or usernames were found in the PR commit diffs; generic redacted log paths and WebAssembly virtual filesystem paths are retained.
Corrected Korean transfer is now archived separately: three valid native executions completed, with exact Undo/Redo and unchanged frozen assets. Rewrite and translation retain checked facts; summary omits Mei Tan and the technical team. Combined coverage is 21 valid task executions (24 attempts including three original setup failures), with 11 narrow passes, six failures and four uncertain manual reviews. Six of seven summaries fail. These are correlated versions of one source, not an independent accuracy estimate; no model adoption follows. Evidence and model-index verifiers pass.
Thinking-mode feasibility contrast completed with process exit 1: both actual loads used four CPU threads. Baseline reproduced the Korean summary actor omissions; thinking timed out at 240 seconds. The timeout catch path did not retain SDK request/intermediate output or native document snapshots, so no request-parity or semantic assertion is made for that row. The original failed receipt and explicit evidence limits are archived; no production model/settings change follows.
Current preserved-build Excel offline lifecycle passed: online warmup and process close, offline restart with service-worker homepage and cached default CPU model, exact B2 tool write, native Undo/Redo, actual editor Save, independent XLSX CRC and sheet1 B2 shared-string validation, then another offline browser restart/file-chooser reopen with the same native B2 value. Aborted resource requests are retained. Scope is one explicit operation in warmed desktop Chromium with offline emulation; no fresh-install, physical-disconnection, PWA/mobile/WebKit or generative-quality claim.
Actual backend date diagnostic ended with exit 1 and both contexts closed. Metal reproduced the prior date-22-to-unrelated-address error after correctly copying date-23. SwiftShader loaded the same model record but its first completion timed out at 300 seconds; no output comparison or Metal-causality claim is justified. Worker-level adapter/progress snapshots and prelaunch artifact hashes were not captured. Known-source date copying is diagnostic only; no model or writing-quality acceptance.
Frozen chat-state contrast completed all three sequences: date22 cold, after successful date23, and after successful date23 plus awaited resetChat(false). Each date22 result was the same unrelated address; date23 remained correct. Worker GPU vendor was apple, and frozen client/Worker/driver/request bytes matched prelaunch hashes. Prior-request state reuse is therefore not necessary for this known error, and reset does not repair it. No production reset or model promotion is added; the root cause and full quality/device acceptance remain unresolved.
Pinned compiled-library contrast completed: actual Apple Metal Worker runs loaded verified upstream base and sg32 WASM bytes. With identical requests/model configuration except model_lib, both copied date23 correctly and produced the same unrelated address for date22. Switching to the subgroup-enabled library is not a sufficient repair. Frozen served client/Worker bytes match prelaunch hashes; weight cache bytes were not independently rehashed. No production library or quality-acceptance change follows.
Frozen original-policy placement diagnostic
The Qwen3-4B-Instruct-2507 CPU development screen completed all 21 tasks with process exit 0, closed contexts/browser, unchanged frozen artifact hashes and full request/native-history verification. Six rows pass bounded manual checks and fifteen fail, including wrong-language, ISO-date, name-spelling, sentence-count and introduced grammar failures. These are development cases without native-speaker certification, not a population accuracy estimate or seven-language quality acceptance. Full raw receipts, hash-bound manual reviews, analysis and bindings are archived. The pre-frozen six-row known-source contrast is now running after that process closed: move the original writing preamble verbatim from user to system, preserving the original task JSON, model, schema and sampling. Its verifier rejects altered policy content. No production prompt, model or default changes are adopted from this diagnostic.
Native CPU readiness and compound lifecycle follow-up
The official pinned Qwen2.5 7B shards passed full byte/hash verification, but all four browser attempts failed native CPU buffer allocation before any completion request. These are setup failures, not semantic test results. A resolved SDK load could still show a loaded note with zero vocabulary/context. The provider now requires a native token-count/context preflight before reporting readiness and releases initialization failures through the existing cleanup path. The new regression fails before the fix and passes after; production build, TypeScript and all 141 test files / 4,533 tests pass. Real Chromium shows failed 7B initialization, zero completion calls and one runtime exit, while the ordinary .6B model still loads and chats.
A supported literal Excel read-range-then-set-cell sequence also passes warmed offline restart, exact whole-table Undo/Redo, native Save, independent OOXML CRC/shared-string inspection and separate offline process reopen. This does not certify arbitrary multi-step planning or physical network disconnection.
Signed/parameterized model-URL persistence remains open: both settings and SDK cache metadata can retain original URLs. A storage/cache policy is awaiting user choice. No seven-language model adoption or complete device/offline/privacy acceptance is claimed.
Fixed CPU memory ceiling and higher-precision diagnostic
The exact matched default WASM imports shared memory64 with a maximum of 65,536 pages (4 GiB); compatibility memory32 has the same maximum. The official 7B run requested a single 4,677,120,000-byte model buffer, exceeding the entire configured memory before context/other allocations. This identifies a current artifact ceiling, not insufficient physical host RAM. No larger runtime has been adopted.
A known-source Q4/Q6 Instruct-2507 contrast is frozen with original production messages/schema and matched sampling. The first Q6 download truncated at 841,791,698 bytes and failed full integrity validation; it was never loaded. Explicit HTTP range continuation is in progress after an exact single-byte 206 response preflight. Only a matching final whole-file hash permits browser loading. This diagnostic has no completion or quality-acceptance claim.
Latest diagnostic evidence
Latest completed regressions and GPU model diagnostic
Completed larger GPU screen and current recovery fix
Pinned Qwen2.5-7B loaded on an actual Apple GPU and completed all 21 original-product Word development tasks. Bounded manual review: 9 narrow passes, 11 failures, 1 uncertain. Missing actors/current status, changed recipients/actions/objects, wrong language/date formats and target-language issues prevent default promotion. Native writes and Undo/Redo pass separately; this is not unused-source transfer or native-speaker/full-device acceptance.
Seven generic loading-failure translations now provide accurate retry/model-switch guidance without implying a first-download network context. Related tests (78) and TypeScript pass. The current production browser probe confirms failed CPU allocation cleans up, then the same panel loads the small model and completes chat; a separate small-model control also passes.
Additional GPU diagnostics (042a14a)
Completed fixed-example rewrite contrast (71866d1)
The original diagnostic exited 1 before generation because its Worker interception omitted production isolation headers. Its timeout is preserved. A separately bound correction retains the production response headers and completes all 21 SDK requests with actual pinned-library cache-read hashes and no Worker errors. Bounded manual review: product 4 narrow passes/2 failures/1 uncertain; minimal policy 3/4/0; fixed examples 5/2/0. Fixed examples produce Chinese from an English rewrite and omit the Japanese operating-team actor. No variant passes the seven known-source rewrites; no product default changes or full acceptance are claimed. A read-only cached-weight audit is in progress; its driver is committed, with no completed audit result claimed yet.
Late GPU failure feedback repair (448a9ad / ab22fba)
Actual device-loss testing found a stale loaded note after interrupted streaming. A failing idle-invalidation regression confirms that request finalization alone misses later Worker failure notifications. The provider now forwards its existing failure signal through an optional callback, and the panel updates readiness under controller-generation ownership. Full suite: 141 files/4538 tests pass; types, changed-file lint and production build pass. Current production native test destroys an actual GPUDevice after streaming starts: loaded status clears, Load model is shown, partial text is retained, input unlocks, document stays unchanged and explicit reload/new greeting succeeds. Original failed probes and pre-fix stale-status evidence remain preserved. This does not certify spontaneous loss, CPU replay monitoring or the full device matrix.
The read-only Qwen2.5-7B cache audit also completes: all 88 shard sizes/MD5s match the pinned manifest (4,284,263,424 bytes), process exit 0. This is post-inference current-cache integrity, not exact earlier consumption or semantic quality. Full PR added-content privacy scan passes; seven-language writing acceptance and broader device/compound-operation gates remain open.
Latest pushed evidence (2026-10-07, 509085d)
Latest model validation evidence (2026-10-07)
Locally compiled Qwen2.5 7B and 14B WebGPU libraries loaded on the observed Apple Metal adapter. The 14B structured request timed out with its default configuration; an explicit 2048-token context completed the same request in diagnostic controls. This does not change the product defaults.
The completed Chinese native writing triplet contains two failures (ISO date format changed in rewrite; delivery-team actor omitted in summary) and one narrowly passing translation sample. The broader language screen is still in progress and is not claimed as passing. Physical Windows/mobile coverage and full offline acceptance remain incomplete. This PR remains experimental and [NEEDS_CHANGE].
All added diff content was scanned before this push for the local username and macOS/Windows user-home paths; no matches were found. Local model binaries, toolchains, profiles, and raw scratch logs remain excluded.
Latest bounded-read repair and validation (2026-10-07)
Fixes WebKit cached model Blob failures by assembling worker buffers from sequential reads of at most 8 MiB. Exact requested bytes and failed-subread cleanup are regression-tested. Related Wllama tests: 10 files / 65 tests passed; TypeScript, changed-test lint/format and production build passed. Native artifact pairs remain unchanged and the client patch can be reversed to the original accepted hash.
On isolated Playwright 1.65.0-alpha-2026-10-07 / WebKit 27.2, the actual rebuilt product completed default CPU model loading, online-page closure, offline new-page navigation, cached-model restoration and exact WEBKIT_OFFLINE_OK reply with unchanged document. Offline spelling-script/update request failures remain recorded. This is one desktop page-close lifecycle, not physical Safari/mobile, process restart, save/reopen or complete offline acceptance. A corrected explicit tools-mode run now verifies offline insertion, Undo and Redo; the original chat-mode observer and correction are preserved. Project Playwright dependencies are unchanged.
The completed diagnostic Qwen2.5-14B 21-task screen gives 7 narrow passes, 12 failures and 2 uncertain outputs. A known-source policy-placement contrast does not repair actor/date/summary failures; no model/default is adopted. Authenticated model URL persistence and directory-query propagation gaps are documented and remain unresolved. The PR remains draft and [NEEDS_CHANGE]. Added content was scanned again for local username and user-home paths before push; no matches.
Offline saved-file reopen repair (2026-10-07)
A fresh WebKit offline save/reopen preserved native text but exposed an uncached font-menu sprite and page load errors. This PR now precaches the ten binary locale/density UI sprites (4,448,211 bytes) in the existing vendor-versioned cache. A regression failed before the repair and passed afterward. Full unit suite: 142 files / 4541 tests passed (asynchronous rejection-handling warnings were observed); TypeScript, changed-file lint/format and production build passed.
Fresh rebuilt-product WebKit 27.2 verified all ten cached sprites, default CPU restoration offline, exact insertion in tools mode, native Undo/Redo, DOCX save with ZIP/XML checks, and native reopening with exact text and no page errors. Original failure evidence is retained. This remains a same-context page lifecycle: process restart, physical Safari/mobile, broader operation coverage and seven-language writing fidelity are not certified. Service-worker update and spelling-script offline request failures remain recorded. No model/default or semantic acceptance was changed.
The complete added diff was scanned for personal username and macOS/Windows user-home paths before push. This PR remains draft and [NEEDS_CHANGE].
Cached WebKit browser restart and writing diagnostics (2026-10-07)
WebKit 27.2 now verifies closure/disconnection of the first browser, a new persistent-context launch, offline navigation served by the service worker, default CPU restoration, exact chat, native insertion/history, DOCX save and reopening with exact text and no page errors. A separate read-only probe found existing same-origin OPFS model bytes in a new configured persistent directory before model initialization; this proves cached restoration, not cold-cache isolation. Physical Safari/mobile, PWA installation and the complete offline/device matrix remain incomplete.
Four known-source fact-ledger writing contrasts yielded two narrow passes and two failures. The same narrow improvements also occur with the extra rendition prompt alone, so extraction adds no demonstrated benefit on those cases. Three protected-date rewrite contrasts restore ISO dates exactly, but Korean loses a permission actor present in both controls. No writing pipeline or model/default was adopted; seven-language factual writing remains unresolved. Raw outputs, mechanical request comparisons, limited manual reviews and failed evidence are preserved.
Full added-content scan before push: 950,683 lines, no local username or macOS/Windows user-home path matches. The draft remains [NEEDS_CHANGE].
Latest verification and evaluation (2026-10-07)
connect-src data:. Script and worker policies remain restrictive. Native browser controls verified image reads and blocked data scripts/workers. Both Excel frozen-pane E2E tests passed against the rebuilt product.Archive integrity and update observation repair (2026-10-07)
currentThemeId; this remains under investigation. Latest commits require a fresh remote CI result. Draft NEEDS_CHANGE remains in effect, including multilingual writing and device/privacy acceptance gaps.Theme readiness and local writing candidates (2026-10-07)
Version-query cleanup and activation controls (2026-10-07)
Quoted cell text and desktop WebKit verification (2026-10-07)
Quoted English/Chinese single-cell assignments now constrain the model plan to the exact address, literal and text type; mismatches are rejected before execution. This repairs the observed leading-zero conversion. Unquoted numeric entry retains native parsing. Full unit suite: 143 files / 4,557 tests passed; type/lint, format and production build passed. Existing asynchronously handled rejection warnings remain.
Desktop WebKit 27.2 process restart with offline enabled before navigation now passes CPU chat plus exact XLSX/PPTX edits, Undo/Redo, native Save and offline reopen. Saved XML independently preserves
00123and the PPT literal. These scoped checks do not establish physical/mobile Safari, arbitrary operations, full font/layout coverage or seven-language writing acceptance. Paired Hy-MT2 literal protection still exhibits semantic failures and is not adopted. Original failed receipts and the successful rerun remain hash-bound in the evaluation archive.