fix: index in the background, coalesce streaming, fix toasts, recover truncation (#176, #177, #178, #179) - #180
Merged
Conversation
…the background (#176) Local ONNX inference ran in the Electron main process, so a book-sized import held the event loop for the whole tokenize + forward pass and the window stalled. Ingestion was also a single long IPC call that resolved only after parse, chunk, embed and persist. - add embeddingWorker (worker_threads) with batching, per-batch progress and per-batch retry inside the worker - add WorkerEmbeddingBackend; EmbeddingService selects it in production and keeps the in-process backend for the injected test loader - build the worker as a second main-process entry (electron.vite.config.ts) - add IngestionQueue and route knowledge:add-files / add-folder through it: documents are registered as `pending` and returned immediately, indexing continues over knowledge:index-progress - broadcast index progress to all windows instead of the original sender - reload the document list when a background import completes or fails
… its own turn (#177) A streaming answer wrote to the store on every snapshot, remapped the whole message array, and re-rendered every historical MessageItem — each with a Markdown parse, syntax highlight and KaTeX pass. - hold the latest snapshot per message and commit on a 40ms timer; flush synchronously before an outcome so terminal state is never delayed - keep the per-message sequence counter out of the store and write it only when a gap is actually detected - guard the pending→streaming transition so a chunk that changes nothing does not commit - render live content and reasoning from the message's own turn, and memoize MessageItem so unchanged answers do not re-render
<Toaster /> lived inside NoteEditor, and sonner replays every still-active toast to a late subscriber. Import, save and excerpt toasts fired while no note was open were therefore held and released together the next time a note opened. - mount a single <Toaster /> at the app root and remove it from NoteEditor - give repeatable toasts a stable id so a later call updates the existing toast instead of stacking a duplicate - add a guard test that pins the single mount and the keyed call sites
A reasoning model spends the output budget on thinking first, so the visible answer can end with finishReason 'length'. The notice was correct; there was no way out of it. - add the 截断时自动继续 setting and a continuation bound (default 2); on a truncated attempt the manager continues in place, seeding the next call with the accumulated answer, and stops at the bound - apply the same bound to the manual 继续生成 chain, recorded in the message metadata so it survives a reload - show the truncation notice and the recovery actions as one unit under the answer - extend the model connection with contextWindow, reasoning, reasoningEffort and reasoningBudget, and translate them per protocol; a connection that declares none sends exactly the request it did before
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Four commits, one per issue. They keep the app responsive while it works and give a truncated answer a way out:
#176— local ONNX inference moves off the Electron main process, and imports are registered immediately and indexed in the background.#177— streaming snapshots are coalesced and each message renders from its own turn instead of a per-chunk remap of the whole transcript.#178— one<Toaster />at the app root, with stable ids for repeatable toasts.#179— a truncated answer can continue in place, bounded by a setting, and connections can declare reasoning effort.Why?
#176: importing a book-sized document froze the window; ingestion was also one long IPC call.#177: every streamed chunk wrote to the store and re-rendered every historicalMessageItem(Markdown + highlight + KaTeX).#178: sonner replays active toasts to a late subscriber, so toasts fired with no note open were held and released in a burst.#179: the truncation notice was correct but there was no recovery, and reasoning models spend the output budget on thinking first.Related issue
Fixes #176
Fixes #177
Fixes #178
Fixes #179
What changed?
embeddingWorker(worker_threads) with batching, progress and per-batch retry inside the worker;WorkerEmbeddingBackend;IngestionQueue;knowledge:add-files/add-folderregisterpendingrows and return; progress broadcasts to all windows; the document list reloads when a background import ends.MessageItemis memoized and reads live text from its ownturn.<Toaster />;NoteEditorno longer mounts one; keyed toasts (note-saved:<id>,excerpt:<id>,import-failed:<name>).autoContinueOnTruncation+maxAutoContinueAttempts(default 2); auto-continue in the manager; the manual continue chain shares the bound via message metadata; notice + Continue/Retry rendered as one unit;contextWindow/reasoning/reasoningEffort/reasoningBudgeton the connection with per-protocol translation.How was this tested?
npm run typecheck— passes (node, web, test).npm test— 469 pass / 0 fail.npm run check:design— no violations.npx eslint src test— 0 errors.npx electron-vite build— succeeds;out/main/embeddingWorker.jsis emitted.Not tested here: packaged
worker_threadsloading fromapp.asar(needsnpm run build:unpack && npm run smoke:packaged), real ONNX inference in the worker, the React Profiler claim, and a live provider for the reasoning-effort effect. NoFixesclaim on those.Checklist
npm run typecheckpasses.npm run buildpasses.Desktop / build changes
npm run build:unpackpasses.npm run smoke:packagedpasses.