Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 63 additions & 0 deletions docs/adr/0002-library-source-reuse.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# ADR 0002: Library sources and notebook memberships (#99)

Status: accepted for implementation, recorded before code changes.

## Schema and identity

Add `library_sources` as the durable library identity and snapshot (canonical text,
parser structure, original URI and local file). Add nullable `documents.sourceId`
referencing it, unique within a notebook. Existing `documents` rows remain notebook
memberships and retain their IDs: historical citations and retrieval scopes must
not change during migration. Existing imports are backfilled one-to-one (do not
merge unrelated sources merely because titles or paths match).

The existing document shape is a compatibility projection of the snapshot plus
membership-local indexing state. Keep the current text/structure fields as cached
projections for the existing retrieval/readers; reuse performs no parsing or file
copy. A library snapshot is immutable while shared. Explicit file refresh creates
a new library snapshot for that membership (copy-on-write); it never changes the
text or page offsets underneath another notebook's citations. Reindex uses the
persisted snapshot, not the original mutable file.

## Embedding-space rule

Chunks, chunk/block mappings, ingestion attempts, indexing status, and vectors
remain per membership/notebook. Citation document IDs identify the membership;
its `sourceId` identifies the shared library snapshot. This preserves every
existing citation's notebook context and page/span contract.

Reuse is an explicit operation. If a donor membership is indexed and its stored
space identity and vector width match the target notebook (and the configured
embedding space), copy its chunks, provenance and vectors with new membership-
local IDs. Never call the embedding provider in the reuse operation. An empty
target may adopt the matching donor space. If there is no compatible indexed
donor, add a pending membership and show that indexing is required; only the
user's explicit reindex action may produce embeddings. Same dimensions alone
are not proof of compatibility. Copying never invalidates other target sources.

A remote embedding connection cannot report its vector width before the first
embedding, so its space id is dimension-agnostic while `notebook_embedding_spaces.dimensions`
carries the width measured during indexing. Reuse therefore compares the exact
space id and, when the current width is known, that width; the donor's persisted
width is separately checked against its own vector table and against every chunk's
embedding row and vector row, and a vector copy that does not affect exactly one
row fails the attach instead of producing a half-indexed membership.

## Removal and lifecycle

Removing a source removes only that notebook's membership and derived index.
Keep the library snapshot and file, including after the last membership is
removed; it remains available for later reuse. Permanent library deletion is a
separate, explicitly confirmed operation and is refused while memberships exist.
Notebook deletion follows the same detach-only semantics. A notebook must never
unlink a file still used by a library source.

## UI and validation

Add “From library” to the existing source-import UI, list snapshots not already
attached, and distinguish ready-to-reuse sources from those needing explicit
indexing. Add confirmed permanent deletion for unused library sources. Provide
both English and Chinese strings. Regression tests cover legacy migration,
two-notebook reuse, no embedding calls on attach, embedding-space mismatch,
page/span preservation, independent removal, notebook deletion and confirmed
last-copy deletion. No automatic deduplication of new imports is introduced.
17 changes: 15 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,13 +153,15 @@ A document is an identity. Everything derived from it can be rebuilt:
```text
PDF / DOCX / URL / Note
↓
documents ← source identity (id, content, structure, localFilePath)
library_sources ← durable library snapshot (#99): content, structure, file
↓
documents ← per-notebook membership: id, status, cached projection
↓
document_blocks ← derived: page / heading level / bbox / char span
↓
chunks + chunk_blocks ← derived: char offsets anchored to documents.content
↓
embeddings + vectors ← derived
embeddings + vectors ← derived, per notebook
```

Two invariants hold across the whole chain:
Expand All @@ -172,6 +174,17 @@ Two invariants hold across the whole chain:
structure is persisted on the source row so a rebuild does not need to re-parse a
file that may have changed on disk.

`#99` splits identity in two. `library_sources` is the library-level snapshot — the
canonical text, its parse structure and the library-owned copy of the original file —
and `documents` is one notebook's membership of it, which keeps its own id and derived
index. Attaching an existing snapshot to a second notebook copies a compatible donor's
chunks, provenance and vectors with new membership-local ids instead of re-parsing or
re-embedding; when the embedding space does not match, the membership lands `pending`
and only an explicit re-index may embed it. Removing a membership never deletes the
library snapshot or its file. A shared snapshot is immutable: refresh is copy-on-write,
so one notebook's file change cannot move another's page and span citations. See
[`adr/0002-library-source-reuse.md`](adr/0002-library-source-reuse.md).

## Retrieval seam

`src/main/services/retrieval/` is the boundary between "how do we search" and
Expand Down
25 changes: 25 additions & 0 deletions src/main/db/migrations/0023_sleepy_luckman.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
CREATE TABLE `library_sources` (
`id` text PRIMARY KEY NOT NULL,
`title` text NOT NULL,
`type` text NOT NULL,
`source_uri` text,
`local_file_path` text,
`content` text,
`structure` text,
`content_hash` text,
`mime_type` text,
`file_size` integer,
`metadata` text,
`created_at` integer NOT NULL,
`updated_at` integer NOT NULL
);
--> statement-breakpoint
ALTER TABLE `documents` ADD `source_id` text REFERENCES library_sources(id);--> statement-breakpoint
CREATE UNIQUE INDEX `idx_documents_notebook_source` ON `documents` (`notebook_id`,`source_id`);--> statement-breakpoint
-- 数据迁移(#99):把现有 documents 一次性回填成一比一的 library_sources。
-- 不按标题/路径合并不同来源;id 取 'lib_' || documents.id,保证一一对应、结果确定。
INSERT INTO `library_sources` (`id`, `title`, `type`, `source_uri`, `local_file_path`, `content`, `structure`, `content_hash`, `mime_type`, `file_size`, `metadata`, `created_at`, `updated_at`)
SELECT 'lib_' || `id`, `title`, `type`, `source_uri`, `local_file_path`, `content`, `structure`, `content_hash`, `mime_type`, `file_size`, `metadata`, `created_at`, `updated_at`
FROM `documents`;--> statement-breakpoint
-- 每个 membership 指向它自己的 snapshot;此后重新索引读的就是这份快照。
UPDATE `documents` SET `source_id` = 'lib_' || `id` WHERE `source_id` IS NULL;
Loading
Loading