Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 26 additions & 1 deletion .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,21 @@ on:
permissions:
contents: read
jobs:
source-snapshot:
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@v4
- run: git archive --format=zip --output=laas-ci-source.zip HEAD
- uses: actions/upload-artifact@v4
with:
name: laas-ci-source-${{ github.sha }}
path: laas-ci-source.zip
retention-days: 3
tests:
timeout-minutes: 15
strategy:
fail-fast: false
matrix:
os: [windows-latest, ubuntu-latest]
python: ['3.11', '3.13']
Expand All @@ -18,7 +31,19 @@ jobs:
with:
python-version: ${{ matrix.python }}
- run: python -m pip install -r requirements-dev.txt
- run: python -m pytest -q
- name: Run tests with bounded diagnostics
timeout-minutes: 8
env:
PYTHONUNBUFFERED: "1"
run: python -X faulthandler -m pytest -vv --durations=20 -o faulthandler_timeout=45 --junitxml=runtime/ci-results.xml
- name: Preserve test diagnostics
if: always()
uses: actions/upload-artifact@v4
with:
name: test-results-${{ matrix.os }}-${{ matrix.python }}
path: runtime/ci-results.xml
if-no-files-found: ignore
retention-days: 7
windows-artifact:
runs-on: windows-latest
needs: tests
Expand Down
8 changes: 7 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
__pycache__/
__pycache__/
*.py[cod]
*$py.class
*.log
Expand All @@ -24,3 +24,9 @@ secrets.*.yaml
*.pfx
/Мои хотелки к первой версии.txt
/Предложения для второй версии от чатгпт.txt

# Private managed-conversation journals (including SQLite sidecars).
*.sqlite3
*.sqlite3-wal
*.sqlite3-shm
*.sqlite3-journal
157 changes: 157 additions & 0 deletions docs/COORDINATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
# Managed coordination: durable conversation first

Status: **M1 implemented; M2a adds an opt-in Qwen outbox/client and external Tool
Guard v1 provider with synthetic HTTP tests. Native Desktop/Telegram ingress and
live result observation are not wired yet.** This does not change the installed
alpha.1 EXE. See [Qwen M2a implementation](QWEN_MANAGED_COORDINATION.md).

## Why a summary is not the source of truth

The user's task evolves during a conversation. A model may omit a new restriction
from `PROJECT_STATE.md`, misunderstand it, or disappear before writing any handoff.
Therefore an LLM is never responsible for persisting user input. The input boundary
must write each original message/edit and its attachments before returning a
successful receipt to the frontend, queue or agent. Failure to commit is a failed
send, not permission to continue without a journal.

`src/coordination/JournalStore` provides this boundary. SQLite is local to the
coordinator PC, with WAL, FULL synchronous writes, transactions and foreign keys.
This protects acknowledged data against process failure within SQLite/filesystem
semantics, not against storage hardware failure, malicious local access or loss of
the entire machine. Backups remain necessary. Do not place a live database on SMB.

A future integration opens `data_dir()/coordination/journal.sqlite3` explicitly.
Importing the package creates nothing. The library uses only Python's standard
library, starts no processes and contacts no model servers. Runtime databases,
attachments, conversations and local raw logs must never be committed to Git.

## Implemented contract

1. `open_project`: register an explicit workspace. A second project ID cannot bind
the same normalized workspace and bypass the project-level single-writer check.
2. `append_user`: exact UTF-8 text, stable source/message ID and optional stored
attachment digests. Return is a durable receipt, including event ID and revision.
Repeated delivery of the same source ID/content returns the same event. Reusing
an ID with different content is an error. Retries must keep the original ID.
3. Edits append a new `user.edit` pointing to the old event ID; no overwrite. Every
human message conservatively advances the project revision, even a clarification
or acknowledgement. No model classifies a message as 'unimportant' for storage.
4. `put_attachment`: content-addressed immutable bytes + media type, stored in the
same database; a screenshot is not merely a path to a disappearing temporary file.
Limits: 20 MiB per attachment, 16 digests per user event, 1 MiB per event payload.
Unsupported size is an explicit error; never silently truncate an input.
5. `record_runtime_event`: exact assistant chunks/messages, tool output and notices
from trusted adapters. These remain evidence, not user commands. Late output from
revoked attempts can be retained without restoring the old attempt's authority.
6. `create_task` and `start_attempt`: each task has a monotonic attempt epoch and
expiring lease. Explicit replacement or expiry revokes the old attempt. Writers
are exclusive within a project; multiple read-only attempts are permitted.
7. `prepare_delivery`: a persisted/replayable packet contains **all original user
events**, including edits, plus the current task's history. A model-authored
summary cannot replace them. The adapter supplies a local tokenizer/budget
counter which must account for the actual chat template, tools and images.
Overflow raises `ContextOverflow` instead of truncating or silently compressing.
8. `acknowledge_delivery`: the trusted transport acknowledges a particular packet
ID/digest/revision. An input arriving in transit invalidates the old receipt.
This is evidence of delivery, **not proof of semantic comprehension**. It is not
a tool which the LLM can call to claim it has read its own instructions.
9. `begin_action`: before a tool is launched, require the current attempt, unexpired
lease, delivered current revision, appropriate write permission and no unresolved
previous action. Persist the intent first. Duplicate tool IDs raise an explicit
conflict rather than running again. The adapter, not the model, classifies tools.
10. `finish_action`: preserve output even if the model/lease/revision became stale.
Such output is `needs_review`, never automatic success for the new attempt.
11. `cancel_attempt`: revoke future admissions. This does NOT terminate an OS process
and does NOT remove uncertain running actions. Reassignment waits for an operator
to verify process/files and call `reconcile_action` with a note/current revision.
12. `finish_attempt`: require the current revision and resolved tools, then propose
`awaiting_review`. A model's 'done' is not a claim that tests or requirements pass.
13. `record_summary`: optional derived text with an explicit source high-water mark.
It never advances a user's revision or an attempt's delivery acknowledgement.
14. `events` / `task_status`: ordered audit stream and task state for future UI.

Example: user says 'use SQLite', later 'keep the existing schema'. The second message
is committed as revision 2 even if the model writes a summary mentioning only SQLite.
An attempt which saw revision 1 cannot start another tool. After failover, the new
attempt's delivery packet still contains both exact messages. A delayed tool request
from the old attempt is refused. If an old tool was already running, automatic
replay stops for reconciliation rather than executing its side effect twice.

## Integration boundaries: do not overclaim

- This is not a new coding harness or a chat UI. External agents still perform
reasoning and tools. Existing Station model/GPU/frontend profiles stay independent.
- The current `QwenCodeAdapter` reports `task_control=False`. The new M2a explicit ingress/Guard APIs
do not install a native GUI/steering hook, stream subscription or live tool gate.
Opening the existing Qwen Desktop therefore does NOT enable these guarantees.
- A model-API proxy alone is insufficient: it sees model requests, not necessarily a
user edit/queued message immediately when entered in the agent UI. The input adapter
must observe the authenticated human-message boundary *before* dispatch/acknowledge.
- If that hook is unavailable in an installed runtime, mark capture as observational
or unsupported and disable automatic authority transfer; do not invent an endpoint,
edit its private session database or ask the model to remember to save messages.
- The library refuses stale **admission**. It does not sandbox arbitrary file access,
prevent an external agent bypassing it, or undo a command already executing when
a correction arrives. A future local tool executor must check current admission,
enforce workspace/approval rules, track owned processes and stop/drain safely.
- Read-only attempts must not receive an unrestricted shell labelled 'read-only'.
A source role ('human' vs assistant) comes from a trusted adapter, never from text.
- SQLite and the API are not an authorization boundary against malicious code running
as the same user. Keep DB/private files in the user's data directory with suitable
permissions; no network control port, secrets headers or tokens in audit metadata.
- End-to-end exactly-once model/tool execution is not claimed. Packets are replayable,
duplicate ingress is idempotent; ambiguous external actions block automatic retry.
- SQLite leases use an injected UTC clock for persistence. Significant clock rollback
needs explicit reconciliation; a process-restart scheduler must not trust old OS
process ownership based only on a PID or assume a lease means a process was killed.
- Project-level revision invalidation is deliberately conservative. Later task-scoped
routing must not silently omit project-wide changes. A new smaller-context leader
currently receives all raw inputs or fails on overflow; selective source-linked
handoff and user-approved scope reduction are later work, not a hidden lossy fallback.
- Backup via SQLite's backup API or a closed database; copying only a live .sqlite3
without its WAL can omit committed records. Retention, encryption at rest and
explicit user deletion/export UX are follow-up work.

## Verification

Run from source:

```text
python -m pytest -q tests/test_coordination_journal.py
python tools/coordination_smoke.py
```

The smoke creates only temporary synthetic data. It demonstrates an omitted
constraint in a model summary, a newer human message, failover and stale-attempt
rejection. It does not start a model or modify the installed Station.

Tests cover retries/conflicts, exact Unicode and edits, attachments, transaction
rollback, concurrent ingress, one writer, lease expiry, immutable-record triggers,
replay after restart, context overflow, new input during delivery and during a tool,
late results, cancellation, operator reconciliation and untrusted runtime evidence.

## Next milestones (not implemented by M1)

**M2b: finish the real Qwen runtime adapter, fail-closed.** M2a outbox and external
Guard provider are implemented; native ingress/output proof remains. Pin/probe installed protocol;
intercept every new/edited/queued human input, store it before acknowledgement,
subscribe to output with persistent source IDs/cursors, and gate every tool through
managed admission. Preserve raw input independently of model summaries. Expose
capture coverage and delivered revision in Station. Unsupported channels stay
explicitly unprotected; no automatic failover there.

**M3: deterministic node/task dispatcher.** Reuse Station model/service registries;
resource-pool slots (not one worker per alias), enabled/paused/ready/loading/busy/offline,
allowed trust boundary per project, one leader, priority and backoff, task-scoped
handoff with provenance, no silent cloud fallback or context shrink, process-aware
cancellation and reconciliation. Do not modify GPU modes or remote services merely
because a model becomes unavailable.

**M4: optional Station controls.** One project/task view with raw conversation,
revisions, leader and worker sessions, queued user corrections, source links,
conflicts and explicit approvals. Use official agent sessions, not a second coding
agent. Requirements edits remain raw events regardless of UI or selected model.

First production acceptance: safe fault injection in a disposable repository,
user edits during generation/tool execution, delayed output after failover and
application restart. No repeat of the multi-model GPU benchmark is required.
38 changes: 38 additions & 0 deletions docs/HANDOFF.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,44 @@
Updated: 2026-09-13. **3.0.0-alpha.1 released**.
Read AGENTS.md, then this file and ignored handoff-local/README.md when available.

## Current work — M2b1 transport hardening, 2026-09-15

The durable journal, Qwen outbox/Guard and SSE observer exist, but live native
input capture, full-packet delivery and owned-process drain are still NOT wired.
`task_control=False`, `automatic_failover=False`, `live_runtime_verified=False`.
Existing Desktop/Telegram conversations are not automatically protected.

This checkpoint fixes reproducible HTTP lifetime problems before M2b2: one owned
exchange deadline now covers capabilities preflight, response headers and streaming
body; explicit response/socket cleanup includes partial and Connection: close bodies.
Cancellation cannot hide concurrent journal/protocol failures or turn uncertain
prompt POSTs into retries. See HTTP_TRANSPORT_LIFETIME.md for precise boundaries.
No new dependency, daemon startup, remote-model access or production change.

Validation: 34 new loopback/SQLite tests; full coordination 254 passed locally.
All three synthetic coordination smokes PASS. Broader non-GUI regression is run
with customtkinter-dependent tests excluded; see VALIDATION.md for exact totals.
The archived fda25d3 source was overlaid with connected-repository ab238e9 files;
Git blob hashes for all affected predecessor modules/tests were checked.

CI blocker: ab238e9 run 34897467428 passed Linux but both Windows jobs exceeded
six hours and were cancelled. Logs have only progress dots, no exact hung test.
The exact Windows root cause remains unconfirmed; do not claim these fixes prove it.
Diagnostics commit d3ebf9c adds verbose test names, stack dumps and strict deadlines.
Runs 34927013300 and 34927010021 failed before runner assignment (empty steps,
runner_id=0); the available API does not establish why. No tests ran in those jobs.
Do not repeatedly rerun or alter repository billing/security to bypass this.

Next: obtain one bounded Windows/Linux run of this transport fix when runners are
available; inspect named test/stack on any hang. Then qualify an installed Qwen in
a disposable project for M2b2: native new/edit/queue/steer capture BEFORE forwarding,
full packet receipt/token budget, foreground tool ownership and queue/process drain.
Do not infer delivery from HTTP 202, replay_complete, model text or turn_complete.
No M3 automatic scheduling until the complete boundary is accepted.

PR #1 remains the review boundary. No main, release, installed EXE, GPU/model,
startup, user setting or remote service was changed.

## Current state

The first early release is published at
Expand Down
74 changes: 74 additions & 0 deletions docs/HTTP_TRANSPORT_LIFETIME.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Bounded Qwen HTTP lifetime — M2b1 hardening

Status: implemented and tested with local loopback HTTP/SQLite fixtures on Linux.
Not an installed Qwen Desktop, Windows driver or remote-model qualification.
`task_control`, `automatic_failover` and `live_runtime_verified` remain false.

## Reproducible defects addressed

1. The SSE deadline previously started after `/capabilities`. A slow capabilities
response ignored the caller's stop event and the stated observation budget.
2. Stopping at an event limit closed `HTTPConnection`, not the retained partial
`HTTPResponse`. A response owns a socket file separately, including when the
connection detaches its socket for `Connection: close`.
3. A stop/deadline racing a SQLite or protocol exception could convert that error
into a normal return. Private storage I/O failures must remain failures as well.

The new private `http_lifetime.bounded_response` owns one connection and response.
One absolute deadline includes connect, headers and the body consumed by its caller.
A caller-owned stop event is checked before connect and after socket attachment.
A watcher only shuts down the owned socket; the request thread closes the response
and connection and joins the watcher. It never closes a buffered reader from a
second thread, executes a tool, kills another process or logs a token/payload.

The SSE receiver shares this budget with the capabilities preflight and limits
reconnect backoff to the remaining time. Only transport interruption is a normal
stop. Event projection/storage errors latch `resync_required` and propagate, even
when cancellation arrives concurrently. No cursor is committed for an incomplete
frame. Prompt POSTs still have at most one attempt; ambiguous admission blocks the
outbox rather than replaying side effects.

A socket read timeout and a wall-clock deadline are different: regular trickle bytes
can prevent a read timeout but do not extend the exchange deadline. Connect remains
bounded by the socket timeout; a stop before a socket exists is rechecked immediately
once connect returns. This network budget is not a guarantee that arbitrary storage
or application code can be preempted. SQLite/application failure is not swallowed.

## Tests and scope

34 new tests cover cancellation/deadlines in capabilities headers/body/trickle,
SSE headers/idle/partial frames/heartbeats, response cleanup including detached
connections, pre-attachment cancellation, non-retried uncertain POSTs, bounded
backoff, invalid budgets, watcher cleanup and concurrent storage/protocol errors.

Commands:

```text
python -m pytest -q tests/test_coordination_http_lifetime.py
python -m pytest -q tests/test_coordination_journal.py tests/test_coordination_qwen.py tests/test_coordination_qwen_events.py tests/test_coordination_http_lifetime.py
```

Final local results: 254 coordination tests passed; the 34-test lifetime suite
passed five additional complete repetitions. All three existing synthetic smokes
passed. Broader non-GUI run: 402 passed, 1 skipped, 1 deselected, with the two
customtkinter-dependent modules excluded. Full GUI collection cannot run in this
container because that dependency is absent; no stub or fake PASS was substituted.
A selected regression subset failed against the exact previous transport files,
then passed with the fix. No actual agent/server/user environment was contacted.

## Windows CI is still a separate acceptance gate

The previous ab238e9 run 34897467428 passed Linux. Both Windows jobs exceeded the
six-hour Actions limit; their logs contain progress dots but no exact stuck test or
stack. These transport fixes must not be advertised as a proven root-cause fix for
that particular Windows hang without a new bounded Windows run.

Diagnostics commit d3ebf9c adds `-vv`, Python faulthandler stack dumps, test-step
and job deadlines, and retained JUnit results. Follow-up runs 34927010021 and
34927013300 failed before runner assignment (empty steps, runner_id=0). The
available API does not establish the infrastructure/account cause. Do not change
billing/security or repeatedly retry jobs to hide this limitation.

Next: one bounded Windows/Linux run when runners are available, then M2b2 native
human ingress, actual full-packet delivery and owned-process draining. Automatic
handoff is not enabled by merely hardening this transport.
Loading
Loading