Skip to content

NetBelt 1.3.1: fix the release workflow failure - #6

Merged
gomanizm merged 8 commits into
mainfrom
fix/v1.3.1-review3
Sep 25, 2026
Merged

gomanizm merged 8 commits into
mainfrom
fix/v1.3.1-review3

Conversation

@gomanizm

Copy link
Copy Markdown
Owner

Fixes what made the release workflow for the v1.3.1 tag fail, so that v1.3.1 can be tagged again.

  • updater.bat now starts Windows PowerShell with its own module path. When NetBelt was started from PowerShell 7, the updater's Windows PowerShell 5.1 inherited PowerShell 7's module path, could not load Get-FileHash, and every update failed.
  • Tests that relied on the OS language, the OS send buffer size, or objects left behind by earlier tests now set up those conditions themselves.

🤖 Generated with Claude Code

gomanizm and others added 8 commits September 25, 2026 16:39
The 1.3.1 release workflow (run 36097480601, windows-latest) stopped at
the test step: every test that runs updater.bat for real failed at [4/6]
with "The term 'Get-FileHash' is not recognized ..." and "ZIP expansion
failed", with a plain TEMP as well as with brackets in it. The runner
executes the test step inside PowerShell 7, which puts its own module
folders at the front of PSModulePath. The powershell that updater.bat
starts (Windows PowerShell 5.1) inherits that, autoloads the 7.x
Microsoft.PowerShell.Utility manifest it finds first, and loses
Get-FileHash, which 5.1 ships as a function in its own psm1 rather than
in the cmdlet DLL both manifests name. Expand-Archive likewise comes from
the first Microsoft.PowerShell.Archive on the path. A user who starts
NetBelt from a PowerShell 7 terminal would see every update fail the
same way.

Reproduced here without PowerShell 7 by putting manifests shaped like
PowerShell 7's in-box ones in front of PSModulePath: 5bdb573's
updater.bat printed the workflow's message word for word, exited 1 and
left the old exe in place. A manifest 5.1 refuses to load
(PowerShellVersion 7.0) stops it the same way.

updater.bat now points PSModulePath at the Windows PowerShell 5.1 system
module folder before its first powershell call, so the hash check, the
expansion and the lock checks always use the in-box modules (the bracket
escaping for Expand-Archive relies on Archive 1.0.1.0, so folders the
user added are left out as well). The caller's value is saved and put
back just before NetBelt is restarted, so the restarted app keeps the
environment it had. The new lines sit after :run, so the byte offsets
that older updaters land on when they overwrite themselves are
unchanged; applying a ZIP with this updater.bat through the published
v1.1.0 / v1.1.1 / v1.2.0 / v1.3.0 updaters gives the same result as
with 5bdb573's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1.3.1 release workflow (windows-latest, English Windows, Python
3.11) failed test_closing_the_tab_while_stalled_stops_the_watcher and
test_reconnecting_after_a_stalled_disconnect_sends_nothing_old on the
precondition "見張りが動いていた" (thread was None). This machine
(Japanese Windows, Python 3.12) passes. _stalled_paste pasted 192KB to a
peer that stopped reading, pumped a fixed 0.6 s, and never checked that
the send had actually stalled. Whether it stalls depends on how much the
OS send buffer absorbs, and Windows sizes that buffer per connection
(dynamic send buffering) unless the socket sets SO_SNDBUF. Measured on
localhost against a non-reading peer with SO_RCVBUF 4096, 512-byte
writes: 73,728 bytes fit by default, 12,288 with SO_SNDBUF 4096 and
1,056,768 with SO_SNDBUF 1MB. Setting NetBelt's Telnet socket to 1MB
right after connect reproduces the CI failure here with the same
message.

Fix the test only (no product change): set SO_SNDBUF 4096 on NetBelt's
socket right after connecting, which turns dynamic send buffering off,
then pump until the drain watcher thread is running (up to 10 s) instead
of a fixed 0.6 s, and assert that as the precondition so a missing
stall fails loudly instead of skipping the check. The peer's receive
buffer stays at SO_RCVBUF 4096 set before listen.

The new test_paste_stall_on_any_send_buffer.py runs the four stalled
paste tests with NetBelt's send buffer forced to 1MB; it failed with the
CI message before this change and passes now. With the product fix
removed (src from b8722f3^ or ae5c3d5^) the adjusted tests still fail
on "相手が詰まっただけでセッションが切れた".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1.3.1 release workflow (windows-latest, English Windows, Python
3.11) failed test_send_returns_while_the_reader_cannot_write_its_replies
on the precondition "応答で送信バッファが埋まらない": within 15 s
NetBelt's socket never stayed unwritable for 0.3 s. This machine passes.
The test relied on 90KB of IAC WONT ECHO replies exceeding NetBelt's
send buffer plus the peer's 4096-byte receive buffer, but Windows sizes
the send buffer per connection (dynamic send buffering) unless the
socket sets SO_SNDBUF. On localhost against a non-reading peer with
SO_RCVBUF 4096, 16KB writes fit 65,536 bytes by default (only 1.4x
below 90KB), 32,768 with SO_SNDBUF 4096 and 1,048,576 with 1MB. Forcing
NetBelt's socket to 1MB right after connect reproduces the CI failure
here with the same message.

Fix the test only (no product change): set SO_SNDBUF 4096 on NetBelt's
socket right after connect, which turns dynamic send buffering off, and
send 100,000 negotiations (300KB of replies) instead of 30,000 so the
margin over what the socket absorbs is about 9x.

The new test_telnet_reply_stall_on_any_send_buffer.py runs the test with
NetBelt's send buffer forced to 1MB; it failed with the CI message
before this change and passes now. With the product fix removed (src
from 9901d8f^) the adjusted test still fails on the GUI thread waiting
3.00 s in send_command.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 1.3.1 release workflow (windows-latest, English Windows, Python
3.11) failed test_unreadable_file_still_refuses_every_host with
"HostKeyStoreError not raised". This machine (cp932) passes, and the
product behaves correctly. Every test in the file assumes that
paramiko's HostKeys.load, which opens the file with the default text
encoding, raises UnicodeDecodeError on the UTF-8 annotation and so takes
NetBelt's own load_known_hosts path. With cp1252 or Python's UTF-8 mode
the annotation decodes, that path is never taken, and the patched
load_known_hosts (PermissionError) is never called. Measured: "ラボ" and
"予備機(待機)" fail only in cp932; "# 検証用ルータ ★重要★" fails in cp932
and cp1252; all decode in UTF-8. So on the runner three other tests
passed without exercising the fix at all.

Fix the test only (no product change): in setUp, replace the open that
paramiko.hostkeys uses with one that opens text without an explicit
encoding as cp932, the Japanese Windows default, and assert right after
writing known_hosts that paramiko's load raises UnicodeDecodeError, so
a broken precondition fails loudly.

The new test_known_hosts_non_ascii_comment_any_locale.py runs the file
with paramiko reading as cp1252 in-process and under python -X utf8 in a
subprocess; both failed with the CI message before this change and pass
now. With the fix removed (src from f990ddc^, or only its except
UnicodeDecodeError fallback removed at HEAD) four tests fail even in
UTF-8 mode, and swallowing errors in the fallback fails
test_unreadable_file_still_refuses_every_host.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test_telnet_negotiation_reply_keeps_reading.py passes on the release
runner, but relies on the same assumption that broke
test_telnet_send_never_waits_for_reply_writer.py there: that the OS
send buffer fills with a 256KB paste or 90KB/300KB of replies. Windows
sizes the send buffer per connection (dynamic send buffering) unless
the socket sets SO_SNDBUF. With NetBelt's Telnet socket forced to 1MB
right after connect, test_a_reply_into_a_full_send_buffer_does_not_stop
_receiving fails on "前提: 送信バッファが埋まる" and
test_piled_up_replies_pause_receiving_and_all_arrive_in_order fails on
"前提: 書き出せない応答が上限まで溜まる" (peak 3 bytes). The queued-reply
test passes, but with a large buffer the reader writes each reply at
once, so the queue it waits for is only non-empty for a moment.

Fix the test only (no product change): in _session, set SO_SNDBUF 4096
on NetBelt's socket right after connect, which turns dynamic send
buffering off; what fits without waiting then stays at 32,768 bytes for
16KB writes and 12,288 for 512-byte writes, several times below the
paste and reply sizes, and replies stay queued once it is full.

The new test_telnet_negotiation_stall_on_any_send_buffer.py runs the
class with NetBelt's send buffer forced to 1MB; it failed with the two
precondition messages before this change and passes now. With the
product fix removed (src from f088505^) the full-buffer test still
fails on "受信が止まった".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test_telnet_paste_to_slow_reader.py passes on the release runner, but
its two stalled-reader tests rely on the assumption that broke
test_paste_to_slow_device_keeps_session.py there (a 192KB paste fills
NetBelt's send buffer), and they checked the drain watcher only under
"if thread is not None:", so without a stall they skipped that check
silently. Windows sizes the send buffer per connection (dynamic send
buffering) unless the socket sets SO_SNDBUF. With NetBelt's Telnet
socket forced to 1MB right after connect, both tests still passed but
the drain watcher thread never ran (it runs once per test with this
machine's default buffer), so stopping it was never checked.

Fix the tests only (no product change): a small _stall helper sets
SO_SNDBUF 4096 on NetBelt's socket (dynamic send buffering off; 512-byte
writes stop at 12,288 bytes) and pumps until the watcher thread runs
(up to 10 s). Both tests then assert the watcher is running as a
precondition and check it stops without the silent condition.

The new test_telnet_slow_reader_stall_on_any_send_buffer.py runs the two
tests with NetBelt's send buffer forced to 1MB and requires that each
one actually started the watcher; it failed with 0 watcher runs for
both before this change and passes now. With the product fix removed
(src from c5abd0f^) the stalled-reader test still fails on
"送信エラー: timed out".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The inspector of the CI fixes found that
test_a_negotiation_reply_waits_for_room_instead_of_disconnecting still
depended on the OS send buffer: with a Windows-grown buffer of 2-8 MB the
slow reader cannot drain it within the test's 20 s and the test fails
(it passed on the current runner only because its buffer stayed small).
Two paste tests (test_a_large_paste_reaches_a_slow_reader_whole_and_in_order,
test_telnet_paste_to_a_slow_reader_arrives_whole) never stalled with a
large buffer, so they did not exercise the back-pressure at all.

Pin SO_SNDBUF to 4096 on NetBelt's socket in all three, as the other
stalled tests now do. With telnet_connection.py from before c5abd0f all
three fail; with the current one they pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In a one-process run of the whole suite about 900,000 objects from earlier
tests stay alive, and a full collection over them stops the event loop. Run
locally, the two longest turns (0.181 s and 0.124 s) were exactly the two
collections (0.152 s and 0.103 s); every other turn took about 0.005 s. The
GitHub runner is slower and failed at 0.27 to 0.29 s. Collect and freeze what
exists before the window is created, and unfreeze afterwards, so the window
and everything the test creates are still collected as usual. The 0.25 s
limit is unchanged, and drawing the whole backlog in one turn still fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@gomanizm
gomanizm merged commit 441ea02 into main Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant