Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions .github/workflows/build-ladybug.yml
Original file line number Diff line number Diff line change
Expand Up @@ -216,6 +216,19 @@ jobs:
${{ matrix.python != 'cp314t' && 'pandas~=2.2 polars~=1.30' || 'pytz' }}
# test_fsm.py fails the same way with upstream's own x86_64 0.19.1 wheel. The deselected
# test_json.py tests INSTALL the json extension from upstream's server, which has no riscv64 build.
#
# 0.20.0 only: test_bytes_param (test_blob_parameter.py) re-executes the same
# parameterized CREATE string through the implicit prepared-statement cache, which
# SIGSEGVs on the second execution (FactorizedTable::clear() derefs a null block
# collection for a write statement's empty-schema result) - ladybugdb/ladybug commit
# d09008e, "Fix SIGSEGV re-executing a parameterized write query string (#862)", not
# released until 0.20.1. test_async_prepare_and_execute_concurrent re-executes a
# parameterized read ($1) concurrently through the same cached-plan fast path and fails
# the same way on 0.20.0 only; ladybugdb/ladybug commit 47443cd, "Fix two state-reuse
# bugs on the cached-physical-plan fast path", lands in the same 0.20.0..0.20.1 range.
# Both are confirmed fixed upstream (0.20.2/0.20.3 build and test clean on this
# fleet) rather than riscv64-specific, so deselect them for 0.20.0 only instead of
# masking them for every version.
CIBW_TEST_COMMAND: >-
python -m pytest -vv tools/python_api/test
--ignore=tools/python_api/test/test_fsm.py
Expand All @@ -224,6 +237,7 @@ jobs:
--deselect=tools/python_api/test/test_json.py::test_get_as_df_json_extract
--deselect=tools/python_api/test/test_json.py::test_get_as_df_json_list
${{ matrix.python == 'cp314t' && '--ignore=tools/python_api/test/test_arrow.py --ignore=tools/python_api/test/test_arrow_memory_backed_table.py --ignore=tools/python_api/test/test_df.py --ignore=tools/python_api/test/test_networkx.py --ignore=tools/python_api/test/test_scan_pandas.py --ignore=tools/python_api/test/test_scan_pandas_pyarrow.py --ignore=tools/python_api/test/test_scan_polars.py --ignore=tools/python_api/test/test_udf.py -k "not test_get_as_df_json"' || '' }}
${{ matrix.version == '0.20.0' && '--deselect=tools/python_api/test/test_blob_parameter.py::test_bytes_param --deselect=tools/python_api/test/test_async_connection.py::test_async_prepare_and_execute_concurrent' || '' }}

- name: Check the extension module and licences made it into the wheel
run: |
Expand Down
7 changes: 7 additions & 0 deletions docs/packages/ladybug.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -23,3 +23,10 @@ versions:
- filename: ladybug-0.19.1-cp314-cp314t-manylinux_2_39_riscv64.whl
sha256: 6f365ce0a83d6fcd6b1abc76e78f34759de07a1a7b823f4f3e4119b2f6382537
requires-python: <3.15,>=3.10
- version: 0.20.0
- version: 0.20.1
- version: 0.20.2
- version: 0.20.3
- version: 0.20.4
- version: 0.21.0
- version: 0.21.1
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Fri, 25 Sep 2026 00:00:00 +0000
Subject: [PATCH] storage: shrink the VMRegion reservation until it fits

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

Every Database reserves its buffer-manager region with one mmap of
max_db_size bytes, and max_db_size defaults to 1 << 43 (8TB). riscv64
Sv39 gives a process 256GB of user address space (the T-Head C910/C920
cores the riscv64 runners use only implement Sv39), so the reservation
fails with ENOMEM and every Database() created with the default
settings throws "Mmap for size 8796093022208 failed." The same happens
under a 39-bit VA arm64 kernel, or on x86-64 under
`ulimit -v 268435456`. Still present in v0.20.4.

On ENOMEM, halve the reservation until it fits. Where the full region
fits nothing changes; elsewhere the database is capped at the largest
power-of-two region the address space can hold, the limit
max_db_size already expresses.
---
diff --git a/src/storage/buffer_manager/vm_region.cpp b/src/storage/buffer_manager/vm_region.cpp
index 429bc61..d102eef 100644
--- a/src/storage/buffer_manager/vm_region.cpp
+++ b/src/storage/buffer_manager/vm_region.cpp
@@ -13,6 +13,8 @@
#else
#include <sys/mman.h>
#include <unistd.h>
+
+#include <cerrno>
#endif

#include "common/assert.h"
@@ -61,6 +63,13 @@ VMRegion::VMRegion(PageSizeClass pageSizeClass, uint64_t maxRegionSize) : numFra
// backed by any file, and its content are initialized to zero.
region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ // The default 8TB region does not fit in a smaller user address space (256GB under riscv64
+ // Sv39, 512GB under a 39-bit VA arm64 kernel), so halve the reservation until it does.
+ while (region == MAP_FAILED && errno == ENOMEM && maxNumFrameGroups > 1) {
+ maxNumFrameGroups /= 2;
+ region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ }
if (region == MAP_FAILED) {
throw BufferManagerException(
"Mmap for size " + std::to_string(getMaxRegionSize()) + " failed.");
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Sat, 27 Sep 2026 00:00:00 +0000
Subject: [PATCH] test: repeat the interrupt until the query stops

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

test_connection_interrupt starts a long query on a thread, sleeps 5s,
calls conn.interrupt() once and expects the thread to end within 100s.
Binding folds each RANGE(1, 1000000) into a million-element list
literal, and ClientContext::executeNoLock() calls resetActiveQuery(),
which clears the interrupted flag, only once compilation is done. On
the riscv64 runners compiling the query takes longer than 5s, so the
interrupt lands during compilation, is wiped, and the query keeps
running. The fixture teardown's close() then waits on it for about 4
hours, and the query thread segfaults freeing its FactorizedTable
after the database is gone. The same loss reproduces on x86-64 with
upstream's 0.19.1 wheel when the sleep is shorter than the compile.
Still present on ladybug-python main.

Re-issue the interrupt every second until the thread ends, within the
same 100s budget. On a fast machine the first interrupt still sticks
and the test behaves as before.
---
diff --git a/tools/python_api/test/test_connection.py b/tools/python_api/test/test_connection.py
index dcc8ee5..aeb2d54 100644
--- a/tools/python_api/test/test_connection.py
+++ b/tools/python_api/test/test_connection.py
@@ -46,6 +46,10 @@ def test_connection_interrupt(conn_db_readwrite: ConnDB) -> None:
execute_thread = threading.Thread(target=run_long_query, args=(conn,))
execute_thread.start()
time.sleep(5)
- conn.interrupt()
- execute_thread.join(timeout=100)
+ # Compiling this query can outlast the sleep on a slow machine, and the interrupt flag is
+ # cleared when execution starts, so an interrupt sent during compilation is lost; repeat it.
+ deadline = time.monotonic() + 100
+ while execute_thread.is_alive() and time.monotonic() < deadline:
+ conn.interrupt()
+ execute_thread.join(timeout=1)
assert not execute_thread.is_alive()
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Sun, 27 Sep 2026 00:00:00 +0000
Subject: [PATCH] extension: report riscv64 as its own platform

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

getArch() starts from "amd64" and only overrides it for x86 and arm64,
so on riscv64 getPlatform() returns "linux_amd64". INSTALL then
downloads the x86-64 build of an extension from
extension.ladybugdb.com into ~/.lbdb/extension/<ver>/linux_amd64/,
and LOAD fails with "cannot open shared object file: No such file or
directory", which is how glibc's dlopen reports an ELF for another
machine. Seen in test_json.py's INSTALL json; LOAD json on the riscv64
runners. Still present on main.

Return "riscv64" there, so INSTALL looks for linux_riscv64 builds
(upstream publishes none yet) and the extension cache is keyed by the
right platform.
---
diff --git a/src/extension/extension.cpp b/src/extension/extension.cpp
index e87004f..e9e2891 100644
--- a/src/extension/extension.cpp
+++ b/src/extension/extension.cpp
@@ -144,6 +144,8 @@ std::string getArch() {
arch = "x86";
#elif defined(__aarch64__) || defined(__ARM_ARCH_ISA_A64)
arch = "arm64";
+#elif defined(__riscv) && __riscv_xlen == 64
+ arch = "riscv64";
#endif
return arch;
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Fri, 25 Sep 2026 00:00:00 +0000
Subject: [PATCH] storage: shrink the VMRegion reservation until it fits

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

Every Database reserves its buffer-manager region with one mmap of
max_db_size bytes, and max_db_size defaults to 1 << 43 (8TB). riscv64
Sv39 gives a process 256GB of user address space (the T-Head C910/C920
cores the riscv64 runners use only implement Sv39), so the reservation
fails with ENOMEM and every Database() created with the default
settings throws "Mmap for size 8796093022208 failed." The same happens
under a 39-bit VA arm64 kernel, or on x86-64 under
`ulimit -v 268435456`. Still present in v0.20.4.

On ENOMEM, halve the reservation until it fits. Where the full region
fits nothing changes; elsewhere the database is capped at the largest
power-of-two region the address space can hold, the limit
max_db_size already expresses.
---
diff --git a/src/storage/buffer_manager/vm_region.cpp b/src/storage/buffer_manager/vm_region.cpp
index 429bc61..d102eef 100644
--- a/src/storage/buffer_manager/vm_region.cpp
+++ b/src/storage/buffer_manager/vm_region.cpp
@@ -13,6 +13,8 @@
#else
#include <sys/mman.h>
#include <unistd.h>
+
+#include <cerrno>
#endif

#include "common/assert.h"
@@ -61,6 +63,13 @@ VMRegion::VMRegion(PageSizeClass pageSizeClass, uint64_t maxRegionSize) : numFra
// backed by any file, and its content are initialized to zero.
region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ // The default 8TB region does not fit in a smaller user address space (256GB under riscv64
+ // Sv39, 512GB under a 39-bit VA arm64 kernel), so halve the reservation until it does.
+ while (region == MAP_FAILED && errno == ENOMEM && maxNumFrameGroups > 1) {
+ maxNumFrameGroups /= 2;
+ region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ }
if (region == MAP_FAILED) {
throw BufferManagerException(
"Mmap for size " + std::to_string(getMaxRegionSize()) + " failed.");
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Sat, 27 Sep 2026 00:00:00 +0000
Subject: [PATCH] test: repeat the interrupt until the query stops

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

test_connection_interrupt starts a long query on a thread, sleeps 5s,
calls conn.interrupt() once and expects the thread to end within 100s.
Binding folds each RANGE(1, 1000000) into a million-element list
literal, and ClientContext::executeNoLock() calls resetActiveQuery(),
which clears the interrupted flag, only once compilation is done. On
the riscv64 runners compiling the query takes longer than 5s, so the
interrupt lands during compilation, is wiped, and the query keeps
running. The fixture teardown's close() then waits on it for about 4
hours, and the query thread segfaults freeing its FactorizedTable
after the database is gone. The same loss reproduces on x86-64 with
upstream's 0.19.1 wheel when the sleep is shorter than the compile.
Still present on ladybug-python main.

Re-issue the interrupt every second until the thread ends, within the
same 100s budget. On a fast machine the first interrupt still sticks
and the test behaves as before.
---
diff --git a/tools/python_api/test/test_connection.py b/tools/python_api/test/test_connection.py
index dcc8ee5..aeb2d54 100644
--- a/tools/python_api/test/test_connection.py
+++ b/tools/python_api/test/test_connection.py
@@ -46,6 +46,10 @@ def test_connection_interrupt(conn_db_readwrite: ConnDB) -> None:
execute_thread = threading.Thread(target=run_long_query, args=(conn,))
execute_thread.start()
time.sleep(5)
- conn.interrupt()
- execute_thread.join(timeout=100)
+ # Compiling this query can outlast the sleep on a slow machine, and the interrupt flag is
+ # cleared when execution starts, so an interrupt sent during compilation is lost; repeat it.
+ deadline = time.monotonic() + 100
+ while execute_thread.is_alive() and time.monotonic() < deadline:
+ conn.interrupt()
+ execute_thread.join(timeout=1)
assert not execute_thread.is_alive()
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Sun, 27 Sep 2026 00:00:00 +0000
Subject: [PATCH] extension: report riscv64 as its own platform

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

getArch() starts from "amd64" and only overrides it for x86 and arm64,
so on riscv64 getPlatform() returns "linux_amd64". INSTALL then
downloads the x86-64 build of an extension from
extension.ladybugdb.com into ~/.lbdb/extension/<ver>/linux_amd64/,
and LOAD fails with "cannot open shared object file: No such file or
directory", which is how glibc's dlopen reports an ELF for another
machine. Seen in test_json.py's INSTALL json; LOAD json on the riscv64
runners. Still present on main.

Return "riscv64" there, so INSTALL looks for linux_riscv64 builds
(upstream publishes none yet) and the extension cache is keyed by the
right platform.
---
diff --git a/src/extension/extension.cpp b/src/extension/extension.cpp
index e87004f..e9e2891 100644
--- a/src/extension/extension.cpp
+++ b/src/extension/extension.cpp
@@ -144,6 +144,8 @@ std::string getArch() {
arch = "x86";
#elif defined(__aarch64__) || defined(__ARM_ARCH_ISA_A64)
arch = "arm64";
+#elif defined(__riscv) && __riscv_xlen == 64
+ arch = "riscv64";
#endif
return arch;
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Fri, 25 Sep 2026 00:00:00 +0000
Subject: [PATCH] storage: shrink the VMRegion reservation until it fits

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

Every Database reserves its buffer-manager region with one mmap of
max_db_size bytes, and max_db_size defaults to 1 << 43 (8TB). riscv64
Sv39 gives a process 256GB of user address space (the T-Head C910/C920
cores the riscv64 runners use only implement Sv39), so the reservation
fails with ENOMEM and every Database() created with the default
settings throws "Mmap for size 8796093022208 failed." The same happens
under a 39-bit VA arm64 kernel, or on x86-64 under
`ulimit -v 268435456`. Still present in v0.20.4.

On ENOMEM, halve the reservation until it fits. Where the full region
fits nothing changes; elsewhere the database is capped at the largest
power-of-two region the address space can hold, the limit
max_db_size already expresses.
---
diff --git a/src/storage/buffer_manager/vm_region.cpp b/src/storage/buffer_manager/vm_region.cpp
index 429bc61..d102eef 100644
--- a/src/storage/buffer_manager/vm_region.cpp
+++ b/src/storage/buffer_manager/vm_region.cpp
@@ -13,6 +13,8 @@
#else
#include <sys/mman.h>
#include <unistd.h>
+
+#include <cerrno>
#endif

#include "common/assert.h"
@@ -61,6 +63,13 @@ VMRegion::VMRegion(PageSizeClass pageSizeClass, uint64_t maxRegionSize) : numFra
// backed by any file, and its content are initialized to zero.
region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ // The default 8TB region does not fit in a smaller user address space (256GB under riscv64
+ // Sv39, 512GB under a 39-bit VA arm64 kernel), so halve the reservation until it does.
+ while (region == MAP_FAILED && errno == ENOMEM && maxNumFrameGroups > 1) {
+ maxNumFrameGroups /= 2;
+ region = static_cast<uint8_t*>(mmap(NULL, getMaxRegionSize(), PROT_READ | PROT_WRITE,
+ MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1 /* fd */, 0 /* offset */));
+ }
if (region == MAP_FAILED) {
throw BufferManagerException(
"Mmap for size " + std::to_string(getMaxRegionSize()) + " failed.");
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
From 0000000000000000000000000000000000000000 Mon Sep 17 00:00:00 2001
From: Ludovic Henry <git@ludovic.dev>
Date: Sat, 27 Sep 2026 00:00:00 +0000
Subject: [PATCH] test: repeat the interrupt until the query stops

Upstream-Status: To upstream [not yet submitted; python-wheels does not open issues/PRs on third-party repos]

test_connection_interrupt starts a long query on a thread, sleeps 5s,
calls conn.interrupt() once and expects the thread to end within 100s.
Binding folds each RANGE(1, 1000000) into a million-element list
literal, and ClientContext::executeNoLock() calls resetActiveQuery(),
which clears the interrupted flag, only once compilation is done. On
the riscv64 runners compiling the query takes longer than 5s, so the
interrupt lands during compilation, is wiped, and the query keeps
running. The fixture teardown's close() then waits on it for about 4
hours, and the query thread segfaults freeing its FactorizedTable
after the database is gone. The same loss reproduces on x86-64 with
upstream's 0.19.1 wheel when the sleep is shorter than the compile.
Still present on ladybug-python main.

Re-issue the interrupt every second until the thread ends, within the
same 100s budget. On a fast machine the first interrupt still sticks
and the test behaves as before.
---
diff --git a/tools/python_api/test/test_connection.py b/tools/python_api/test/test_connection.py
index dcc8ee5..aeb2d54 100644
--- a/tools/python_api/test/test_connection.py
+++ b/tools/python_api/test/test_connection.py
@@ -46,6 +46,10 @@ def test_connection_interrupt(conn_db_readwrite: ConnDB) -> None:
execute_thread = threading.Thread(target=run_long_query, args=(conn,))
execute_thread.start()
time.sleep(5)
- conn.interrupt()
- execute_thread.join(timeout=100)
+ # Compiling this query can outlast the sleep on a slow machine, and the interrupt flag is
+ # cleared when execution starts, so an interrupt sent during compilation is lost; repeat it.
+ deadline = time.monotonic() + 100
+ while execute_thread.is_alive() and time.monotonic() < deadline:
+ conn.interrupt()
+ execute_thread.join(timeout=1)
assert not execute_thread.is_alive()
Loading
Loading