Skip to content

Add LINCS/CCMA constraints and accelerate SHAKE with small-group solvers - #73

Merged
xiaoxuan-yu merged 5 commits into
lab/sidereus-aifrom
feature/constraint-solvers-small-groups
Sep 21, 2026
Merged

xiaoxuan-yu merged 5 commits into
lab/sidereus-aifrom
feature/constraint-solvers-small-groups

Conversation

@yuhaosimba

@yuhaosimba yuhaosimba commented Sep 19, 2026 •

Copy link
Copy Markdown
Member

改动

原 SHAKE 默认执行固定 25 轮迭代;默认含氢键约束通常形成很小的独立连通组,重复 kernel 启动与不必要迭代增加了开销。本 PR 为这些小组提供融合求解路径,并新增 LINCS / CCMA 两种可选求解器。

  • SHAKE、LINCS、CCMA 共用约束分组,每组最多 3 条约束 / 4 个原子。若任一本地组超出限制,整批回退通用路径。
  • SHAKE / CCMA 小组按相对键长 tolerance 自动提前停止,默认 1e-4;默认迭代上限分别 25 / 8。通用路径仍采用固定迭代次数。
  • LINCS 使用 order / corrections 控制计算量,默认 8 / 1;参数说明提供普通步长、大步长和 NVE 的起始配置。
  • CCMA 使用初始几何构造局部稀疏近似逆,不是完整复刻 OpenMM CCMA。
  • 接入初态、局部映射、位置/速度与六分量 virial 修正;支持 CPU / GPU 路径。
  • 公共分组实现归入 constrain.h/.cpp,复用现有矩阵 atomicAdd;报错统一使用 SPONGE 接口。移除开发期最终误差检查、迭代统计及异常转换包装。
  • 修正局部约束数组按约束数分配,更新输入 schema 和简明参数文档。

Benchmark 体系与方法

体系 构成 总原子数 SETTLE 后通用约束数
small 2M2C DNA–COU–Na–水;DNA 884、COU 17、Na 26、水 30,735 原子(10,245 个水) 31,662 172
large 蛋白–RNA–离子–水,使用周期性最小化 restart 189,441 6,415

小体系使用仓库 benchmarks/performance/sinkmeta/statics/dna_cou_sinkmeta/2m2c_* 拓扑与 Pmin_coordinate.txt,去掉偏置和 DNA restraint。两个体系都含大量水;小体系水原子占 97.07%,不是低含水体系。它们是两个不同体系,不能作为同一模型严格扩容的 scaling 曲线。

硬件 RTX 4090 / CUDA 13,单 GPU;周期 NVT、PME、10 Å cutoff、默认 middle-Langevin 参数,默认含氢键约束(质量阈值 3.3 Da)。优化阶段每配置 10,000 步,排除前 1,000 步,输出间隔 1,000 步;分别测试 2 fs / 4 fs。

以下耗时均为 每步约束区间(微秒),包含旧键方向记录、SETTLE 和通用约束求解,不是纯 kernel 时间,也不是完整 MD 吞吐。采用隔离计时构建和 CUDA events;初始化与文件输出不计入。每配置一次轨迹,时间块不是独立重复实验。测试、计划及原始日志按开发约定归档在仓库外,未随本 PR 加入。

性能结果

2 fs:相对原始 SHAKE 的历史优化收益

原始固定 25 次 SHAKE 基线:small 139.48 μs/步,large 305.49 μs/步。只做小组融合、保持固定工作量时,SHAKE 为 72.65 / 109.09 μs,同轮加速 1.92× / 2.80×。

随后参数测试中已测较快的配置如下;这部分用于展示优化过程的收益量级:

方法 / 参数 small μs/步 相对原始 SHAKE large μs/步 相对原始 SHAKE
SHAKE:tolerance=1e-4,上限25 32.80 4.25× 59.28 5.15×
LINCS:order=4,corrections=1 53.52 2.61× 86.49 3.53×
CCMA:tolerance=1e-4,上限8 51.99 2.68× 84.42 3.62×

这是历史不同轮次的参数测试,不是最终 commit 的统一 A/B 复测;此阶段的停止判据和矩阵求解器最终检查随后有所调整,不能把这些倍率宣称为最终版本的精确速度保证。对应本地记录:OPTIMIZATION_REPORT.md、DEFAULT_TOLERANCE_REPORT.md、LINCS_ORDER_REPORT.md。

4 fs:最终停止判据下的参数性能

SHAKE / CCMA 使用最终的相对键长区间判据,不再扣除盒尺寸相关误差预算;三种方法均不做额外最终运行时误差检查。后续提交前的代码整理没有重新进行长程性能测量。LINCS 数值计算未随 tolerance 改变,其两行来自相同阶数/修正次数的不同测试运行。

目标容差 方法 / 参数 small μs/步 large μs/步
1e-4 SHAKE:tolerance=1e-4,上限25 34.75 61.70
1e-4 LINCS:order=4,corrections=2 47.51 76.92
1e-4 CCMA:tolerance=1e-4,上限8 39.73 68.62
1e-5 SHAKE:tolerance=1e-5,上限25 35.74 86.29
1e-5 LINCS:order=4,corrections=2 47.27 76.89
1e-5 CCMA:tolerance=1e-5,上限8 41.22 73.10

没有原始 SHAKE 的同条件 4 fs 基线,因此不将 2 fs 基线用于计算这张表的加速比。这是给定参数的耗时比较,不是相同实际精度下的排名或穷尽搜索得到的全局最优参数。

目标容差不保证输出误差:最终 1e-4 测试中 large SHAKE / CCMA 保存帧最大相对误差约 1.00934e-4 / 1.00001e-4;1e-5 测试三种方法均未满足所有选中键(包括 SETTLE)的严格离线阈值。对应本地记录:TOLERANCE_FIX_REPORT.md、NO_CHECK_4FS_REPORT.md、TOLERANCE_1E5_REPORT.md。

复杂耦合与更严格容差的解释

  • 对本次默认含氢键的小组,1e-4 下优化 SHAKE 最快。
  • 在 4 fs、1e-5 的大体系参数测试中,LINCS / CCMA 比 SHAKE 耗时分别低约 11% / 15%;小体系仍是 SHAKE 更快。这说明收益依赖体系和计算量,但由于实际误差并不相同,不能据此宣称同精度优势。
  • 早期 small DNA 全键约束探索(0.2 fs、2,000 步,尚未完成本 PR 后续优化)测得 SHAKE / LINCS / CCMA 为 208.7 / 131.5 / 60.7 μs,提示矩阵方法在更大耦合网络上可能有价值;该结果来自不同步长、参数和实现阶段,不代表当前版本的标准步长结果。对应本地记录:SMALL_PBC_REPORT.md。
  • 因此,复杂耦合约束或更高精度要求(更小的 tolerance)是 LINCS / CCMA 值得进一步验证的适用方向,而不是本 PR 已证实的普遍性能优势。更大的 tolerance 表示更宽松的精度要求,通常也更有利于 SHAKE / CCMA 提前停止。

验证与边界

  • CUDA 生产构建通过;CPU 约束源文件语法检查通过。
  • 最终整理后 SHAKE / CCMA / LINCS 各 100 步 smoke test 通过,输出有限;LINCS 使用 Andersen 路径覆盖速度投影。
  • 开发阶段 CPU / CUDA 独立双精度非线性参考测试覆盖位置、速度、六分量 virial、PBC、质量差异、小组及通用回退、局部重映射;这些历史测试没有作为最终提交重新全量运行。
  • LINCS / CCMA 的耦合约束必须留在同一 MPI rank;尚不支持跨 rank 耦合求解,也未验证多 rank 接受性。
  • 4 fs 记录仅为 40 ps NVT 测试,未建立长期稳定性、NVE 能量守恒或 NPT 系综正确性。早期长程 LINCS 曾触发保护,短程完成不能替代长期验证。

CI follow-up: Windows H5 runtime memory errors

Commit 282b451 fixes existing memory errors exposed while investigating the win-64 runtime benchmark (0xC0000374). The failures occur in rerun tests rather than constraint integration:

  • SW allocated N floats for an energy buffer used as N+1 (sum plus per-atom energies). CPU storage aliases the host allocation, so the reset wrote past the allocation; GPU copy also read one float past it. Both native and legacy initialization now allocate N+1.
  • Molecule periodicity storage aliased a temporary vector on CPU, leaving a dangling pointer used by trajectory output. It now owns an allocated copy.
  • The synthetic single-element EAM fixture used out-of-range type 1 in legacy/sidecar inputs. The two runtime tests now normalize those inputs to {0,0}, matching their native cases and the existing smoke-matrix treatment; assertions remain unchanged.

Linux AddressSanitizer reproduced the SW heap-buffer-overflow, EAM out-of-bounds read and molecule periodicity use-after-free. After the fixes both test_h5_restart_load_runtime_closure and test_h5_reaxff_edip_runtime_parity pass with ASan (leak checking disabled); ordinary CPU tests and CUDA build also pass. Native Windows confirmation is pending the new CI run. Earlier CI follow-ups also added the required source BOM and linked constrain.cpp into the standalone velocity projection probe.

@yuhaosimba

Copy link
Copy Markdown
Member Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-20T06:32:20.021764Z 48836af Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 48836af5e7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@xiaoxuan-yu
xiaoxuan-yu merged commit 57e291c into lab/sidereus-ai Sep 21, 2026
42 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants