Repository navigation
Add LINCS/CCMA constraints and accelerate SHAKE with small-group solvers - #73
Merged
xiaoxuan-yu merged 5 commits intoSep 21, 2026
Merged
Conversation
Member
Author
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Codex Review: Didn't find any major issues. Swish! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
xiaoxuan-yu
approved these changes
Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
改动
原 SHAKE 默认执行固定 25 轮迭代;默认含氢键约束通常形成很小的独立连通组,重复 kernel 启动与不必要迭代增加了开销。本 PR 为这些小组提供融合求解路径,并新增 LINCS / CCMA 两种可选求解器。
tolerance自动提前停止,默认1e-4;默认迭代上限分别 25 / 8。通用路径仍采用固定迭代次数。order/corrections控制计算量,默认 8 / 1;参数说明提供普通步长、大步长和 NVE 的起始配置。constrain.h/.cpp,复用现有矩阵atomicAdd;报错统一使用 SPONGE 接口。移除开发期最终误差检查、迭代统计及异常转换包装。Benchmark 体系与方法
小体系使用仓库
benchmarks/performance/sinkmeta/statics/dna_cou_sinkmeta/2m2c_*拓扑与Pmin_coordinate.txt,去掉偏置和 DNA restraint。两个体系都含大量水;小体系水原子占 97.07%,不是低含水体系。它们是两个不同体系,不能作为同一模型严格扩容的 scaling 曲线。硬件 RTX 4090 / CUDA 13,单 GPU;周期 NVT、PME、10 Å cutoff、默认 middle-Langevin 参数,默认含氢键约束(质量阈值 3.3 Da)。优化阶段每配置 10,000 步,排除前 1,000 步,输出间隔 1,000 步;分别测试 2 fs / 4 fs。
以下耗时均为 每步约束区间(微秒),包含旧键方向记录、SETTLE 和通用约束求解,不是纯 kernel 时间,也不是完整 MD 吞吐。采用隔离计时构建和 CUDA events;初始化与文件输出不计入。每配置一次轨迹,时间块不是独立重复实验。测试、计划及原始日志按开发约定归档在仓库外,未随本 PR 加入。
性能结果
2 fs:相对原始 SHAKE 的历史优化收益
原始固定 25 次 SHAKE 基线:small 139.48 μs/步,large 305.49 μs/步。只做小组融合、保持固定工作量时,SHAKE 为 72.65 / 109.09 μs,同轮加速 1.92× / 2.80×。
随后参数测试中已测较快的配置如下;这部分用于展示优化过程的收益量级:
这是历史不同轮次的参数测试,不是最终 commit 的统一 A/B 复测;此阶段的停止判据和矩阵求解器最终检查随后有所调整,不能把这些倍率宣称为最终版本的精确速度保证。对应本地记录:
OPTIMIZATION_REPORT.md、DEFAULT_TOLERANCE_REPORT.md、LINCS_ORDER_REPORT.md。4 fs:最终停止判据下的参数性能
SHAKE / CCMA 使用最终的相对键长区间判据,不再扣除盒尺寸相关误差预算;三种方法均不做额外最终运行时误差检查。后续提交前的代码整理没有重新进行长程性能测量。LINCS 数值计算未随 tolerance 改变,其两行来自相同阶数/修正次数的不同测试运行。
没有原始 SHAKE 的同条件 4 fs 基线,因此不将 2 fs 基线用于计算这张表的加速比。这是给定参数的耗时比较,不是相同实际精度下的排名或穷尽搜索得到的全局最优参数。
目标容差不保证输出误差:最终 1e-4 测试中 large SHAKE / CCMA 保存帧最大相对误差约 1.00934e-4 / 1.00001e-4;1e-5 测试三种方法均未满足所有选中键(包括 SETTLE)的严格离线阈值。对应本地记录:
TOLERANCE_FIX_REPORT.md、NO_CHECK_4FS_REPORT.md、TOLERANCE_1E5_REPORT.md。复杂耦合与更严格容差的解释
SMALL_PBC_REPORT.md。验证与边界
CI follow-up: Windows H5 runtime memory errors
Commit
282b451fixes existing memory errors exposed while investigating the win-64 runtime benchmark (0xC0000374). The failures occur in rerun tests rather than constraint integration:Linux AddressSanitizer reproduced the SW heap-buffer-overflow, EAM out-of-bounds read and molecule periodicity use-after-free. After the fixes both
test_h5_restart_load_runtime_closureandtest_h5_reaxff_edip_runtime_paritypass with ASan (leak checking disabled); ordinary CPU tests and CUDA build also pass. Native Windows confirmation is pending the new CI run. Earlier CI follow-ups also added the required source BOM and linked constrain.cpp into the standalone velocity projection probe.