Skip to content

SevenNet+D3 with TorchSim batched MD can produce NaNs while ASE remains finite #316

Description

@sek1ro-yuzzz

Summary

I am seeing reproducible NaNs when running SevenNet v0.13.0 + D3(BJ)-PBE through the TorchSim batched-MD path. The same structures and MD settings remain finite when run through the ASE SevenNet+D3 path.

I also opened a TorchSim-side issue here:

TorchSim/torch-sim#579

I am reporting this here as well in case the SevenNet TorchSim wrapper, accelerator configuration, or SevenNet+D3 combination has known caveats or recommended diagnostics.

Setup

The TorchSim call follows the official batched-MD style:

ts.integrate(system=list_of_atoms, ...)

where list_of_atoms contains 16 independent ASE Atoms objects.

System:

  • 16 independent amorphous Na16Ta16Cl96 structures
  • 128 atoms per structure
  • Initial structures look reasonable; initial nearest-neighbor distances are about 2.26 Angstrom
  • SevenNet OMNI checkpoint
  • modal="omat24"
  • D3(BJ)-PBE enabled

MD settings:

  • Fixed-cell NVT
  • Nose-Hoover chain
  • 300 K
  • 2 fs timestep
  • 200 fs thermostat damping
  • 200 warmup steps
  • 1000 measured MD steps

Environment

Primary test:

  • GPU: NVIDIA RTX 5090
  • sevenn 0.13.0
  • torch 2.11.0+cu128
  • torch-sim-atomistic 0.6.0
  • ase 3.28.0

Repeat test:

  • GPU: NVIDIA RTX 4090
  • Same package versions: sevenn 0.13.0, torch 2.11.0+cu128, torch-sim-atomistic 0.6.0, ase 3.28.0

Observed behavior

On the RTX 5090 machine, for the same batch of 16 structures:

  • TorchSim + SevenNet+D3 + cueq: MD completes, but some batch elements end with NaN energy/temperature.
  • TorchSim + SevenNet+D3 + flash: MD completes, but some batch elements end with NaN energy/temperature.
  • ASE + SevenNet+D3 + flash: all 16 structures remain finite under the same structures, potential, and MD settings.

On the RTX 4090 machine, repeating TorchSim + SevenNet+D3 + cueq also produces NaN energy/temperature in some batch elements. This suggests the issue is not specific to a single RTX 5090/Blackwell machine.

Static single-point comparison

For the same structure, TorchSim SevenNet+D3 and ASE SevenNet+D3 agree closely in static single-point calculations:

  • Energy difference: about 0.000679 meV/atom
  • Max force-component difference: about 1.62e-5 eV/Angstrom
  • Max stress-component difference: about 1.13e-8 eV/Angstrom^3

So the problem appears more likely to be in the MD / batched integration path than in the static SevenNet+D3 energy/force evaluation itself.

Question

Do you have any suggestions for diagnosing this from the SevenNet side? In particular:

  • Are there known limitations for SevenNet v0.13.0 with TorchSim batched MD and D3?
  • Are cueq / flash expected to be numerically equivalent for MD use here?
  • Is there a recommended minimal diagnostic to separate a SevenNet wrapper issue from a TorchSim integration/state-update issue?

I can provide the benchmark config, structures, or a smaller reproducer if that would be useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions