Skip to content

Uni-HamGNN: clarifications on training-set composition and the OpenMX label protocol (for accurate description and consistent reference calculations) #114

Description

@hjunxi22

Hi, and thank you for releasing Uni-HamGNN. We are using the released predictor.pkl (Zenodo 15568557) with OpenMX-generated graph_data.npz inputs for inorganic halide perovskites. We are preparing a methods section that describes the pretrained model, and we also run our own OpenMX reference calculations for comparison, which we would like to match the training protocol as closely as possible. A few details are not stated in the paper/SI (arXiv:2504.19586 / Nat. Mach. Intell. 2026) and we could not find them in this repository, so we would be grateful for clarification.

Training set

  1. Structure selection from the Materials Project: were any filters applied (e.g., maximum atoms per cell, energy above hull, exclusion of magnetic or f-element systems), and which MP release was used?
  2. Sec. 2.2 mentions 44,000 non-SOC structures for H0, while the Introduction mentions 40,000. Which is the exact number? Is the ~10,000-structure SOC set a subset of the non-SOC set, and was it drawn randomly or weighted toward heavy elements?
  3. Does each structure contribute a single relaxed MP geometry (i.e., no perturbed or MD configurations)?
  4. Roughly, what is the distribution of atoms per cell in the training set (median / maximum)? If a list of MP IDs for the training and test splits could be shared, that would let us check overlap with our own test systems.
  5. In Fig. 3b ("Element count in the whole training dataset"), are the numbers atom counts or structure counts?

OpenMX label protocol

  1. Exchange-correlation functional (we assume GGA-PBE, as for the CPL 2024 non-SOC model), OpenMX version, and the PAO/VPS database (DFT_DATA19?).
  2. Was the 6×6×6 Monkhorst-Pack grid used for all cells regardless of cell size, or scaled with the cell dimensions?
  3. scf.criterion: the paper states 1.0e-8 Hartree, while the readme template uses 1.0e-7. Which was used for the training labels? Likewise, were scf.ElectronicTemperature 100 and rmm-diis mixing from the readme SOC template the training settings?
  4. Can we confirm that the training labels used exactly the PAO/VPS choices in DFT_interfaces/openmx/utils.py (PAO_dict / PBE_dict, i.e., basis_def_26 including the "quick" bases for P/S/Cl/Ar discussed in Basis_def_26 inconsistent with OpenMX official #63), with scf.SpinPolarization nc and scf.SpinOrbit.Coupling on for the SOC labels?

Thanks again for your time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions