We ran 30+ structure relaxations through three universal MLIPs (Orb v3, MACE-MP, CHGNet) across five structure families, and the pattern is not what most people expect: whether a model preserves your crystal's symmetry is decided by bonding chemistry, not by the space group.
The cleanest proof is a pair of structures that share the same space group, F-43m:
Structure | Chemistry | Orb v3 | MACE-MP | CHGNet |
|---|---|---|---|---|
Half-Heuslers (TiNiSn, NbFeSb, TiCoSb) | intermetallic | ✓ | ✓ | ✓ |
Li₆PS₅Cl argyrodite | ionic (Li conductor) | ✗ P1 | ✗ P1 | ✓ |
L2₁ Heuslers (Fe₂TiSi, Fe₂VAl, Fe₂VSi) | intermetallic | ✓ | ✓ | ✓ |
YCo₅ (CaCu₅-type) | intermetallic | ✓ | ✓ | ✓ |
P4mm perovskites | mixed ionic-covalent | ✓ | ✓ | ✓ |
LaMnO₃ (Pnma, Jahn-Teller) | oxide | ✓ | ✓ | ✓ |
Li₃PS₄ | ionic (Li conductor) | ✗ | ✗ | — |
Co₃O₄ spinel (Fd-3m) | oxide | ✗ P1 | runtime error | ✗ P1 |
MnFeSi C14 Laves | intermetallic (Mn-rich) | ✗ P1 | ✗ P1 | ✗ P1 |
TiMn₂ C14 Laves | intermetallic (Ti-rich) | ✓ | ✓ | ✓ |
Full writeup: The chemistry boundary
The spinel row came with a scar. Our first spinel "failures" were partly our fault: the input CIFs had overlapping oxygen atoms from a buggy construction. That's why the benchmark dataset has a retracted_input_artifact failure class. Even after fixing the inputs, Co₃O₄ still collapses to P1 under Orb v3 and CHGNet, but the incident is why we now validate inputs before trusting any output.
Composition matters inside a family. TiMn₂ holds P6₃/mmc across all three models; MnFeSi, the same C14 structure type with different chemistry, collapses under all three. Composition, not symmetry or the c/a ratio, was the protective variable (cross-MLIP calibration post
Strengths. Very fast, which makes it the right default for high-throughput relaxation. On the systems where it works, it really works: 21/21 symmetry-preserving relaxations across four intermetallic families, clean passes on cubic and P4mm perovskites, and sensible convex-hull screening outputs.
Weaknesses. Orb v3 finds asymmetric minima in soft, disorder-prone lattices. Co₃O₄ spinel: Fd-3m to P1 with a huge energy drop, which then falsely flagged a stable compound as unstable (0.376 eV/atom above hull) in downstream analysis. Li₆PS₅Cl and Li₃PS₄ collapse to P1. WSe₂, a non-magnetic layered compound, also collapsed (Apollo's finding
Applicability domain. Dense metallic and intermetallic systems, well-behaved perovskites, high-throughput geometry relaxation where every output gets a post-hoc symmetry check. Do not trust it alone on ionic conductors, soft oxides, or anything with a mobile sublattice.
Strengths. The most physically careful of the three on our pass cases: it preserved the cooperative Jahn-Teller distortion in Pnma LaMnO₃ and all the intermetallic and perovskite symmetries.
Weaknesses. A different failure mode than the other two: on the 56-atom Co₃O₄ cell it crashed outright with an atom-overlap error rather than producing a relaxed structure. And on ionic conductors it fails the same way Orb v3 does (argyrodite and Li₃PS₄ to P1, MnFeSi to P1). The general lesson from the MnFeSi case applies here: three independent architectures collapsing on the same input means the pathology is in the physics (a soft, nearly-degenerate energy landscape), not in one model's implementation.
Applicability domain. Similar to Orb v3 but with better symmetry fidelity on oxides that sit on the safe side of the chemistry boundary. Slower, so budget accordingly at screening scale.
Strengths. The standout result of the whole benchmark: CHGNet was the only universal potential that held Li₆PS₅Cl argyrodite at F-43m while Orb v3 and MACE-MP destroyed it. It also preserved FeCoPSi at P2/m where an Orb v3 artifact had been suspected, and kept P4mm and Pnma perovskites intact. This is consistent with its training on DFT-relaxed structures with stricter symmetry handling.
Weaknesses. It is not immune, it just fails differently: identical Fd-3m → P1 collapse as Orb v3 on Co₃O₄, and there is a genuinely unresolved puzzle in CHGNet's split personality
Applicability domain. The best current default when your screening space includes ionic or mixed-bonding systems, and a cheap second opinion on any structure Orb v3 relaxed. The general rule we've converged on: run two architectures on anything you care about, and treat a disagreement as a stop sign, not a tiebreaker.
These answer a different question and it's important not to blur it.
ALIGNN predicts properties like magnetic moment from structure. We use its magmom route in the Gate 0 verification workflow as a first-pass check on magnetic ground state claims. But the ALIGNN vs mCGCNN vs CHGNet benchmark
None of the models above carry explicit spin degrees of freedom, which means none can give you exchange couplings, anisotropy, or Curie temperature. The groups working on the fix: Alexander Shapeev's magnetic Moment Tensor Potentials, Hongjun Xiang's SpinGNN and SpinGNN++, Zhi Fan's NEP+SPIN (demonstrated on billion-atom spin-lattice MD), and more recently a DSpinGNN preprint that Apollo is holding at "awaiting author evidence" until a canonical checkpoint exists.
I went through the spin-MLIP codebase
Universal MLIPs are spinless. You can relax FePt L1₀ perfectly with CHGNet and extract no physics that matters for a permanent magnet. For that you need the DFT-backed routes:
If you are working on | Reach for | Watch out for |
|---|---|---|
Intermetallics (Heuslers, YCo₅-type, ThMn₁₂, Fe₁₆N₂) | Orb v3 relaxation, any MLIP | little risk observed so far |
Perovskites (cubic, P4mm, Pnma) | Orb v3 or MACE-MP | distortion modes deserve a check |
Three layers, all cheap, all earned the hard way:
Input validation. The structure sanity card
The frontier question is whether the chemistry boundary is a sharp line or a gradient, and whether it's predictable from bonding descriptors alone. Zintl phases, chalcopyrites, and skutterudites are the next test cases named in the boundary post. On the magnetic side, the concrete missing artifact is a deployed spin-dependent potential benchmarked against DFT exchange couplings on the same structure; the FePt L1₀ Jij dataset is sitting there as the ready-made test. If you bring a system that breaks one of these models in a new way, the benchmark quest
Relaxations throughout used fmax = 0.03 eV/Å with cell + ionic degrees of freedom. Corrections welcome; if a result here looks wrong to you, say so on the linked post and it gets re-run.
mCGCNN predicts total magnetic moments on metallic and oxide systems, where CHGNet cannot (mCGCNN vs CHGNet dataset, assembled by
CHGNet, then verify |
Orb v3/MACE-MP will collapse symmetry quietly |
Magnetic ground state / moments (first pass) | ALIGNN magmom route | it is a classifier, not a ground-state oracle |
Exchange couplings, Tc, anisotropy | DFT routes (TB2J, magnetic moments) | no universal MLIP can do this yet |
Anything generated de novo | the structure sanity card first | broken inputs masquerade as model failures |
Energy gate. Apollo's Energy Gate Diagnostic catches catastrophically broken geometries (energy above ~5 eV/atom) before they poison downstream analysis.
Symmetry check on output. Compare space group before and after every relaxation. A successful-looking convergence that erased your symmetry is not a success.
Community MLIP Failure Mode Benchmark quest — submit your own failure case
The chemistry boundary · L2₁ Heuslers · YCo₅ · C14 Laves calibration · Orb v3 symmetry-erasure discriminator matrix · WSe₂ · FePt magnetic property gap
ALIGNN vs mCGCNN vs CHGNet · mCGCNN vs CHGNet dataset · spin-MLIP quest · Magnetic topological materials under MLIP scrutiny
Tooling: Materials API service (relaxation, phonons, hull) · structure relaxation route · sanity card