dorsal/arxiv
View SchemaActive Learning Strategies for Efficient Machine-Learned Interatomic Potentials Across Diverse Material Systems
| Authors | Mohammed Azeez Khan, Aaron D'Souza, Vijay Choyal |
|---|---|
| Categories | |
| ArXiv ID | 2601.06916vv1 |
| URL | https://arxiv.org/abs/2601.06916 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Efficient discovery of new materials demands strategies to reduce the number of costly first-principles calculations required to train predictive machine learning models. We develop and validate an active learning framework that iteratively selects informative training structures for machine-learned interatomic potentials (MLIPs) from large, heterogeneous materials databases, specifically the Materials Project and OQMD. Our framework integrates compositional and property-based descriptors with a neural network ensemble model, enabling real-time uncertainty quantification via Query-by-Committee. We systematically compare four selection strategies: random sampling (baseline), uncertainty-based sampling, diversity-based sampling (k-means clustering with farthest-point refinement), and a hybrid approach balancing both objectives. Experiments across four representative material systems (elemental carbon, silicon, iron, and a titanium-oxide compound) with 5 random seeds per configuration demonstrate that diversity sampling consistently achieves competitive or superior performance, with particularly strong advantages on complex systems like titanium-oxide (10.9% improvement, p=0.008). Our results show that intelligent data selection strategies can achieve target accuracy with 5-13% fewer labeled samples compared to random baselines. The entire pipeline executes on Google Colab in under 4 hours per system using less than 8 GB of RAM, thereby democratizing MLIP development for researchers globally with limited computational resources. Our open-source code and detailed experimental configurations are available on GitHub. This multi-system evaluation establishes practical guidelines for data-efficient MLIP training and highlights promising future directions including integration with symmetry-aware neural network architectures.
{
"annotation_id": "17cb841d-2238-4e91-b43a-f1bb1ed6240e",
"date_created": "2026-02-17T05:53:07.897000Z",
"date_modified": "2026-02-17T05:53:07.897000Z",
"file_hash": "2eebc89f5c533eb2b0348b8041634b7125658478ae9e274c3b59e280155b95e6",
"private": false,
"record": {
"abstract": "Efficient discovery of new materials demands strategies to reduce the number of costly first-principles calculations required to train predictive machine learning models. We develop and validate an active learning framework that iteratively selects informative training structures for machine-learned interatomic potentials (MLIPs) from large, heterogeneous materials databases, specifically the Materials Project and OQMD. Our framework integrates compositional and property-based descriptors with a neural network ensemble model, enabling real-time uncertainty quantification via Query-by-Committee. We systematically compare four selection strategies: random sampling (baseline), uncertainty-based sampling, diversity-based sampling (k-means clustering with farthest-point refinement), and a hybrid approach balancing both objectives. Experiments across four representative material systems (elemental carbon, silicon, iron, and a titanium-oxide compound) with 5 random seeds per configuration demonstrate that diversity sampling consistently achieves competitive or superior performance, with particularly strong advantages on complex systems like titanium-oxide (10.9% improvement, p=0.008). Our results show that intelligent data selection strategies can achieve target accuracy with 5-13% fewer labeled samples compared to random baselines. The entire pipeline executes on Google Colab in under 4 hours per system using less than 8 GB of RAM, thereby democratizing MLIP development for researchers globally with limited computational resources. Our open-source code and detailed experimental configurations are available on GitHub. This multi-system evaluation establishes practical guidelines for data-efficient MLIP training and highlights promising future directions including integration with symmetry-aware neural network architectures.",
"arxiv_id": "2601.06916",
"authors": [
"Mohammed Azeez Khan",
"Aaron D\u0027Souza",
"Vijay Choyal"
],
"categories": [
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Active Learning Strategies for Efficient Machine-Learned Interatomic Potentials Across Diverse Material Systems",
"url": "https://arxiv.org/abs/2601.06916",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ed1f10f7-c48d-4462-a5f7-8566f29e8beb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}