Search NASA⌕ Search

SEARCH · Search NASA

Results for “Algorithms and data structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dynamics and Formation of Antiferromagnetic Textures in MnBi 2 Te 4 Single Crystal

We report coherent X-ray imaging of antiferromagnetic (AFM) domains and domain walls in MnBi 2 Te 4 , an intrinsic AFM topological insulator. This technique enables direct visualization of domain morphology without reconstruction algorithms, allowing us to resolve antiphase domain walls as distinct dark lines arising from the A-type AFM structure. The wall width is determined to be 550(30) nm, in good agreement with earlier magnetic force microscopy results. The temperature dependence of the AFM order parameter extracted from our images closely follows previous neutron scattering data. Remarkably, however, we find a pronounced hysteresis in the evolution of domains and domain walls: upon cooling, dynamic reorganizations occur within a narrow ∼1 K interval below 𝑇 𝑁 , whereas upon warming, the domain configuration remains largely unchanged until AFM order disappears. These findings reveal a complex energy landscape in MnBi 2 Te 4 , governed by the interplay of exchange, anisotropy, and domain-wall energies, and underscore the critical role of AFM domain-wall dynamics in shaping its physical properties. These sharply defined and hysteretically evolving walls may provide a controllable AFM texture in MnBi 2 Te 4 , hinting at potential use in low-power spintronic devices based on domain-wall dynamics.

36 MATERIALS SCIENCE↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Capturing thin structures in VOF simulations with two-plane reconstruction

A novel interface reconstruction strategy for volume of fluid (VOF) methods is introduced that represents the liquid-gas interface as two planes that co-exist within a single computational cell. In comparison to the piecewise linear interface calculation (PLIC), this new algorithm greatly improves the accuracy of the reconstruction, in particular when dealing with thin structures such as films. The placement of the two planes requires the solution of a non-linear optimization problem in six dimensions, which has the potential to be overly expensive. Further, an efficient solution to this optimization problem is presented here that exploits two key ideas: an algorithm for extracting multiple plane orientations from transported surface data, and an efficient and mass-conserving distance-finding algorithm that accounts for two planes with arbitrary orientation. Additionally, a simple and robust strategy is presented to accurately represent the surface tension forces produced at the interface of subgrid-thickness films. The performance of this new VOF reconstruction is demonstrated on several test cases that illustrate the capability to handle arbitrarily thin films.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Portable Parallel Algorithms and Frameworks for Exascale Graph Analytics

Graphs (or networks) are a tool used to model the interactions among various entities. Efficiently processing large graphs has recently attracted significant attention due to the applications of graphs in various domains, such as biology, chemistry, and cyber-security. Analyzing the structure and properties of these graphs is an important component of many scientific computing pipelines. With the explosion in the volume of data, graphs have become very large and can contain hundreds of billions of vertices and trillions of edges. Therefore, it is crucial to develop high-performance methods to enable graph analysis to be done quickly and energy-efficiently. Furthermore, these solutions should be highly parallel in order to take advantage of modern parallel machines. However, designing efficient solutions is not enough. With the wide variety of computing environments available, each with different programmability and performance characteristics, it is necessary to develop solutions that are portable in terms of both performance (i.e., provide theoretical guarantees) and programmability (i.e., provide high level abstractions).

97 MATHEMATICS AND COMPUTING↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

Constraints on f ( R ) gravity from thermal-Sunyaev-Zel’dovich-effect-selected SPT galaxy clusters and weak lensing mass calibration from DES and HST

We present constraints on the f ( R ) gravity model using a sample of 1005 galaxy clusters in the redshift range 0.25–1.78 that have been selected through the thermal Sunyaev-Zel’dovich effect from South Pole Telescope data and subjected to optical and near-infrared confirmation with the multicomponent matched filter algorithm. We employ weak gravitational lensing mass calibration from the Dark Energy Survey Year 3 data for 688 clusters at z < 0.95 and from the Hubble Space Telescope for 39 clusters with 0.6 < z < 1.7 . Our cluster sample is a powerful probe of f ( R ) gravity, because this model predicts a scale-dependent enhancement in the growth of structure, which impacts the halo mass function (HMF) at cluster mass scales. To account for these modified gravity effects on the HMF, our analysis employs a semianalytical approach calibrated with numerical simulations. Combining calibrated cluster counts with primary cosmic microwave background temperature and polarization anisotropy measurements from the Planck 2018 release, we derive robust constraints on the f ( R ) parameter f R 0 . Our results, log 10 | f R 0 | < − 5.32 at the 95% credible level, are the tightest current constraints on f ( R ) gravity from cosmological scales. This upper limit rules out f ( R ) -like deviations from general relativity that result in more than a ∼ 20 % enhancement of the cluster population on mass scales M 200 c > 3 × 10 14 M ⊙ . Published by the American Physical Society 2025

79 ASTRONOMY AND ASTROPHYSICS↗

Angular Dependence of Drell-Yan in the SeaQuest Experiment using Deep Neural Network-Based Reconstruction

The SeaQuest and SpinQuest experiments at Fermilab were designed to probe the internal dynamics of protons and neutrons via the angular dependence of muons created by colliding 120 GeV protons into stationary targets. To do this, it is necessary to translate the detector data into coherent information about the particles detected in the experiment. This dissertation describes the development and implementation of QTracker, a neural network-based algorithm designed to reconstruct muons. The application of QTracker to experimental data yields improved reconstruction of muon tracks, allowing for a more precise investigation of the transverse momentum distributions and angular modulations in the Drell-Yan process. This work highlights the potential of neural network-based methods to advance particle tracking and enhance our understanding of nucleon structure.

Conover, Arthur↗

Quantitative x-ray scattering of free molecules

Advances in x-ray free electron lasers have made ultrafast scattering a powerful method for investigating molecular reaction kinetics and dynamics. Accurate measurement of the ground-state, static scattering signals of the reacting molecules is pivotal for these pump-probe x-ray scattering experiments as they are the cornerstone for interpreting the observed structural dynamics. Here, this article presents a data calibration procedure, designed for gas-phase x-ray scattering experiments conducted at the Linac Coherent Light Source x-ray Free-Electron Laser at SLAC National Accelerator Laboratory, that makes it possible to derive a quantitative dependence of the scattering signal on the scattering vector. A self-calibration algorithm that optimizes the detector position without reference to a computed pattern is introduced. Angle-of-scattering corrections that account for several small experimental non-idealities are reported. Their implementation leads to near quantitative agreement with theoretical scattering patterns calculated with ab-initio methods as illustrated for two x-ray photon energies and several molecular test systems.

74 ATOMIC AND MOLECULAR PHYSICS↗

Ecosystem Science with NISAR: Final Preparations in The Pre-Launch Period

The NISAR mission which in its most recent round of launch preparations was set to launch in the spring of 2024, and now delayed until later in the fall or early spring of 2025, will serve as an unprecedented resource for the Remote Sensing of Ecosystems Science community. The two frequency, L- and S-band will full-polarimetric capability over a 250 km wide swath using the SweepSAR technique [1] will collect reliable set of observations (60 per year; 30 each for ascending and descending passes) on a continuing basis that will allow for the modeling and observation of time-varying processes that are prevalent in the living environment broadly described as Ecosystems. Among the prime science goals of the NISAR Ecosystems disciplines are in the characterization of agriculture, disturbance, biomass, forest structure and water dynamics seen in the world’s rivers, coasts, and permafrost regions. In this paper we provide an overview of the Ecosystem science that will be enabled by the NISAR mission and give a status of the basic algorithms that are being used to provide a basic set of tools to the community to make use of the data that NISAR will provide.

Siqueira, Paul↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

Ab Initio Study of the Beryllium Isotopes 7 Be to 12 Be

We present a systematic ab initio study of the low-lying states in beryllium isotopes from 7 Be to 12 Be using nuclear lattice effective field theory with the N 3 ⁢LO interaction. Our calculations achieve good agreement with experimental data for energies, radii, and electromagnetic properties. We introduce a novel, model-independent method to quantify nuclear shapes, uncovering a distinct pattern in the interplay between positive and negative parity states across the isotopic chain. By combining Monte Carlo sampling of the many-body density operator with a novel nucleon-grouping algorithm, the prominent two-center cluster structures, the emergence of one-neutron halo, complex nuclear molecular dynamics such as 𝜋 orbital and 𝜎 orbital, emerge naturally.

binding energy & masses↗

An adaptive and stability-promoting layerwise training approach for sparse deep neural network architecture

This work presents a two-stage adaptive framework for progressively developing deep neural network (DNN) architectures that generalize well for a given training data set. In the first stage, a layerwise training approach is adopted where a new layer is added each time and trained independently by freezing parameters in the previous layers. We impose desirable structures on the DNN by employing manifold regularization, sparsity regularization, and physics-informed terms. We introduce a ε – δ – stability-promoting concept as a desirable property for a learning algorithm and show that employing manifold regularization yields a ε – δ stability-promoting algorithm. Further, we also derive the necessary conditions for the trainability of a newly added layer and investigate the training saturation problem. In the second stage of the algorithm (post-processing), a sequence of shallow networks is employed to extract information from the residual produced in the first stage, thereby improving the prediction accuracy. Numerical investigations on prototype regression and classification problems demonstrate that the proposed approach can outperform fully connected DNNs of the same size. Moreover, by equipping the physics-informed neural network (PINN) with the proposed adaptive architecture strategy to solve partial differential equations, we numerically show that adaptive PINNs not only are superior to standard PINNs but also produce interpretable hidden layers with provable stability. As a result, we also apply our architecture design strategy to solve inverse problems governed by elliptic partial differential equations.

42 ENGINEERING↗

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE↗

Combined Machine Learning and Molecular Dynamics Reveal Two States of Hydration of a Single Functional Group of Cationic Polymeric Brushes

The state of hydration of a macromolecular system regulates a plethora of different properties of such a system. In this article, we develop a novel machine learning (ML) approach, based on the unsupervised clustering algorithm, for probing the hydration behavior of the {N(CH 3 ) 3 } + functional group of the PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] polyelectrolyte (PE) brush system. The PE brushes and the brush-supported water molecules and counterions (chloride ions) are first described using all-atom molecular dynamics (MD) simulations. The simulation data is subsequently used in our ML framework to identify that (1) the {N(CH 3 ) 3 } + functional groups of the PMETAC brushes have two distinct hydration states with one state (state 1) being characterized by less structured water molecules and the other state (state 2) being characterized by more structured water molecules and (2) an enhancement in the brush grafting density leads to the progressive dissapparenace of state 2. An increase in the grafting density increases the number of chloride counterions in a given volume around the {N(CH 3 ) 3 } + functional group and increases the number of shared water molecules between the {N(CH 3 ) 3 } + and Cl - . The chloride counterions are associated with a hydration layer with much less structured water molecules. Therefore, with an increase in the grafting density, an increase in the percentage of shared water molecules leads to the prevalence of the hydration state [of the {N(CH 3 ) 3 } + moiety] with less structured water molecules. Finally, we explain how the present findings are commensurate with two key previous related results, namely a significantly large chloride ion mobility inside the PMETAC brush layer and the {N(CH 3 ) 3 } + -Cl - average distance remaining independent of the PMETAC brush grafting density. Furthermore, we anticipate that the combined ML-MD-simulation approach proposed in this study can be adapted to probe other soft matter systems to reveal new insights of the underlying mechanisms of emergent phenomenon.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advanced Polymer Characterization: Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)

Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry encodes structural information across diverse homo- and copolymer ensembles, yet decrypting these spectra requires a systematic analytical approach. We introduce Modular Operations for Spectral Alignment by Iterative Compression (MOSAIC)─a general cipher algorithm that applies modular arithmetic to filter monomer-derived mass contributions and cluster MALDI peaks by nonconstitutional repeating units (non-CRUs). MOSAIC performs sequential modular operations using monomer mass differences as base units to compress complex spectral data, revealing end-group distributions and comonomer incorporation. As a demonstration, we applied MOSAIC to five copolymers formed by two different polymerization mechanisms. Furthermore, the resulting remainder–mass plots clearly resolve polymer homologs with distinct non-CRUs into visually apparent clusters, enabling intuitive assignment of mass spectral features.

Wang, Hanlin M. [University of Illinois at Urbana−↗