Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Reward Driven Workflows for Unsupervised Explainable Analysis of Phases and Ferroic Variants From Atomically Resolved Imaging Data

Rapid progress in aberration corrected electron microscopy necessitates development of robust methods for the identification of phases, ferroic variants, and other pertinent aspects of materials structure from imaging data. While unsupervised methods for clustering and classification are widely used for these tasks, their performance can be sensitive to hyperparameter selection in the analysis workflow. In this study, the effects of descriptors and hyperparameters are explored on the capability of unsupervised ML methods to distill local structural information, exemplified by the discovery of polarization and lattice distortion in Sm − dopped BiFeO 3 (BFO) thin films. It is demonstrated that a reward-driven approach can be used to optimize these key hyperparameters across the full workflow, where rewards are designed to reflect domain wall continuity and straightness, ensuring that the analysis aligns with the material's physical behavior. This approach allows the discovery of local descriptors that are best aligned with the specific physical behavior, providing insight into the fundamental physics of materials. The reward driven workflow is further extended to disentangle structural factors of variation via an optimized variational autoencoder (VAE). Lastly, the importance of well-defined rewards is explored as a quantifiable measure of the success of the workflow.

Barakati, Kamyar [University of Tennessee, Knoxvil↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Giant Magnetostriction in Ferrimagnetic SmFe 5 As 3

Magnetostrictive materials are of interest not only from a fundamental perspective but also for their potential applications, spanning spintronics to energy harvesting. A new magnetostrictive material—SmFe 5 As 3 —reveals a complex interplay between magnetostriction, magnetic properties, and the crystal structure behavior. The ground state of SmFe 5 As 3 is ferrimagnetic, as evidenced by magnetic susceptibility data, band-structure calculations, and XANES measurements. At T m1 = 28 ± 4 K, part of the Fe sublattice reorients, transitioning into a ferromagnetic state. This is followed by an entrance into the paramagnetic state at T m2 = 76 ± 4 K. The effects observed in the magnetic measurements are accompanied by structural phase transitions. All three phases—below T m1 , between T m1 and T m2 , and above T m2 —are described by the same structural motif of the UCr 5 P 3 type (monoclinic space group P2 1 /m), differing only in the degree of deformation of the Fe–As framework. The new material exhibits giant magnetostriction of 2500 x 10 -6 . Dilatometry measurements on a SmFe 5 As 3 single crystal indicate strongly diverse behavior, exhibiting not only negative, but also zero, as well as positive thermo-elastic effects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploiting correlations in multi-coincidence Coulomb explosion patterns for differentiating molecular structures using machine learning

Coulomb explosion imaging (CEI) is a powerful technique for capturing the real-time motion of individual atoms during ultrafast photochemical reactions. CEI generates high-dimensional data with naturally embedded correlations that allow mapping the coordinated motion of nuclei in molecules. This enables reliable separation of competing reaction pathways and makes this approach uniquely suited for characterizing weak reaction channels. However, rich information contained in experimental CEI patterns remains largely underexploited due to challenges in visualizing correlations between multiple observables in multi-dimensional parameter space. Here we present a new approach to CEI of intermediate-sized polyatomic molecules, detecting up to eight ionic fragments in coincidence and leveraging machine-learning-based analysis to identify patterns and correlations in the resulting high-dimensional momentum-space data, enabling robust molecular structure identification and differentiation. Our approach provides high-dimensional background-free data encoding exceptionally rich structural information and establishes an automated, scalable framework for extracting insightful information from the data. As a demonstration, we apply this method to image and distinguish dichloroethylene isomers, showcasing its potential for broader applications in molecular imaging. Our results pave the way for channel-specific analysis of ultrafast structural dynamics in chemically relevant systems, particularly for disentangling mixed reaction pathways and detecting contributions from weak channels and minority species.

Chemical Physics (physics.chem-ph)↗

Autonomous phase mapping of gold nanoparticles synthesis with differentiable models of spectral shape

Autonomous experimentation–or self-driving labs–offers a systematic approach to accelerate materials discovery by integrating automated synthesis, characterization, and data-driven decision-making. We present a closed-loop workflow for the on-demand synthesis and structural characterization of colloidal gold nanoparticles, enabling direct mapping from composition to nanoscale structure. Our framework leverages differentiable models of spectral shape to address two central tasks in self-driving labs: (a) phase mapping, or identifying compositional regions with distinct structural behavior; and (b) material retrosynthesis, or optimizing compositions for target structure. Using functional data analysis, we develop a data-driven model with generative pre-training, active learning, and high-throughput experiments to predict spectral responses across composition space. We demonstrate the approach on seed-mediated growth of gold nanoparticles, showcasing its ability to extract design rules, reveal secondary interactions, and efficiently navigate morphology space. Gradient-based optimization of the models enables inverse design, making this a unified platform.

36 MATERIALS SCIENCE↗

Basin-Scale Structural Features Database

The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.

basin scale↗

Tropical Cyclone Wind Shear-Relative Asymmetry in Reanalyses

Abstract While tropical cyclones (TCs) are axisymmetric vortices to the first order, they often exhibit noteworthy structural asymmetries. These often result from environmental vertical wind shear, which tilts the vortex and induces a wavenumber 1 pattern in the circulation and precipitation fields. Reanalyses and climate models have improved in representing the TC structure and climatology, but their relatively coarse resolution and dependence on parameterized physics cast doubt on their ability to capture the asymmetric TC structure. We perform the most comprehensive process-oriented assessment of TC asymmetry to date in reanalyses. Specifically, we analyze the composite shear-relative TC structure in ERA5 and Climate Forecast System Reanalysis (CFSR), which vary in their resolutions, physical parameterization suites, and data assimilation techniques. These structures are compared with aircraft reconnaissance radar observations. In agreement with the observations, the strongest tangential winds are usually found left-of-shear, while inner core rainfall, ascent, vortex tilt, and low-level inflow are favored directly downshear or in the downshear-left quadrant. Outer rainband convection generally peaks in the downshear-right quadrant. Thermodynamic asymmetries are also apparent, with anomalous low-level moisture right-of-shear, midlevel warmth in the upshear-right quadrant (uptilt), and cloud properties suggestive of a realistic precipitation life cycle from growth to fallout. We also decompose rainfall contributions from the convective parameterization and large-scale cloud schemes and highlight the roles of vorticity advection, buoyancy advection, and diabatic processes in driving asymmetric vertical motions in the inner core and outer rainband regions. Our results suggest that process-level studies of TC asymmetry and TC–wind shear interaction under future warming are viable using climate models. Significance Statement Asymmetries are common in tropical cyclones (TCs), influencing their intensity, track, and hazards. Vertical wind shear often plays a leading-order role in causing these asymmetries. It is uncertain how well asymmetric structures and processes are captured in reanalyses and global climate models (GCMs) with grid spacings of 0.25° and coarser. In this study, we first evaluate TC asymmetry in reanalyses, which have the benefit of being forced by observations. This helps to assess whether the resolutions associated with GCMs sufficiently capture asymmetric structures and processes and motivates upcoming work with free-running GCMs to study how TC asymmetry may change in a warming climate.

Carstens, Jacob D.↗

Experimental and Computational Insights into the Structural Dynamics of the Fc Fragment of IgG1 Subtype from Biosimilar VEGF‐Trap

The constant fragment (Fc) of the immunoglobulin G1 (IgG1) subtype is increasingly recognized as a crucial scaffold in the development of advanced therapeutics due to its enhanced specificity, efficacy, and extended half‐life. A prime example is VEGF‐Trap (Aflibercept), a recombinant fusion protein that merges the Fc region of the IgG1 subtype with the binding domains of vascular endothelial growth factor receptors (VEGFR)‐1 and VEGFR‐2. The Fc region's role in N‐glycosylation is particularly important, as it significantly influences protein stability. Herein, the first near‐physiological temperature structures of the N‐glycan‐bound Fc fragment of IgG1 subtype from a biosimilar VEGF‐Trap are presented, determined using the SPring‐8 Angstrom Compact free electron LAser (SACLA) and the Turkish Light Source (Turkish DeLight). Comparative analysis with cryogenic structures, including existing data, reveals alternate conformations within the glycan‐binding pocket. Furthermore, molecular dynamics simulations indicate the presence of a high degree of structural plasticity, explaining how the protein adapts its structure through conformational changes. The observed structural fluctuations/conformational changes demonstrate the effect of N‐glycans on protein stability. These findings offer new insights into the molecular basis of Fc‐mediated functions and provide valuable information for the design of next‐generation therapeutics.

60 APPLIED LIFE SCIENCES↗

A Bayesian desmearing algorithm for Bonse–Hart USANS with anisotropic scattering

Ultra-small-angle neutron scattering (USANS) using Bonse–Hart optics provides micrometer-scale structural insights but suffers from severe slit-geometry smearing. While well-established for isotropic systems, quantitative desmearing of anisotropic data remains a challenge because conventional corrections break down for non-radial scattering. In this work, we address this by developing a resolution-aware Bayesian framework that explicitly incorporates anisotropy via an affine deformation to the scattering pattern, guided by the principle of parsimony. This results in orientation-resolved point-spread functions that enable a self-consistent determination of both the resolution and deformation parameters. Using Gaussian process regression with uncertainty quantification and a probabilistic correction for multiple scattering, we demonstrate the framework’s effectiveness through numerical benchmarks and experimental studies of a stretched polymer melt. Our approach enables the seamless integration of SANS and USANS data, facilitating quantitative structural analysis of deformed materials at nanometer to micrometer scales.

36 MATERIALS SCIENCE↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Ba 4 RuMn 2 O 10 : A Noncentrosymmetric Polar Crystal Structure with Disordered Trimers

Phase-pure polycrystalline Ba 4 RuMn 2 O 10 was prepared and determined to adopt the noncentrosymmetric polar crystal structure (space group Cmc2 1 ) based on results of second harmonic generation, convergent beam electron diffraction, and Rietveld refinements using powder neutron diffraction data. The crystal structure features zigzag chains of corner-shared trimers, which contain three distorted face-sharing octahedra. The three metal sites in the trimers are occupied by disordered Ru/Mn with three different ratios: Ru1:Mn1 = 0.202(8):0.798(8), Ru2:Mn2 = 0.27(1):0.73(1), and Ru3:Mn3 = 0.40(1):0.60(1), successfully lowering the symmetry and inducing the polar crystal structure from the centrosymmetric parent compounds Ba 4 T 3 O 10 (T = Mn, Ru; space group Cmca). The valence state of Ru/Mn is confirmed to be +4 according to X-ray absorption near-edge spectroscopy. Ba 4 RuMn 2 O 10 is a narrow bandgap (~0.6 eV) semiconductor exhibiting spin-glass behavior with strong magnetic frustration and antiferromagnetic interactions.

36 MATERIALS SCIENCE↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

Influence of Alkyne Precursor Structure on Carbon Nanotube Chiral Distribution: Data-Dense Analysis Across Multiple Catalyst Types

Carbon nanotubes (CNTs) are a desirable material in the field of optoelectronics and semiconductors due to electronic properties (e.g., bandgap) that are dependent upon their chirality, defined by their diameter and lattice angle. Unfortunately, industrial-scale syntheses have yet to realize growth of a single desired chirality and instead rely on postsynthetic separation techniques to refine a chiral mixture, which increases process complexity and cost. Here, we studied the influence of precursor structure on chiral distribution, using a series of terminal alkyne precursors (acetylene, methylacetylene, vinylacetylene, 1-butyne, two enantiomers of 3-butyn-2-ol and a racemic mixture thereof) to grow CNTs across five transition-metal catalysts (Fe, FeMo, and three proportions of CoMo). Multiwavelength Raman spectroscopy on 5,145 spots (5 catalysts, 7 precursors, 3 lasers, and 49 distinct substrate locations on each) determined that acetylene grew the smallest diameter CNTs, while vinylacetylene produced fewer subnanometer CNTs. Though precursor structure did not dictate a uniform chiral shift, it was shown to broaden or narrow chiral distribution, while catalyst structure played a dominant role. In conclusion, this is consistent with metal-precursor binding occurring through unsaturated bonds in the hydrocarbons via the alkyne polymerization mechanism.

Carbon nanotubes↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

Spectroscopy-guided discovery of three-dimensional structures of disordered materials with diffusion models

Spectroscopy techniques such as x-ray absorption near edge structure (XANES) provide valuable insights into the atomic structures of materials, yet the inverse prediction of precise structures from spectroscopic data remains a formidable challenge. In this study, we introduce a framework that combines generative artificial intelligence models with XANES spectroscopy to predict three-dimensional atomic structures of disordered systems, using amorphous carbon (a-C) as a model system. In this work, we introduce a new framework based on the diffusion model, a recent generative machine learning method, to predict 3D structures of disordered materials from a target property. For demonstration, we apply the model to identify the atomic structures of a-C as a representative material system from the target XANES spectra. We show that conditional generation guided by XANES spectra reproduces key features of the target structures. Furthermore, we show that our model can steer the generative process to tailor atomic arrangements for a specific XANES spectrum. Finally, our generative model exhibits a remarkable scale-agnostic property, thereby enabling generation of realistic, large-scale structures through learning from a small-scale dataset (i.e. with small unit cells). Our work represents a significant stride in bridging the gap between materials characterization and atomic structure determination; in addition, it can be leveraged for materials discovery in exploring various material properties as targeted.

36 MATERIALS SCIENCE↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗