Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Utility-Scale Solar, 2024 Edition: Analysis of Empirical Plant-level Data from U.S. Ground-mounted PV, PV+battery, and CSP Plants (exceeding 5 MWAC)

Berkeley Labs "Utility-Scale Solar", 2024 Edition presents analysis of empirical plant-level data from the U.S. fleet of ground-mounted photovoltaic (PV), PV+battery, and concentrating solar-thermal power (CSP) plants with capacities exceeding 5 MWAC. While focused on key developments in 2023, this report explores trends in deployment, technology, capital and operating costs, capacity factors, the levelized cost of solar energy (LCOE), power purchase agreement (PPA) prices, wholesale market value, net value, and interconnection queue data.

analysis↗

High‐speed 4‐dimensional scanning transmission electron microscopy using compressive sensing techniques

Abstract Here we show that compressive sensing allows 4‐dimensional (4‐D) STEM data to be obtained and accurately reconstructed with both high‐speed and reduced electron fluence. The methodology needed to achieve these results compared to conventional 4‐D approaches requires only that a random subset of probe locations is acquired from the typical regular scanning grid, which immediately generates both higher speed and the lower fluence experimentally. We also consider downsampling of the detector, showing that oversampling is inherent within convergent beam electron diffraction (CBED) patterns and that detector downsampling does not reduce precision but allows faster experimental data acquisition. Analysis of an experimental atomic resolution yttrium silicide dataset shows that it is possible to recover over 25 dB peak signal‐to‐noise ratio in the recovered phase using 0.3% of the total data. Lay abstract : Four‐dimensional scanning transmission electron microscopy (4‐D STEM) is a powerful technique for characterizing complex nanoscale structures. In this method, a convergent beam electron diffraction pattern (CBED) is acquired at each probe location during the scan of the sample. This means that a 2‐dimensional signal is acquired at each 2‐D probe location, equating to a 4‐D dataset. Despite the recent development of fast direct electron detectors, some capable of 100kHz frame rates, the limiting factor for 4‐D STEM is acquisition times in the majority of cases, where cameras will typically operate on the order of 2kHz. This means that a raster scan containing 256^2 probe locations can take on the order of 30s, approximately 100‐1000 times longer than a conventional STEM imaging technique using monolithic radial detectors. As a result, 4‐D STEM acquisitions can be subject to adverse effects such as drift, beam damage, and sample contamination. Recent advances in computational imaging techniques for STEM have allowed for faster acquisition speeds by way of acquiring only a random subset of probe locations from the field of view. By doing this, the acquisition time is significantly reduced, in some cases by a factor of 10‐100 times. The acquired data is then processed to fill‐in or inpaint the missing data, taking advantage of the inherently low‐complex signals which can be linearly combined to recover the information. In this work, similar methods are demonstrated for the acquisition of 4‐D STEM data, where only a random subset of CBED patterns are acquired over the raster scan. We simulate the compressive sensing acquisition method for 4‐D STEM and present our findings for a variety of analysis techniques such as ptychography and differential phase contrast. Our results show that acquisition times can be significantly reduced on the order of 100‐300 times, therefore improving existing frame rates, as well as further reducing the electron fluence beyond just using a faster camera.

Robinson, Alex W.↗

The advancing wave front on a sloping channel covered by a rod canopy following an instantaneous dam break

The drag coefficient Cd for a rigid and uniformly distributed rod canopy covering a sloping channel following the instantaneous collapse of a dam was examined using flume experiments. The measurements included space x and time t high resolution images of the water surface h(x, t) for multiple channel bed slopes So and water depths behind the dam Ho along with drag estimates provided by sequential load cells. Using these data, an analysis of the Saint-Venant equation (SVE) for the front speed was conducted using the diffusive wave approximation. An inferred Cd=0.4 from the h(x, t) data near the advancing front region, also confirmed by load cell measurements, is much reduced relative to its independently measured steady-uniform flow case. This finding suggests that drag reduction mechanisms associated with transients and flow disturbances are more likely to play a dominant role when compared to conventional sheltering or blocking effects on Cd examined in uniform flow. The increased air volume entrained into the advancing wave front region as determined from an inflow–outflow volume balance partly explains the Cd reduction from unity.

Mechanics↗

Expanding NSI searches at NOvA

NOvA is an accelerator-based long-baseline neutrino experiment with two functionally equivalent detectors, designed to study neutrino oscillations. NOvA has also been able to look for signals of new physics like non-standard interactions with matter, setting constraints on the parameters governing that beyond-standard neutrino-physics phenomenon. With data collection progressing, and an upgraded analysis including new data samples and improved understanding of the systematics, we are able to further advance our quest of constraining new physics. Here we will present an update on the analysis status of NOvA on the NSI parameters when the addition of more data and a set of low-energy electron neutrino events not considered in our previous analysis.

Acero Ortega, Mario Andres [U. Atlantico, Barranqu↗

Statistics and sensitivity of axion wind detection with the homogeneous precession domain of superfluid helium-3

The homogeneous precession domain (HPD) of superfluid He 3 has recently been identified as a detection medium which might provide sensitivity to the axion-nucleon coupling g a N N competitive with, or surpassing, existing experimental proposals. In this work, we make a detailed study of the statistical and dynamical properties of the HPD system in order to make realistic projections for a full-fledged experimental program. We include the effects of clock error and measurement error in a concrete readout scheme using superconducting qubits and quantum metrology. This work also provides a more general framework to describe the statistics associated with the axion gradient coupling through the treatment of a transient resonance with a nonstationary background in a time-series analysis. Incorporating an optimal data-taking and analysis strategy, we project a sensitivity approaching g a N N ∼ 10 − 12 GeV − 1 across a decade in axion mass. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Constraints on cosmology and baryonic feedback with joint analysis of Dark Energy Survey Year 3 lensing data and ACT DR6 thermal Sunyaev-Zel'dovich effect observations

We present a joint analysis of weak gravitational lensing (shear) data obtained from the first three years of observations by the Dark Energy Survey and thermal Sunyaev-Zel'dovich (tSZ) effect measurements from a combination of Atacama Cosmology Telescope (ACT) and Planck data. A combined analysis of shear (which traces the projected mass) with the tSZ effect (which traces the projected gas pressure) can jointly probe both the distribution of matter and the thermodynamic state of the gas, accounting for the correlated effects of baryonic feedback on both observables. We detect the shear$~\times~$tSZ cross-correlation at a 21$\sigma$ significance, the highest to date, after minimizing the bias from cosmic infrared background leakage in the tSZ map. By jointly modeling the small-scale shear auto-correlation and the shear$~\times~$tSZ cross-correlation, we obtain $S_8 = 0.811^{+0.015}_{-0.012}$ and $\Omega_{\rm m} = 0.263^{+0.023}_{-0.030}$, results consistent with primary CMB analyses from Planck and P-ACT. We find evidence for reduced thermal gas pressure in dark matter halos with masses $M < 10^{14} \, M_{\odot}/h$, supporting predictions of enhanced feedback from active galactic nuclei on gas thermodynamics. A comparison of the inferred matter power suppression reveals a $2-4\sigma$ tension with hydrodynamical simulations that implement mild baryonic feedback, as our constraints prefer a stronger suppression. Finally, we investigate biases from cosmic infrared background leakage in the tSZ-shear cross-correlation measurements, employing mitigation techniques to ensure a robust inference. Our code is publicly available on GitHub.

Pandey, S. [Johns Hopkins U.; Columbia U.] (ORCID:↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

HOD-dependent systematics in Emission Line Galaxies for the DESI 2024 BAO analysis

The Dark Energy Spectroscopic Instrument (DESI) will provide precise measurements of Baryon Acoustic Oscillations (BAO) to constrain the expansion history of the Universe and set stringent constraints on dark energy. Therefore, precise control of the global error budget due to various systematic effects is required for the DESI 2024 BAO analysis. In this work, we estimate the level of systematics induced in the DESI BAO analysis due the assumed Halo Occupation Distribution (HOD) model for the Emission Line Galaxy (ELG) tracer. We make use of mock galaxy catalogs constructed by fitting various HOD models to early DESI data, namely the One-Percent survey data. Our analysis includes typical HOD models for the ELG tracer used in the literature as well as extensions to the baseline models. Among the extensions, we consider various recipes for galactic conformity and assembly bias. We use 25 AbacusSummit simulations under the ΛCDM cosmology for each HOD model and perform independent analyses in Fourier space and in configuration space. To recover the BAO signal from our mocks we perform BAO reconstruction and apply the control variates technique to reduce sample variance noise. Our BAO analyses can recover the isotropic BAO parameter α iso within 0.1% and the Alcock Paczynski parameter α AP within 0.3%. Overall, we find that the systematic error due to the HOD dependence is below 0.17%, with the Fourier space analysis being more robust against the HOD systematics. We conclude that our analysis pipeline is robust enough against the HOD systematics for the ELG tracer in the DESI 2024 BAO analysis, for the assumptions made.

79 ASTRONOMY AND ASTROPHYSICS↗

Temperature‐Dependent Crystallization in Two‐Step Perovskite Deposition Revealed by In Situ GIWAXS and Machine Learning‐Guided Analysis

The performance and stability of perovskite solar cells are strongly governed by the crystallization behavior of their active layer. In two-step sequential deposition, early-stage film formation plays a decisive role in determining final phase purity and device quality. Guided by a data-driven analysis of nearly 39 000 devices in the FAIR perovskite database, we identified solvent-mediated quenching and thermal processing as key variables affecting power conversion efficiency (PCE), particularly in two-step fabrication. Here, to investigate these effects in real time, we designed and implemented a custom-built, temperature-controlled spin-coating system, enabling precise thermal modulation during precursor deposition. Using this platform, we performed in situ GIWAXS measurements to study the crystallization dynamics of FA 0.5 MA 0.5 PbI 3 films over a temperature range of 30°C–90°C. Our results reveal a non-monotonic relationship between spin-coating temperature and α-phase formation, governed by the interplay between precursor interdiffusion, PbI 2 crystallinity, and δ-phase suppression. The custom thermal control enabled us to isolate and quantify these competing effects during the earliest stages of film formation, providing mechanistic insight into how spin-coating temperature governs both phase purity and kinetic pathways in two-step perovskite systems. Temperature-dependent SEM and photovoltaic device measurements further demonstrate that early-stage crystallization pathways directly translate into differences in morphology, charge-transport continuity, and device performance. These findings inform targeted strategies for optimizing deposition protocols to balance rapid nucleation, phase stability, and device performance.

Saadawy, Ahmed [King Fahd University of Petroleum ↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

EV Profile Capture

NextGen Profiles' EV profile capture efforts aimed to explore the variance in performance and evaluate how different operational conditions influence production EV charging behavior. Data were collected at a frequency of 10 Hz from both the EV and EVSE during each charge session. These charge session parameters were then entered into a time-series database for further analysis. The data were gathered under different operational conditions to examine the effects of various factors such as battery state of charge, battery temperature, vehicle condition, smart charge management, and EVSE limitations. The EV profile capture dataset includes extensive high-power charging data from 16 different EVs—comprising light-, medium-, and heavy-duty vehicles—along with EVSE from various suppliers. To protect confidentiality, the EV and EVSE metadata are anonymized, and the publicly released datasets are aggregated to 0.1-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Schema Elements for Granta Annual Report: FY2024

Granta: Materials Intelligence (Granta: MI) is a commercial database software distributed by Ansys, Inc. that is utilized by the Nuclear Security Enterprise (NSE) to organize and store relevant materials data. Lack of standard and well-documented database schema is the primary obstacle to an NSE materials data management solution, so the objective of this project is to create and document such a schema. In FY21, an approach for designing, documenting, and managing a standard database schema was described based on the creation of schema elements (collections of attributes used to describe particular aspects of the data) to be used as building blocks for creating various database tables without duplication. In FY22, these methods were applied through a multi-site collaboration to create and document the schema elements necessary to build a thermogravimetric analysis (TGA) testing table. In FY23 the schema was expanded to include elements for a differential scanning calorimetry (DSC) table, along with schema for supporting metadata tables including Instruments, Projects, Documents, and Testing Series. In FY24 the following progress was made, again through multi-site collaboration: • The existing schema elements were modified to accommodate thermomechanical analysis (TMA) data, and a table, Test Data: TMA, was created for managing TMA data. • The elements necessary for the following additive manufacturing (AM) data tables (directed at data specific to selective laser sintering AM technology) were created: • AM Builds • AM Processes • AM Part Designs • Built AM Parts • AM Feedstock Materials • AM Feedstock Material Batches • The elements necessary for creating a Calibrated Material Models table were created, and the Calibrated Material Models table was created. In FY25 the existing schema will be deployed on the production enterprise Granta instance on the enterprise secure network. Schema elements will be appended, and new elements created as necessary, to allow the creation of tables specifically to support materials testing, AM process development, and design and analysis for modernization programs.

36 MATERIALS SCIENCE↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Electromagnetic Transient Modeling of Large Data Centers for Grid-Level Studies: Beta Release

The magnitude and complexity of electricity usage patterns from large data centers are having significant impacts on the operation and dynamics of the power grid; grid operators and planners require a range of specialized data center models to properly evaluate these impacts and specify technical solutions as needed. Towards addressing this need, Pacific Northwest National Laboratory (PNNL) has developed a library of electromagnetic transient (EMT) models for grid-level studies of data centers called the data center model library (DML). This report describes how the DML was created and how it may properly be used. This report details the DML’s beta release, completed in July 2026. This is a revision and expansion of the alpha release, which was made available in January 2026 The models present in the DML are generic models; subject matter expertise and additional technical data are needed to modify these models before they can represent any real data center. However, they will significantly reduce the level of effort required to develop site-specific models and can serve as a common starting point to guide industry towards a more refined consensus. Most of the models within DML are dedicated to representing the power electronics interfaces commonly used in modern data centers, such as double-conversion uninterruptible power supplies and single-phase power factor correction converters. These models are intended for use in grid-level studies and are a simplified aggregation of many small components. That said, background material on the physical and electrical design of large data centers is provided as companion material so that users can be aware of many of the details which have been omitted or streamlined as a matter of practical necessity. Additionally, guidance on the application of EMT analysis for data center interconnection studies is provided, which aids users in identifying when the DML is necessary and what sort of additional model development may be necessary for conducting real-world studies.

electromagnetic transients↗

Machine learning inversion of interatomic force constants from single-crystal inelastic neutron scattering

Atomic vibrations govern many macroscopic properties of materials, but experiments to comprehensively probe them remain challenging. Inelastic neutron scattering (INS) is a powerful technique to map phonon dispersions in crystals, especially when leveraging modern time-of-flight (ToF) spectrometers with large detectors. However, efficiently and robustly extracting interatomic force constants (FCs) parameterizing phonon dynamics from experimental spectra remains a bottleneck due to the complexity and high dimensionality of ToF INS datasets. Here, we present a machine learning approach for the direct inversion of FCs from single-crystal INS measurements. The framework leverages synthetic training data generated using universal machine-learned force fields and an efficient physics-based forward model. We benchmark two neural architectures–one emphasizing structured latent representation learning and the other direct, supervised spectral regression–across simulated datasets for two materials under idealized and noisy conditions. The latent-representation model is subsequently applied to experimental single-crystal INS data on germanium. The model is shown to reproduce FCs derived from both first-principles simulations and from iterative optimization, and furthermore achieves reliable inference even from sparse, single-orientation measurements representing short data acquisitions. Analysis of the learned latent space reveals semantically continuous and physically interpretable encodings that support strong cross-domain generalization. By bridging theoretical and experimental domains, we establish a path toward rapid inversion of experimental spectra and data-driven interpretation of temperature-dependent lattice dynamics.

42 ENGINEERING↗