Search NASASearch

SEARCH · Search NASA

Results for “Factorization machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Mechanism of Antiferroelectricity in Polycrystalline ZrO 2

The size and electric field dependent induction of polarization in antiferroelectric ZrO 2 is the key to several technological applications that are unimaginable a decade ago. However, the lack of a deeper understanding of the mechanism hinders progress. Molecular dynamics simulations of polycrystalline ZrO 2 , based on machine-learned interatomic forces with near ab initio quality, shed light on the fundamental mechanism of the size effect on the transition fields. Stress in the oxygen sublattice is the most important factor. The so constructed interatomic forces allow the calculation of the transition fields as a function of the ZrO 2 film thickness and predict the ferroelectricity at large thickness. The simulation results are validated with electrical and piezo response force microscopy measurements. The results allow a clear interpretation of the properties of the double-hysteresis loops as well as the construction of the free energy landscape of ZrO 2 grains.

36 MATERIALS SCIENCE

Developing a Prototype Methodology to Rank CO2-EOR Wells and Assess Their Reuse Potential for Geologic Carbon Storage

This paper presents a prototype methodology to assess the possible transition of Class II carbon dioxide-enhanced oil recovery (CO2-EOR) wells to Class VI wells. The focus is on wellbore construction materials—casing, cement, tubing, and the packer—and includes comprehensive workflows to evaluate these materials, with primary emphasis on compliance with Environmental Protection Agency (EPA) Class VI well construction and conversion guidelines. These workflows systematically assess material properties and performance criteria to ensure regulatory compliance and optimize long-term wellbore integrity and functionality. Utilizing Python scripts and JavaScript Object Notation (JSON) representations, the study automates checks on digitized Texas Railroad Commission (TRRC) data to rank wells based on workflow criteria. By emphasizing critical factors such as casing integrity, cementing techniques, tubing compatibility, and packer selection, the methodology helps well owners and operators prioritize wells for potential reuse as CO2 injection wells. Given limitations in digitized data, manual user verification is required in some sections. Future improvements include integrating non-digitized data through web scraping and machine learning techniques. This research serves as a practical guide for stakeholders, supporting environmental compliance and sustainable well operations.

geologic carbon sequestration

Exploring new frontiers in type 1 diabetes through advanced mass-spectrometry-based molecular measurements

Type 1 diabetes (T1D) is a devastating autoimmune disease for which advanced mass spectrometry (MS) methods are increasingly used to identify new biomarkers and better understand underlying mechanisms. For example, integration of MS analysis and machine learning has identified multimolecular biomarker panels. In mechanistic studies, MS has contributed to the discovery of neoepitopes, and pathways involved in disease development and identifying therapeutic targets. However, challenges remain in understanding the role of tissue microenvironments, spatial heterogeneity, and environmental factors in disease pathogenesis. Recent advancements in MS, such as ultra-fast ion-mobility separations, and single-cell and spatial omics, can play a central role in addressing these challenges. Here, in this work, we review recent advancements in MS-based molecular measurements and their role in understanding T1D.

60 APPLIED LIFE SCIENCES

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi

A moment-conserving discontinuous Galerkin representation of the relativistic Maxwellian distribution

Kinetic simulations of relativistic gases and plasmas are critical for understanding diverse astrophysical and terrestrial systems, but the accurate construction of the relativistic Maxwellian, the Maxwell–Jüttner distribution, on a discrete simulation grid is challenging. Difficulties arise from the finite velocity bounds of the domain, which may not capture the entire distribution function, as well as errors introduced by projecting the function onto a discrete grid. Here, we present a novel scheme for iteratively correcting the moments of the projected distribution applicable to all grid-based discretizations of the relativistic kinetic equation. In addition, we describe how to compute the needed nonlinear quantities, such as Lorentz boost factors, in a discontinuous Galerkin scheme through a combination of numerical quadrature and weak operations. The resulting method accurately captures the distribution function and ensures that the moments match the desired values to machine precision.

astrophysical plasmas

Boron Coordination in Multicomponent Glasses: Analytical Models and Machine Learning With Uncertainty

Borosilicate glasses are extensively used in a variety of applications from kitchenware to nuclear waste immobilization due to the strong network formed by the Si-O-B bond that makes it resistant to chemical corrosion and gives it a low thermal expansion. Boron, however, exists in both trigonal BO3 and tetrahedral BO4 bonds in glass systems, which impacts the chemical durability and thermal resistance of the glass, amongst other properties. Boron coordination (N4), or the ratio of the amount of BO4 to BO3 within a glass, may aid in predicting these properties but is difficult to derive without experimental data due to the complexity of impacts from varied glass compositions and processing factors. For this reason, compositional models have been developed to predict boron coordination, but the models typically include a limited number of glass components. To help fill this gap in the models, in this work, a diverse multicomponent glass dataset of 809 glasses is compiled from a literature search, and then a number of analytical and machine learning (ML) models are trained on the dataset. Previously developed modified Bernstein and modified Du Stebbins analytical models were fitted to update parameters with the new dataset. Then, partially Bayesian neural networks, Gaussian process regressor, and heteroskedastic deterministic neural networks were evaluated. The ML models examined all have different strategies to overcome the potential for overfitting as a result of a limited training dataset, and return results that account for model uncertainty, which can be valuable for understanding model reliability. For the first time, cooling rate is introduced as an input parameter for ML models, showing consistent improvements in performance and solidifying the importance of including parameters outside of composition alone for N4 prediction. The machine learning models examined here show promise in accurate predictions of boron coordination in borosilicate glasses, all achieving R2 values of 0.91.

boron coordination

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

Mapping tree height in complex terrain of northern China using ultra-high-resolution images

Tree height is a key parameter for estimating forest biomass and carbon sequestration. In recent years, notable progress has been made in mapping tree height using satellite imagery. However, existing tree height products show low accuracy in mountainous and complex terrains, and few studies typically addressed tree height estimations in mountain areas. This study examines the Mentougou district of Beijing, China, characterized by complex terrain and mountainous landscapes. We analyzed two methods for estimating tree height: one using only spectral features and another combining spectral features with topographic factors (elevation, slope, aspect). We used 3-m resolution PlanetScope 8-band multispectral imagery, with 710 field-measured individual tree heights averaged to obtain 471 pixel-level tree height values as ground-truth, to develop tree height prediction models using eXtreme Gradient Boosting (XGBoost), Random Forest (RF), and Gradient Boosting Machine (GBM) models. The results show that the XGBoost model consistently presented the highest accuracy for both methods evaluated. Specifically, the XGBoost model that combined spectral data with elevation and slope variables with an R² of 0.75 and an RMSE of 2.69 m. Using the XGBoost model, we generated the tree height map for the Mentougou area at 3 m resolution, showing tree heights ranging from 0.5 to 30.4 m, and the model’s prediction error standard deviations ranged from 2.50 to 4.71 m, indicating reliable performance across varied terrain. Additionally, we compared and evaluated the global tree height products, identifying limitations in the accuracy within complex terrains. This study demonstrates the potential for accurately predicting tree heights by combining high-resolution multispectral satellites with a terrain factor modeling approach.

Complex terrain

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling

Entire four-graviton EFT from the duality between color and kinematics

The Bern-Carrasco-Johansson (BCJ) double-copy construction reveals a fundamental structural connection between gauge and gravity theories. At its core, the BCJ double copy is directly due to a duality between the algebraic relations of a color root and those of a kinematic root. We generalize this principle beyond the conventional Lie algebra structure of tree-level Yang-Mills theory. By demanding color-kinematics duality for the complete basis of four-point color structures—including those involving the symmetric 𝑑 𝑎⁢𝑏⁢𝑐 constants—we define the universal double copy. We systematically classify the bases of all such parity-even generalized gauge-theory numerators and, independently, the space of all parity-even four-graviton higher-derivative operators. We demonstrate that our universal double-copy construction precisely spans the entire tower of parity-even four-graviton amplitudes in any dimension, except for the Lovelock 𝑅 3 contribution in 𝐷 > 6 which we can express in terms of a particularly simple universal triple-copy involving gauge theories coupled to scalars. Explicit machine-readable expressions for the complete basis of gauge-theory numerators and fundamental gravitational building blocks are provided in the Supplemental Material. This establishes that all possible four-point gravitational interactions can be factorized into products of gauge-theory building blocks governed by this universal notion of color-kinematics duality.

Carrasco, John Joseph M. [Northwestern Univ., Evan

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING

Dynamics and lipid membrane coupling of the RAS-RAF complex revealed via multiscale simulations

To gain molecular and mechanistic insights into initiation of the RAS-RAF signaling cascade, we developed and used a combination of multiscale simulation and experimental approaches. The influence and impact of the membrane on RAS and RAF proteins is a factor we are just beginning to understand and appreciate in more detail. Molecular simulation is an ideal methodology to further study this complicated relationship between the membrane and associated proteins. Our previous work using Multiscale Machine-learned Modeling Infrastructure investigated different lipid compositions solely around the KRAS4b protein and the interplay between protein behavior and these membrane environments. Multiscale Machine-learned Modeling Infrastructure uses machine learning to couple adjacent simulation scales and has been efficiently scaled across some of the world’s largest high-performance computers. Recently, we have expanded this multiresolution framework to include the all-atom simulation scale and to incorporate the RAF RBDCRD domains. Here, we present the overall analysis results from this new simulation campaign comprising a mixture of RAS and RAF RBDCRD proteins. Approximately 35,000 coarse-grained and 10,000 all-atom molecular dynamics simulations were completed, sampled from a variety of protein/lipid composition configurations that were generated from a micron-scale continuum simulation containing hundreds of copies of the proteins. Our studies suggest that orientations of the RAS-RBDCRD complex on the membrane occupy distinct configurational states, and the spatial patterns of lipid arrangements around these different protein states are unique to each state. The extent and size of lipid “fingerprints” imposed on the membrane by the RAS-RBDCRD protein complex are significantly larger than observed for just the RAS protein on its own. These protein complexes strongly associate, but we do not observe statistically significant preferred protein-protein orientations. These observations indicate that spatial colocalization of RAS-RBDCRD proteins in the same vicinity may be assisted by specific membrane environments, acting to increase the probability of signaling complex formation.

Carpenter, Timothy S. [Lawrence Livermore National

Developing a Prototype Methodology to Rank CO2-EOR Wells and Assess Their Reuse Potential for Geologic Carbon Sequestration

This study presents a prototype methodology for evaluating the reuse potential of Class II CO₂-enhanced oil recovery (CO₂-EOR) wells as Class VI wells for geologic carbon sequestration. The approach focuses on assessing wellbore construction materials—casing, cement, tubing, and packers—based on U.S. Environmental Protection Agency (EPA) Class VI well conversion guidelines. Utilizing Python scripts and JSON representations, the methodology automates checks on digitized Texas Railroad Commission (TRRC) data to rank wells based on regulatory and integrity criteria. Key factors include casing integrity, cementing techniques, tubing compatibility, and packer selection. Due to limitations in digitized data, manual verification is required for certain sections. Future enhancements include incorporating non-digitized data via web scraping and machine learning. This research provides a practical framework for well owners and regulators, supporting informed decision-making for sustainable CO₂ storage.

carbon sequestration

FEDERATED LEARNING ON STOCHASTIC NEURAL NETWORKS

Federated learning is a machine learning paradigm that leverages edge computing on client devices to optimize models while maintaining user privacy by ensuring that local data remain on the device. However, since all data are collected by clients, federated learning is susceptible to latent noise in local datasets. Factors such as limited measurement capabilities or human errors may introduce inaccuracies in client data. To address this challenge, we propose the use of a stochastic neural network as the local model within the federated learning framework. Stochastic neural networks not only facilitate the estimation of the true underlying states of the data but also enable the quantification of latent noise. We refer to our federated learning approach, which incorporates stochastic neural networks as local models, as federated stochastic neural networks. In this work we will present numerical experiments demonstrating the performance and effectiveness of our method, particularly in handling nonindependent and identically distributed data.

97 MATHEMATICS AND COMPUTING

Position Papers for Inverse Methods for Complex Systems under Uncertainty Workshop

The ability to solve inverse problems – inferring unknown parameters, structures, or states of a system from observed data – is essential for advancing scientific discovery and innovation capabilities for the DOE mission. Basic research needs and challenges are particularly acute in emerging areas such as the interactive, data-driven, modeling and simulation of digital twins; decision support for experiments at DOE scientific user facilities; and for other complex systems and workflows. Inverse problems are at the heart of understanding and controlling complex systems due to factors such as observational data with varying modalities and fidelities, inherent uncertainties in physical measurements and numerical models, and the computational demands of rapid and high-fidelity simulations. The convergence of recent scientific computing trends – scientific machine learning, artificial intelligence, and computing advances such as exascale computing – is creating unprecedented opportunities. These advancements offer the potential to revolutionize how we approach inverse problems to extract actionable insights with the required level of accuracy and computational efficiency. This workshop and the Call for Position Papers are vital steps in bringing together experts to collectively explore and identify the new computational and mathematical directions needed in inverse methods for complex systems under uncertainty.

97 MATHEMATICS AND COMPUTING

Evolution and Degradation Patterns of Electrochemical Cells Based on the Analysis of Interfacial Phenomena at Li Metal Anode/Electrolyte Interfaces

In this work, we report the results of a theoretical–computational analysis of the solid electrolyte interphase (SEI) growth and degradation dynamics occurring in lithium metal batteries during cycling. We use ab initio-kinetic Monte Carlo simulations to generate a synthetic data set, which is analyzed by machine learning methods. We aim to determine: (i) how modifications in interfacial interaction energies between solid electrolyte interphase (SEI) blocks and between Li ions and SEI facets impact the Coulombic efficiency (CE) of the battery and (ii) what factors, including reactions, microscopic transport, and other interfacial events, may lead to cell performance “failure” during prolonged charge and discharge cycles, signaled as a sharp decay in the CE over cycling. The demonstration of our approach is done on a cell including a Li metal surface interfacing with a previously introduced state-of-the-art electrolyte, and the idea can be applied to any electrochemical system. Outcomes include the identification of the leading chemical, physical, and structural variables causing cell failure and relating them to the electrolyte formulation, thus paving the way to future more refined analysis and electrolyte design.

batteries

Deterministic High-Fidelity Neutronics Simulation of Pebble Bed Reactors Using Pebble Tracking Transport

The pebble tracking transport (PTT) algorithm offers a high-fidelity deterministic approach for neutron transport for pebble bed reactors (PBRs). This approach requires the mesh for the active-core region to consist exclusively of tetrahedral elements, where each node in the pebble-packing region represents a pebble centroid. This paper investigates the application of PTT for full-scale PBRs, considering both the isothermal and the temperature-dependent core conditions. Macroscopic cross sections are generated using Serpent 2 full-core eigenvalue simulations where pebbles are grouped into disjoint subsets using machine learning. To minimize the need for individual cross-section sets for each pebble in the core, K-means clustering is used to group pebbles by temperature and neutronic environment parameters. Here, we compare the multiplication factor and power rate distributions between PTT simulations using the Griffin reactor physics software and reference solutions from Serpent 2. Our analysis shows that a full-core, high-fidelity PTT calculation produces accurate results with minimal local (pebblewise) errors. Additionally, timing results indicate that PTT simulations converge rapidly on modern supercomputing platforms.

Griffin

Recent advances in plasma control and physics research in the Large Helical Device

The Large Helical Device (LHD), the largest superconducting helical system in the world, is equipped with advanced heating and diagnostic tools, facilitating plasma control and physics research. Data assimilation was employed for electron temperature control using a real-time Thomson scattering system and real time prediction code. A virtual LHD environment enabled visualization of escaping high-energy tritium ions and demonstrated that these ions impact the rear side of the divertor plate. Pioneering results crucial to plasma control have also been achieved. Real-time wall conditioning using Lithium granule dropping improved bulk ion energy and particle transport while simultaneously enhancing the heavy impurity transport. Progress has also been made in the investigation of turbulence-driven transport. At the confinement bifurcation, ion-scale turbulence decreased, while electron-scale turbulence increased. A change in the anisotropy of turbulent eddies was also observed at the confinement bifurcation. Coexistence of local and non-local turbulence was identified in electron-scale turbulence. Non-local turbulence exhibited the rapid spatial propagation of perturbations throughout the plasma, while local turbulence followed the temperature gradient. A transition between drift-wave turbulence and magnetohydrodynamics (MHD) turbulence was observed with the turbulence minimized at the transition condition. Machine learning analysis was employed to evaluate the temperate and density conditions of this turbulence transition. Then, real-time control of fueling and heating was applied to maintain the turbulence transition condition, improving the energy confinement enhancement factor by 20%. In addition, evidence was obtained for collisionless ion heating by energetic-ion-driven geodesic acoustic modes and MHD bursts. These achievements represent unique contributions to the development of fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY