Search NASASearch

SEARCH · Search NASA

Results for “Factorization machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Identifying Anomalous DESI Galaxy Spectra with a Variational Autoencoder

The tens of millions of spectra being captured by the Dark Energy Spectroscopic Instrument (DESI) provide tremendous discovery potential. In this work we show how Machine Learning, in particular Variational Autoencoders (VAE), can detect anomalies in a sample of approximately 200,000 DESI spectra comprising galaxies, quasars and stars. We demonstrate that the VAE can compress the dimensionality of a spectrum by a factor of 100, while still retaining enough information to accurately reconstruct spectral features. We then detect anomalous spectra as those with high reconstruction error and those which are isolated in the VAE latent representation. The anomalies identified fall into two categories: spectra with artefacts and spectra with unique physical features. Awareness of the former can help to improve the DESI spectroscopic pipeline; whilst the latter can lead to the identification of new and unusual objects. To further curate the list of outliers, we use the Astronomaly package which employs Active Learning to provide personalised outlier recommendations for visual inspection. In this work we also explore the VAE latent space, finding that different object classes and subclasses are separated despite being unlabelled. We demonstrate the interpretability of this latent space by identifying tracks within it that correspond to various spectral characteristics. For example, we find tracks that correspond to increasing star formation and increase in broad emission lines along the Balmer series. In upcoming work we hope to apply the methods presented here to search for both systematics and astrophysically interesting objects in much larger datasets of DESI spectra.

Nicolaou, C. [University Coll. London] (ORCID:0000

A Near-Real-Time Model for Predicting Electricity Disruptions in Texas During Winter Storms

There has been an increase in extreme weather events, posing a threat to power grid systems, potentially influenced by factors such as population growth, changes in ecosystems, land cover, and land use in the service area, as well as the growth of certain vegetation types. This research seeks to develop a predictive model to mitigate potential damages caused by future winter storms. This research utilizes the Light Gradient Boosting Machine (LightGBM), incorporating the number of power outages experienced at the county level, geographic details, weather information, and lagged outage and lagged weather data. The developed models were broadly divided into two groups, with six models in each group - one group without optimization and another with optimization, totaling 12 trained models. For model optimization, Bayesian optimization was employed using Root Mean Squared Error (RMSE) as the objective function. In results, when comparing Group 2 (the optimized group) with Group 1 (the non-optimized group), it was found that optimization did not always lead to a reduction in RMSE and Mean Absolute Error (MAE). However, in terms of Mean Directional Accuracy (MDA), while all results in Group 1 were below the baseline accuracy of 0.33, all results in Group 2 exceeded 0.33, with some cases showing an increase of more than three times the baseline. The results indicated that, in the optimized model group, Population and Pressure were the most influential factors when using current weather data and geographical information. When using lagged data, lagged recorded outages and lagged Pressure emerged as the most significant factors. Among the 12 developed models, the L-1-2-O model showed the lowest RMSE and MAE, as well as the highest accuracy, with values of 390.62 households and 168.13 households, respectively. To normalize the RMSE and MAE values, each metric was divided by the average number of households among the counties in Texas. For the L-1-2-O model, the scaled RMSE was 0.88% and the scaled MAE was 0.38%. In terms of MDA, which indicates the accuracy of the prediction direction, the L-1-O model achieved the highest score of 0.41. Although this study focused on Texas, which suffered the greatest impact from the winter storms in 2021, with additional validation, the methodology used in this research could be applied to other regions.

Lee, Jangjae [Texas A & M Univ., College Station,

Side-by-Side Comparison of Subhourly Clipping Models

Over the past several years there have been numerous attempts at quantifying the inherent power clipping of inverters due to subhourly irradiance variability that is not captured in hourly PV performance models. Different models have been proposed to correct for these clipping losses in PV performance estimates, including matrix lookup models, distribution modeling of the PV power performance within a given hour, and machine learning methods. To date, there have been few comprehensive quantitative comparisons of these inverter clipping correction modeling approaches to evaluate the effectiveness of these approaches in predicting the actual behavior of PV system inverter clipping. In this study, we perform such a comparison, evaluating the Allen and Walker correction loss modeling approaches recently implemented in the System Advisor Model (SAM) against clipping losses modeled with 1-minute climate data. These comparisons were performed across a variety of climate locations and inverter loading ratios to thoroughly analyze the effectiveness of these modeling approaches relative to each other. Results from this analysis reveal that both clipping correction approaches improve annual energy accuracy to within 2% of 1-minute modeled energy yield. The two models predict annual clipping loss more accurately than simple hourly power limit clipping, with the Allen method typically being slightly more accurate at typical ILR values and the Walker method often being slightly more accurate at high ILR values The models can improve accuracy over the status quo clipping approach up to 3 percentage points in systems with ILR of 2.0, showing the importance of this modeling factor in energy yield estimates.

accuracy

Quantitative Analysis and Prediction of Thermal Runaway Metrics of High-Nickel Oxide Cathodes by Machine Learning Models

The pursuit of higher energy density in lithium-ion batteries has made high-nickel (Ni) layered oxides leading cathode candidates for next-generation electric vehicles. However, their poor thermal stability, particularly at Ni contents ≥ 90%, increases the risk of cathode-initiated thermal runaway. Furthermore, we present a data-driven framework combining linear and nonlinear machine learning models to predict key thermal runaway descriptors from a high-throughput differential scanning calorimetry database. With cathode composition and state of charge (SOC) as input features, the ensemble model accurately predicts peak temperature, heat release, and peak heat flow. SHAP analysis identifies Ni content and SOC as the dominant factors controlling thermal runaway temperature, while SOC primarily governs heat release and peak heat flow. Al, Mg, and Mn improve thermal stability by strengthening metal–oxygen bonding and delaying structural transformation, whereas B mainly reduces heat release through surface passivation. Validation with a new cathode composition confirms accurate prediction of SOC-dependent thermal runaway behavior and critical SOC.

25 ENERGY STORAGE

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics

CAMELSH: A Large-Sample Hourly Hydrometeorological Dataset and Attributes at Watershed-Scale for CONUS

We present CAMELSH (Catchment Attributes and Hourly HydroMeteorology for Large-Sample Studies), the first large-sample hydrometeorological dataset at the hourly scale for the contiguous United States. CAMELSH intergrates hourly meteorological time series, catchment attributes and boundaries from GAGES-II and HydroATLAS for 9,008 catchments across diverse climatic, hydrological, and anthropogenic conditions. In addition, hourly streamflow time series is provided for 3,166 catchments. The dataset spans 45 years (1980–2024) with 11 meteorological variables from the NLDAS-2 forcing dataset, from which we compute nine climate indices related to precipitation, evapotranspiration, seasonality, and snow fraction. Additionally, CAMELSH includes two sets of catchment attributes: 439 from GAGES-II and 195 derived from HydroATLAS. These attributes include factors related to climate, geology, hydrology, river/stream morphology, landscape, nutrient, soil, topography, and anthropogenic influences. Developed in accordance with FAIR (Findability, Accessibility, Interoperability, and Reusability) principles, CAMELSH is the first large-sample dataset at an hourly timescale, supporting machine learning applications for short-term streamflow (flood) prediction and advancing data-driven hydrological research across multiple timescales.

54 ENVIRONMENTAL SCIENCES

Interpretable Machine Learning for Characterizing Electric Vehicle Charging Behavior: Insights from Real-World Data

As electric vehicle (EV) adoption rises globally, concerns about the impact on aging electrical grids grow, particularly regarding the charging behavior of EV drivers. This study analyzes real-world driving and charging data from Ford battery electric vehicles (BEVs) collected between 2018 and 2019 to develop interpretable models that characterize charging behavior and quantify influencing factors. Prior research has relied on assumptions regarding driver behavior, often overlooking actual charging patterns. By employing generalized linear mixed models (GLMMs), this work offers insights into how various elements, such as next trip distance and state of charge (SOC), influence charging decisions. The dataset comprises over three million park-trip pairs from 1,997 vehicles, revealing that features related to driving behavior significantly dictate charging behavior, while infrastructure and regional factors have lesser impacts. The findings suggest that existing simulation models may oversimplify EV charging behavior assumptions. This work utilizes real-world EV driving and charging data to train interpretable models that describe charging behavior and quantify the factors most associated with how drivers use charging infrastructure. This research underscores the need for interpretable, data-driven methodologies to inform future EV infrastructure planning and grid management.

29 - ENERGY PLANNING, POLICY AND ECONOMY

Machine learning insights into microstructural origins of transport and mechanical properties in porous microstructures

Multifunctional porous materials are increasingly needed across various fields, but their complex microstructures create significant challenges due to the intricate microstructure-property relationships. This complexity, combined with limitations of traditional analysis methods, hinders efforts to understand and optimize microstructure–property relationships. Here, to address this, we integrate physics-based mesoscale modeling with interpretable machine learning (ML) to uncover how microstructural features govern effective diffusivity and elastic modulus. At constant porosity, we show diffusivity varies by over 150 × and modulus by ∼50 ×, highlighting the power of microstructure engineering. Statistical analysis reveals bimodal behavior in diffusivity and unimodal in modulus. ML identifies connectivity as the dominant factor, while modulus is also sensitive to domain size and feature interactions. Controlled simulations further highlight domain shape as a critical feature for modulus. This framework enables efficient exploration of microstructure-property correlations, offering new insights to guide the design of advanced porous materials.

Bicontinuous microstructure

Error and Correction Analysis for the FFA@CEBAF Energy Upgrade

An energy upgrade design for the Continuous Electron Beam Accelerator Facility (CEBAF) is under development, using fixed field alternating gradient (FFA) return arcs to recirculate electron beam up to an additional five times through the accelerating structures at CEBAF. A necessary component of any large accelerator is a beam steering and optical correction system. Small environmental changes and system errors can lower beam quality or even shut down the machine; and in pursuit of the scientific mission of JLab, high quality electron beams must be delivered to the experimental halls on a predictable schedule. Correction in the novel FFA arcs of the current upgrade design is complicated by several factors. These complexities inform the choice of correction algorithm structure and parameter values. A baseline algorithm in addition to diagnostic and correction hardware configuration is presented. The effect of this correction protocol is shown with respect to estimated errors, and several possible extensions of the algorithm are discussed. This work presents an important proof of concept for the FFA@CEBAF design effort, and provides a functional correction strategy which may be simply adjusted and optimized for future design changes.

Coxe, Alex [Old Dominion Univ., Norfolk, VA (Unite

Micromechanical Surrogate Machine Learning Model for Creep Deformation Modeling

Process variability during the manufacture of gas turbine engine hot section components can significantly affect the material’s resulting microstructure. In casting, for instance, geometric variation within a component (thin sections versus thick sections, radial location) influences cooling rates and the resulting grain size. The high temperature creep response is known to be sensitive to grain size owing to a diffusional creep mechanism which occurs more readily along grain boundaries. Microstructural variation correspondingly drives mechanical behavior which propagates into component scale performance uncertainty. These factors are essential when planning inspection, maintenance, and repair strategies within a reliability framework. These benefits provide opportunities to increase overall energy efficiency through refined margins. Critically, there is an opportunity to bolster existing data-driven reliability models using physics-driven process-structure-property relations. Here we present recent work establishing a framework for evaluating the probabilistic creep performance of high-temperature materials. A novel microstructure-sensitive crystal plasticity finite element model is established that captures both grain boundary and crystallographic deformation effects. The computationally expensive physics model is calibrated using a statistical approach and this high-fidelity model is subsequently used to train a computationally efficient machine learning surrogate model. The surrogate model is essential for sampling a large ensemble of simulated structure-property pair results. The ensemble data are then mined to extract salient trends to be incorporated into a microstructure-sensitive reliability model. The proposed approach represents a novel way to capture microstructure-sensitive trends from physics-based models within a modern reliability framework.

Fernandez-Zelaia, Patxi [ORNL]

Causal machine learning uncovers conditions for convective intensification driven by organic and sulfate aerosols

Aerosols are often hypothesized to invigorate deep convective clouds (DCCs), but observational evidence remains limited and inconclusive. Clarifying this hypothesis is critical for regions vulnerable to thunderstorms and flooding, particularly highly polluted coastal cities. Leveraging a novel causal discovery–inference pipeline and high-resolution observations near Houston, TX, we identify multiple causal pathways among aerosols (mostly organic and sulfate), DCCs, and meteorological factors. However, a direct causal link from aerosols to DCCs is found to be uncommon, occurring in less than 35% of analyzed scenarios, and is characterized by strong conditionality and nonlinearity. When aerosol impacts on DCCs do occur, they can be substantial, enhancing DCC core heights by approximately 1.7 km, with 92% of this effect concentrated in warmer-phase cloud regions. Notably, the presence of sea breezes and the inclusion of all measured aerosol particles each enhance DCCs in over 95% of aerosol-sensitive cases.

54 ENVIRONMENTAL SCIENCES

Leveraging design of experiments to build chemometric models for the quantification of uranium (VI) and HNO3 by Raman spectroscopy

Partial least squares regression (PLSR) and support vector regression (SVR) models were optimized for the quantification of U(VI) (10–320 g L −1 ) and HNO 3 (0.6–6 M) by Raman spectroscopy with optimized calibration sets chosen by optimal design of experiments. The designed approach effectively minimized the number of samples in the calibration set for PLSR and SVR by selecting sample concentrations with a quadratic process model, despite complex confounding and covarying spectral features in the spectra. The top PLS2 model resulted in percent root mean square errors of prediction for U(VI), HNO 3 , and NO 3 − of 3.7%, 3.6%, and 2.9%, respectively. PLS1 models performed similarly despite modeling an analyte with a majority linear response (i.e., uranyl symmetric stretch) and another with more covarying vibrational modes (i.e., HNO 3 ). Partial least squares (PLS) model loadings and regression coefficients were evaluated to better understand the relationship between weaker Raman bands and covarying spectral features. Support vector machine models outperformed PLS1 models, resulting in percent root mean square error of prediction values for U(VI) and HNO 3 of 1.5% and 3.1%, respectively. The optimal nonlinear SVR model was trained using a similar number of samples (11) compared with the PLSR model, even though PLS is a linear modeling approach. The generic D-optimal design presented in this work provides a robust statistical framework for selecting training set samples in disparate two-factor systems. This approach reinforces Raman spectroscopy for the quantification of species relevant to the nuclear fuel cycle and provides a robust chemometric modeling approach to bolster online monitoring in challenging process environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Bridging the Gap Between Modern UX Design and Particle Accelerator Control Room Interfaces

Accelerator control systems often represent relatively complex and safety-sensitive human-machine interfaces within process control industries. These systems are technically robust and reflect the cumulative integration of solutions built and adapted across decades. One of the regular, unfortunate casualties of provisional accelerator control system updates is their human-system interfaces (HSIs) which often lag behind modern usability and design standards. An additional challenge is that although there is a multitude of established human factors (HF), and user experience (UX) principles for everyday digital applications, there are very few (if any) established principles for complex and safety-critical applications for an accelerator. This paper argues for the importance of established HF and UX principles (herein referred to as human-centered design principles) into the development of accelerator HSIs, emphasizing the need for clarity, consistency, responsiveness, and cognitive accessibility. Drawing from HF/UX best practices and human-centered design, this paper discusses how these approaches can enhance operator performance, reduce human error, and improve accelerator personnel collaboration. Case studies from Accelerator Control Operations Research Network (ACORN) at Fermilab are explored to demonstrate how interfaces built with human-centered design principles can scale with system complexity while remaining intuitive and efficient for diverse user roles including operators, machine experts, and engineers. By bridging the gap between traditional control system design and modern human-centered design methods, this paper provides a roadmap for evolving accelerator HSIs into more usable, maintainable, and effective tools.

Hill, Rachael [Idaho Natl. Lab.]

Machine Learning Based Prediction of Airflow Maldistribution in A-Type Heat Exchangers

Airflow maldistribution is one of the primary causes of performance degradation in air-to refrigerant heat exchangers (HX) and has been shown to decrease heat transfer by as much as 35%. As a result, many units are oversized to meet the target capacity, resulting in increased system cost and refrigerant charge. Several studies have explored how characteristics like package type and HX geometry impact the flow profile, but results are restricted to a limited range of parameters and cannot be extrapolated to new designs. In this work, a machine learning (ML) model is trained to predict the inlet flow profile of dry air entering A-type HXs across a broad range of geometries and conditions. Flow profiles are generated using a porous media CFD model and used to train an Artificial Neural Network (ANN) which exhibits maximum and average relative L2 norm errors of 0.48 and 0.05. Additionally, these predictions take less than a second to generate resulting in a speed up factor of 2.42E5 compared to CFD. Component-level simulations are conducted to determine the performance degradation resulting from the predicted airflow maldistribution profiles. The new ML model will enable rapid and accurate prediction of performance degradation resulting from airflow maldistribution in A-type HXs, allowing for more accurate and cost-effective HX design.

42 ENGINEERING

Uncovering Structure–Conductivity Relationships in Anion Exchange Membranes (AEMs) Using Interpretable Machine Learning

Anion exchange membranes (AEMs) play a vital role in the performance of water electrolyzers and fuel cells, yet their discovery and optimization remain challenging due to the complexity of structure–property relationships. In this study, we introduce a machine learning framework that leverages conditional graph neural networks (cGNNs) and descriptor-based models and a hybrid graph neural network (HGARE) to predict and interpret ionic conductivity. The descriptor-based pipeline employs principal component analysis (PCA), ablation, and SHAP analysis to identify factors governing anion conductivity, revealing electronic, topological, and compositional descriptors as key contributors. Beyond prediction, dimensionality reduction and clustering are performed by employing t-SNE and KMeans as well as SOM, which reveal distinct membranes clusters, some of which were enriched with high anion conductivity. Among graph-based approaches, the graph convolutional (GCN) achieved strong predictive performance, while the Hybrid Graph Autoencoder-Regressor Ensemble (HGARE) achieved the highest accuracy. Additionally, atom-level saliency maps from GCN provide spatial explanations for conductive behavior, revealing the importance of polarizable and flexible regions. This work contributes to the accelerated and data-driven design of high-performance AEMs.

Naghshnejad, Pegah [Department of Chemical Enginee

On the hardness of learning ground state entanglement of geometrically local Hamiltonians

Characterizing the entanglement structure of ground states of local Hamiltonians is a fundamental problem in quantum information. In this work we study the computational complexity of this problem, given the Hamiltonian as input. Our main result is that to show it is cryptographically hard to determine if the ground state of a geometrically local, polynomially gapped Hamiltonian on qudits (d=O(1)) has near-area law vs near-volume law entanglement. This improves prior work of Bouland et al. (arXiv:2311.12017) showing this for non-geometrically local Hamiltonians. In particular we show this problem is roughly factoring-hard in 1D, and LWE-hard in 2D. Our proof works by constructing a novel form of public-key pseudo-entanglement which is highly space-efficient, and combining this with a modification of Gottesman and Irani's quantum Turing machine to Hamiltonian construction. Our work suggests that the problem of learning so-called "gapless" quantum phases of matter might be intractable.

Computational Complexity (cs.CC)

Improving Neutrino Energy Reconstruction with Machine Learning

Faithful energy reconstruction is foundational for precision neutrino experiments like DUNE, but is hindered by uncertainties in our understanding of neutrino--nucleus interactions. Here, we demonstrate that dense neural networks are very effective in overcoming these uncertainties by estimating inaccessible kinematic variables based on the observable part of the final state. We find improvements in the energy resolution by up to a factor of two compared to conventional reconstruction algorithms, which translates into an improved physics performance equivalent to a 10-30% increase in the exposure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY