Search NASASearch

SEARCH · Search NASA

Results for “Factorization machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Randomized Algorithms for Symmetric Nonnegative Matrix Factorization

Symmetric Nonnegative Matrix Factorization (SymNMF) is a technique in data analysis and machine learning that approximates a matrix with a product of a nonnegative, low-rank matrix and it transpose. To design faster and more scalable algorithms for SymNMF we develop two randomized algorithms for its computation. The first method uses randomized matrix sketching to compute an initial low-rank approximation to the input matrix and proceeds to uses this as a low-rank input to rapidly compute a SymNMF. The second methods uses randomized leverage score sampling to approximately solve constrained least squares problems. Many successful methods for SymNMF rely on (approximately) solving sequences of constrained least squares problems. Here, we prove theoretically that leverage score sampling can approximately solve constrained least squares problems to e-accuracy. Finally we demonstrate both methods work in practice by applying them to graph clustering tasks on large real world data sets. These experiments show that our methods approximately maintain solution quality and achieve significant speed ups for both large dense and large sparse problems.

97 MATHEMATICS AND COMPUTING

Correlating processing variables to material properties in recycled polypropylene: A data‐driven approach

Abstract Polypropylene (PP) is one of the most widely used plastics, yet its recycling remains limited, with less than 1% of solid waste PP being reprocessed. Mechanical recycling through extrusion is the most practical method, but inconsistent reprocessing conditions introduce variability in material properties. While temperature, screw speed, and residence time influence the thermomechanical stress applied during reprocessing, there are no standardized guidelines for optimizing these parameters. This study examines how these factors shape the properties of recycled PP, using conditions designed to mimic post‐industrial recycled (PIR) scrap. Residence time was measured using colorimetric tracking and correlated with molecular weight, viscosity, and mechanical properties over multiple extrusion cycles. Data‐driven modeling, including response surface methodology, support vector machines, and artificial neural networks, identified processing temperature as the dominant factor in material degradation, followed by residence time. Mechanical properties remained stable, while viscosity decreased predictably with increasing residence time. By linking reprocessing conditions to property evolution, this study provides a method to optimize processing parameters and reduce variability in recycled PP. These findings help manufacturers improve process control, making recycled PP more predictable for reuse in manufacturing. Highlights Study of PIR‐quality PP without additives or compatibilizers. Residence time analysis shows processing temperature drives PP property changes. Mark‐Houwink enables quick molecular weight checks for quality control. Models predict mechanical and rheological shifts in reprocessing. Optimized processing parameters minimize property degradation in recycling.

Estela‐García, John E. [Polymer Engineering Center

Agricultural practices influence soil microbiome assembly and interactions at different depths identified by machine learning

Agricultural practices affect soil microbes which are critical to soil health and sustainable agriculture. To understand prokaryotic and fungal assembly under agricultural practices, we use machine learning-based methods. We show that fertility source is the most pronounced factor for microbial assembly especially for fungi, and its effect decreases with soil depths. Fertility source also shapes microbial co-occurrence patterns revealed by machine learning, leading to fungi-dominated modules sensitive to fertility down to 30 cm depth. Tillage affects soil microbiomes at 0-20 cm depth, enhancing dispersal and stochastic processes but potentially jeopardizing microbial interactions. Cover crop effects are less pronounced and lack depth-dependent patterns. Machine learning reveals that the impact of agricultural practices on microbial communities is multifaceted and highlights the role of fertility source over the soil depth. Machine learning overcomes the linear limitations of traditional methods and offers enhanced insights into the mechanisms underlying microbial assembly and distributions in agriculture soils.

60 APPLIED LIFE SCIENCES

Universal Nuclear Accident Dosimeter

The Lawrence Livermore National Laboratory (LLNL) Universal Nuclear Accident Dosimetry (UNAD) project is a four-year initiative aimed at advancing nuclear accident dosimetry methods. This article presents an overview of the research, key findings, and the progress made throughout the project. The primary goals included a background into the history of nuclear accident dosimetry, consolidating current dosimetry techniques within the NNSA/DOE complex, fostering collaboration among subject matter experts, and exploring novel technologies for potential implementation. The technical focus centered on investigating new and novel technologies, instrumentation methods, and analysis methods to develop recommendations for a potential nuclear accident dosimeter (NAD) to be universally deployed through the DOE complex. A multilaboratory and multinational Usergroup was established, conducting periodic meetings to facilitate knowledge exchange. The UNAD team has participated in two international nuclear accident dosimetry intercomparison exercises and one characterization exercise, where the existing LLNL NAD and a prototype alanine electron paramagnetic dosimeter NAD were deployed. Ongoing improvements are being made to the prototype NAD based on results from the exercises, laboratory studies, and collaboration with other laboratories. A machine learning algorithm to optimize the geometry and conversion factors of the current LLNL NAD is being implemented, and the resulting design will be tested in the next exercise. In conclusion, key lessons learned and future directions for the project are discussed.

Electron paramagnetic resonance spectroscopy

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories

Characterization of Fuel Cladding Chemical Interaction on a High Burnup U-10Zr Metallic Fuel via Electron Energy Loss Spectroscopy Enhanced by Machine Learning

Fuel cladding chemical interaction (FCCI) is one of the main performance limiting factors for metallic nuclear fuels. The interaction destabilizes the martensitic microstructure and deteriorates mechanical properties of HT-9 cladding. The detection of low atomic number elements (Z<10) and overlapping of elemental peaks can be problematic in interpreting energy dispersive X-ray spectroscopy (EDS) data. Electron energy loss spectroscopy (EELS) provides precise elemental edge energy values and can detect elements with a low atomic number. This work utilizes EELS to study the distribution of lanthanides and light elements at the interaction region. The sample was prepared from the FCCI region of a U-10Zr (wt.%) solid fuel with HT-9 cladding, irradiated to a burnup of 13.2 at.%. Processing the EELS data included three major steps: 1) enhance the signal to noise ratio by denoising the spectrum with principal component analysis (PCA) method, removing background and performing deconvolution; 2) identify chemical elements with core energy loss edges; 3) confirm different phases using a popular machine learning method, K-means. This work presents qualitative assessment of lanthanides and light elements like carbon (C) and oxygen (O) enhanced by the application of machine learning algorithms. By comparing with EDS elemental maps, EELS provides higher resolution chemical maps, reveals the distribution of carbon at the interaction region supporting the formation of zirconium carbide, a rind-like microstructure feature that was proposed to mitigate the chemical interaction. Furthermore, the plasmon peak map was also found to indicate an energy shift associated with the formation of phases/compounds. K-means clustering method was used on the processed electron energy loss (EEL) spectrum to automatically reveal different phases. The resulting clustered maps from K-means clustering align well with elemental maps confirming certain phases, especially Fe-Ce and Zr-C, in the FCCI region.

EELS

Modeling Single-Crystal Battery Materials: From Fundamental Understanding to Performance Evaluation

The performance of rechargeable batteries is fundamentally influenced by the physicochemical properties and microstructural features of their key material components. Recent experimental advancements have highlighted the potential of single-crystal (SC) morphologies to address inherent limitations of polycrystalline (PC) electrodes and solid-state electrolytes, offering tunable charge transport kinetics and improved cell cycling performance. Here, this review examines how state-of-the-art computational modeling, from atomistic and mesoscale to continuum-level approaches, including machine learning methodologies, has been utilized to investigate the critical factors governing the electrochemical behavior of SC battery materials. We explore how predictive modeling can elucidate the processing–structure–property–performance relationships of SC cathodes, anodes, and solid-state electrolytes, with a focus on unique SC characteristics such as crystallographic anisotropy, size effects, and facet-dependent properties. Additionally, we identify limitations in commonly used modeling techniques and discuss strategies to address these challenges. By integrating high-fidelity simulations with experimental insights, this review aims to outline a clear path for the rational design and optimization of SC battery components, paving the way for accelerated advancements in energy storage technologies.

Materials science

Machine learning inversion from small-angle scattering for charged polymers

We develop Monte Carlo simulations for uniformly charged polymers and a machine learning algorithm to interpret the intra-polymer structure factor of the charged polymer system, which can be obtained from small-angle scattering experiments. The polymer is modeled as a chain of fixed-length bonds, where the connected bonds are subject to bending energy, and there is also a screened Coulomb potential for charge interaction between all joints. The bending energy is determined by the intrinsic bending stiffness, and the charge interaction depends on the interaction strength and screening length. All three contribute to the stiffness of the polymer chain and lead to longer and larger polymer conformations. The screening length also introduces a second length scale for the polymer besides the bending persistence length. To obtain the inverse mapping from the structure factor to these polymer conformation and energy-related parameters, we generate a large data set of structure factors by running simulations for a wide range of polymer energy parameters. We use principal component analysis to investigate the intra-polymer structure factors and determine the feasibility of the inversion using the nearest neighbor distance. We employ Gaussian process regression to achieve the inverse mapping and extract the characteristic parameters of polymers from the structure factor with low relative error.

36 MATERIALS SCIENCE

Advancing Artificial Intelligence with Liquid Argon Neutrino Experiments (Technical Report)

The grant allowed two main contributions: 1) The development of a first successful demonstration of the employment of Optimal Transport in liquid argon time projection chamber neutrino detectors. Optimal Transport, used in other contexts and specifically with LHC calorimetric data, was adapted to address a key particle identification challenge in LArTPCs: the separation of pi0 backgrounds from single-electrons produced in charged-current electron neutrino interactions. The work, leveraging ML methods such as k-nearest-neighbor (kNN) and support-vector-machine (SVM), showed an increase in background rejection of a factor of two or more. Work is now ongoing to incorporate this development in physics analyses for LArTPC experiments and more broadly expand the use of OT in LArTPC detectors including DUNE. This work was done in collaboration with the phenomenology group led by Nathaniel Craig at UCSB. 2) The deployment of NuGraph2, a graph neural network developed for LArTPC reconstruction, in the MicroBooNE experiment. NuGraph2 uses novel graph-neural-network methods on the rather simple LArTPC inputs of reconstructed hits, greatly simplifying the workflow compared to the use of waveform or signal-deconvolved wire ROIs. The network performed particle classification and was shown to address many challenging problems in LArTPC imaging including track-shower separation and the identification of protons and charged pions from primary muons. Our group collaborated with Giuseppe Cerati (FNAL scientist) who is one of the core developers of NuGraph2 to integrate this tool in MicroBooNE’s analysis framework. This consisted in tow key contributions: a) Studying performance on real data, which came with several months of iterations because the MC-trained version of the network was found to show significant bias that our group investigated and addressed. b) Integrating the output hit labeling of NuGraph2 into the existing particle tracking and shower reconstruction code. As a result of this work led by our team NuGraph2 is now enabling a suite of new analyses which benefit from enhanced capabilities and thus broader physics reach. The grant supported primarily the salary of UCSB graduate student Chuyue “Michaelia” Fang as well as partial summer salary support for PI Caratelli. Some funds were used for travel by Michaelia to ML related schools and conferences.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Securing The Future: 2026 Manufacturing & Critical Infrastructure Threat Landscape

This report outlines the current state of manufacturing weaknesses introduced by the complexities of modern environments, including cloud services and Internet of Things (IoT) devices, with particular attention paid to the unique vulnerabilities encountered by SMMs. It also highlights CyManII’s strategic initiatives and collaborative solutions to mitigate these risks and strengthen the cybersecurity posture of the manufacturing ecosystem. Utilizing data from 2025 to inform forward-looking mitigation strategies, this report provides manufacturers with a clear understanding of both current and emerging cybersecurity threats, as well as practical opportunities to strengthen their cyber ecosystems. The following sections detail key vulnerabilities and threat vectors, along with actionable mitigation strategies, many of which have been developed or piloted through CyManII-led efforts. A thorough understanding of these risks and mitigation strategies is essential for manufacturers seeking to strengthen the security and resilience of their manufacturing operations.

3D Printing

Harnessing the Power of Machine Learning and Omics to Identify Environmental Regulation on Microbial Functional Composition for Soil C, N, and P Cycling

Microbial enzyme-mediated soil organic matter (SOM) decomposition regulates many key ecosystem functions, such as elemental cycling, soil carbon sequestration, and soil fertility. However, representing microbial processes in Earth system models (ESMs) remains challenging due to a limited understanding of the spatial patterns of diverse microbial functions responsible for soil carbon (C), nitrogen (N), and phosphorus (P) cycling as well as the underlying mechanisms regulating their relative abundances across various environments. We collected published metagenomics data across the continental US (CONUS) to identify hundreds of microbial genes involved in soil C, N, and P cycling and grouped them into eight enzyme functional classes (EFCs). Each EFC represented a group of gene-encoded potential enzymes that decompose similar soil compounds. By integrating the abundances of omics-informed EFCs with the corresponding environmental information, we trained a machine learning (ML) model to identify key edaphic, climate, and vegetation factors regulating the abundances of each EFC. Quantitative analysis of effects of these factors revealed that the spatial distribution of eight EFCs for soil C, N, and P cycling across CONUS reflected potential resource optimization strategies of microbial communities under nutrient limitation, preferential organic-mineral associations, and climatological stresses. This insight, together with the interpreted ML tool and the CONUS-level benchmark for EFCs abundances, paves the way for parameterizing environmental-regulated microbial functional dynamics in biogeochemical models.

machine learning

Multiplexed profiling of transcriptional regulators in plant cells

Transcriptional regulators play key roles in plant growth, development and environmental responses; however, understanding how their regulatory activity is encoded at the protein level has been hindered by a lack of multiplexed large-scale methods to characterize protein libraries in planta. Here we present enrichment of nuclear trans-elements reporter assay in plants with sequencing (ENTRAP-seq), a high-throughput method that introduces protein-coding libraries into plant cells to drive a nuclear magnetic sorting-based reporter, enabling multiplexed measurement of regulatory activity from thousands of protein variants. Using ENTRAP-seq and machine learning, we screen 1,495 plant viruses and identify hundreds of putative transcriptional regulatory domains found in structural proteins and enzymes not associated with gene regulation. In addition, we combine ENTRAP-seq with machine-guided design to engineer the activity of a plant transcription factor in a semirational fashion. Our findings demonstrate how scalable protein function assays deployed in planta will enable the characterization of natural and synthetic coding diversity in plants.

Alamos, Simon

Heavy-Duty Nonroad Material Handler Electrification Part 1: Real-World Drive Cycle Development

Knowing a detailed operating cycle is critical for developing and testing equipment. Operating cycles can be separated by two clear distinctions: (1) regulatory or non-regulatory and (2) application at the engine-only or full machine level. The Environmental Protection Agency’s (EPA) Nonroad Transient Cycle (NRTC) may be a good representation of engine use in many types of equipment, but there is a gap in standardized and validated drive cycles specifically for nonroad material handlers. Lacking a standardized drive cycle makes it difficult to accurately benchmark machine performance and validate new powertrain technologies. The objective of this investigation is to illustrate the development of a custom drive cycle augmented with real-world customer use data that serves multiple purposes: (1) understand the range of operation and utilization that formulated inputs for electrified architecture analysis and (2) develop a repetitive and consistent maneuver to establish baseline energy consumption enabling equivalent comparison to future electrified prototype builds. This article presents a solution specifically for a 23-ton nonroad material handler in which material handling, machine transport, and extended idle were homologated to form representative short cycles defined by machine velocity and hydraulic cylinder position. The most intensive material handling short cycles had a load factor of 40% and an average fuel rate of 16 L/h. Combined with a visual aid, the short cycles exhibited low variability, having less than 5% root mean square (RMS) error in lift and reach position with respect to the average. The machine’s performance on these short cycles at the Advanced Power Systems Research Center (APSRC) was compared to results from two real-world customer locations operating the instrumented test machine in a cyclical manner, and for similar ground conditions were found to be comparable in fuel consumption.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Improving tropical cyclone rapid intensification forecasts with satellite measurements of sea surface salinity and calibrated machine learning

Forecasting rapid intensification (RI) of tropical cyclones (TC) is a mission known for large errors. One under-researched factor that affects TC intensification is salinity, which is important for density stratification in certain ocean regions and can affect the surface enthalpy flux under a strengthening hurricane. To investigate the impact and efficacy of using salinity information in state-of-the-art forecasting, we use a statistical model consisting of a variety of machine learning (ML) methods. For salinity data, we use satellite measurements of pre-storm sea surface salinity (SSS) as a proxy for the salinity stratification. We train and test the model on various ocean basins, including the Atlantic, eastern North Pacific and western North Pacific. A calibrator is trained on top of the ML models to correct and enhance probability forecasts. The calibrator significantly improves probability forecasts relative to recent works. The ML model performance is improved with the addition of SSS in the Eastern North Pacific, western North Pacific, and the Caribbean subregion of the North Atlantic, and the overall model performance is better than previous studies. SSS decreases model skill for a model trained on the full Atlantic basin. In the Indian Ocean, SSS is also notably correlated with RI occurrence, but the TC samples are not sufficient to train ML models.

hurricane

Discovery of unconventional and nonintuitive self-assembling peptide materials using experiment-driven machine learning

Prediction of peptide secondary structure is challenging because of complex molecular interactions, sequence-specific behavior, and environmental factors. Traditional design strategies, based on hydrophobicity and structural propensity, can be biased and could indeed prevent discovery of interesting, diverse, and unconventional peptides with desired nanostructure assembly. Using β sheet formation in pentapeptides as a case study, we used an integrated high-throughput experimental workflow and an artificial intelligence–driven active learning framework to improve prediction accuracy of self-assembly. By focusing on sequences where machine learning (ML) predictions deviate from conventional design strategies, we synthesized and tested 268 pentapeptides, successfully finding 96 forming β sheet assemblies, including unconventional sequences (e.g., ILFSM, LMISI, MITIY, MISIW, and WKIYI) not predicted by traditional methods. Our ML models outperformed conventional β sheet propensity tables, revealing useful chemical design rules. A web interface is provided to facilitate community access to these models. This work highlights the value of ML-driven approaches in overcoming the limitations of current peptide design strategies.

Talluri, Y. Nissi [Indian Inst. of Technology (IIT

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data driven investigation to understand the influence of total solids on biological biogas upgrading

In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.

In situ biogas upgrading

Aerosol Influences on Cloud Water: Insights From ARM EPCAPE Observations With Explainable Machine Learning

This study employs an explainable machine learning (ML) framework (XGBoost-SHapley Additive exPlanations analysis) to investigate controlling factors on cloud liquid water path (LWP) using EPCAPE observations near the California coast. Aerosols are found to be the dominant factor explaining LWP variability, surpassing meteorological factors (MFs). By isolating aerosol effects from meteorological influences, the ML reveals a negative linear relationship between LWP and cloud droplet number concentration (Nd) in log space, likely driven by entrainment drying via evaporation-entrainment feedback. This aligns with the negative regime of the inverted-V relationship reported in previous studies, while no positive LWP responses are found due to a limited number of precipitating cases in EPCAPE. Furthermore, the sensitivity of LWP to Nd shows a non-linear dependence on MFs like moisture contrast between surface and free troposphere and lower-tropospheric stability. This occurs due to the interplay between the MFs' direct effects on entrainment drying and indirect effects through LWP adjustments.

EPCAPE observations