Search NASASearch

SEARCH · Search NASA

Results for “multivariate experimental data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

MethodOpt: a Shiny-based graphical user interface for multivariate optimization of sampling and analytical instrumentation

Method optimization is an important step in producing useful data in various experimental settings involving the use of sampling and analytical instrumentation, such as gas-chromatography mass-spectrometry or other analytical techniques. However, traditional optimization techniques often lack the sophistication of more modern optimization techniques developed in areas of applied mathematics. A graphical user interface has been developed that implements a multivariate, multi-objective optimization technique for spectra-generating sampling and analytical instrumentation, which saves substantial time and resources compared to the more traditional approaches to method development.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]

Sensitive Detection of Structural Differences using a Statistical Framework for Comparative Crystallography

Chemical and conformational changes underlie the functional cycles of proteins. Comparative crystallography can reveal these changes over time, over ligands, and over chemical and physical perturbations in atomic detail. A key difficulty, however, is that the resulting observations must be placed on the same scale by correcting for experimental factors. We recently introduced a Bayesian framework for correcting (scaling) X-ray diffraction data by combining deep learning with statistical priors informed by crystallographic theory. To scale comparative crystallography data, we here combine this framework with a multivariate statistical theory of comparative crystallography. By doing so, we find strong improvements in the detection of protein dynamics, element-specific anomalous signal, and the binding of drug fragments.

Hekstra, Doeke R. [Harvard Univ., Cambridge, MA (U

Uncertainty quantification of a physics-informed model based on sparse identification of a Thermal Energy Distribution System

Integrated energy systems (IES)s are crucial for enhancing the economy and efficiency of power generation sources (e.g., nuclear energy) necessary to unleash American energy dominance. These systems can be integrated with thermal energy storage (TES) and intermittent renewable energies to optimize overall energy use, peak-load regulation, and demand-side responses. However, the stabilization of energy generation, transport, and utilization introduces operational complexities that exceed the challenges of managing each sub-component individually. Currently, though IESs rely on human operators for efficiency and stability, reducing human error risk and enhancing performance through automation is highly desirable. Recent advances at Idaho National Laboratory have demonstrated successful control of the Thermal Energy Distributed System (TEDS). However, the automatic control system depends on a deterministic Sparse Identification of Nonlinear Dynamics with Control (SINDyC) model, which are trained based on simulation data from physics-based simulations. Because of uncertainties in physics-based simulation, SINDyC model results in large discrepancies against experimental data and cannot be reliably used in automatic control. In this paper, we present an innovative approach to address these discrepancies by quantifying uncertainties and developing a more robust model. We first generated trajectories by using first-principles physics codes to encapsulate the experiment. Next, we trained thousands of models by randomly sampling these trajectories. We then collapsed all those models into one probabilistic SINDyC by fitting a multivariate Gaussian distribution onto the resulting coefficient’s distribution. Despite its simplicity, our approach successfully produced 95% confidence intervals that captured the experimental trajectories. It even did so with a higher probability and better U-pooling score across six of the seven relevant quantities of interest (QoIs), as compared to other classical approaches. In conclusion, ongoing research is focusing on generating new experimental trajectories to validate this approach, and on employing Bayesian calibration to refine parametric uncertainties and guide future model development efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Optimizing high energy density sulfur cathodes: A multivariate approach to electrode formulation and processing

Lithium-sulfur (Li-S) batteries involve complex solid-liquid-solid phase transformations during both discharging and charging processes, where cathode materials, formulation, and structure play a crucial role. Here, a design of experiments (DoE) methodology and an empirical model are developed to systematically explore the interactions and trade-offs among cathode factors and process variables, and to obtain generalizable effects estimates for the multivariate system. Compared to the conventional one-factor-at-a-time (OFAT) approach, this work demonstrates advantages in both efficiency and accuracy by allowing the data to guide future research and decisions. Further, an optimized cathode formulation and processing parameters are predicted and validated experimentally, achieving over 1000 mAh g -1 in discharge capacity and improved cycling under practical lean electrolyte (4 µL mg -1 S) and high S-loading cathodes (>4 mg cm -2 ) conditions. The optimized cathode was scaled up and assembled into Li-S pouch cells, achieving 316 Wh kg -1 in cell-level energy, proving that the comprehensive and rigorous framework for optimizing complex systems with DoE leads to improved performance in a practical pouch cell system.

25 ENERGY STORAGE

Investigating lab-scaled offshore wind aerodynamic testing failure and developing solutions for early anomaly detections

As offshore wind systems become more complex, the risk of human error or equipment malfunction increases during experimental testing. This study investigates a lab-scale incident involving a 1 : 50 scale 5 MW wind turbine, where a generator failure led to rotor overspeed and a blade–tower strike. To improve early fault detection, we propose a data-driven method based on multivariate long short-term memory (LSTM) models. High-frequency measurements are projected onto principal components, and anomalies are identified using reconstruction error and its time derivative. Two models are trained on different healthy datasets and tested using single- and multi-principal component (1PC and MPC) variations. Results show that combining both error and error derivative improves detection accuracy. The 1PC model detects faults faster, has a higher recall rate, and achieves a 43 % improvement in anomaly detection accuracy, while the MPC model yields higher precision. This approach provides a simple and effective tool for early anomaly detection in lab-scale experiments, helping to reduce the risk of future failures during the testing of new technologies.

17 WIND ENERGY

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Search for Light Pseudoscalar Bosons, Pair-Produced in Higgs Boson Decays in the Four-Electron Final State in Proton-Proton Collisions at $\sqrt{s}=13$ TeV

A search for pairs of light neutral pseudoscalar bosons (𝐴) resulting from the decay of a Higgs boson is performed. The search is conducted using LHC proton-proton collision data at $\sqrt{s}=13$ TeV, collected with the CMS detector in 2016–2018 and corresponding to an integrated luminosity of 138 fb −1 . The 𝐴 boson decays into a highly collimated electron-positron pair. A novel multivariate algorithm using tracks and calorimeter information is developed to identify these distinctive signatures, and events are selected with two such merged electron-positron pairs. No significant excess above the standard model background predictions is observed. Upper limits on the branching fraction for 𝐻 → 𝐴⁢𝐴 → 4⁢𝑒 are set at 95% confidence level, for masses between 10 and 100 MeV and proper decay lengths below 100 μ⁢m, reaching branching fraction sensitivities as low as 10 −5 . This is the first search for Higgs boson decays to four electrons via light pseudoscalars at the LHC. It significantly improves the experimental sensitivity to axionlike particles with masses below 100 MeV.

Hayrapetyan, A. [Yerevan Physics Institute]

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition

Search for a light charged Higgs boson in $t \rightarrow H^{\pm } b$ decays, with $H^{\pm } \rightarrow cs$, in $pp$ collisions at $\sqrt{s}={13}\hbox { TeV}$ with the ATLAS detector

A search for a light charged Higgs boson produced in decays of the top quark, $t \rightarrow H^{\pm } b$ with $H^{\pm } \rightarrow cs$, is presented. This search targets the production of top-quark pairs $t\bar{t} \rightarrow WbH^{\pm } b$, with $W \rightarrow ℓv(ℓ = e, μ)$, resulting in a lepton-plus-jets final state characterised by an isolated electron or muon and at least four jets. The search exploits b-quark and c-quark identification techniques as well as multivariate methods to suppress the dominant $t\bar{t}$ background. The data analysed correspond to 140 fb -1 of $pp$ collisions at $\sqrt{s}$ = 13 TeV recorded with the ATLAS detector at the LHC between 2015 and 2018. Observed (expected) 95% confidence-level upper limits on the branching fraction $\mathscr{B}(t \rightarrow H^{\pm } b)$, assuming $\mathscr{B}(t \rightarrow Wb) + \mathscr{B}(t \rightarrow H^{\pm }(\rightarrow cs)b$, are set between 0.066% (0.077%) and 3.6% (2.3%) for a charged Higgs boson with a mass between 60 and 168 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Search for the production of a Higgs boson in association with a single top quark in pp collisions at $\sqrt{s}=13$ TeV with the ATLAS detector

A search for the production of a Higgs boson in association with a single top quark, tH, is presented. The analysis uses proton-proton collision data corresponding to an integrated luminosity of 140 fb −1 at a centre-of-mass energy of 13 TeV, collected by the ATLAS detector at the LHC. The search targets Higgs-boson decays into $b\bar{b}$, WW * , ZZ * , and ττ, accompanied by an isolated lepton (electron or muon) from the top-quark decay. Multivariate techniques are employed to enhance the separation between signal and background processes. The observed signal strength, μ tH , defined as the ratio between the measured cross-section and the predicted Standard Model value, is μ tH = 8.1 ± 2.6 (stat.) ± 2.0 (syst.). The significance of the observed (expected) signal above the background-only expectation is 2.8 (0.4) standard deviations. The corresponding observed (expected) upper limit at the 95% confidence level on the tH cross-section is found to be 13.9 (6.1) times the value predicted by the Standard Model. An interpretation with an inverted sign of the top-quark Yukawa coupling is performed, and the signal strength and corresponding limit are reported.

Hadron-Hadron Scattering