Search NASA⌕ Search

SEARCH · Search NASA

Results for “data reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values

Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Quantifying health benefits of sustainable aviation fuels: Modeling decreased ultrafine particle emissions and associated impacts on communities near the Seattle-Tacoma International Airport

Exposure to ultrafine particles (UFP, ≤100 nm) is an emerging health concern linked to premature mortality, with jet fuel combustion identified as a significant source of UFPs near airports. Sustainable aviation fuel (SAF) adoption has the potential to reduce aviation-related UFPs and may particularly benefit populations who reside nearby. However, assessing aviation-specific impacts on health remains challenging due to the lack of tools capable of addressing: fine-scale exposure evaluation, novel ambient pollutants, and groups with increased exposure or susceptibility. We develop and apply a method to estimate reductions in mortality associated with aviation-related UFP reductions at the Seattle-Tacoma (SEA-TAC) International Airport under SAF adoption scenarios, with a focus on near-airport communities. Using UFP exposure surfaces generated from AERMOD modeling, flight count data, and UFP measurements, we evaluated UFP reductions under various control scenarios. We estimated mortality reductions by combining this with population data, baseline mortality, and a hazard ratio of 1.012 (95 % confidence interval: 1.010, 1.015) per interquartile range increment of 2723 particles/cm 3 . Our analysis included 412 census tracts representing almost 1.5 million adults. Baseline aviation-related UFP exposures averaged 1145 (SD: 277) particles/cm 3 . The highest baseline concentrations and subsequent reductions under SAF scenarios were near SEA-TAC. Mortality case reductions averaged between 3.1 (95 % range: 2.5–3.7) for a 5 % UFP reduction to 31.0 (24.6–37.4) for a 50 % reduction, with corresponding mortality rate reductions of 0.2 (0.2–0.3) to 2.1 (1.7–2.5) cases per 100,000 people per year. Mortality rate reductions were larger among populations residing closer to SEA-TAC, including those that were Hispanic or Latino, below-poverty, and did not identify as White. Reducing aviation-related UFPs through SAF adoption could lead to lower mortality, particularly in near-airport communities. This reproducible approach can be adapted to other settings to evaluate health benefits from aviation-related UFP reductions.

Aviation-related air pollution↗

Combining High-Throughput Experiments and Active Learning to Characterize Deep Eutectic Solvents

The high tunability of deep eutectic solvents (DESs) stems from the ease of changing their precursors and relative compositions. However, measuring the physicochemical properties across large composition and temperature ranges, necessary to properly design target-specific DESs, is tedious and error-prone and represents a bottleneck in the advancement and scalability of DES-based applications. As such, active learning (AL) methodologies based on Gaussian processes (GPs) were developed in this work to minimize the experimental effort necessary to characterize DESs. Owing to its importance for large-scale applications, the reduction of DES viscosity through the addition of a low-molecular-weight solvent was explored as a case study. A high-throughput experimental screening was initially performed on nine different ternary DESs. Then, GPs were successfully trained to predict DES viscosity from its composition and temperature, showcasing the ability of these stochastic, nonparametric models to accurately describe the physicochemical properties of complex mixtures. Finally, the ability of GPs to provide estimates of their own uncertainty was leveraged through an AL framework to minimize the number of data points necessary to obtain accurate viscosity modes. This led to a significant reduction in data requirements, with many systems requiring only five independent viscosity data points to be properly described.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder↗

A transfer learning approach to energy-efficient control of small and medium-sized commercial buildings

Model-free reinforcement learning (RL) provides a data-driven and adaptive approach to optimize building energy use while satisfying occupant comfort. This powerful tool does not need any prior knowledge about the environment and system it is optimizing and can adapt its policy based on the changes in captures. Like any other data-driven tool, it faces high training costs due to the extensive agent-environment interactions required to capture long-term building dynamics and user comfort. Transfer learning, particularly policy distillation, offers a promising way to accelerate training by leveraging pretrained RL agents in different building and system types. Here, this study investigates online student distillation, in which the student model updates its neural network weights using outputs from teacher models. The work introduces a student distillation strategy designed for efficient knowledge transfer, along with a teacher selection method that ensures high-quality guidance. The approach is validated using a highly calibrated whole building energy model for a small/medium commercial building test facility. Results show substantial reductions in training time and data requirements while surpassing the performance of ASHRAE Guideline 36, an advanced rule-based control strategy. The distilled RL model required 45% less data and achieved 20% higher cumulative rewards than a state-of-the-art RL model, with faster convergence and lower energy consumption. These outcomes demonstrate that effective transfer learning enables a scalable and data-efficient energy management solution for commercial buildings.

ASHRAE guideline 36↗

cnor_pub: R code and data for nitrous oxide synthesis by purified bacterial cNOR

This zipped archive of a GitHub repository includes the experimental data (csv files) collected for the reduction of NO to N2O by purified Paracoccus denitrificans cytochrome c nitric oxide reductase (cNOR) and the R code (qmd files) used to analyze these data. A link to the corresponding GitHub repository is also provided.

Hegg, Eric L. [GLBRC - Michigan State University]↗

Machine Learning to Select Experiments Driven by Fundamental Science and Applications for Targeted Nuclear Data Improvement

This work describes a blueprint for a process that accelerates progress in science by quantitatively answering the following question: What is the optimal combination of fundamental-science and application-driven experiments to maximally reduce pertinent data uncertainties? Answering this question entails solving a high-dimensional and complex optimization problem that is best solved with advanced statistic techniques often classified as machine learning. We apply this process within the framework of nuclear data with the aim to select an experiment combination that will reduce uncertainties in 239 Pu nuclear data for neutron energies between 1 and 600 keV. In this field, fundamental-physics driven data, called differential, look at one nuclear physics observable at a time. They are contrasted to application-driven, integral, data where one or few resulting values inform a broad set of nuclear data across several nuclides and energies. The candidates for integral experiments are criticality measurements that were refined by a genetic algorithm to be maximally sensitive to 239 Pu fission cross sections in the desired energy range. Twenty-three candidate differential experiments were investigated and span multiple nuclear physics observables (e.g., total, capture cross sections) for isotopes appearing in the integral experiments. The optimal combination among these candidate experiments was investigated via generalized least squares fitting, augmented with Gaussian processes to ameliorate statistical irregularities in data, and the D-optimality criterion. The latter evaluates for each pair of candidates the joint reduction in uncertainties of all 12200 nuclear data appearing in the integral experiments compared to the knowledge we have from 168 past experiments, theory, and nuclear data. We chose as differential measurements those that investigate 63 Cu and 239 Pu total cross sections, based on D-optimality rank and feasibility constraints. Two integral (criticality) experiments were selected: An experiment with Al 2 ⁢O 3 and graphite interleaved with Pu and a thick Cu reflector explores 1–30 keV, while we target the 30–600 keV range with an experiment that swaps boron in place of graphite with a different geometry.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗

Tabulated data for Lignin-Derived Phenolic Compounds and Water Are Effective Cosolvents for Reductive Catalytic Fractionation

Reductive catalytic fractionation (RCF) is an effective lignin-first biorefining method to extract lignin as a stabilized oil from lignocellulosic biomass. To realize RCF at scale, process modeling has shown that minimizing the use of exogenous organic solvents is critical. To this end, here we investigate the ability of lignin-derived monomers to act as either solvents or cosolvents for RCF. We begin by examining the influence of lignin-derived aromatic compounds (4-propylguaiacol, 4-propylphenol, and propylbenzene) on RCF monomer yields and subsequently extend our analysis to mixtures of 4-propylguaiacol and either methanol or water. We demonstrate that 4-propylguaiacol is an effective solvent for lignin extraction and depolymerization during RCF, especially when used in combination with water as a cosolvent. Cosolvent mixtures of 4-propylguaiacol and water enable up to 81% lignin extraction, monomer yields up to 25 wt %, and postreaction phase separation. However, unlike methanol, water as a cosolvent fails to inhibit aromatic ring hydrogenation when conducted over Ru/C as a catalyst, potentially leading to excess hydrogen consumption in a process utilizing this approach. Nonetheless, these results suggest a promising strategy for eliminating external organic solvents from RCF by utilizing mixtures of lignin-derived compounds and water as alternative extraction solvents. This is the tabulated data for the paper

biorefining↗

Addressing the impact of Lyman opacity in inference of divertor plasma conditions with 2D spectroscopic camera analysis of Balmer emission during detachment in JET L-mode plasmas

The impact of re-absorption of the deuterium Lyman series emission was addressed in inferring divertor plasma conditions from Balmer series emission with 2D spectroscopic camera analysis during detachment in JET L-mode plasmas. The previously presented methodology was amended by modifying the standard photon emission coefficients and ionization and recombination rate coefficients of the ADAS database to consider the re-population of excited states due to Lyman opacity. This resulted in the estimate for the atomic density near the outer strike point to decrease by up to 75% at the onset of detachment at strike point temperatures of T e,osp ≈ 1.0–3.0 eV with respect to the strongly overestimated previously obtained values, whereas the estimated electron temperature and density were unaffected by the opacity correction within the scatter of the data and only a moderate reduction by up to 20% was observed in the estimate for the molecularly induced fraction of the Balmer emission. No noticeable change was seen in the ionization rate, calculated from the estimated outer strike point conditions, due to the decrease in the atomic density estimate compensating for the increased values of the opacity-corrected ADAS rate coefficients for ionization. In detached conditions at T e,osp ≈ 0.5–1.0 eV, 25%–35% lower recombination rates were provided by the opacity-corrected model. The observed effects on the experimental analysis were supported by a corresponding synthetic analysis based on EDGE2D-EIRENE simulations.

Balmer emission↗

Optimization of La 2 NiO 4+δ Electrolysis Cell Oxygen Electrode through Surfactant-Enabled LaCoO 3±δ Nanocatalyst Deposition

Lanthanum nickelate (LNO) has shown promise as a Cr-resistant air electrode material for SOECs but has suboptimal surface oxygen exchange properties. Nanocoating of the LNO surface with lanthanum cobaltite (LCO) was chosen to improve cell performance as a surface oxygen conductor. The work focused on the implementation of a two-step nano-LCO film deposition utilizing catechol molecules in a porous LNO electrode. The subgoals of the work were to maintain nanosized LCO particles/ grains to increase active surface area and to control the regularity/ homogeneity of the coating across the microstructure. To achieve these goals, a novel surfactant-enhanced liquid infiltration method was utilized, where nucleation sites were spread across the electrode structure to control the location and size of LCO particles. Various catechol surfactant compositions were evaluated for their ability to control the kinetics of nanoparticle deposition and the homogeneity of the coating. Chelated LCO was characterized by X-ray diffraction (XRD), which found a substantial improvement in LCO formation with surfactant addition and determined polymerized norepinephrine to be the best-performing surfactant, with 88.4% pure LCO formed at low temperature. X-ray photoelectron spectroscopy (XPS) confirmed LCO nanostructures formed by the two-step infiltration process, showing no impurities and a stable perovskite structure. Deposition kinetics were analyzed using atomic force microscopy (AFM), correlating infiltration times and solution molarity to nanoparticle size and distribution, the results of which were confirmed in symmetrical cell samples by scanning electron microscopy (SEM). Electrochemical impedance spectroscopy (EIS) testing demonstrated substantial improvements in polarization resistance, where the nanocoating reduced the resistance by ∼55% to 0.152 Ω·cm 2 at 700 °C and 0.039 Ω·cm 2 at 800 °C. Electrical conductivity relaxation (ECR) at this temperature confirmed an improved surface oxygen exchange coefficient of the LCO + LNO heterostructure predicted by the Bode data from EIS, alongside a reduction in activation energy by about 30%.

Deposition↗

A New Track Trigger for Characterization of the Antiproton-Induced Background in the Mu2e Experiment

The Mu2e experiment at Fermilab will enable the search for the neutrinoless muon to electron conversion in the field of an Al nucleus, a charged lepton flavor violating process. If observed, there would be a clear indication of physics beyond the Standard Model. Mu2e aims to reach a single event sensitivity of $3 /times 10^{-17}$, improving from the previous limit by 4 orders of magnitude. This improvement relies on the development of trigger selection systems, designed to discard data from background-induced events by placing kinematic, topological cuts on a particle’s reconstructed track. One of the largest sources of background Mu2e faces is proton-antiproton annihilation. These annihilations produce a 2 GeV shower of particles, among which there could be an electron that mimics the conversion electron signal, with an expected number of 0.010 ± 0.010. The large uncertainty on this number is dominated by the systematic uncertainty associated with the theoretical production model. To better characterize this background, we have developed an antiproton trigger selection by taking advantage of the track multiplicity and topology of these events. We discuss the steps taken in this development and the first performance study of this trigger, evaluating the signal efficiency and background rate. This trigger is essential to enable a data-driven analysis targeting the reduction of the systematic uncertainty of the antiproton-induced background in the Mu2e experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety reliability for hydrogen facilities by integrating risk reduction methodologies and component reliability data taxonomies that support hydrogen infrastructure failure rate analysis.

availability↗

Framework of compressive sensing and data compression for 4D-STEM

Four-dimensional Scanning Transmission Electron Microscopy (4D-STEM) is a powerful technique for high-resolution and high-precision materials characterization at multiple length scales, including the characterization of beam-sensitive materials. However, the field of view of 4D-STEM is relatively small, which in absence of live processing is limited by the data size required for storage. Furthermore, the rectilinear scan approach currently employed in 4D-STEM places a resolution- and signal-dependent dose limit for the study of beam sensitive materials. Improving 4D-STEM data and dose efficiency, by keeping the data size manageable while limiting the amount of electron dose, is thus critical for broader applications. Here we introduce a general method for reconstructing 4D-STEM data with subsampling in both real and reciprocal spaces at high fidelity. The approach is first tested on the subsampled datasets created from a full 4D-STEM dataset, and then demonstrated experimentally using random scan in real-space. The same reconstruction algorithm can also be used for compression of 4D-STEM datasets, leading to a large reduction (100 times or more) in data size, while retaining the fine features of 4D-STEM imaging, for crystalline samples.

4D-STEM↗