Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

A universal implementation of radiative effects in neutrino event generators

Due to the similarities between electron-nucleus (eA) and neutrino-nucleus scattering (νA), eA data can contribute key information to improve cross-section modeling in eA and hence in νA event generators. However, to compare data and generated events, either the data must be radiatively corrected or radiative effects need to be included in the event generators. We implemented a universal radiative corrections program that can be used with all reaction mechanisms and any eA event generator. Our program includes real photon radiation by the incident and scattered electrons, and virtual photon exchange and photon vacuum polarization diagrams. It uses the “extended peaking” approximation for electron radiation and neglects charged hadron radiation. This method, validated with GENIE, can also be extended to simulate νA radiative effects. This work facilitates data-event-generator comparisons used to improve νA event generators for the next-generation of neutrino experiments. Program Title: emMCRadCorr CPC Library link to program files:https://doi.org/10.17632/hmsxg82vnf.1 Developer's repository link:https://github.com/e4nu/emMCRadCorr Licensing provisions: AGPLv3 Programming language:C++ Nature of problem: Radiative effects can significantly modify the event kinematics and the resulting cross-sections. Such effects must be accounted for when comparing event generators to eA data. Existing radiative correction codes are tailored to specific processes and topologies, and are limited to a restricted phase space defined by the spectrometer acceptance. Therefore, a more general approach is required to apply radiative corrections to semi-inclusive and exclusive eA measurements. Solution method: Our program incorporates real photon radiation from both the incident and scattered electrons, as well as virtual photon exchange and photon vacuum polarization effects. It employs the “extended peaking” approximation for electron radiation while neglecting contributions from charged hadron radiation. The code is fully decoupled from event generator codes and can be used for all event generators in the market.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An In Situ , Automated High-Explosives Aging Method Utilizing Two-Dimensional Gas Chromatography–Mass Spectrometry

Understanding chemical changes that occur in high explosives as they age is of great importance to the safe employment and storage of these compounds. Traditional methods of aging high explosives even under accelerated aging conditions are time intensive with durations on the order of months to years. The nature of traditional aging analyses reduces each sample to a snapshot data point often separated widely in time, requiring many assumptions as to how the degradation products develop. Further complicating matters, several analytical techniques are typically employed for each sample analysis in order to ascertain an entire picture of the decomposition pathways. To address these shortcomings with existing methods, a new method of accelerated aging of high explosives utilizing comprehensive two-dimensional gas chromatography coupled to high-resolution mass spectrometry (GC × GC-HRMS) was developed using 2,4,6,8,10,12-hexanitro-2,4,6,8,10,12-hexaazaisowurtzitane (CL-20) as a model compound for method development. This in situ automated method reduces the time scale of aging to a matter of hours using the inlet of the GC × GC as the aging vessel. GC × GC in combination with HRMS allowed for the collection of both evolved gases and other decomposition products produced during the entire aging process in real time with HRMS providing far greater certainty in identification of explosives aging products. Additionally, this method allowed for a higher throughput of samples with greatly simplified sample preparation. Chemometric analysis of the GC × GC-HRMS data set via the alteration analysis (ALA) enabled discovery of statistically significant chemical changes providing insight into the variation of decomposition pathways with varying aging temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CMPLE: Correlation Modeling to Decode Photosynthesis Using the Minorize–Maximize Algorithm

In plant genomic experiments, correlations among various biological traits (phenotypes) give new insights into how genetic diversity may have tuned biological processes to enhance fitness under diverse conditions. Consequently, knowing how the correlations are affected by genetic (G) and environmental (E) factors helps develop climate-resilient plants. However, the current literature lacks any method for assessing the effect of predictors on pairwise correlations among multiple phenotypes together with easily interpretable model parameters. To address this need, we propose to model pairwise correlations directly in terms of G and E and develop a computationally efficient inference procedure. Two major novelties in our methodology are (1) the use of a composite pairwise likelihood method to avoid the positive definiteness restriction on the correlation matrix and (2) the use of a novel Minorize–Maximize (MM) algorithm for the efficient estimation of a large number of parameters. The proposed method shows excellent numerical performance on synthetic datasets. Here, the analysis of the motivating data on cowpea reveals that the rates of solar energy storage by photosynthesis (the aggregate trait) are differentially affected by different genetic loci through two distinct processes: “photoinhibition” which results from photodamage caused by excess light, and “photoprotection” which protects plants from photodamage but also results in energy loss.

Correlation modeling↗

Precision Measurements of the Neutron Magnetic Form Factor to High Momentum Transfer using Durand’s Method

Protons and neutrons, collectively known as nucleons, along with electrons, constitute the fundamental building blocks of the visible universe. Understanding their internal structure is crucial for addressing key scientific questions about our origin and existence. Elastic electron-nucleon scattering provides insights into the spatial distributions of charge and current within nucleons through their electromagnetic form factors. Accurate knowledge of these form factors over a broad range of Q2, the squared four-momentum transfer in the scattering process, reveals details about the nucleon's internal structure. However, high-Q2 data of the nucleon electromagnetic form factor is scarce due to the challenges associated with such measurements. This thesis reports preliminary results from high-precision measurements of the neutron magnetic form factor (GMn) to unprecedented Q2 using Durand's method, also known as the "ratio" method. Systematic errors are greatly reduced by ext

Datta, Provakar↗

HostSub_GP: Precise Galaxy Background Subtraction in Transient Long-slit Spectroscopy with Gaussian Processes

We present a novel host galaxy subtraction technique in long-slit spectroscopy for extragalactic transients. Unlike classic methods which generally estimate the background using simple interpolation of local galaxy flux in the 2D spectrum, our approach leverages multi-band archival images of the host galaxies to model the background emission from the galaxy in the 2D spectrum. Such imaging encodes the wavelength-dependent galaxy profile along the slit, and is readily accessible through wide-field imaging surveys. We construct a smooth prior for the 2D galaxy profile with a Gaussian process (GP) based on these reference images, and use another GP to model the correlated deviations from the prior in the observed spectrum. This enables accurate inference of the galaxy flux blended with the transient. On synthetic long-slit data of a spiral galaxy extracted from a Multi Unit Spectroscopic Explorer hyper-spectral cube, the GP method remains robust as long as the host galaxy is spatially resolved and consistently outperforms classic methods. We apply the method to archival Keck spectra of two real transients, SN 2019eix and AT 2019qiz, to further demonstrate how the method uniquely recovers weak spectral features amid strong galaxy contamination, enabling refined constraints on the properties of both transients. We have released the software implementation, HostSub_GP, a scalable toolkit that leverages JAX, with an MIT license.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Investigating Temperature Uniformity and Accuracy in PV Module Lamination: A Verification Study

This study investigates the temperature uniformity and accuracy of a photovoltaic (PV) module lamination process by addressing inconsistencies identified in 2017 data where irregular temperature changes were observed across setpoints. The 2017 data showed a notable drop in temperature upon bladder initiation, except for the 145 degrees Celsius profile. This inconsistency indicated potential inaccuracies in manual data recording methods. To address this concern, a verification experiment was conducted to evaluate temperature uniformity across the 2014 Bent River SPL2828 laminator platen and within test samples. Thermocouples, paired with Omega data acquisition software, were deployed to measure temperatures at multiple platen locations and within test samples. The experiment compared lamination temperatures of polyethylene-co-vinyl acetate (EVA) encapsulant when paired with solite glass or TPE backsheets. The methodology included verifying temperature uniformity directly on the platen and by using a large glass/EVA/glass sample using multiple thermocouples. Smaller samples were built with glass/EVA/glass and glass/EVA/backsheet configurations with one centered thermocouple to verify and compare sample temperatures. This verification aims to refine lamination temperature profiles, enhance data accuracy and provide insights into optimal process control for uniform module lamination. Ensuring consistent and uniform lamination may improve the accuracy and reliability of research outcomes.

14 SOLAR ENERGY↗

On the Training and Generalization of Deep Operator Networks

Here, we present a novel training method for deep operator networks (DeepONets), one of the most popular neural network models for operators. DeepONets are constructed by two subnetworks, namely the branch and trunk networks. Typically, the two subnetworks are trained simultaneously, which amounts to solving a complex optimization problem in a high dimensional space. In addition, the nonconvex and nonlinear nature makes training very challenging. To tackle such a challenge, we propose a two-step training method that trains the trunk network first and then sequentially trains the branch network. The core mechanism is motivated by the divide-and-conquer paradigm and is the decomposition of the entire complex training task into two subtasks with reduced complexity. Therein the Gram–Schmidt orthonormalization process is introduced which significantly improves stability and generalization ability. On the theoretical side, we establish a generalization error estimate in terms of the number of training data, the width of DeepONets, and the number of input and output sensors. Numerical examples are presented to demonstrate the effectiveness of the two-step training method, including Darcy flow in heterogeneous porous media.

deep operator networks↗

Permafrost Region Greenhouse Gas Budgets Suggest a Weak CO 2 Sink and CH 4 and N 2 O Sources, But Magnitudes Differ Between Top-Down and Bottom-Up Methods

Large stocks of soil carbon (C) and nitrogen (N) in northern permafrost soils are vulnerable to remobilization under climate change. However, there are large uncertainties in present-day greenhouse gas (GHG) budgets. We compare bottom-up (data-driven upscaling and process-based models) and top-down (atmospheric inversion models) budgets of carbon dioxide (CO 2 ), methane (CH 4 ) and nitrous oxide (N 2 O) as well as lateral fluxes of C and N across the region over 2000–2020. Bottom-up approaches estimate higher land-to-atmosphere fluxes for all GHGs. Both bottom-up and top-down approaches show a sink of CO 2 in natural ecosystems (bottom-up: -29 (-709, 455), top-down: -587 (-862, -312) Tg CO 2 -C yr -1 ) and sources of CH 4 (bottom-up: 38 (22, 53), top-down: 15 (11, 18) Tg CH 4 -C y -1 ) and N 2 O (bottom-up: 0.7 (0.1, 1.3), top-down: 0.09 (-0.19, 0.37) Tg N 2 O-N yr -1 ). The combined global warming potential of all three gases (GWP-100) cannot be distinguished from neutral. Over shorter timescales (GWP-20), the region is a net GHG source because CH 4 dominates the total forcing. The net CO 2 sink in Boreal forests and wetlands is largely offset by fires and inland water CO 2 emissions as well as CH 4 emissions from wetlands and inland waters, with a smaller contribution from N 2 O emissions. Priorities for future research include the representation of inland waters in process-based models and the compilation of process-model ensembles for CH 4 and N 2 O. Discrepancies between bottom-up and top-down methods call for analyses of how prior flux ensembles impact inversion budgets, more and well-distributed in situ GHG measurements and improved resolution in upscaling techniques.

54 ENVIRONMENTAL SCIENCES↗

Precision Measurements of the Neutron Magnetic Form Factor to High Momentum Transfer using Durand's Method

Protons and neutrons, collectively known as nucleons, along with electrons, constitute the funda- mental building blocks of the visible universe. Understanding their internal structure is crucial for addressing key scientific questions about our origin and existence. Elastic electron-nucleon scatter- ing provides insights into the spatial distributions of charge and current within nucleons through their electromagnetic form factors. Accurate knowledge of these form factors over a broad range of Q2, the squared four-momentum transfer in the scattering process, reveals details about the nucleon’s internal structure. However, high-Q2 data of the nucleon electromagnetic form factor is scarce due to the challenges associated with such measurements. This thesis reports preliminary results from high-precision measurements of the neutron magnetic form factor (Gn M ) to unprecedented Q2 using Durand’s method, also known as the “ratio” method. Systematic errors are greatly reduced by extracting Gn M from the ratio of neutron-coincident (D(e, e'n)) to proton-coincident (D(e, e'p)) quasi-elastic electron scattering from deuteron. The scattered electrons were detected in the BigBite spectrometer, which features multiple Gas Elec- tron Multiplier (GEM) layers with large active area for high-precision tracking at very high rates. Simultaneous nucleon detection was performed by the Super BigBite spectrometer, which utilizes a dipole magnet with large solid angle acceptance at forward angles and a novel hadron calorimeter with very high and comparable detection efficiencies for both protons and neutrons. This setup could handle very high luminosity, making high-Q2 measurements feasible. Data were collected at five Q2 points: 3, 4.5, 7.4, 9.9, and 13.6 (GeV/c)2. Preliminary results are reported for all, with the lowest two Q2 points in good agreement with existing world data, while the higher points significantly extend the Q2 range in which Gn M is known accurately. The precision of the highest Q2 point is expected to remain unmatched for years to come.

Datta, Provakar↗

Bayesian And Human Reliability Analysis (hra)-aided Method For The Reliability Analysis Of Software (bahamas)

The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.

Wang, Congjian (0000000207789927)↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

Transfer learning nonlinear plasma dynamic transitions in low dimensional embeddings via deep neural networks

Deep learning algorithms provide a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. Development of novel, data-driven model reduction methods, coupled with detection of abnormal modes with plasma physics, opens a unique opportunity to identify plasma instabilities through automated construction of parsimonious models that can be tuned to balance accuracy and cost. Our fusion transfer learning (FTL) model demonstrates success in rapidly reconstructing nonlinear kink mode structures by learning from a limited amount of nonlinear simulation data. The knowledge transfer process leverages a pre-trained neural encoder–decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL’s capacity to capture transitional behaviors and dynamical features in plasma dynamics—a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics modes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

DNN-based Signal Processing for Liquid Argon Time Projection Chambers

We investigate a deep learning-based signal processing for liquid argon time projection chambers (LArTPCs), a leading detector technology in neutrino physics. Identifying regions of interest (ROIs) in LArTPCs is challenging due to signal cancellation from bipolar responses and various detector effects observed in real data. We approach ROI identification as an image segmentation task, and employ a U-ResNet architecture. The network is trained on samples that incorporate detector geometry information and include a range of detector variations. Our approach significantly outperforms traditional methods while maintaining robustness across diverse detector conditions. This method has been adopted for signal processing in the Short-Baseline Neutrino program and provides a valuable foundation for future experiments such as the Deep Underground Neutrino Experiment.

Bhat, Avinay [Chicago U.]↗

Comparison of three measurement modalities for 3D characterization of manufactured features and process-induced porosity in titanium alloy additively manufactured parts

Nondestructive characterization of internal features and defects within complex components is vital for many industrial applications, particularly with the advent of additive manufacturing (AM) technologies. However, community understanding of the limitations of nondestructive methods such as X-ray Computed Tomography (CT) can be limited in certain industrial sectors as these may be emergent applications. In this paper, we investigate the limits of X-ray CT measurements and compare extracted data with mechanical polishing serial sectioning (MPSS) and confocal laser scanning microscopy (CLSM). The test object is an additively manufactured titanium alloy disk that contains both process-induced porosity and machined features, including focused ion beam milled features designed to probe the resolution limits of X-ray CT. Results show that each of these characterization techniques has advantages and disadvantages. We compare data acquisition times, spatial resolution, geometric measurement accuracy and defect visualization fidelity across these modalities to establish a practical framework.

Additive manufacturing↗

Positron emission tomography harmonization in the Alzheimer's Disease Neuroimaging Initiative: A scalable and rigorous approach to multisite amyloid and tau quantification

Abstract INTRODUCTION A key goal of the Alzheimer's Disease NeuroImaging Initiative (ADNI) positron emission tomography (PET) Core is to harmonize quantification of β‐amyloid (Aβ) and tau PET image data across multiple scanners and tracers. METHODS We developed an analysis pipeline (Berkeley PET Imaging Pipeline, B‐PIP) for ADNI Aβ and tau PET images and applied it to PET data from other multisite studies. Steps include image pre‐processing, refacing, magnetic resonance imaging (MRI)/PET co‐registration, visual quality control (QC), quantification of tracer uptake, and standardization of Aβ and tau standardized uptake value ratios (SUVrs) across tracers. RESULTS Measurements from 10,105 cross‐sectional and longitudinal Aβ and tau PET scans acquired in several studies between 2010 and 2024 can be processed, harmonized, and directly merged across tracers and cohorts. DISCUSSION The B‐PIP developed in ADNI is a scalable image harmonization approach used in several observational studies and clinical trials that facilitates rigorous Aβ and tau PET quantification and data sharing. Highlights Quantitative results from ADNI Aβ and tau PET data are generated using a rigorous, scalable image processing pipeline This pipeline has been applied to PET data from several other large, multisite studies and trials Quantitative outcomes are harmonizable across studies and are shared with the scientific community

Neurosciences & Neurology↗

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION↗