Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Hot or Not? An Evaluation of Methods for Identifying Hot Moments of Nitrous Oxide Emissions From Soils

Abstract Effectively quantifying hot moments of nitrous oxide (N 2 O) emissions from agricultural soils is critical for managing this potent greenhouse gas. However, we are challenged by a lack of standard approaches for identifying hot moments, including (a) determining thresholds above which emissions are considered hot moments, and (b) considering seasonal variation in the magnitude and frequency distribution of net N 2 O fluxes. We used one year of hourly N 2 O flux measurements from 16 autochambers that varied in flux magnitude and frequency distribution in a conventionally tilled maize field in central Illinois, USA, to compare three approaches to identify hot moment thresholds: standard deviations (SD) above the mean, 1.5x the interquartile range (IQR), and isolation forest (IF) identification of anomalous values. We also compared these approaches on seasonally subdivided data (early, late, and non‐growing seasons) versus the whole year. Our analyses revealed that 1.5x IQR method best identified N 2 O hot moments. In contrast, using 2 or 4 SD both yielded hot moment threshold values too high, and IF yielded threshold values too low, leading to missed N 2 O hot moments or low net N 2 O fluxes mischaracterized as hot moments, respectively. Furthermore, seasonally subdividing the data set not only facilitated identification of smaller hot moments in the late‐ and non‐growing seasons when N 2 O hot moments were generally smaller but it also increased hot moment threshold values in the early growing season when N 2 O hot moments were larger. Consequently, of the methods evaluated here, we recommend using the 1.5x IQR method on whole year data sets to identify N 2 O hot moments.

Stuchiner, Emily R. [Institute for Sustainability,↗

Development of Machine Learning Algorithm for Pebble Bed Modular Reactor Misuse Detection

The objective of this work was to develop a machine learning ensemble that could assist pebble bed reactor verification by evaluating whether a given pebble circulating through a PBR was normal or anomalous using gamma spectroscopy measurements from a notional PBR burnup measurement system. Using a PBR reference design, data sets of synthetic gamma spectra representative of BUMS measurements of normal and anomalous pebbles that may be used to produce special fissile material were generated to train and test an ML anomaly detection ensemble on two reference scenarios – substitution of normal pebbles with target pebbles for production of Pu or 233 U. The ML ensemble correctly identified all anomalous pebbles in the testing data set, and while perfect ensemble performance is normally indicative of overfitting, it was concluded that significantly lower photon intensity of target pebbles produced distinctly less intense photon spectra to where perfect ensemble performance was expected.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Characterization of uncertainties in electron-argon collision cross sections

Abstract The predictive capability of a plasma discharge model depends on accurate representations of electron-impact collision cross sections, which determine the corresponding reaction rates and electron transport properties. The values of cross sections can be known only approximately either through experiments or simulations and are thus subject to uncertainties. Quantifying the uncertainties in plasma simulations allows us to assess the reliability of simulations and to provide a basis for interpreting discrepancies between simulations and experiments. For such uncertainty quantification of plasma simulations, it is essential to quantify the uncertainties of the underlying cross sections. Although much effort has been committed to calibrate the cross section values, their uncertainties are not well investigated. We characterize uncertainties in electron-argon atom collision cross sections using a Bayesian framework. Six collision processes—elastic momentum transfer, ionization, and four excitations—are characterized with semi-empirical models, which effectively capture the features important to the macroscopic properties of the plasma. A probability model for the uncertain parameters of these semi-empirical models is developed. Specifically, a Gaussian-process likelihood model is proposed to capture discrepancies among data sets, as well as the model-form inadequacies of the semi-empirical models. Two other likelihood models are compared with the proposed Gaussian-process model, to illustrate the importance of the choice of the likelihood model. The cross section models are calibrated using the electron-beam experiments and ab-inito quantum simulations. The resulting calibrated uncertainties capture well the scattering among the data sets. The calibrated cross section models are further validated against swarm-parameter experiments and zero-dimensional Boltzmann equation simulations of widely used cross section datasets.

Chung, Seung Whan (ORCID:0000000302501549)↗

Automated and High-Throughput Phase Separation Control for Supramolecular Polymer Blends Enabled by Machine Learning

Supramolecular polymer blends (SPBs) offer tunable morphologies that dictate their macroscopic properties, yet their rational design is limited by the absence of predictive structure−morphology models. Here, we introduce a data-driven highthroughput workflow that integrates modular polymer synthesis, robotic formulation, automated morphology characterization, and machine learning (ML) for accelerated SPB discovery. Using a plug-and-play synthetic strategy, 33 hydrogen-bonding endfunctional homopolymers were prepared and orthogonally combined to generate 260 SPBs in 1 day. A fully automated atomic force microscopy (AFM) pipeline enabled systematic imaging, producing 2340 morphology data sets with minimal human intervention. Domain spacings were extracted through complementary imageprocessing methods and used to train ML models. A support vector regression (SVR) model accurately predicted target phase-separation sizes (50, 100, and 150 nm), which were experimentally validated. This work demonstrates the power of coupling high-throughput experimentation with ML to accelerate morphology discovery and provides one of the first large-scale experimental data sets for supramolecular polymer systems.

ML-guided polymer design↗

CLEAP Project: OR-SAGE Analysis for MT, UT, and CO States

The OR-SAGE tool is designed to use industry-accepted practices in screening sites and then employ the proper array of data sources through the considerable computational capabilities of GIS technology available at ORNL. The tool was developed to screen the potential for NPP siting on a national and regional basis. However, because of the tool granularity, it is often focused specifically on the immediate area around user sites of interest. If data center siting parameters can be added to OR-SAGE, the ability to evaluate data center siting on a localized scale will be beneficial.1 More than 60 data sets have been collected and processed by ORNL to develop exclusionary, avoidance, and suitability criteria for screening sites for a variety of power generation types, including nuclear power plants. Available site evaluation parameters include population density, slope, seismic activity, proximity to cooling-water sources, proximity to hazard facilities, avoidance of protected lands and floodplains, susceptibility to landslide hazards, and many others. All siting parameters should be considered as flags to inform siting decisions and should not be used to rule in or rule out any NPP site. Once data center siting parameters are identified, appropriate data sets will be collected and processed. The OR-SAGE process is very versatile. Essentially, OR-SAGE is a visual, relational database. The database partitions the contiguous United States, a total of 720 million hectares (~1.8 billion acres), into 100-m by 100-m (1 hectare or ~2.5 acre) cells. The database is tracking just under 700 million individual land cells. Successive suitability criterion is applied to each cell in the database. User-specified thresholds can be applied to each siting parameter data layer. In this manner, a variety of scenarios can be quickly and thoroughly evaluated. Data can be added and/or revised within OR-SAGE to address user interests. Siting security assessment capability is currently being added to OR-SAGE. Security is expected to be of concern at data centers whether it is collocated with a nuclear power generating technology or not. If data center is collocated with a nuclear power generating source, the security threat attractiveness level of both will likely increase. It will be of additional benefit if a potential data center site is also assessed for security vulnerability.

97 MATHEMATICS AND COMPUTING↗

SPRUCE: Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2024

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2024. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored at Oak Ridge National Laboratory and may be available for further analysis. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B00VQ. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR. Note: Only dried and ground material from C Cores are available for new analysis.

Birkebak, Joshua [ORNL] (ORCID:0009000955611494)↗

Rheology Investigations with Sludges from Metro Vancouver

Rheological investigation were performed with primary and secondary waste water treatment sludge. Using a rheometer equipped with a high pressure/temperature cell, flow curves were generated over shear rate at 0 to 1000 s-1. Temperature sweeps spanning 25 to 300 C were also performed are are reported here. The original release, PNNL-SA-185826, is being revised. The revision include a revision table, a disclaimer, and the underlying data set is being added.

sludge wastewater treatment plant sludge character↗

A Mountain Glacier Perspective on the Bipolar Seesaw

A global record of mountain glacier terminations during the last deglaciation (∼19–11 ka) dated by a large, uncurated data set of cosmogenic-nuclide exposure ages highlights a statistically significant asynchrony in termination ages between the Northern and Southern Hemispheres. This interhemispheric offset in the timing of glacier terminations is consistent with previously correlated ice core records that show a systematic interhemispheric lag in the timing of abrupt climate events, with the Southern Hemisphere leading the Northern Hemisphere by ∼300–3,500 years. Our analysis (a) aggregates cosmogenic-nuclide exposure ages from a global data set of moraines to discern climatically driven peaks in moraine emplacement events, and (b) utilizes a Monte Carlo simulation based on a null hypothesis that moraine emplacement is interhemispherically synchronous to estimate the statistical significance of the observed offset. The observed lag of Northern Hemisphere emplacement events compared to the Southern Hemisphere is statistically significant and is consistent with the “bipolar seesaw” pattern observed in ice core records.

58 GEOSCIENCES↗

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology↗

Simulations suggest offshore wind farms modify low-level jets

Abstract. Offshore wind farms are scheduled to be constructed along the East Coast of the US in the coming years. Low-level jets (LLJs) – layers of relatively fast winds at low altitudes – also occur frequently in this region. Because LLJs provide considerable wind resources, it is important to understand how LLJs might change with turbine construction. LLJs also influence moisture and pollution transport; thus, the effects of wind farms on LLJs could also affect the region’s meteorology. In the absence of observations or significant wind farm construction as yet, we compare 1 year of simulations from the Weather Research and Forecasting (WRF) model with and without wind farms incorporated, focusing on locations chosen by their proximity to future wind development areas. We develop and present an algorithm to detect LLJs at each hour of the year at each of these locations. We validate the algorithm to the extent possible by comparing LLJs identified by lidar, constrained to the lowest 200 m, to WRF simulations of these very low LLJs (vLLJs). In the NOW-WAKES simulation data set, we find offshore LLJs in this region occur about 25 % of the time, most frequently at night, in the spring and summer months, in stably stratified conditions, and when a southwesterly wind is blowing. LLJ wind speed maxima range from 10 m s−1 to over 40 m s−1. The altitude of maximum wind speed, or the jet “nose”, is typically 300 m above the surface, above the height of most profiling lidars, although several hours of vLLJs occur in each month in the data set. The diurnal cycle for vLLJs is less pronounced than for all LLJs. Wind farms erode LLJs, as LLJs occur less frequently (19 %–20 % of hours) in the wind farm simulations than in the no-wind-farm (NWF) simulation (25 % of hours). When LLJs do occur in the simulation with wind farms, their noses are higher than in the NWF simulation: the LLJ nose has a mean altitude near 300 m for the NWF jets, but that nose height moves higher in the presence of wind farms, to a mean altitude near 400 m. Rotor region (30–250 m) wind veer is reduced across almost all months of the year in the wind farm simulations, while rotor region wind shear is similar in both simulations.

17 WIND ENERGY↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Modeling Partial Reflection Paths for Infrasound Analysis

Numerical methods enabling simulation of scattered and partially reflected infrasonic propagation paths produced by interaction with fine-scale structure in the middle atmosphere have been implemented in the infraGA ray tracing software. This capability enables simulation of ensonification in the classical stratospheric “shadow zone” that has been observed during the Humming Roadrunner and LSECE surface explosion campaigns as well as in other data sets. In the case of LSECE, a pair of stations roughly 140 kilometers east of the source location observed arrivals with celerities (horizontal group velocities) slightly slower than observed stratospheric paths at similar azimuths. The arrivals exhibited increasing trace velocity later in the wavetrain indicating a steepening of the arrival path for longer or slower propagation paths. Simulation of partially reflected paths using the updated infraGA software methods finds good agreement between observed and predicted infrasonic ensonification at these locations within the stratospheric shadow zone. Further development of the partial reflection physics and comparison with other data sets is needed to more robustly understand how such anomalous infrasonic signals can be predicted; however, the demonstration of this capability is a promising first step in such analyses.

97 MATHEMATICS AND COMPUTING↗

AERO-MAP: a data compilation and modeling approach to understand spatial variability in fine- and coarse-mode aerosol composition

Abstract. Aerosol particles are an important part of the Earth climate system, and their concentrations are spatially and temporally heterogeneous, as well as being variable in size and composition. Particles can interact with incoming solar radiation and outgoing longwave radiation, change cloud properties, affect photochemistry, impact surface air quality, change the albedo of snow and ice, and modulate carbon dioxide uptake by the land and ocean. High particulate matter concentrations at the surface represent an important public health hazard. There are substantial data sets describing aerosol particles in the literature or in public health databases, but they have not been compiled for easy use by the climate and air quality modeling community. Here, we present a new compilation of PM2.5 and PM10 surface observations, including measurements of aerosol composition, focusing on the spatial variability across different observational stations. Climate modelers are constantly looking for multiple independent lines of evidence to verify their models, and in situ surface concentration measurements, taken at the level of human settlement, present a valuable source of information about aerosols and their human impacts complementarily to the column averages or integrals often retrieved from satellites. We demonstrate a method for comparing the data sets to outputs from global climate models that are the basis for projections of future climate and large-scale aerosol transport patterns that influence local air quality. Annual trends and seasonal cycles are discussed briefly and are included in the compilation. Overall, most of the planet or even the land fraction does not have sufficient observations of surface concentrations – and, especially, particle composition – to characterize and understand the current distribution of particles. Climate models without ammonium nitrate aerosols omit ∼ 10 % of the globally averaged surface concentration of aerosol particles in both PM2.5 and PM10 size fractions, with up to 50 % of the surface concentrations not being included in some regions. In these regions, climate model aerosol forcing projections are likely to be incorrect as they do not include important trends in short-lived climate forcers.

Mahowald, Natalie M. (ORCID:000000022873997X)↗

Analyzing the impact of design factors on solar module thermomechanical durability using interpretable machine learning techniques

Solar modules in utility-scale systems are expected to maintain decades of lifetime to rival conventional energy sources. However, cyclic thermomechanical loading often degrades their long-term performance, highlighting the importance of effective design to mitigate thermal expansion mismatches between module materials. Given the complex composition of solar modules, isolating the impact of individual components on overall durability remains a challenging task. In this work, we analyze a comprehensive data set that comprises bill-of-materials (BOM) and thermal cycling power loss from 251 distinct module designs to identify the predominant design factors and their impacts on the thermomechanical durability of modules. The methodology of our analysis combines machine learning modeling (random forest) and Shapley additive explanation (SHAP) to correlate design factors with power loss and interpret the model’s decision-making. The interpretation reveals that silicon type (monocrystalline or polycrystalline), encapsulant thickness, busbar numbers, and wafer thickness predominantly influence the degradation. With lower power loss of around 0.6% on average in the SHAP analysis, monocrystalline cells present better durability than polycrystalline cells. This finding is further substantiated by statistical testing on our raw data set. The SHAP analysis also demonstrates that while thicker encapsulants lead to reduced power loss, further increasing their thickness over around 0.6 to 0.7 mm does not yield additional benefits, particularly for the front side one. In addition, other important BOM features such as the number of busbars are analyzed. This study provides a blueprint for utilizing explainable machine learning techniques in a complex material system and can potentially guide future research on optimizing the design of solar modules.

14 SOLAR ENERGY↗

Quantifying motion blur by imaging shock front propagation with broadband and narrowband X-ray sources

Time-integrated radiography using MeV Bremsstrahlung X-ray sources is the norm for imaging during system-level testing of components and structures under dynamic condition. One source of error in the analysis of the time-integrated radiography data sets stems from motion blur which smears out sharp interfaces to a greater degree with longer exposure times, which become necessary to provide sufficient signal-to-noise with low X-ray penetration of objects of interest. To quantify motion blur, a 1D shock wave through PMMA was investigated experimentally at The Dynamic Compression Sector at The Advanced Photon Source (DCS@APS) with tapered broadband and 25.46 ± 1.06 keV narrowband X-rays. Four cameras with different exposure times were used for each experiment to compare the effect that exposure time has on motion blur. In addition, our methodology to accurately simulate motion blur in terms of transmission and shape is presented and compared to our experimental results and quantified. There is a high level of agreement between the experimental and simulation results across the range of data sets investigated in this study with a percent difference range of 0.29–1.31% for the four shots. The methodology of this work serves as a steppingstone towards a physically validated model that could be used in conjunction with experimental results to deconvolve physical parameters, densities, and interfaces of interest in a way that would not be possible with experimental results alone.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Source apportionment of aerosols at the White River IMPROVE site near the SAIL site

This data set contains source apportionment results at the White River IMPROVE site (39.1536, -106.8209), which is about 30 km north of the Surface Atmosphere Integrated Field Laboratory (SAIL) Campaign site. The IMPROVE network (Malm et al. 1994) collected 24-hour aerosol filter samples every three days over several decades at this site. Chemical concentrations in the PM2.5 fraction of 19 elements (Al, As, Br, Ca, Cl, Cr, Cu, Fe, K, Mg, Mn, Na, Ni, Pb, Se, Si, Ti, V, and Zn), along with nitrate, sulfate, elemental carbon (EC), organic carbon (OC), and calculated coarse mass concentrations (PM10−PM2.5 mass concentrations), from 2014 to 2023, were used as input for the PMF analysis. PMF was performed using EPA PMF 5.0 (Norris et al. 2014). A five-factor solution was chosen as the optimal solution. These factors were identified as coarse dust, fine dust, biomass burning, sulfate-dominated, and nitrate-dominated sources. This data set is useful for understanding aerosol sources and their long-term variability near this region.

biomass burning↗

3 He +𝛼 resonances in 7 Be

Resonances in 7 Be which decay into the 3 He+α exit channel have been measured with improved precision using preexisting data sets. The energy and width of the J π =7/2 - state have been extracted from an invariant-mass study of projectile-breakup products originating from interactions of an E/A=10.7-MeV 10 C beam on Be and C targets. The excitation energy of this state (from the pole of the S-matrix) is determined to be E*=4.545(6)~MeV, a factor of 8 improvement in precision as compared to the ENSDF value. This improvement is enabled by fine tuning the detector calibrations using calibration resonances in 6 Li, 7 Li, 6 Be, 9 B, and 12 C whose decay energies are known to high precision. The J π =7/2- resonance in 7 Be can now itself be used as a calibration resonance in invariant-mass experiments. This utility is particularly helpful for 3 He energy calibrations of CsI(Tl) detectors which are often used in detector arrays employed for measurements with fast beams. This utility is demonstrated with a data set associated with E/A=70 MeV 7 Be beams which are inelastically excited to the 7/2 - and 5/2 - 1 resonances. Again using the pole of the S-matrix as the definition of the resonance parameters, the fitted excitation energy of the 5/2 - 1 resonance is 6.376(17) MeV, approximately 300 keV lower than the ENSDF value. Finally, its width of 565(4) keV is roughly half of the ENSDF value.

energy levels↗