Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

The Early Data Release of the Dark Energy Spectroscopic Instrument

The Dark Energy Spectroscopic Instrument (DESI) completed its 5 month Survey Validation in 2021 May. Spectra of stellar and extragalactic targets from Survey Validation constitute the first major data sample from the DESI survey. This paper describes the public release of those spectra, the catalogs of derived properties, and the intermediate data products. In total, the public release includes good-quality spectral information from 466,447 objects targeted as part of the Milky Way Survey, 428,758 as part of the Bright Galaxy Survey, 227,318 as part of the Luminous Red Galaxy sample, 437,664 as part of the Emission Line Galaxy sample, and 76,079 as part of the Quasar sample. In addition, the release includes spectral information from 137,148 objects that expand the scope beyond the primary samples as part of a series of secondary programs. Here, we describe the spectral data, data quality, data products, Large-Scale Structure science catalogs, access to the data, and references that provide relevant background to using these spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

Joint Analysis of Small-scale Galaxy Clustering and Galaxy–Galaxy Lensing from BOSS Galaxies

We present a joint analysis of galaxy clustering and galaxy–galaxy lensing measurements from BOSS galaxies using a simulation-based emulation method combined with a halo occupation distribution model. Our emulators are constructed with the Aemulus ν simulations, a suite of wνCDM N-body simulations with massive neutrinos as independent particle species. We combine small-scale analysis of clustering from 0.1 to 60.2 h −1 Mpc and lensing from 1.7 to 60.2 h −1 Mpc to perform cosmological constraints. We split the BOSS galaxies into three redshift bins to measure their clustering and employ galaxies from Dark Energy Camera Legacy Survey and Hyper Suprime-Cam as source galaxies to measure lensing separately. We find that the addition of lensing significantly improves the constraining power on $S_8 = σ_8(Ω_m/0.3)^{0.5}$, with a weak improvement for fσ 8 . Our results of fσ 8 indicate tensions of around 1σ−4σ below the results of the cosmic microwave background observations of Planck. For S 8 , our results are also lower than Planck, and the tension can be mitigated when considering possible systematics in lensing measurement. As a by-product, our analysis prefers a nonzero neutrino mass but without strong significance, with the constraining power dominated by the clustering. Given the accuracy and precision of our model and the observational data, it is anticipated that larger and higher-quality spectroscopic data sets will improve the constraints on this fundamental property in the near future.

Gao, Wenhao [Shanghai Jiao Tong University (China)↗

Status of the CERBERUS Evaluation for the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook

Modeling & Simulation (M&S) tools are used to analyze advanced reactor designs and the safety of current nuclear operations. As computers continue to improve, we are able to enhance resolution in our calculations. Therefore, the limitations of simulation capability are in the quality of data that is being used, including our ability to quantify the uncertainty and sensitivity of that data. In order to model systems of interest with increasing accuracy, the industry must improve key nuclear data measurements. The International Criticality Safety Benchmark Evaluation Project (ICSBEP) compiles and evaluates experiment data in a handbook that can be used by criticality safety engineers and others to validate computer codes and cross section libraries at nuclear facilities. Both critical and subcritical experiments are included in the handbook. These experiments, along with differential measurements, can help improve the quality of nuclear data. Concerns regarding the accuracy of Cu nuclear data have been published. The large values and trend of C-E for the Zeus intermediate energy benchmark, being one of the primary examples. Furthermore, very few experiments have been designed to be sensitive to Cu (as shown in Figure 1), so an integral, critical experiment is needed to help resolve these differences.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

WFIP3 - NOAA SHIP site - NREL Ceilometer (Vaisala CL51) / Derived Data

NOAA SHIP ceilometer: netCDF L3 data files have level 3 (L3) data that have gone through the calculation service and contain all the data from the algorithms, including mixing layer height values, and quality index data. L3 default files contain L3 data that use the default preset for a live plot. File naming schema: L3_DEFAULT_ _YYYYMMDDHHMM_ _ .nc Name Description: L3 Identification of the data level DEFAULT Identification of the L3 file type CUSTOM OFFLINE STATION_NUMBER WMO station number, if defined YYYYMMDDHHMM UTC time ParameterKey Identification of the advanced algorithm settings. See the table below for an explanation. FREE_FORMAT File suffix, if defined

17 WIND ENERGY↗

WFIP3 - CACO site - NREL Ceilometer (Vaisala CL51) / Derived Data

CACO ceilometer: netCDF L3 data files have level 3 (L3) data that have gone through the calculation service and contain all the data from the algorithms, including mixing layer height values, and quality index data. L3 default files contain L3 data that use the default preset for a live plot. File naming schema: L3_DEFAULT_ _YYYYMMDDHHMM_ _ .nc Name Description: L3 Identification of the data level DEFAULT Identification of the L3 file type CUSTOM OFFLINE STATION_NUMBER WMO station number, if defined YYYYMMDDHHMM UTC time ParameterKey Identification of the advanced algorithm settings. See the table below for an explanation. FREE_FORMAT File suffix, if defined

17 WIND ENERGY↗

Quantifying uncertainty in regional-scale seismic moment tensors

We examine the ability of three different inversion methods: 1) first motion (FM) inversions, 2) amplitude inversions, and 3) full waveform (FW) time variable moment tensor (TVMT) inversions, to recover an accurate source mechanism for simulated data for an earthquake as well as an explosion. The ability of inversion methods, such as the ones described above, to recover accurate models representative of the data depends on both a priori information as well as the quality of data being inverted. Therefore, we examine the effect that station geometry, geologic model, and noise level has on the inversion and estimated source mechanism. We find that FM data can provide more robust solutions than amplitude data and are not as sensitive to inaccurate earth models, especially in low-noise cases, and overall FW inversions are the most accurate out of the methods examined, but can still be biased by inaccurate earth models.

58 GEOSCIENCES↗

Fast event-based electron counting for small-molecule structure determination by MicroED

Electron counting helped realize the resolution revolution in single-particle cryoEM and is now accelerating the determination of MicroED structures. Its advantages are best demonstrated by new direct electron detectors capable of fast (kilohertz) event-based electron counting (EBEC). This strategy minimizes the inaccuracies introduced by coincidence loss (CL) and promises rapid determination of accurate structures. We used the Direct Electron Apollo camera to leverage EBEC technology for MicroED data collection. Given its ability to count single electrons, the Apollo collects high-quality MicroED data from organic small-molecule crystals illuminated with incident electron beam flux densities as low as 0.01–0.045 e − /Å 2 /s. Under even the lowest flux density (0.01 e − /Å 2 /s) condition, fast EBEC data produced ab initio structures of a salen ligand (268 Da) and biotin (244 Da). Each structure was determined from a 100° wedge of data collected from a single crystal in as few as 50 s, with a delivered fluence of only ∼0.5 e − /Å 2 . Fast EBEC data collected with a fluence of 2.25 or 3.33 e − /Å 2 also facilitated a 1.5 Å structure of thiostrepton (1665 Da). While refinement of these structures appeared unaffected by CL, a CL adjustment applied to EBEC data further improved the distribution of intensities measured from the salen ligand and biotin crystals. However, CL adjustment only marginally improved the refinement of their corresponding structures, signaling the already high counting accuracy of detectors with counting rates in the kilohertz range. Overall, by delivering low-dose structure-worthy data, fast EBEC collection strategies open new possibilities for high-throughput MicroED.

EBEC↗

End-to-end deep learning pipeline for real-time Bragg peak segmentation: from training to large-scale deployment

X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.

Wang, Cong↗

Biophysical model of eelgrass and water quality in Coos Bay, OR shows greater mitigation potential for ocean acidification than hypoxia

Seagrass beds provide important ecosystem services and are valued, in part, for their potential to mediate stressors such as ocean acidification and hypoxia (OAH) for sensitive species. However, the susceptibility of seagrasses to anthropogenic impacts and recent declines motivate the need to better understand the drivers of seagrass and the water quality consequences that occur with variation in seagrass abundance. To meet this need, we leveraged existing monitoring data (water quality and seagrass), hydrodynamic circulation model, and biogeochemical model framework with seagrass submodel, to produce a biophysical model of Coos Bay estuary, Oregon, U.S. The model includes biogeochemical processes involving water quality, plankton, seagrass, and sediment-water interactions. Ecosystem models like this are useful for evaluating complex estuarine systems because they allow us to extend our understanding of system dynamics beyond existing observations and perform experiments to identify the processes driving observed patterns. We used the biophysical model of Coos Bay to evaluate the dynamics of water quality and native eelgrass (Zostera marina) under three eelgrass abundance scenarios (zero eelgrass, current extent, and maximum observed extent) to elucidate the relationship between eelgrass and OAH. Including eelgrass in the Coos Bay model produced results that more closely resembled water quality observations - dissolved oxygen (DO) and pH were more dynamic in simulations with eelgrass, often having both higher highs and lower lows. While there were some areas of the estuary where DO improved with the addition of eelgrass to the model there was overall a small net increase in harmful DO conditions (based on a salmon physiological threshold). In contrast, ocean acidification conditions, pH and calcium carbonate saturation state for aragonite (Ω), were improved (based on oyster requirements) with the addition of eelgrass - although the magnitude of improvement differed seasonally and spatially. Our new model represents a useful tool - one which accounts for and controls the relevant physical and biogeochemical processes - to evaluate conditions that confer resilience or enhance vulnerability to OAH in an important Pacific Northwest coastal estuary and results can inform the OAH-related dynamics occurring in other eastern boundary current estuaries.

FVCOM-ICM↗

Computational Workflows for Uncertainty-Quantified Nuclear Reactions: From Nuclear Theory Inputs to Astrophysical Reaction Rates

Reactions on unstable nuclei, particularly those on the neutron-rich side of stability, are important for both fundamental and applied physics. For fundamental science, the most prevalent use case is astrophysi cal nucleosynthesis by rapid neutron capture—the r-process—by which heavy nuclei are formed in extreme astrophysical environments, such as in supernovae and neutron star mergers; see, e.g., Refs. [1–3]. For ap plications, these processes are relevant for the interpretation of radiochemical data from historic nuclear tests, which contribute to our ability to certify the enduring stockpile in the absence of nuclear testing [4]; see Ref. [5] for a broader discussion of applications. However, reaction cross sections involving unsta ble species are generally poorly understood, for the simple reason that useful data become scarce as one moves away from stability. While there are avenues for improving the amount and quality of data for these species [6], one is fundamentally reliant on nuclear theory to make progress on these fields of study.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ecological Insights from Transferable Plant Biomass Mapping across the Arctic using High-resolution Structure-from-Motion and LiDAR Data

Warmer temperatures, permafrost thaw, and increased wildfire activity are driving rapid ecological change across the Arctic, significantly altering plant productivity and aboveground biomass (AGB). These rapid changes highlight the urgent need to improve monitoring of vegetation dynamics in the Earth’s northern ecosystems, where high spatiotemporal heterogeneity occurs at scales finer than those captured by traditional satellite observations. The growing use of Unoccupied Aerial Systems (UASs) presents an opportunity to overcome this limitation. Yet, the diversity of UAS platforms, sensors, and data collection and processing workflows presents challenges for developing standardized, generalizable approaches. To address this challenge, we compiled 672 AGB plots co-located with 183 UAS-based Structure-from-Motion (SfM) or Light Detection and Ranging (LiDAR) surveys collected across the Arctic. Here, we: (1) evaluated the generalizability of UAS-derived canopy structure derived from high-resolution SfM and LiDAR for estimating AGB, (2) assessed scaling errors and their sources in two recent satellite-based AGB products derived from Landsat and MODIS, and (3) demonstrated the use of high-resolution AGB maps to quantify biomass variation across tundra plant functional types (PFTs) and to monitor post-fire recovery. Our results show that both SfM and LiDAR accurately captured AGB and its variability across tundra PFTs using a Random Forest (RF) model (overall RMSE: 0.336 kg/m2), with mapping performance varying slightly by region and data source. Using UAS-derived AGB maps as a benchmark, we identified systematic biases in satellite-derived AGB products, largely attributable to the magnitude of AGB and structural heterogeneity within coarse-resolution pixels. Applying our model to repeat UAS surveys following a tundra fire on Seward Peninsula, we observed rapid AGB recovery in non-shrub patches, with biomass recovering to pre-fire levels within 2 years. In contrast, shrub patches recovered more slowly, with AGB gains continuing over 2–4 years through both in-patch growth and lateral expansion (via dispersal) into remaining burned areas. Overall, these findings demonstrate the generalizability of UAS-based SfM and LiDAR data for estimating tundra AGB and highlight the potential of our approach to be broadly applied to generate high-quality AGB data for ecological monitoring and model benchmarking across the Arctic.

Yang, Daryl [ORNL] (ORCID:0000000317057823)↗

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics↗

ENVnet provides a global molecular resource of dissolved organic matter

Dissolved organic matter (DOM) is an important component of Earth's carbon cycle and one of the planet's most chemically diverse pools, yet the molecular structures of its constituents remain largely unresolved. This limitation has hindered our ability to link DOM composition to microbial processes and ecosystem function. Here we present ENVnet, a global molecular repository built from tandem mass spectrometry data collected across 13 terrestrial and aquatic environment types, including 419 newly generated samples that expand publicly available DOM metabolomics data and cover previously underrepresented environments. By computationally deconvolving chimeric mass spectra, a longstanding challenge in environmental metabolomics, we recover high-quality fragmentation data for >22,000 distinct molecular features (defined by a specific precursor mass and fragmentation pattern). Using ENVnet, we uncover conserved and environment-specific molecular patterns in DOM composition and underlying biogeochemical processes. We also use molecular features encoded in ENVnet to train predictive models of DOM persistence, allowing molecular-level assessment of microbial turnover in independent systems.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the Nation's Pipeline Infrastructure with NETL's Advanced Infrastructure Integrity Model (AIIM)

This poster is a part of BIL-EDX4CCS Task 36: Advanced Infrastructure Integrity Modeling to Evaluate Existing Energy Infrastructure Reusability and Risk, the goal of which is to produce a smart tool that will assess existing energy infrastructure reusability and risk using the Advanced Infrastructure Integrity Model (AIIM). This model forecasts lifespan and potential risk using a multitude of factors such as incidents reports, structural characteristics, and the surrounding environment. The project aims to provide scientific insights for a better understanding of carbon storage (CS), potential to support CS stakeholder needs, national decarbonization, and mitigating climate change. AIIM will utilize an energy infrastructure database as its input, developed by acquiring publicly available data as well as NETL derived products. These resources include incidents, geohazards, and infrastructure variables. Soil data in the form of rasters and pipeline incident reports were processed and a script was developed to count the number of times features such as roads, railroads, and rivers intersected with pipeline segments which were then converted to points. Distance to oil and natural gas wells, petroleum ports, intermodal freight facilities, and geologic structures were also calculated. After data preparation and quality control was completed, the data was integrated into the pipeline points. Once models are complete, a smart tool will be created in the form of an online dashboard.

Malay, Caleb↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 8: Marine Renewable Energy Data and Information Systems

As the marine renewable energy (MRE) sector grows, large amounts of environmental and technical data and information are being collected. When these data and information are openly available, they can be used to guide research and development, inform responsible siting and consenting of projects, and increase stakeholder understanding through transparency. For example, quality environmental data collected during the siting, consenting, construction, operation, and decommissioning of MRE projects can all play key roles in better characterizing baseline conditions, developing effective monitoring and mitigation strategies, and retiring environmental risks through data transferability (see Chapter 6). Ensuring that these data and information are easily discoverable and accessible will help the MRE sector make informed decisions and coexist in an increasingly busy ocean environment.

16 TIDAL AND WAVE POWER↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

The Dynamic Networks Experiments: Virtual Experiments to Quantify Gains in Nuclear Explosion Monitoring

We describe an ongoing series of virtual experiments conducted collaboratively by four United States National Laboratories: Sandia National Laboratories, Los Alamos National Laboratory, Lawrence Livermore National Laboratory, and Pacific Northwest National Laboratory. These Dynamic Network Experiments (DNEs) provide an experimental framework to evaluate the potential impact of new research tools on nuclear explosion monitoring. The second DNE (DNE2), completed in 2024, exploited waveform data (seismic, infrasound, and electromagnetic) that was recorded by multi-modal sensors within and near the Nevada National Security Site and synthetic radionuclide signatures over multiple time periods. During the execution of DNE2, we processed and analyzed data through a multi-stage event processing pipeline that ingested raw data, performed quality control, detected signals, built events from these signals, located these events, and characterized the events’ source types and sizes. For each stage and over the entire event processing pipeline, we evaluated performance changes by comparing the performance of new data processing methods, models, and algorithms against a baseline. We also performed an additional execution phase to assess event processing pipeline function, speed, and efficiency against that of an expert analyst, including computational and manual efforts. Finally, we assessed the impact and effort of modern computing infrastructure on the monitoring pipeline. This paper describes key elements of the DNEs, from formulation through execution, as demonstrated in DNE2. The DNEs introduce several novel concepts to quantitatively measure the potential impact of new methods on explosion monitoring, including the collaborative design of multi-modal datasets, performance and logistical metrics, and integrated analyses.

42 ENGINEERING↗