Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty

Inverse problems, which aim to infer unknown properties of a system using experimental and observational data, are central to addressing many of the U.S. Department of Energy’s (DOE) most critical scientific and engineering challenges. Accurate, computationally efficient, and data-efficient solutions to inverse problems are essential for advancing DOE mission-critical science drivers, including analyzing data from large-scale experimental facilities, optimizing fusion reactor performance, accelerating materials discovery, enhancing geophysical imaging, improving wildfire predictions, and enabling autonomous systems and digital twins. However, these problems are becoming increasingly complex, often involving nonlinear, highdimensional, and interconnected systems and models that span multiple physics and scales, while relying on data with varying quantity, quality, and information content. Compounding these challenges is the uncertainty inherent in DOE-relevant systems, where errors in inputs, noise in data, incompleteness of data, and discrepancies between models and reality constrain the accuracy and precision of solutions. At the same time, the convergence of recent scientific computing trends—scientific machine learning, artificial intelligence, and computing advances such as exascale computing—is creating unprecedented opportunities for tackling these challenges. The cross-cutting nature of inverse problems, combined with their growing complexity and rapidly evolving data and algorithmic demands, strongly motivates the formulation of a prioritized research agenda to maximize their capabilities and impact. In response to this need, DOE’s Advanced Scientific Computing Research (ASCR) program in the Office of Science convened the Workshop on Basic Research Needs for Inverse Problems for Complex Systems Under Uncertainty in June 2025. This workshop brought together experts across disciplines to identify grand challenges and major opportunities in the field. Through collaborative discussions, the workshop defined transformative research directions aimed at addressing the mathematical, statistical, and computational challenges posed by inverse problems under uncertainty. As a result of these efforts, four priority research directions (PRDs) were identified to guide future research and development in this area. These PRDs, summarized below, represent a roadmap for advancing the foundational science and mathematics of inverse problems, enabling robust, scalable, and uncertainty-aware solutions that are critical for DOE applications.

97 MATHEMATICS AND COMPUTING↗

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Online energy consumption forecast for battery electric buses using a learning-free algebraic method

Accurately predicting the energy consumption plays a vital role in battery electric buses (BEBs) route planning and deployment. Based on the algebraic derivative estimation, we present a novel method to forecast the energy consumption in real time. In contrast to the mainstream machine-learning-based methods, the proposed method does not require access to the historical energy consumption data. It eliminates the time-consuming and computationally expensive offline training. Consequently, its prediction performance is not constrained by the quantity and quality of the training data. Moreover, the method can swiftly adapt to new situations not included in the previous driving cycles, which makes it especially suitable for emerging transport modes, e.g., on-demand transit services. In addition, its online execution only involves algebraic calculations, yielding superior calculation efficiency. Using real-world data, we comprehensively compare the performance of the proposed learning-free algebraic method with multiple representative machine-learning-based methods. Finally, the advantages and limitations of the proposed method are discussed in detail.

33 ADVANCED PROPULSION SYSTEMS↗

NCAR/EOL ISFS Data for LASSO-CACTI Overview Paper

5 minute averages of surface meteorology and flux data collected by the NCAR/EOL Integrated Surface Flux System (ISFS) at 15 sites during the RELAMPAGO field campaign. These data have been quality-controlled and are available in NetCDF format. Winds reported by the sonic anemometers have been tilt corrected and rotated into geographic coordinates. Data providence, citation, and acknowledgement This ARM data set is a copy of v2.0 of the NCAR data set obtained on 6-Jun-2024 from https://doi.org/10.26023/ZPHJ-JW9W-2B0Y. The citation for the original data source is: NCAR/EOL In-situ Sensing Facility, Oncley, S. 2021. NCAR/EOL ISFS Surface Meteorology and Flux Products, 5-minute. Version 2.0. UCAR/NCAR - Earth Observing Laboratory. https://doi.org/10.26023/ZPHJ-JW9W-2B0Y Accessed 06 Jun 2024. In addition to the citation reference and any other acknowledgements, please acknowledge NCAR/EOL in your publications with text such as: "Data provided by NCAR/EOL under the sponsorship of the National Science Foundation. https://data.eol.ucar.edu/"

atmosphere: surface↗

Meteorological Variables and Energy Fluxes at the Pumphouse Site, Crested Butte, CO 2017-2019

This data contains output from the pumphouse eddy covariance tower that includes shortwave radiation, longwave radiation, net radiation, air temperature, relative humidity, as well as sensible, latent, and ground heat fluxes. Also included is calculated evapotranspiration from the latent heat flux and the latent heat of vaporization. All data are on a daily timestep and displayed in Mountain Time. The data has been processed, and Quality Assurance / Quality Control (QA/QC) was done, but any daily gaps in the data have not been filled in. This research was funded by the Department of Energy and performed as part of the Watershed Function Scientific Focus Area. This research aimed to constrain evapotranspiration in a high-elevation catchment.The dataset includes one comma-separated values (CSV) data file (EddyCovariance_MeteorlogicalVariables_CrestedButtePumphouse.csv). Additionally, three metadata CSV files are included: (1) location metadata file (locations.csv), which contains location metadata and coordinates; (2) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (3) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Solar Resource Measurements in Eugene, OR: Cooperative Research and Development Final Report, CRADA Number CRD-07-00252

Site-specific, long-term, continuous, and high-resolution measurements of solar irradiance are important for developing renewable resource data. These data are used for several research and development activities consistent with the NLR mission: establish a national 3-year climatological database of measured solar irradiances; provide high quality ground-truth data for satellite remote sensing validation; support development of radiative transfer models for estimating solar irradiance from available meteorological observations; provide solar resource information needed for technology deployment and operations. Data acquired under this agreement will be available to the public through NLR's Measurement & Instrumentation Data Center – MIDC (http://www.nlr.gov/midc) Or the Renewable Resource Data Center - RReDC (http://rredc.nlr.gov). The MIDC offers a variety of standard data display, access, and analysis tools designed to address the needs of a wide user audience (e.g., industry, academia, and government interests).

14 SOLAR ENERGY↗

PIPES (Pipeline for Integrated Projects in Energy Systems) [SWR-24-89]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. https://github.com/nrel-pipes/pipes-api https://github.com/nrel-pipes/pipes-web https://github.com/nrel-pipes/nrel-pipes

Gu, Jianli↗

Pipeline for Integrated Projects in Energy Systems (PIPES): A Tool for Integrated System Planning [Slides]

The Pipeline for Integrated Projects in Energy Systems (PIPES) is a comprehensive project, data, and workflow management tool designed for integrated modeling teams. PIPES facilitates the management of data requirements, tasks, and progress tracking, serving as a higher-level integration layer that works across various data and modeling software. This tool integrates models, data, and tools to perform large-scale, integrated analysis work at scale. PIPES is designed to streamline integrated modeling projects, enhance collaboration, and ensure the quality and efficiency of data management and workflow processes. This presentation introduces PIPES a multi-model tool for integrated system planning; it describes the underlying architecture, deep dives into common user workflows, and outlines the upcoming development roadmap beyond its current alpha state.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗

Best Practices Handbook for the Collection and Use of Solar Resource Data for Solar Energy Applications: Fourth Edition

As the world increasingly seeks low-carbon energy solutions, solar power emerges as the most abundant resource on our planet. However, the challenge of effectively harnessing this energy is crucial in the coming years. Solar energy applications such as photovoltaics, solar heating and cooling, and concentrating solar power use different technologies to capitalize on sunlight. Each system has unique capabilities and requirements, underscoring the need for reliable information about solar resources across diverse installations, from residential rooftops to large-scale power plants. This is especially important for substantial projects, often exceeding $1 billion in construction costs. Before embarking on such ventures, it is imperative to obtain accurate data concerning solar resource quality and reliability at specific sites. Developers require detailed historical information, including seasonal, daily, hourly, and, ideally, subhourly variability to effectively predict a power plant's annual performance. Without these vital data, financial analyses fall short. Moreover, with the growing adoption of distributed photovoltaics, integrating these generation sources becomes critical to maintaining grid reliability and stability. By accurately forecasting generation patterns, utilities and system operators can facilitate greater integration of solar energy, thus ensuring the operational stability of the grid. The complexity and importance of these issues have prompted the foremost experts in the field to collaborate under the auspices of the International Energy Agency's (IEA's) Photovoltaic Power Systems Programme (PVPS) Task 16 to publish this handbook, which summarizes state-of-the-art information about all these topics. The efforts focus on providing reliable data and insights that can help shape our investments in solar energy and drive a sustainable future.

14 SOLAR ENERGY↗

Automatic Generation of Algorithms for High-Speed Reliable Lossy Data Compression (Final Report)

Fast reliable data compression is urgently needed for many leading-edge scientific instruments and for exascale high-performance computing applications because they produce vast amounts of data at extremely high rates. The goal of this project has been to develop a framework named LC that is able to automatically generate high-speed lossless and reliable lossy compression and decompression algorithms that can be customized for different kinds of data. The resulting LC framework is freely available on GitHub. To achieve high-speed operation, LC outputs optimized and parallelized CPU and GPU implementations of the generated algorithms. To ensure the quality of lossily compressed data, LC guarantees the user-provided error bound. To be able to customize the compression algorithm to various use cases, LC can synthesize millions of different algorithms and automatically search for the one that works best for the given data. We have already employed LC to create state-of-the-art lossless and lossy compressors for scientific data as well as leading lossless compressors for images. We hope that LC and the customized, fast, reliable, and CPU/GPU-compatible compression algorithms that it can generate will greatly benefit the many scientific applications that need not only high trustworthiness but also high performance.

97 MATHEMATICS AND COMPUTING↗

Precision Plant Biomass Characterization in Agriculture: Harnessing Machine Learning and Hyperspectral Imaging [Slides]

Efficient Biomass Separation Object detection of anatomical parts (Cob, Stalk, Husk) in IR images enables precise separation, improving preprocessing (e.g., drying, grinding) for biofuel production. Detailed Biomass Characterization with Hyperspectral Data Hyperspectral imaging captures spectral signatures of biomass, allowing for the identification of specific traits like moisture content, lignin levels, and nutrient composition, leading to optimized treatments for each biomass part. Enhanced Feedstock Quality By leveraging hyperspectral data, feedstock can be processed based on its chemical composition, improving conversion efficiency and biofuel yield. Automation for Large-Scale Operations Automated object detection and hyperspectral data analysis reduce manual labor, ensuring accurate sorting and faster processing, making large-scale biofuel production more efficient. Maximized Biomass Utilization Accurate identification of biomass properties minimizes waste and ensures that each part is processed according to its highest biofuel potential.

09 BIOMASS FUELS↗

Comparing Outdoor to Indoor Performance for Bifacial Modules Affected by Polarization-Type Potential-Induced Degradation

Bifacial photovoltaic (PV) modules have the advantage of using light reflected off of the ground to contribute to power production. Predicting the energy gain is challenging and requires complex models to do so accurately. Often, module degradation over time is neglected in models for the sake of simplicity or is underestimated. Comparing outdoor and indoor current–voltage (I–V) performance for bifacial modules is more challenging than for monofacial modules, as there are additional variables to consider such as rear albedo non-uniformity, cell mismatch, and their effects on temperature. This challenge is compounded when heterogeneous degradation modes occur, such as polarization-type potential-induced degradation (PID-p). To examine the effects of PID-p on I–V predictions using an empirical data-driven approach, 16 bifacial PERC modules are installed outdoors on racks with different albedo conditions. A subset is exposed to high-voltage biases of −1500 V or +1500 V. Outdoor data are traced at irradiance ranges of 150–250 W/m 2 , 500–600 W/m 2 , and 900–1000 W/m 2 . These curves are corrected using control module temperature, wire resistivity, and module resistance measured indoors. We examine several methods to transform indoor I–V curves to accurately, and more simply than existing methods, approximate outdoor performance for bifacial modules without and with varying levels of PID-p degradation. This way, bifacial performance modeling can be more accessible and informed by fielded, degraded modules. Distributions of percent errors between indoor and outdoor performance parameters and Mean Absolute Percent Errors (MAPEs) are used to assess method quality. Results including low-irradiance data (150–250 W/m 2 ) are discussed but are filtered for quantifying method quality as these data introduce substantial errors. The method with the most optimal tradeoff between low MAPE and analysis simplicity involves measuring the front side of a module indoors at an irradiance equal to plane-of-array irradiance plus the product of module bifaciality and albedo irradiance. This method gives MAPE values of 1–6.5% for non-degraded and 1.6–5.9% for PID-p degraded module performance.

14 SOLAR ENERGY↗

U-Pu and Ba-Cs isotopic measurements on Trinitite by laser ablation sampling on the Neoma MC-ICP-MS

In this study we present the results of combined U-Pu and Ba-Cs isotope measurements obtained by laser ablation (LA) sampling of two glassy debris fragments (‘Trinitite’) from the world's first atomic bomb detonation conducted in New Mexico on July 16, 1945. Our primary goal in conducting these measurements was to understand whether examination of the U-Pu and Ba-Cs systematics by direct sampling (e.g. without any chemical separation or purification prior to isotope ratio measurement) could yield meaningful information that would differentiate the Trinitite fragments from glassy material lacking a nuclear fission signature. These measurements were conducted on a ThermoFisher Scientific Neoma multi collector – inductively coupled plasma – mass spectrometer (MC-ICP-MS), which is a relatively new MC-ICP-MS platform, so we also examine the behavior of these isotope systems in standards sampled in solution and via LA. Unsurprisingly, the measurements made on purified solutions of the U, Pu, and Ba isotopic standards produce high precision isotope ratios. Furthermore, this extends to the U-Pu measurements made by LA sampling, with the expected degradation in precision and accuracy related to matrix effects and signal intensity fluctuation. However, the Ba-Cs data acquired by LA is of low precision across all of the matrices examined and bears evidence of complex mass fractionation that will require further investigation to resolve. In total, our results indicate that the observed U-Pu isotope data are of sufficient quality to accurately constrain the U and Pu isotopic composition of glass containing sub-ppm levels of these elements which in turn could be used to differentiate glass containing anthropogenic fission products from natural glass whereas the Ba-Cs LA data cannot be used for this purpose until further methodological refinement is performed.

Ba-Cs↗