Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Indoor Air Quality (IAQ) Monitoring for Space Farming Institute [Slides]

Through the U.S. Department of Energy's Energy to Communities (E2C) program, NREL, other national laboratory experts, and select organizations provide Expert Match - free, short-term technical assistance to address near-term energy challenges and questions. Expert Match is for community stakeholders who have decision-making power or influence in their community but need access to additional energy expertise to inform key upcoming decisions. This Expert Match request supported the Space Farming Institute, a nonprofit organization located in Anchorage, AK, with an indoor air quality analysis. The NREL team analyzed indoor air quality data provided by the Space Farming Institute, which experts at PNNL used to design an indoor bioreactor to grow Ulva algae for indoor air quality mitigation purposes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

On High-Temperature Dynamometer Test Stand Development

RePED is a novel prototype alternator that has been designed and built to operate at 35 kW under 250 degrees C ambient temperature conditions. The operating efficiency will be measured using a 225-kW motor to drive the machine in a back-to-back configuration. The goal is to determine if the RePED alternator can deliver 35 kW at a rotational speed of 1,000 rpm while operating in 250 degrees C ambient temperature and have a minimum efficiency of 60%. Dynamometer testing of a high-temperature alternator for geothermal drilling presents several unique challenges, especially around maintaining and managing the various sources and types of energy passing through the system without impacting data quality. The objectives of this report are to (1) describe the test stand in detail, the relative orientation of each major component, and the data acquisition system and (2) discuss the behavior of the test stand components such that this report can serve as a reference document for future projects.

15 GEOTHERMAL ENERGY↗

Quality-Controlled Meteorological Data from the Flood Control District of Maricopa County (FCDMC) Network, Phoenix, Arizona (1987-2024)

This dataset contains 15- or 30-minute interval meteorological data from the Flood Control District of Maricopa County (FCDMC), Arizona, USA, covering eight key variables across multiple sensor stations between 1987 and 2024. Each variable is stored as a separate CSV file, containing time-series data that have undergone rigorous quality control (QC) procedures and, where appropriate, short-gap interpolation for consistency. The quality control (QC) pipeline consisted of four sequential tests: (1) a range test to ensure all values fall within physically realistic limits, (2) a step test to identify abrupt and implausible changes between consecutive records, (3) a proximity test that validates flagged values from step test using data from nearby stations and exceedance probability thresholds, and (4) a persistence test to detect and remove periods of unrealistically constant readings. These thresholds were calibrated to Arizona’s environmental conditions and sensor specifications. After QC, short gaps (≤2 hours) were linearly interpolated to ensure consistent temporal resolution, except for wind variables. Due to a major upgrade in FCDMC’s data transmission system, only ALERT-2 protocol data (2016–2024) for wind variables are included; earlier ALERT-1 data were excluded because of irregular sampling and high missing rates. This dataset supports regional climate and infrastructure resilience studies by providing standardized, high-resolution meteorological data for the greater Phoenix metropolitan area.

54 ENVIRONMENTAL SCIENCES↗

Accelerating discoveries at DIII-D with the Integrated Research Infrastructure

DIII-D research is being accelerated by leveraging high performance computing (HPC) and data resources available through the National Energy Research Scientific Computing Center (NERSC) Superfacility initiative. As part of this initiative, a high-resolution, fully automated, whole discharge kinetic equilibrium reconstruction workflow was developed that runs at the NERSC for most DIII-D shots in under 20 min. This has eliminated a long-standing research barrier and opened the door to more sophisticated analyses, including plasma transport and stability. These capabilities would benefit from being automated and executed within the larger Department of Energy Advanced Scientific Computing Research program’s Integrated Research Infrastructure (IRI) framework. The goal of IRI is to empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. For transport, we are looking at producing flux matched profiles and also using particle tracing to predict fast ion heat deposition from neutral beam injection before a shot takes place. Our starting point for evaluating plasma stability focuses on the pedestal limits that must be navigated to achieve better confinement. This information is meant to help operators run more effective experiments, so it needs to be available rapidly inside the DIII-D control room. So far this has been achieved by ensuring the data is available with existing tools, but as more novel results are produced new visualization tools must be developed. In addition, all of the high-quality data we have generated has been collected into databases that can unlock even deeper insights. This has already been leveraged for model and code validation studies as well as for developing AI/ML surrogates. The workflows developed for this project are intended to serve as prototypes that can be replicated on other experiments and can be run to provide timely and essential information for ITER, as well as next stage fusion power plants.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Ten questions concerning low-cost indoor air quality sensors: Perspectives from research and practice

Low-cost indoor air quality (IAQ) sensors are increasingly being used in homes and commercial and public buildings, driven by growing concerns about the impact of air on health, cognitive performance, and occupant wellbeing. These sensors offer a potentially transformative opportunity to increase spatial and temporal coverage of IAQ monitoring at a fraction of the cost of conventional reference instruments. However, their widespread use raises questions around accuracy, calibration, placement, data handling and interpretation, and integration into existing standards and workflows. This paper presents ten critical questions concerning the use of low-cost IAQ sensors in buildings, drawing on the latest empirical research, field deployments, and emerging practice. It discusses potential frameworks for deployment and evaluation, examines current sensor capabilities for measuring common pollutants, identifies methodological gaps in validation and uncertainty quantification, and outlines the extent to which existing IAQ standards can accommodate sensor-based evidence. The paper also explores how monitoring needs and deployment models vary by building type, the potential of real-time IAQ data to support building operations, and the ethical and legal implications of widespread sensor use. While significant challenges remain in ensuring data quality and building stakeholder trust, new applications are emerging through open data initiatives and advances in analytics and visualization. As the technology, science, and standards co-evolve, low-cost IAQ sensors are poised to become integral to routine building operation, building science, and environmental health research.

Parkinson, Thomas↗

ECLEIRS: Exact conservation law embedded identification of reduced states for parameterized nonlinear conservation laws from sparse and noisy data

Multi-query applications such as parameter estimation, uncertainty quantification and design optimization for parameterized partial differential equation (PDE) systems are expensive. While reduced/latent state dynamics approaches for parameterized PDEs offer a viable alternative, these approaches rely on high-quality data and struggle with highly sparse spatiotemporal noisy measurements typically obtained from experiments. Furthermore, there is no guarantee that these models satisfy governing physical conservation laws. In this article, we propose a reduced state dynamics approach, referred to as ECLEIRS, that embeds exact conservation in the solution and flux representation by utilizing a space-time divergence-free neural network formulation. We compare ECLEIRS with other reduced state dynamics approaches, those that do not enforce any physical constraints and those with physics-informed loss functions, for three shock-propagation problems: 1-D advection, 1-D Burgers and 2-D Euler equations. In conclusion, the numerical experiments conducted in this study demonstrate that ECLEIRS provides the most accurate prediction of dynamics for unseen parameters even in the presence of highly sparse and noisy data.

97 MATHEMATICS AND COMPUTING↗

Regulatory Testing of WTP HL W Glasses for Compliance with Delisting Requirements, VSL-03R3780-1, Rev. 1

(Part of the data collected for this work and discussed in this report was subject to data quality and bias issues. All the affected tests have subsequently been repeated, as directed by the Waste Treatment Plant Project, and new statistical analyses have been performed on the revised data set. A subsequently-issued report supercedes this report and describes the revised data, together with the revised composition-property models (Kot et al. 2004). The reader should refer to the new report for discussion of the revised data set and composition-property (TCLP cadmium release) models. None of the data for the spike glasses, designed for Case 1 and Case 2 COPCs testing, were affected by this issue.)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Data Sharing as a Catalyst for Expanding the Energy Frontier

As the energy landscape evolves to include technologies such as geothermal energy, comprehensive data become essential for driving innovation and scalability, particularly with the growing use of tools like machine learning and artificial intelligence. In emerging sectors, the cost of gathering high-quality data across large spatial areas can present a significant barrier. A key solution is leveraging existing data from well-established industries like oil and gas. However, the proprietary nature of data in these industries often hinders collaboration. This paper explores how cultivating a culture of data sharing can act as a catalyst for progress, fueling breakthroughs across both conventional and renewable energy sectors. Practical compromises that protect business interests while enabling data access are proposed, and real-world success stories are highlighted, demonstrating how collaboration has accelerated advancements in geothermal, carbon capture, and other innovative technologies.

15 GEOTHERMAL ENERGY↗

Detecting outbreaks using a spatial latent field

In this paper, we present a method for estimating the infection-rate of a disease as a spatial-temporal field. Our data comprises time-series case-counts of symptomatic patients in various areal units of a region. We extend an epidemiological model, originally designed for a single areal unit, to accommodate multiple units. The field estimation is framed within a Bayesian context, utilizing a parameterized Gaussian random field as a spatial prior. We apply an adaptive Markov chain Monte Carlo method to sample the posterior distribution of the model parameters condition on COVID-19 case-count data from three adjacent counties in New Mexico, USA. Our results suggest that the correlation between epidemiological dynamics in neighboring regions helps regularize estimations in areas with high variance (i.e., poor quality) data. Using the calibrated epidemic model, we forecast the infection-rate over each areal unit and develop a simple anomaly detector to signal new epidemic waves. Our findings show that anomaly detector based on estimated infection-rates outperforms a conventional algorithm that relies solely on case-counts.

Safta, Cosmin [Sandia National Laboratories (SNL-C↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

The state of the art for neutron irradiation experiments from the perspective of the High Flux Isotope Reactor (HFIR)

Irradiation experiment campaigns are critical to advancing nuclear energy technologies by providing data on material performance under relevant radiation conditions. Successful irradiation experiments require integrated design efforts that balance technical goals with facility constraints. Here, this paper presents an expert-informed overview of irradiation experiment design at the High Flux Isotope Reactor. It addresses the nuclear materials research and irradiation experiment communities to guide them toward developing technically sound, facility-compatible campaigns. The High Flux Isotope Reactor is a multipurpose reactor supporting isotope production, neutron scattering, and materials testing. Its high, steady-state neutron flux is ideal for irradiation experiments, but successful execution demands coordinated thermal, structural, and reactor physics analyses. The paper outlines the complete development workflow from concept definition and design optimization to safety qualification and post-irradiation examination. Standardized capsule platforms are also discussed in terms of flexibility, specimen capacity, and thermal performance. Common failure modes such as unanticipated geometric variations, can impact temperature-dose profiles and compromise data reliability. Therefore, detailed thermal modeling and accurate as-built characterization are essential for meaningful post-irradiation data interpretation. Key recommendations include early engagement all stakeholders, clearly defined design expectations, and alignment of specimen geometries with post-irradiation examination capabilities. This approach reduces design iterations, enhances data quality, and supports more efficient use of irradiation resources. Strategic and well-planned irradiation testing not only improves individual campaign success but also accelerates the deployment of advanced nuclear technologies. By closing critical data gaps and reducing development risks, the nuclear materials community can more effectively contribute to the future of clean, resilient energy systems.

Experiments↗

guppy i : a code for reducing the storage requirements of cosmological simulations

ABSTRACT As cosmological simulations have grown in size, the permanent storage requirements of their particle data have also grown. Even modest simulations present a major logistical challenge for the groups which run these boxes and researchers without access to high performance computing facilities often need to restrict their analysis to lower quality data. In this paper, we present guppy, a compression algorithm and code base tailored to reduce the sizes of dark matter-only cosmological simulations by approximately an order of magnitude. guppy is a ‘lossy’ algorithm, meaning that it injects a small amount of controlled and uncorrelated noise into particle properties. We perform extensive tests on the impact that this noise has on the internal structure of dark matter haloes, and identify conservative accuracy limits which ensure that compression has no practical impact on single-snapshot halo properties, profiles, and abundances. We also release functional prototype libraries in C, Python, and Go for reading and creating guppy data.

79 ASTRONOMY AND ASTROPHYSICS↗

The Linear Point Standard Ruler with DESI DR1 and DR2 Data

The linear point, a purely geometric feature in the monopole of the two-point correlation function, has been proposed as an alternative standard ruler. Compared to the peak in the correlation function, it is more robust to late-time nonlinear effects at the percent level. In light of improved simulations and high quality data, we revisit the robustness of the linear point and use it as an alternative to template-based fitting approaches typically used in BAO analyses. We present the linear point measurements on galaxy samples from the first and second data releases (DR1 and DR2) of the DESI survey. We convert the linear point into a dimensionless parameter $α_{iso,LP}$, defined as the ratio of the linear point in the fiducial cosmology and the observed value, analogous to the isotropic BAO scaling parameter $α_{iso}$ used in previous BAO measurements. Using the 2nd generation of AbacusSummit mock catalogs, we find that linear point measurements are more precise when calculated in the post-reconstruction regime with 15-60% smaller uncertainties than those pre-reconstruction. We find a systematic shift in the linear point measurements compared against the isotropic BAO measurements in mocks; we attribute this to the isotropic damping parameter responsible for smearing the linear point in the nonlinear regime. We propose a sample-dependent correction that mitigates the impact of late-time nonlinear effects. While this introduces a cosmology dependence in an otherwise model-independent measurement, this is necessary given the sub-percent precision dictated by current cosmological surveys. Comparing $α_{iso,LP}$ with isotropic BAO measurements made on the DESI DR1 and DR2 galaxy samples, we find excellent agreement after applying this correction, particularly post-reconstruction. We discuss future scope regarding cosmological inference with linear point measurements.

Uberoi, N. [Yale U.] (ORCID:0000000275179629)↗

Synthesis of ARM User Facility Surface Rainfall Datasets to Construct a Best Estimate Value Added Product (PrecipBE)

Surface precipitation measurements are essential for Earth system model (ESM) evaluation and understanding cloud processes. An ever-growing need for robust, temporally evolving, and easy-to-use statistical datasets provides motivation for a baseline ground-based precipitation properties data product. The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility operates an extensive suite of precipitation instruments with various sensitivities and operating mechanisms, which render the decision of which instrument to use based on one or more fixed thresholds challenging and prone to errors and bias. Using a long-term instrument inter-comparison from a unique per-precipitation event perspective, rather than instantaneous sample comparison, we demonstrate that ARM rainfall-measuring instruments are generally consistent with each other at the statistical level. Inter-instrument deviations at the single event level can be large, especially for specific rainfall event properties such as maximum precipitation rates. A machine-learning (ML) analysis using a random forest regressor indicates that in some cases, depending on instrument, local site climatology, and/or specific deployment configuration, certain atmospheric state variables influence the measured quantities in an unpredictable manner. Thus, a-priori weighting of different instruments does not necessarily lead to more accurate and less biased synthesis of instrument data. These results motivate the design of the ARM precipitation best-estimate (PrecipBE) value-added product, which incorporates all valid precipitation data while considering data quality and other instrument limitations. PrecipBE consists of time series and tabular statistics datasets in an easy-to-use and insightful per-precipitation event format. It provides a large set of precipitation event properties supplemented with ancillary data from ARM datasets that correspond to the detected precipitation events. We describe the PrecipBE algorithm and demonstrate its use via the examination of a single-day output as well as a long-term trend analysis of precipitation events at the ARM Southern Great Plains (SGP) site, covering more than 30 years of data. The trend analysis tentatively suggests a long-term temporal tendency for mainly shorter and less intense precipitation events at the SGP site, but a long-term increase in annual rainfall by more than 36 mm (5 %) per decade. This rainfall trend is catalyzed primarily by more extreme event properties of relatively rare, intense precipitation events, with event total and 1 min maximum precipitation rate at a 1 year timeframe increasing up to 5 mm and 9 mm h −1 (several percent) per decade, respectively. While the currently available PrecipBE datasets (at https://adc.arm.gov/discovery/, last access: 8 December 2025) cover rainfall from multiple ARM deployments up to March 2025, PrecipBE is planned to be expanded to include solid-phase precipitation and will soon become an operational product with a several-day lag from real-time. We invite the ARM user community to leverage this new product and welcome user feedback to further enhance the dataset.

Silber, Israel [Pacific Northwest National Laborat↗

Factorization Machine‐Based Active Learning for Functional Materials Design with Optimal Initial Data

The optimization of functional materials is important to enhance their properties, but their complex geometries pose great challenges to optimization. Data-driven algorithms efficiently navigate such complex design spaces by learning relationships between material structures and performance metrics to discover high-performance functional materials. Surrogate-based active learning, continually improving its surrogate model by iteratively including high-quality data points, has emerged as a cost-effective data-driven approach. Furthermore, it can be coupled with quantum computing to enhance optimization processes, especially when paired with a special form of surrogate model (i.e., quadratic unconstrained binary optimization), formulated by factorization machine (FM). However, current practices often overlook the variability in design space sizes when determining the initial data size for optimization. In this work, we investigate the optimal initial data sizes required for efficient convergence across various design space sizes. By employing averaged piecewise linear regression, we identify initiation points where convergence begins, highlighting the crucial role of employing adequate initial data in achieving efficient optimization. These results contribute to the efficient optimization of functional materials by ensuring faster convergence and reducing computational costs in FM-based active learning.

active learning↗