Search NASA⌕ Search

SEARCH · Search NASA

Results for “data quality analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Physics-informed graph neural networks for predicting cetane number with systematic data quality analysis

Designing alternative fuels for advanced compression ignition engines necessitates a predictive model for cetane number (CN). In this study, the physics-informed graph neural networks are introduced for a reliable CN prediction by considering molecular features pertinent to the physical properties of molecules that affect CN. The reliability of measured data is another key factor to consider for improving the predictive model. Various experimental instruments for measuring CN exist, including standard and non-standard methods. In this regard, a systematic data quality analysis was carried out for the total 630 CNs collected from literature and new measurements in this study using Advanced Fuel Ignition Delay Analyzer (AFIDA). The results from this data curation process were reflected in the model by imposing lower sample weights on the data coming from less reliable measurement techniques. This approach effectively maximized the prediction accuracy while incorporating data from all available sources. Using the sample weights decreased the mean absolute error (MAE) up to 0.8 CN units. The accuracy was also improved by introducing the CN-related physical properties (the number of hydrogen bond donors and acceptors); the test set MAE is 5.74 and 7.01 for the model with and without such properties, respectively. Investigating molecular structural effects on CN was also carried out to gain chemical insights into factors used to design new fuel candidates. The dimensionality reduction analysis of feature vectors showed a clear clustering in terms of functional groups and CN and the structural effect derived from the model was consistent with the physicochemical insights. Finally, this physics-informed model and data curation would be helpful for accurate CN prediction and inform rational fuel design.

97 MATHEMATICS AND COMPUTING↗

Optical Particle Measurements during EPCAPE Field Campaign Report

This campaign requested the deployment of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) User Facility optical particle counter (OPC) at the first ARM Mobile Facility (AMF1) located at the Scripps Pier in La Jolla, California during the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE). The addition of the OPC was requested for two reasons. (1) Close the gap between the scanning mobility particle sizer (SMPS) and aerodynamic particle sizer (APS) size distribution from the Aerosol Observing System (AOS) measurements. (2) Principal investigator Petters has been working with Tracking Aerosol Convection Interaction Experiment (TRACER) data to compute particle fluxes from Doppler lidar (Petters et al. 2024). Briefly, backscatter flux is obtained using the eddy covariance technique using the Doppler vertical velocity and attenuated backscatter. Building upon prior studies, we were able to relate backscatter to particle number concentration by calibrating the lidar retrievals against optical particle counter-measured ground-based aerosol size distribution and radiosonde-interpolated relative humidity at lidar sample height. Performing similar analysis was of interest to EPCAPE to better understand the emissions and vertical transport of large particles into the overlying stratus clouds. However, as stated above, this analysis requires an optical size distribution that covers the 0.3-30-μm-diameter size range. The OPC was deployed between 2023-04-14 and 2024-02-14. The deployment, data quality analysis, and data archiving was handled by the DOE ARM instrument mentor team without additional involvement by the principal investigator. Data quality was marked as “routine” for the majority of the campaign.

54 ENVIRONMENTAL SCIENCES↗

Using Synchrophasor Status Word as Data Quality Indicator: What to Expect in the Field?

Data quality plays a crucial role in successful applications of synchrophasor data in power system operation and control. This paper presents the results of a data quality analysis of a multi-year field-recorded synchrophasor dataset. The analysis has identified several typical data quality issues encountered in the field data. An examination of the PMU status words included with the dataset has revealed several inconsistent implementations and the lack of correlation between the PMU data quality and the status word, which impacts the usefulness of such information. Our investigation has concluded that the status word alone as found in the recorded field dataset could not be used as a reliable indicator of data quality for field-recorded data. Several recommendations are proposed to improve the usefulness of the PMU status word.

Cheng, Zheyuan↗

The Marine and Hydrokinetic Toolkit (Mhkit) for Data Quality Control and Analysis

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

The Marine and Hydrokinetic ToolKit for Data Quality Control and Analysis: Preprint

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

The high level trigger and express data production at STAR

To meet the demands of the Beam Energy Scan phase-II (BES-II) program, the STAR experiment at the Relativistic Heavy Ion Collider (RHIC) developed a dual real-time framework consisting of a High Level Trigger (HLT) and an Express Data Production system (xProduction). The HLT operates online within the Data Acquisition (DAQ) chain on a dedicated multi-core CPU cluster with the option to offload compute-intensive kernels to Xeon Phi coprocessors. It uses parallelized algorithms, such as the Cellular Automaton (CA) Track Finder, to perform rapid tracking, vertexing, and event filtering. This allows it to select events of interest in real time and provide immediate feedback on detector and beam conditions. In contrast, the xProduction workflow runs concurrently and independently of the DAQ loop. It applies near offline-quality calibration and reconstruction within hours of data collection. The xProduction input is the express data stream, whose content can be enriched by HLT trigger/priority selections under DAQ/HLT resource constraints, and it uses the STAR calibration/conditions framework, incorporating online calibration/QA information when available. This enables early preliminary physics analysis, including the reconstruction of rare signals, such as hyperons and hypernuclei. It also provides collaboration-wide access to analysis-ready datasets. Together, the HLT and xProduction systems form a complementary architecture: the HLT performs online event selection while the xProduction chain delivers high-quality results within a short amount of time. This integrated framework has enabled the prompt reconstruction of the $^5_Λ$ He hypernucleus with high statistical significance and the efficient processing of hundreds of millions of heavy-ion collision events. In conclusion, its demonstrated scalability and robustness establish a model for future high-luminosity experiments requiring both online event filtering and rapid access to analysis-quality data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Becoming a 10: A Closer Look at the U.S. Department of Energy Home Energy Score's Updates, Improvements, and Expansion

The U.S. Department of Energy (DOE)'s Home Energy Score provides homeowners, buyers, and renters directly comparable and credible information about a home's estimated energy use and costs. Certified Qualified Assessors conduct low-cost assessments to provide each home a 1-10 score alongside a set of cost-effective upgrades to improve the score. As of February 2022, hundreds of assessors have delivered over 175,000 scores to homes across the country. Originally released in 2012 using DOE2.1e as the modeling backend, after years of effort, a new version of Home Energy Score was released in 2021 utilizing DOE's flagship energy modeling software, EnergyPlus. The updated architecture leverages modeling advancements and enables new building technologies to be added to the Scoring Tool. The new release represents a leap forward in harmonizing modeling assumptions across DOE and industry programs. In this paper we discuss the rigorous approach to model comparison with DOE2 undertaken prior to the update, utilizing test homes and real homes from the Home Energy Score database to strike a balance between consistency and more accurate energy predictions. We also discuss additional new capabilities, including improvements made to the upgrade recommendations methodology, the inclusion of an energy cost estimate metric based on ResStock analysis for use in home appraisals, and improved data analysis for quality assurance. Finally, we look at the impact Home Energy Score has had over the last decade and its future potential as its uptake in state energy plans, local ordinances, utility programs, and real estate data continues to grow.

building energy modeling↗

Adaptable SEC‐SAXS data collection for higher quality structure analysis in solution

Abstract The two major challenges in synchrotron size‐exclusion chromatography coupled in‐line with small‐angle x‐ray scattering (SEC‐SAXS) experiments are the overlapping peaks in the elution profile and the fouling of radiation‐damaged materials on the walls of the sample cell. In recent years, many post‐experimental analyses techniques have been developed and applied to extract scattering profiles from these problematic SEC‐SAXS data. Here, we present three modes of data collection at the BioSAXS Beamline 4–2 of the Stanford Synchrotron Radiation Lightsource (SSRL BL4‐2). The first mode, the High‐Resolution mode, enables SEC‐SAXS data collection with excellent sample separation and virtually no additional peak broadening from the UHPLC UV detector to the x‐ray position by taking advantage of the low system dispersion of the UHPLC. The small bed volume of the analytical SEC column minimizes sample dilution in the column and facilitates data collection at higher sample concentrations with excellent sample economy equal to or even less than that of the conventional equilibrium SAXS method. Radiation damage problems during SEC‐SAXS data collection are evaded by additional cleaning of the sample cell after buffer data collection and avoidance of unnecessary exposures through the use of the x‐ray shutter control options, allowing sample data collection with a clean sample cell. Therefore, accurate background subtraction can be performed at a level equivalent to the conventional equilibrium SAXS method without requiring baseline correction, thereby leading to more reliable downstream structural analysis and quicker access to new science. The two other data collection modes, the High‐Throughput mode and the Co‐Flow mode, add agility to the planning and execution of experiments to efficiently achieve the user's scientific objectives at the SSRL BL4‐2.

Matsui, Tsutomu↗

A Comprehensive Analysis of Real-World Accelerometer Data Quality in a Global Smartphone-based Seismic Network

The proliferation of low-cost sensors in smartphones has facilitated numerous applications; however, large-scale deployments often encounter performance issues. Sensing heterogeneity, which refers to varying data quality due to factors such as device differences and user behaviors, presents a significant challenge. In this research, we perform an extensive analysis of 3-axis accelerometer data from the MyShake system, a global seismic network utilizing smartphones. We systematically evaluate the quality of approximately 22 million 3-axis acceleration waveforms from over 81 thousand smartphone devices worldwide, using metrics that represent sampling rate and noise level. We explore a broad range of factors influencing accelerometer data quality, including smartphone and accelerometer manufacturers, phone specifications (release year, RAM, battery), geolocation, and time. Our findings indicate that multiple factors affect data quality, with accelerometer model and smartphone specifications being the most critical. In addition, we examine the influence of data quality on earthquake parameter estimation and show that removing low-quality accelerometer data enhances the accuracy of earthquake magnitude estimation.

58 GEOSCIENCES↗

Editorial: Resolving atmospheric flow in complex environments: recent experiments in terrain and forest canopies

The characterization of atmospheric flows in complex environments, which may include steep terrain slopes and heterogeneous vegetation and/or forest cover, is a long-standing challenge in boundary-layer meteorology. Atmospheric observations are complicated by the presence of transient, terrain-induced flow features, forest-canopy-atmosphere interactions, and atmospheric stability effects, not to mention the logistical hurdles involved with instrument deployment, data analysis, and quality control. Furthermore, challenges in atmospheric modeling arise due to numerical errors associated with complex terrain flows, as well as reliance on simplified parameterizations for unresolved processes such as turbulent mixing and land-surface or forest-canopy-atmosphere interactions. These modeling challenges are exacerbated in the so-called “gray zone,” wherein features of interest have length scales that are similar to the model grid spacing, or when the principal flow layer is smaller than the grid spacing (e.g., slope flows).

54 ENVIRONMENTAL SCIENCES↗

Evaluation of WRF-Solar Cloud Forecast Using the NSRDB: Preprint

Cloud forecast is a crucial component in predicting solar irradiance from numerical weather prediction (NWP) models. Assessing cloud properties from the NWP models requires significant work due to the need for high-quality data, spatial analysis covering model extent, and detailed analysis of model performance for different types of clouds. This study presents an evaluation of the WRF-Solar cloud forecast using the National Solar Radiation Database (NSRDB). We propose an evaluation framework applied to a single model prediction as well as ensemble-based forecasts. Various cloud detection metrics are calculated when comparing with the satellite-derived dataset. The mismatched clouds from the WRF-Solar model are quantified using nine cloud types classified by cloud top height and cloud optical depth. The results based on the WRF-Solar forecasts covering the entire U.S. for the full year of 2018 shows mismatched cloud frequency in the range of 8% - 46% for thick and high-level (deep convective) to thin and low-level (cumulus) clouds.

cloud mask forecast↗

Stakeholder analysis for designing an urban air quality data governance ecosystem in smart cities

Cities, the world over, are fuelling economic growth. At the same time, rapid urbanization is a root cause of serious environmental damage. Recent WHO global air pollution guidelines highlight air pollution as a critical environmental threat along with climate change. To address these threats, smart cities and clean air programs are on a rise. In smart cities, data and Information and Communication Technologies (ICT) are major drivers of city transformations. The 4th Industrial Revolution (4IR) technologies such as the Internet of Things (IoT), big data, artificial intelligence (AI), and cloud computing have the potential to accelerate these transformations toward urban resilience. However, the success of smart cities and clean air programs depends on cohesive multi-sector stakeholder contributions. This study conducted interdisciplinary participative stakeholder analysis to understand the data, and sectorial challenges, to outline the technological opportunities to facilitate clean air programs in Indian smart cities. The research highlights gaps due to siloed stakeholder operations, lack of data calibration, non-alignment of smart city and air quality management services, non-availability of health exposure data, and difficulty in translating scientific data into implementable actions. Stakeholders expressed potential ‘fit for the purpose’ use of IoT devices, satellites, smartphones, and mobility data augmented by AI methods in bridging these gaps. In conclusion, the analysis points toward a need to develop an easily accessible and ubiquitous urban data governance ecosystem enabling seamless cross-sector data exchanges to build trusting relationships among the stakeholders across the air quality management value chain.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data Fusion to Enhance Quality Control and Analysis with Instruments at the Marine and Coastal Research Laboratory

Deploying environmental monitoring instruments in the marine environment can be challenging, facing challenges around device survivability, biofouling and corrosion, and consistent data collection. This project explores the use of data fusion – the process of integrating multiple data sources to produce more consistent, accurate, and useful information – to build a consistent long-term monitoring system at the Marine and Coastal Research Laboratory (MCRL) in Sequim, Washington. Unused instruments that had been acquired from past projects were inventoried and deployments planned on the MCRL pier and floating dock. A total of 8 instruments were deployed including a tide gauge, hydrophone, acoustic Doppler current profiler (ADCP), photosynthetically active radiation (PAR) sensors, meteorological station, and three water quality sensors. Deployments were planned to be well-protected around the pier structure and a maintenance schedule was created for cleaning and recalibration. An automated data pipeline was created to aggregate data on edge computers that push data to Amazon Web Services (AWS) cloud storage every 15 minutes, performing automated quality control and data transformations using the Time Series Data Analytical Toolkit (TSDAT). Continued efforts are underway to maintain this system into the future, take a data-driven approach to maintenance scheduling, improve the reliability of the system, and share the data with a variety of end-users.

54 ENVIRONMENTAL SCIENCES↗

Los Alamos National Laboratory 2022 Annual Site Environmental Report (Rev. 2)

Los Alamos National Laboratory (Laboratory) annual site environmental reports are prepared each year by the Laboratory’s environmental organizations as required by U.S. Department of Energy Order 231.1B, Administrative Change 1, Environment, Safety, and Health Reporting, and Order 458.1, Administrative Change 4, Radiation Protection of the Public and the Environment. The chapters in this report discuss our compliance with environmental laws, regulations, and orders (Chapter 2, Compliance Summary); how we manage the Laboratory’s environmental performance and assure the quality of data from analysis of environmental samples (Chapter 3, Environmental Programs and Analytical Data Quality); how we monitor for air emissions of radioactive materials and for weather conditions (Chapter 4, Air Quality); how we monitor for effects of Laboratory operations on groundwater quality (Chapter 5, Groundwater Protection); how we monitor the levels of chemicals and radionuclides in storm water runoff and sediment (Chapter 6, Watershed Quality); how we monitor for the levels and effects of chemicals and radionuclides in plants, animals, soil, and vegetation (Chapter 7, Ecosystem Health); and finally, what radioactive dose or risk from chemical exposure that members of the public could experience as a result of Laboratory operations (Chapter 8, Public Dose and Risk Assessment).

54 ENVIRONMENTAL SCIENCES↗

Quality Analysis of Baseline Time-Lapse CSEM Data CarbonSAFE Project, North Dakota

Conference presentation at International Meeting for Applied Geoscience & Energy (IMAGE), Houston, TX, August 28 – September 2, 2022. When a series of time-lapse CSEM surveys are designed to measure the often-small variations in signal seen in CCUS projects, the primary concern must be collecting the highest-quality data and reducing as much noise as possible. This includes optimization during planning and feasibility, careful and consistent quality checks in the field, and transparent and repeatable postprocessing steps. We apply this practice to a baseline (prior to CO 2 injection) CSEM data set collected for a time-lapse survey in Center, North Dakota, as a part of the North Dakota CarbonSAFE project, and describe the rigorous quality control and assessment methods used, including noise removal and data validation with 1D and 3D inversion to tie results to borehole logs. The final result is an accurate and representative CSEM data set and information that stakeholders can use to inform future time-lapse survey costs and designs.

20 FOSSIL-FUELED POWER PLANTS↗