Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

OC6 Phase Ia - Nonlinear hydrodynamic loading validation dataset

Two validation campaigns were examined within the Offshore Code Comparison Collaboration, Continued, with Correlation and unCertainty (OC6) Phase 1 project to examine the modeling tools' underprediction of loads and motion of a floating wind semisubmersible (semi) at their surge and pitch natural frequencies. These campaigns were performed at the Maritime Research Institute Netherlands (MARIN) in 2017 and 2018. The load cases (LC) considered include: LC1 – Load measurements across semi under current loading; LC2 - Load measurements across semi under forced surge oscillation; LC3 – Load measurements across semi under wave loading, while held fixed; LC4 – Free-decay motion measurements in surge, pitch, and heave; and LC5 – Motion measurements under wave loading. Details on the results from the OC6 Phase Ia project can be found in the reference, “OC6 Phase 1: Investigating the underprediction of low-frequency hydrodynamic loads and responses of floating wind turbines”, J Phys: Conf Series 1618 032033.

17 WIND ENERGY↗

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES↗

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab↗

Public water supply infrastructure extensification and diversification in surface waters is insufficient to meet future demands in Texas

The data were developed to evaluate the capacity of existing and potential new surface water supply infrastructure to meet projected public water demands across districts in Texas under multiple future socioeconomic and climate scenarios. The database integrates hydrologic, water quality, infrastructure, energy, cost, demographic, and demand-projection information for candidate surface water supply locations. Candidate sites include stream reaches, waterbodies, reservoir surplus locations, and potential new reservoir sites. Water availability is characterized using historical and projected flow conditions, while site suitability is evaluated using five indicators: Water Availability Index (WAI), Water Quality Index (WQI), Energy Requirement Index (ERI), Water Treatment Cost (WTC), and Water Infrastructure Cost (WIC). The datasets include statewide candidate-site information, district-level demand projections under Shared Socioeconomic Pathways (SSPs), runoff-based allocation constraints, climate-stress metrics, and optimization outputs evaluating alternative infrastructure planning strategies. Optimization results compare Business-as-Usual (BAU) and All Surface Water (AllSW) demand-management approaches under both scaled and fixed cost-cap strategies. Associated validation datasets provide district-level feasibility assessments, infrastructure selection outcomes, cost-cap utilization, demand satisfaction metrics, and constraint diagnostics. Additional datasets quantify projected changes in storage and flow conditions as well as water availability stress for both existing and newly selected intake locations under the SSP5 scenario for mid-century and late-century climate conditions. Together, these datasets support assessment of the extent to which surface-water infrastructure expansion and diversification strategies can satisfy future public water demands while accounting for hydrologic, economic, and planning constraints across Texas. Dataset(s) Description Dataset_preoptimization.xlsx Comprehensive pre-optimization dataset containing candidate water-supply sites and associated hydrologic, water-quality, infrastructure, climate, demographic, runoff, and demand-projection variables used as inputs to the optimization analyses. Includes variable descriptions and the full statewide candidate-site database. District_level_site_selection.zip - Compressed archive containing all SSP-specific district-level optimization result files MESIO_ssp1_results.xlsx District-level site selection results for SSP1 (MESIO). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. MESID_ssp2_results.xlsx District-level site selection results for SSP2 (MESID). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. LCMRD_ssp3_results.xlsx District-level site selection results for SSP3 (LCMRD). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. IRDev-Low_ssp4l_results.xlsx District-level site selection results for SSP4-Low (IRDev-Low). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. IRDev-High_ssp4h_results.xlsx District-level site selection results for SSP4-High (IRDev-High). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. RSIM_ssp5_results.xlsx District-level site selection results for SSP5 (RSIM). Includes variable descriptions, BAU and AllSW site-selection results under scaled and fixed cost strategies, and district-level validation diagnostics. tx_hydrological_stress.xlsx Hydrological stress dataset for existing and newly selected intake locations. Includes projected mid-century and late-century changes, gain/loss classifications, planning strategy information, and accompanying variable descriptions. Also includes water-stress metrics derived from historical and projected low-flow conditions.

Okoye, Perpetua I. (ORCID:0000000215545033)↗

Comparison of CNN-Based Image Classification Approaches for Implementation of Low-Cost Multispectral Arcing Detection

Camera-based sensing has benefited in recent years from developments in machine learning data processing methods, as well as improved data collection options such as Unmanned Aerial Vehicles (UAV) mounted sensors. However, cost considerations, both for the initial purchase of sensors as well as updates, maintenance, or potential replacement if damaged, can limit adoption of more expensive sensing options for some applications. To evaluate more affordable options with less expensive, more available, and more easily replaceable hardware, we examine the use of machine learning-based image classification with custom datasets, utilizing deep learning based-image classification and the use of ensemble models for sensor fusion. Utilizing the same models for each camera to reduce technical overhead, we showed that for a very representative training dataset, camera-based detection can be successful for detection of electrical arcing. We also use multiple validation datasets, based on conditions expected to be of varying difficulty, to evaluate custom data. These results show that ensemble models of different data sources can mitigate risks from gaps in training data, though the system will be less redundant for those cases unless other precautions are taken. We found that with good quality custom datasets, data fusion models can be utilized without specialization in design to the specific cameras utilized, allowing for less specialized, more accessible equipment to be utilized as multispectral camera components. This approach can provide an alternative to expensive sensing equipment for applications in which lower-cost or more easily replaceable sensing equipment is desirable.

convolutional neural networks↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗

Event-Based Energy Impact Tracking and Forecasting with Limited Measurements for Rooftop Units

Packaged air conditioning units and heat pumps, also known as rooftop units (RTUs), are responsible for almost 133 billion kWh of electricity usage annually on site for space cooling U.S. commercial buildings. In addition, the use of heat pumps is a trend we expect to accelerate as buildings transition from fossil fuel-based heating to electricity as a key step for decarbonizing the U.S. commercial buildings sector. However, the operation conditions and energy use of RTUs and heat pumps are usually not well monitored as they are not commonly integrated with building automation systems and lack exposed sensing and control points. To fill this gap, this paper proposes a framework for tracking and forecasting energy impacts resulting from degradation of performance and improved performance for unit servicing using limited data. The proposed framework makes use of a constrained dataset, specifically measurements of the outdoor air temperature and the power demand of individual RTUs, to track and forecast changes in energy use associated with changes in performance over various temporal horizons ranging from days to weeks. Following the detection of an RTU fault, performance degradation, or performance improvement, the framework employs a prediction model to assess the cumulative energy impact. We demonstrate the effectiveness of the method with field-collected data for servicing and degradation examples and compare the predicting accuracy of Gradient Boosting Decision Tree (GBDT) Regression models to Support Vector Regression and Linear Regression models. The results show that GBDT achieved the best accuracy for time-series validation datasets for the servicing and degradation cases, and the prediction model was able to track the cumulative energy impacts of events. The proposed framework can inform building owners of the cumulative change in energy usage of RTUs associated with performance degradation, performance improvement, or a fault.

packaged air conditioners, packaged heat pumps, ro↗

Network Anomaly Detection in Distributed Edge Computing Infrastructure

As networks continue to grow in complexity and scale, detecting anomalies has become increasingly challenging, particularly in diverse and geographically dispersed environments. Traditional approaches often struggle with managing the computational burden associated with analyzing large-scale network traffic to identify anomalies. This paper introduces a distributed edge computing framework that integrates federated learning with Apache Spark and Kubernetes to address these challenges. We hypothesize that our approach, which enables collaborative model training across distributed nodes, significantly enhances the detection accuracy of network anomalies across different network types. We show that by leveraging distributed computing and containerization technologies, our framework not only improves scalability and fault tolerance but also achieves superior detection performance compared to state-of-the-art methods. Extensive experiments on the UNSW-NB15 and ROAD datasets validate the effectiveness of our approach, demonstrating statistically significant improvements in detection accuracy and training efficiency over baseline models, as confirmed by MannWhitney U and Kolmogorov-Smirnov tests (p<0.05).

Marfo, William [University of Texas at El Paso,Dep↗

Entropy-Infused Deep Learning Loss Function for Capturing Extreme Values in Wind Power Forecasting

Extreme scenarios in wind power generation occur with higher frequency and larger magnitude in the recent years due to the ever-increasing extreme meteorological factors. Accurate forecasting of the occurrence of extreme values in wind power generation is of great concern to ensure reliable power system operation. Recently, deep learning models have surged in popularity for wind power forecasting, with the mean squared error (MSE) loss function being commonly used. However, the MSE loss function, being sensitive to extreme values, disproportionately penalizes larger errors, cannot adequately capture the extreme values present in wind energy data, and novel loss functions have seldom been tailored for wind power forecasting. To this end, in this paper, we introduce a novel loss function specifically crafted to capture extreme values in wind power forecasting. The experimental results with four fundamental deep learning methods on open source wind power dataset validate that the new loss function is efficient and superior in all cases compared to MSE in capturing extreme values while maintaining forecasting performance.

17 WIND ENERGY↗

Detecting Masquerade Attacks in Controller Area Networks Using Graph Machine Learning

Modern vehicles rely on a myriad of electronic control units (ECUs) interconnected via controller area networks (CANs) for critical operations. Despite their ubiquitous use and reliability, CANs are susceptible to sophisticated cyberattacks, particularly masquerade attacks, which inject false data that mimic legitimate messages at the expected frequency. These attacks pose severe risks such as unintended acceleration, brake deactivation, and rogue steering. Traditional intrusion detection systems (IDS) often struggle to detect these subtle intrusions due to their seamless integration into normal traffic. This paper introduces a novel framework for detecting masquerade attacks in the CAN bus using graph machine learning (ML). We hypothesize that the integration of shallow graph embeddings with time series features derived from CAN frames enhances the detection of masquerade attacks. We show that by representing CAN bus frames as message sequence graphs (MSGs) and enriching each node with contextual statistical attributes from time series, we can enhance detection capabilities across various attack patterns compared to using graph-based features only. Our method ensures a comprehensive and dynamic analysis of CAN frame interactions, improving robustness and efficiency. Extensive experiments on the ROAD dataset validate the effectiveness of our approach, demonstrating statistically significant improvements in the detection rates of masquerade attacks compared to a baseline that uses graph-based features only as confirmed by Mann-Whitney U and Kolmogorov-Smirnov tests (p < 0.05) .

Marfo, William [Univ. of Texas, El Paso, TX (Unite↗

EMT data generation

The integration of inverter-based resources (IBRs) in power systems is accelerating, bringing with it significant benefits such as reduced greenhouse gas emissions, improved grid resilience, and increased energy independence. Despite these advantages, the widespread adoption of IBRs introduces several challenges, including issues related to grid stability, increased operational complexity, and the need for updated regulatory frameworks. To address these challenges, IEEE released Standard 2800 in 2022, which sets forth the necessary interconnection capabilities and performance criteria for IBRs connected to transmission and sub-transmission systems. This standard outlines the performance requirements to ensure the reliable integration of IBRs into the bulk power system. Furthermore, in 2023, the North American Electric Reliability Corporation (NERC) published a reliability guideline for electromagnetic transient (EMT) modeling of BPS-connected IBRs. This guideline provides recommendations for developing EMT model requirements, performing model quality checks, and implementing verification practices specifically for EMT models representing BPS-connected inverter-based resources in reliability studies conducted by transmission planners and planning coordinators. These standards and guidelines have a profound impact on EMT studies for transmission networks, influencing system stability analyses, grid recovery and resynchronization processes, fault ride-through evaluations, protection and coordination strategies, advanced control methodologies, and the inclusion of IBRs in transient models of transmission networks. As a result, the generation of EMT data is crucial for conducting various transient-based studies to understand the impact of IBRs. EMT data generation use cases serve as the basis for scenarios in event detection and identification use cases, providing comprehensive details about EMT data generation for transmission grids with inverter-based resources. These use cases supply sufficient training and validation datasets for subsequent EMT analysis algorithms.

Xia, Qianxue↗

Integration of GOES Data for Solar Resource Assessment of the Contiguous United States

The National Solar Radiation Database (NSRDB), produced by the National Laboratory of the Rockies (NLR), provides high-resolution solar resource data for the contiguous United States (CONUS) using Geostationary Operational Environmental Satellite (GOES) East and West observations. This study evaluates the integration of multi-satellite data within the GOES-East/West overlap regions, where conventional longitude-based selection methods often produce an artificial boundary seam. Our results demonstrate that an advanced blending algorithm, which incorporates sun-satellite scattering angles and satellite viewing zenith angles, improves NSRDB accuracy and creates a spatially continuous dataset. Validation against ground-based irradiance measurements reveals reductions in both percentage error (PE) and normalized Root Mean Square Error (nRMSE), particularly in the central United States. The dynamical integration of multi-satellite data provides a robust foundation for more precise modeling of solar resource and improved spatiotemporal analysis of solar ramp across the CONUS.

14 SOLAR ENERGY↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Using pile-up collisions as an abundant source of low-energy hadronic physics processes in ATLAS and an extraction of the jet energy resolution

During the 2015–2018 data-taking period, the Large Hadron Collider delivered proton-proton bunch crossings at a centre-of-mass energy of 13 TeV to the ATLAS experiment at a rate of roughly 30 MHz, where each bunch crossing contained an average of 34 independent inelastic proton-proton collisions. The ATLAS trigger system selected roughly 1 kHz of these bunch crossings to be recorded to disk. Offline algorithms then identify one of the recorded collisions as the collision of interest for subsequent data analysis, and the remaining collisions are referred to as pile-up. Pile-up collisions represent a trigger-unbiased dataset, which is evaluated to have an integrated luminosity of 1.33 pb -1 in 2015–2018. This is small compared with the normal trigger-based ATLAS dataset, but when combined with vertex-by-vertex jet reconstruction it provides up to 50 times more dijet events than the conventional single-jet-trigger-based approach, and does so without adding any additional cost or requirements on the trigger system, readout, or storage. The pile-up dataset is validated through comparisons with a special trigger-unbiased dataset recorded by ATLAS, and its utility is demonstrated by means of a measurement of the jet energy resolution in dijet events, where the statistical uncertainty is significantly reduced for jet transverse momenta below 65 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MARIAH PCAP data for Validation Demonstration

This dataset holds simulated PCAP (packet capture) data from the SCEPTRE validation demonstration model as a set of pairwise communications between devices via specific protocols. All connections should be assumed to be symmetric, as this data is an aggregation of the true PCAP. A mapping is also provided associating each IP address with its true device type.

cyber-physical system↗

A network-enabled pipeline for gene discovery and validation in non-model plant species

Identifying key regulators of important genes in non-model crop species is challenging due to limited multi-omics resources. To address this, we introduce the network-enabled gene discovery pipeline NEEDLE, a user-friendly tool that systematically generates coexpression gene network modules, measures gene connectivity, and establishes network hierarchy to pinpoint key transcriptional regulators from dynamic transcriptome datasets. After validating its accuracy with two independent datasets, we applied NEEDLE to identify transcription factors (TFs) regulating the expression of cellulose synthase-like F6 ( CSLF6 ), a crucial cell wall biosynthetic gene, in Brachypodium and sorghum. Our analyses uncover regulators of CSLF6 and also shed light on the evolutionary conservation or divergence of gene regulatory elements among grass species. These results highlight NEEDLE’s capability to provide biologically relevant TF predictions and demonstrate its value for non-model plant species with dynamic transcriptome datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Dark Energy Survey: Modeling strategy for multiprobe cluster cosmology and validation for the Full Six-year Dataset

We introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3x2pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (1) identification of modeling and parameterization choices and impact studies using simulated analyses with each possible model misspecification (2) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3x2pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy--galaxy lensing and cosmic shear two-point statistics) and the CL+3x2pt+BAO+SN (combination of CL+3x2pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗