Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Equation-of-state measured via x-ray phase contrast imaging for Epon 828/DEA epoxy

Epoxies are a broad class of polymer materials often used as adhesive, structural or binding materials. Epon 828 is an epoxy resin that can be polymerized with a variety of curing agents with the choice of curing agent potentially having an effect on the resulting epoxy polymer’s material properties. In this study, the dynamic behavior of Epon 828 epoxy resin cured with diethanolamine (DEA) is investigated through a series of tamped Richtmyer-Meshkov instability (RMI) experiments measured with x-ray phase-contrast imaging. The measured shock and particle velocities are combined with data in the literature to calibrate Mie-Grüneisen equations-of-state (EOS) for portions and combinations of the collective dataset. The calibrated Mie-Grüneisen EOS are validated against particle velocity profiles extracted from published literature using the Eulerian hydrocode CTH. Here, the Mie-Grüneisen EOS fit to only the tamped RMI experimental data presented here most closely follows the particle velocity profile in the published literature.

36 MATERIALS SCIENCE↗

Dynamical Sketching for Enhanced Communication Efficiency in Federated Learning

Federated learning (FL) has revolutionized distributed machine learning by enabling collaborative model training without sharing local data. However, communication efficiency and privacy guarantees remain significant challenges. This paper introduces a dynamic sketching mechanism in FL, optimizing the trade-off between communication efficiency and model accuracy. By dynamically selecting the sketch matrix size, our approach adapts to the evolving characteristics of the data and the model, ensuring optimal performance across diverse scenarios. We leverage Bayesian optimization to systematically tune the sketch parameters, achieving an effective balance between resource efficiency and model performance. Experimental results on the MNIST dataset using a convolutional neural network (CNN) architecture validate the proposed method's efficiency and scalability. Our dynamic sketching approach significantly outperforms fixed-size sketching techniques, achieving higher compression ratios (up to 62x) and providing better privacy guarantees while maintaining high model accuracy. These findings highlight the robustness and versatility of our approach and make it a valuable solution for privacy-preserving, communication-efficient federated learning.

Afrose, Sharmin [ORNL]↗

Data-Driven Digital Twin for Reliability Assessment of DC/DC Buck Converter

In commercial applications, the operation of DC/DC converters significantly impacts overall system performance and long-term reliability. This study introduces a data-driven digital twin (DT) approach for estimating critical degradation parameters of DC/DC BUCK converter under steady-state condition. Initially, a circuit-level MATLAB/Simulink digital model (DM C ) is refined against a hardware prototype’s switching model dataset using offline particle swarm optimization. The optimized digital model’s steady-state response is then verified with its average model response while varying the duty and load. Subsequently, degradation profiles are imposed on the inductor, capacitor, MOSFET in the DMC. A large dataset is generated from this model, allowing training, validation, and testing of machine learning (ML) models for component health regression tasks. The proposed method employs random forest ML models, achieving impressive regression results with a squared R value as high as 0.99978 and a root mean square error of 4.2× 10 –6 . The method is further validated on a medium power level DC/DC BUCK prototype with varying load conditions, and includes the analysis of MOSFET’s on-resistance under degradation conditions. This data-driven DT method shows promise for identifying parasitic degradation and ohmic loss parameters, enhancing converter reliability assessments in a non-invasive, generalized, and computationally efficient manner.

14 SOLAR ENERGY↗

Hardware In the Loop for Demand Flexibility (HIL4DF) v1.0

The software package in question is a collection of simulation models in the Modelica language, representing a variety of mechanical system designs and envelope conditions related to LBL's FLEXLAB facility. The collection of models also features multiple controls sequences that can be simulated with the FLEXLAB model to simulate different demand flexibility scenarios. Additionally, this package will feature datasets from 3 experimental tests, used for calibration, validation and comparison against the Modelica models, this includes weather data that can be used to replicate different scenarios in simulation across the same weather conditions experienced in real experiments. Given FLEXLAB high level of instrumentation and available data, the models are calibrated across multiple measurement points, and thus results from the extension of this model to other climate zones or control sequences, would provide high level of confidence.

Huang, Weiping↗

TEAMER: Performance mapping of Re Vision's Persistance PTO

Contains datasets from experimental measurements that were used to validate Re Vision's Persistence PTO's efficiency and performance. These measurements were obtained using a dynamometer test bench setup. The data includes open-circuit voltage and loss measurements to validate machine characteristics, efficiency mapping tests to determine the generator's performance mapping, and efficiency mapping tests to determine the converter's efficiency over the feasible operating range. This data was collected between June 2023 and September 2023. The data was collected at the National Renewable Energy Laboratory's Flatirons Campus, Colorado, United States. The data was collected using NREL's 5-kW dynamometer test bench, equipped with a torque sensor and various voltage and current sensors fed to a dedicated data acquisition system. Units for the data are included in the data file headers for each data series. A text editor or spreadsheet software such as Excel is required to view the *.csv data. The data are also provided in *.mat files. To view data plots, a Matlab script with *.mat files are provided.

16 TIDAL AND WAVE POWER↗

WFIP3

The Wind Forecasting Improvement Project 3 (WFIP-3) is the first offshore-based wind resource characterization project within the WFIP construct, funded by the U.S. Department of Energy. WFIP-3 will provide a unique field study that will deliver the comprehensive suite of data needed to inform a series of modeling efforts that will develop and evaluate parameterization schemes suited to offshore environments and improved industry-targeted applications. The field study has two goals: (1) detailed sampling of the vertical structure of the Marine Atmospheric Boundary Layer (MABL) at key observational areas, creating a rich dataset that will be used to refine and validate parameterization schemes, and (2) wide-area sampling of the MABL to create a multi-scale array of observations informing and guiding models of resource characterization. We will deploy a multi-platform array of measurements that span the MABL and create a multi-scale observational array stretching south from Marth’s Vineyard across the wind energy areas.

17 WIND ENERGY↗

Automatic Calibration and Health Monitoring of Infrastructure Sensors

Smart transportation infrastructure relies on networks of heterogeneous sensors - cameras, radars, and lidars - continuously monitoring traffic conditions. However, executing the initial spatial calibration of multiple sensors and the subsequent health monitoring presents significant operational challenges. Environmental factors, mechanical vibrations, and gradual drift cause spatial misalignment, degrading fusion performance and tracking accuracy. Traditional calibration approaches require manual intervention with specialized targets or survey equipment, resulting in service interruptions and high maintenance costs. This work presents an automated framework for initial calibration and continuous health monitoring without human intervention or service disruption. Our approach addresses two critical problems: (1) detecting when sensors become miscalibrated during operation, and (2) automatically re-establishing spatial alignment using only operational traffic data. The health monitoring component analyzes measurement innovations - differences between sensor observations and predicted object states - to detect systematic biases indicative of calibration drift. By computing bias magnitude, directional consistency, and rejection rates, the system identifies miscalibrations as small as 0.5 meters. Unlike traditional methods requiring known calibration targets, our diagnostic operates continuously on live traffic observations, enabling early detection before fusion quality degrades. The automatic recalibration algorithm leverages overlapping sensor fields-of-view and temporal correlation of vehicle observations. Using graph-based optimization, the system automatically discovers which sensor pairs observe common regions, estimates pairwise spatial transformations using RANSAC-based robust estimation, and jointly optimizes all sensor poses through bundle adjustment. The framework handles practical deployment challenges, including different sensor sampling rates (1-10 Hz), varying installation positions, unknown orientations, and limited overlap regions (>10%). When approximate sensor positions are available from installation surveys (+/-1m accuracy), the algorithm additionally estimates sensor orientations, refining both position and rotation to sub-meter and sub-degree accuracy. We validate the framework on multi-hour traffic datasets from six heterogeneous sensors with sampling rates ranging from 1 Hz to 10 Hz. Results demonstrate successful calibration even with sparse overlap (<20%) and automatic detection of miscalibrations exceeding 0.8 meters. This work enables a "deploy-and-forget" sensor infrastructure that maintains calibration autonomously, reducing maintenance costs while improving tracking accuracy. The techniques generalize beyond transportation to any multi-sensor monitoring application requiring robust spatial alignment, including smart cities, industrial monitoring, and surveillance systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A robust approach to Gaussian process implementation

Abstract. Gaussian process (GP) regression is a flexible modeling technique used to predict outputs and to capture uncertainty in the predictions. However, the GP regression process becomes computationally intensive when the training spatial dataset has a large number of observations. To address this challenge, we introduce a scalable GP algorithm, termed MuyGPs, which incorporates nearest-neighbor and leave-one-out cross-validation during training. This approach enables the evaluation of large spatial datasets with state-of-the-art accuracy and speed in certain spatial problems. Despite these advantages, conventional quadratic loss functions used in the MuyGPs optimization, such as root mean squared error (RMSE), are highly influenced by outliers. We explore the behavior of MuyGPs in cases involving outlying observations and, subsequently, develop a robust approach to handle and mitigate their impact. Specifically, we introduce a novel leave-one-out loss function based on the pseudo-Huber function (LOOPH) that effectively accounts for outliers in large spatial datasets within the MuyGPs framework. Our simulation study shows that the LOOPH loss method maintains accuracy despite outlying observations, establishing MuyGPs as a powerful tool for mitigating unusual observation impacts in the large data regime. In the analysis of US ozone data, MuyGPs provides accurate predictions and uncertainty quantification, demonstrating its utility in managing data anomalies. Through these efforts, we advance the understanding of GP regression in spatial contexts.

Mukangango, Juliette↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Nantucket (WFIP3 Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars, wind profiling radars, and sonic anemometers across Northeast U.S. coastal/offshore sites during the WFIP3 campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include comprehensive uncertainty estimates. The Nantucket dataset covers February 2024–September 2025, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

A dataset for understanding self-reported patterns influencing residential energy decisions

Household occupant behavior and decision-making dynamics substantially impact technology uptake and residential building energy performance. Although significant research underscores the importance of social science in energy studies, few public data with representative samples on household energy decision-making patterns are available. The dataset (UPGRADE-E: Understanding Patterns Guiding Residential Adoption and Decisions about Energy Efficiency) presents 9,919 responses from U.S. residents of single-family and small multifamily homes. Derived from a national-scale internet survey, the dataset contains 391 variables: demographics, building characteristics, home modifications, willingness to adopt new technologies, motivations for making changes, barriers, program participation, trusted information sources, and energy scenarios. Responses were validated via internal consistency checks and comparison with other U.S. national scale datasets. UPGRADE-E advances knowledge of household energy related decision-making, tying demographics, home modifications, and self-reported cognitive drivers together at a scale and breadth that has not been previously achieved. Policymakers and researchers at local, regional, and national levels may leverage this dataset to understand drivers influencing the adoption of key technologies in U.S. homes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

WINDPROF: Merged Best-Estimate Wind Profile Data – Block Island (WFIP3 Campaign)

WINDPROF provides 10-minute wind and turbulence profiles, integrating Doppler lidars, wind profiling radars, and sonic anemometers across Northeast U.S. coastal/offshore sites during the WFIP3 campaign. Key data include wind speed, direction, vertical velocity, and turbulence parameters, with standardized quality control (e.g., instrument-specific thresholds and inter-instrument validation). Profiles are interpolated to a height grid (20 m spacing below 100 m; 30 m above) and include comprehensive uncertainty estimates. The Block Island dataset covers February 2024–September 2025, offering reproducible methods for atmospheric research, model validation, and wind energy studies.

17 WIND ENERGY↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Demonstration of neutron identification in neutrino interactions in the MicroBooNE liquid argon time projection chamber

A significant challenge in measurements of neutrino oscillations is reconstructing the incoming neutrino energies. While modern fully-active tracking calorimeters such as liquid argon time projection chambers in principle allow the measurement of all final state particles above some detection threshold, undetected neutrons remain a considerable source of missing energy with little to no data constraining their production rates and kinematics. We present the first demonstration of tagging neutrino-induced neutrons in liquid argon time projection chambers using secondary protons emitted from neutron-argon interactions in the MicroBooNE detector. We describe the method developed to identify neutrino-induced neutrons and demonstrate its performance using neutrons produced in muon-neutrino charged current interactions. The method is validated using a small subset of MicroBooNE’s total dataset. The selection yields a sample with 60% of selected tracks corresponding to neutron-induced secondary protons. At this purity, the integrated efficiency is 8.4% for neutrons that produce a detectable proton.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Detecting Neutrons in MicroBooNE

A significant challenge in measurements of neutrino oscillations is reconstructing the incoming neutrino energies. While modern fully-active tracking calorimeters such as liquid argon time projection chambers in principle allow the measurement of all final state particles above some detection threshold, undetected neutrons remain a considerable source of missing energy with little to no data constraining their production rates and kinematics. We present the first demonstration of tagging neutrino-induced neutrons in liquid argon time projection chambers using secondary protons emitted from neutron-argon interactions in the MicroBooNE detector. We describe the method developed to identify neutrino-induced neutrons and demonstrate its performance using neutrons produced in muon-neutrino charged current interactions. The method is validated using a small subset of MicroBooNE's total dataset. The selection yields a sample with $60\%$ of selected tracks corresponding to neutron-induced secondary protons.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗