Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framework to understand, assess, and quantify the potential information leakage associated with machine learning systems. Designing effective MIAs is a challenging task that usually requires extensive manual exploration of model behaviors to identify potential vulnerabilities. In this paper, we introduce AutoMIA -- a novel framework that leverages large language model (LLM) agents to automate the design and implementation of new MIA signal computations. By utilizing LLM agents, we can systematically explore a vast space of potential attack strategies, enabling the discovery of novel strategies. Our experiments demonstrate AutoMIA can successfully discover new MIAs that are specifically tailored to user-configured target model and dataset, resulting in improvements of up to 0.18 in absolute AUC over existing MIAs. This work provides the first demonstration that LLM agents can serve as an effective and scalable paradigm for designing and implementing MIAs with SOTA performance, opening up new avenues for future exploration.

Tran, Toan Viet [Emory University]↗

Measurement of the differential cross section for neutral pion production in charged-current muon neutrino interactions on argon with the MicroBooNE detector

We present a measurement of neutral pion production in charged-current interactions using data recorded with the MicroBooNE detector exposed to Fermilab’s booster neutrino beam. The signal comprises one muon, one neutral pion, any number of nucleons, and no charged pions. Studying neutral pion production in the MicroBooNE detector provides an opportunity to better understand neutrino-argon interactions, and is crucial for future accelerator-based neutrino oscillation experiments. Using a dataset corresponding to 6.86 ×10 20 protons on target, we present single-differential cross sections in muon and neutral pion momenta, scattering angles with respect to the beam for the outgoing muon and neutral pion, as well as the opening angle between the muon and neutral pion. Data extracted cross sections are compared to generator predictions. We report good agreement between the data and the models for scattering angles, except for an over-prediction by generators at muon forward angles. Similarly, the agreement between data and the models as a function of momentum is good, except for an underprediction by generators in the medium momentum ranges, 200–400 MeV for muons and 100–200 MeV for pions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - Simulated Marine Hydrokinetic Tidal Turbine

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis, hydrogen compression and storage, and variable hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset is part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with other energy technologies. This dataset contains inputs and outputs from simulations of a floating marine hydrokinetic turbine over approximately half a tidal cycle (~6.6 hours). Inflow conditions were derived from field measurements in Alaska’s Cook Inlet and represent a tidal environment in which the current speed ramps from near 0 m/s to a peak of 3 m/s and back. The original acoustic doppler current profiler dataset is publicly available on the Marine and Hydrokinetic Data Repository. In a full tidal cycle, the flow reverses and the rotor would reorient; this reversal was not modeled. In the Cook Inlet campaign , turbulence intensity was similar in both directions. Two inflow cases are included. In the first case, labeled “raw” in the files, the measured current time series was used directly in the InflowWind module of OpenFAST. Speed and direction were applied as a function of time and elevation, uniformly in the horizontal direction. With full spatial coherence, this approach captures high turbulent variability and results in pronounced power fluctuations, so it is considered a conservative, near-worst-case representation of loading. In the second case, labeled “average” in the files, a 30-minute moving average was applied to extract the slowly varying mean speed. The residual fluctuations about this mean were used to generate spatially varying, full-field turbulence inputs with TurbSim, giving a more physically realistic representation of the inflow across the rotor disk. Two random realizations were used to produce distinct inflow conditions for two OpenFAST simulations representing a two-turbine array. The same turbulence intensity is applied across the full time series, producing larger fluctuations at the start and end, where the mean speed is low. The second case is the more appropriate framework for performance and power assessment but overpredicts turbulence at lower flow speeds and underpredicts it at higher speeds. As the floating platform moves and the rotor changes its x-position, Taylor’s frozen turbulence hypothesis used by InflowWind assumes a constant rather than a time-varying mean velocity, introducing some inaccuracy in the velocity plane sampling. The turbine modeled is the 500-kW Reference Model 1, a horizontal-axis two-bladed hydrokinetic turbine on a four-column floating semisubmersible substructure . Simulations were performed using OpenFAST v4.1 with the Reference Open Source Controller (ROSCO) v2.10. All input files required to reproduce the simulations are included. The electrolyzer is a 1.25-MW proton exchange membrane type MC250 system manufactured by Nel . This unit supports up to 2.5 MW, but NLR has only a single 1.25-MW stack. The datasets report hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. The system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operating current of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The simulated tidal turbine time series data was translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz. Each zip file represents a single tidal electrolysis experiment and is named: {technology}_{inflow method}_{number of 500 kW tidal turbines connected} For instance, “tidal-500kW-RM1_average_2.zip” is a 6-hour experiment using the 500-kW tidal reference model, scaled by 2x (1-MW) to better match the electrolyzer maximum of 1.25MW, fed with the 30-minute moving average current case. Each zip folder contains the following files: A .csv file of raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production in kilograms per hour, electrolysis power consumption, and input wave power. A .csv file combines all tidal profiles as "combined_tidal_experiments.csv." A separate experiment, “characterization_200.zip,” shows the MC250 electrolyzer steady-state response with 30-minute load steps over 5 hours and is accessible with this entry.

08 HYDROGEN↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

PyJMAK: An Open-Source Python Toolkit for Modeling Solid-State Metallurgical Phase Transformations

Accurate prediction of metallurgical phase transformations is an essential basis for autonomous optimization and rapid part qualification. Several methods can be used to estimate the evolution of phase fractions such as JMAK kinetics-based models, phase-field models, thermodynamic models, and data-driven machine learning models. Thermodynamic and phase-field-based methodologies solve multiphysics equations requiring numerous calibration parameters and significant computational resources. As a result, the computation domain is limited to a point or on order of micron-meters. The data-driven models rely on large datasets from experiments and simulations. While the JMAK model only provides information about phase fraction evolution, it can predict this evolution in near real-time using thermal history and thermodynamic data without restriction on the domain. JMAK models have been popularly used by researchers to model phase transformations occuring during additive manufacturing or over arbitrary temperature profiles. Commercial proprietary software such as Abaqus and Ansys or closed-source in-house implementations offer the ability to model JMAK based kinetics to predict phase transformation. However, these software packages are not open-source or freely available for use and development in conjunction with manufacturing machines, sensors, and machine learning algorithms. In addition, the use of the model is restricted by a license token. In contrast, given temperature profiles at multiple points in the domain, this Python-based PyJMAK model can compute phase evolution in parallel due to its stand-alone modular, voxel-based structure, and it can be executed on high-performance computing resources without any license restrictions.

Prabhune, Bhagya [Oak Ridge National Laboratory (O↗

Active learning for the design of polycrystalline textures using conditional normalizing flows

Generative modeling has opened new avenues for solving previously intractable materials design problems. However, these new opportunities are accompanied by a drastic increase in the required amount of training data. This is in stark juxtaposition to the high expense and difficulty in curating such large materials datasets. In this work, we propose a novel framework for integrating generative models within an active learning loop. Further, this enables the training of generative models with datasets significantly smaller than what has previously been demonstrated, providing a direct route for their application in data constrained environments. The functionality of this framework is then demonstrated by addressing the challenge of designing polycrystalline textures associated with target anisotropic mechanical properties. The developed protocol exhibited a cost reduction between 14 to 18 times over a randomly sampled experimental design.

36 MATERIALS SCIENCE↗

The 4D Camera: An 87 kHz Direct Electron Detector for Scanning/Transmission Electron Microscopy

We describe the development, operation, and application of the 4D Camera—a 576 by 576 pixel active pixel sensor for scanning/transmission electron microscopy which operates at 87,000 Hz. The detector generates data at ~480 Gbit/s which is captured by dedicated receiver computers with a parallelized software infrastructure that has been implemented to process the resulting 10–700 Gigabyte-sized raw datasets. The back illuminated detector provides the ability to detect single electron events at accelerating voltages from 30 to 300 kV. Through electron counting, the resulting sparse data sets are reduced in size by 10--300× compared to the raw data, and open-source sparsity-based processing algorithms offer rapid data analysis. The high frame rate allows for large and complex scanning diffraction experiments to be accomplished with typical scanning transmission electron microscopy scanning parameters.

47 OTHER INSTRUMENTATION↗

Data for "Verification of the kinetic electron role in the microinstabilities in a negative triangularity model equilibrium"

This contains the dataset used in "Verification of the kinetic electron role in the microinstabilities in a negative triangularity model equilibrium" published in Physics of Plasmas in Oct 2024. It consists of plaintext files as well as .bp files (which may be read by ADIOS open-source software) used for creating the plots that appear in the paper. These are derived from simulation outputs from XGC (X-point Gyrokinetic Code). The data mainly consists of growth rate and frequency measurements as well as poloidal cross-section data.

gyrokinetic↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

MapsTorch : automatic differentiation for X-ray fluorescence data analysis

X-ray fluorescence (XRF) is a popular spectroscopy technique for elemental analysis. Spectrum fitting and parameter tuning are at the core of XRF analysis and are conventionally manually intensive, especially for synchrotron experiments involving large amounts of diverse samples. This work introduces the automatic differentiation (AD) technique to XRF and an open-source package called MapsTorch. By transforming an analytical model of the XRF spectrum into a differentiable computation graph with AD, MapsTorch enables robust optimization of parameters and elemental intensities. We evaluate MapsTorch by conducting computational experiments on a large number of historical synchrotron XRF datasets and compare its performance with the currently practiced fitting tool NLopt. The results show that MapsTorch consistently achieves high-quality fits and often leads to better fitting quality than NLopt, particularly in tasks such as initial spectrum fitting and elemental intensity refinement. The robust performance of MapsTorch paves the way for developing automated and high-throughput XRF data analysis workflows to handle the increasing data volumes expected from next-generation synchrotron facilities.

X-ray fluorescence↗

Lab Scale Demonstration of Pipeline Third-Party Damage Classification Using Convolutional Neural Networks

This research aims to propose a simple experiment for third party damage classification problem by generating a dataset of third-party damage events on a laboratory scale utilizing single mode-multi mode-single mode (SMS) fiber acoustic sensor. The sound samples representative of various third-party activities, such as vehicle movements, excavation, and digging, were sourced from open-source databases. These samples were then played through a speaker in proximity to an SMS sensor, and the resultant fiber acoustic vibration data were recorded for each event. This process yielded a collection of 200 samples across 13 distinct third-party events. Convolutional Neural Networks (CNNs) were employed to classify these samples into their respective categories, and an accuracy exceeding 97% was obtained from our results.

Bukka, Sandeep Reddy↗

Lab-Scale Demonstration of Pipeline Third-Party Damage Classification Using Convolutional Neural Networks

This research aims to mitigate the challenges of field tests for classification of third-party damages by generating a dataset of third-party damage events on a laboratory scale utilizing single mode-multi mode-single mode (SMS) fiber acoustic sensor. The sound samples representative of various third-party activities, such as vehicle movements, excavation, and digging, were sourced from open-source databases. These samples were then played through a speaker in proximity to an SMS sensor, and the resultant fiber acoustic vibration data were recorded for each event. This process yielded a collection of 200 samples across 13 distinct third-party events. Convolutional Neural Networks (CNNs) were employed to classify these samples into their respective categories, and an accuracy exceeding 97% was obtained from our results.

Bukka, Sandeep Reddy↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Wind Loading on CSP Collectors

The project significantly enhanced the community's understanding of the fundamental physics drivers underlying the wind-loading experienced by concentrating solar power (CSP) collector structures (i.e., parabolic troughs and heliostats) as well as their support structures. This project had two overarching objectives: (1) detailed measurements to characterize the prevailing wind conditions and resulting operational loads on collector structures, and (2) development and validation of a computationally efficient, high-fidelity modeling tool capable of predicting wind-loading in deep-array installations. Over three years, we conducted comprehensive at-scale field measurements of the atmospheric turbulent wind conditions, and the resulting wind loads on parabolic troughs and heliostats. Two at-scale measurement campaigns yielded first-of-its-kind, high-resolution, long-term datasets that are used to characterize the complex flow field and wind loading on parabolic-troughs and heliostats in operational power plants. The high-resolution measurements collected during these campaigns were used to validate the high-fidelity computational models developed at NREL. These open-source computationally efficient models were shown to be accurate in predicting wind-driven loads on collectors without the need for a large supercomputer.

14 SOLAR ENERGY↗

Modeling Multi-View Impedance-Based Cross-Geometry SOH Estimator for Li-ion Batteries

Abstract: Accurately estimating battery’s State of Health (SOH) remains challenging when models must generalize across cell designs and operating conditions. Most Electrochemical Impedance Spectroscopy (EIS)-based approaches either (i) hand-engineer a few Nyquist-plot features for shallow models—fast but does not generalize across geometries—or (ii) learn directly from Nyquist plots with deep networks, which removes manual feature extraction, yet still limited to a single plot type. As a result, cross-geometry robustness and deployability on constrained Internet of Things (IoT) devices remain open problems. We propose a compact Convolutional Neural Network (CNN) (∼ 10k parameters) that takes multi-representation EIS inputs—Nyquist (real/imaginary) and phase–magnitude (|Z|/ϕ) stacked as four channels, so the model can learn complementary degradation signatures while remaining small enough for fast inference. We build a dataset from cyclic aging of two geometries (LG INR18650MJ1 cylindrical cells and LIR2032 coin cells), acquire EIS every ten cycles from 10 kHz to 10 mHz (10 points/decade), and evaluate with leave-one-cell-out testing strategy. We further study fusion vs. single-representation inputs and assess feasibility for on-device deployment (e.g., NVIDIA Jetson device). The results show that training on multiple EIS representations improves SOH estimation accuracy and cross-geometry generalization compared to single-representation models, which uses only Nyquist or phase–magnitude plots. This design targets accurate, generalizable SOH prediction without manual feature engineering while enabling practical real-time use.

Bakr, Ahmed [The University of Alabama (UA)]↗

Survey of Use Cases and Scenarios on the Open Energy Data Initiative Solar Systems Integration (OEDI SI) Platform

The Open Energy Data Initiative Solar Systems Integration (OEDI SI) Data and Modeling Platform offers a comprehensive set of use cases tailored for power systems analysis. Each use case is centered around a specific power system analysis problem, supported by composite input data and reference algorithms. These composite input datasets are meticulously assembled using OEDI SI's data preprocessing tools, which integrate raw data from various sources. The primary objectives of the OEDI SI Platform include facilitating access to composite input data through widely accepted input/output formats and verified results. This accessibility enables power system network researchers and developers to validate their algorithms and showcase their applications' capabilities to the broader community. Moreover, the platform strives to promote reproducible, robust, replicable, and generalizable solar systems integration research.

14 SOLAR ENERGY↗