Search NASA⌕ Search

SEARCH · Search NASA

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Expanding on the BRIAR Dataset: A Comprehensive Whole Body Biometric Recognition Resource at Extreme Distances and Real-World Scenarios (Collections 1-4)

The state-of-the-art in biometric recognition algorithms and operational systems has advanced quickly in recent years providing high accuracy and robustness in more challenging collection environments and consumer applications. However, the technology still suffers greatly when applied to non-conventional settings such as those seen when performing identification at extreme distances or from elevated cameras on buildings or mounted to UAVs. This paper summarizes an extension to the largest dataset currently focused on addressing these operational challenges, and describes its composition as well as methodologies of collection, curation, and annotation.

Cornett, David [ORNL] (ORCID:0000000222910860)↗

Remark on Algorithm 1012: Computing Projections with Large Datasets

In ACM TOMS Algorithm 1012, the DELAUNAYSPARSE software is given for performing Delaunay interpolation in medium to high dimensions. When extrapolating outside the convex hull of the training set, DELAUNAYSPARSE calls the nonnegative least squares solver DWNNLS to compute projections onto the convex hull. However, DWNNLS and many other available sum-of-squares optimization solvers were not intended for usage with many variable problems, which result from the large training sets that are typical in machine learning applications. Thus, a new PROJECT subroutine is given, based on the highly customizable quadratic program solver BQPD. This solution is shown to be as robust as DELAUNAYSPARSE for projection onto both synthetic and real-world datasets, where other available solvers frequently fail. Although it is intended as an update for DELAUNAYSPARSE, due to the difficulty and prevalence of the problem, this solution is likely to be of external interest as well.

97 MATHEMATICS AND COMPUTING↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

Dataset for the paper titled "Investigating the relationship between bolide entry angle and apparent direction of infrasound signal arrivals"

This dataset includes outputs generated for the journal publication titled: "Investigating the relationship between bolide entry angle and apparent direction of infrasound signal arrivals". The outputs include .csv files with model-generated synthetic trajectories of asteroids entering Earth at a variety of impact and approach (azimuthal) angles. All outputs are based on hypothetical but realistic scenarios.

Herrera, Natalie [Sandia National Laboratories (SN↗

Idaho National Laboratory Quality Of Service Dataset

The code is designed to run tests to generate and collect data from a Wi-Fi network using OPENWRT or a simulated a 5G network using Open5gs and UERANSIM. The tests simulate the network performing downloads or uploads of various files with a varying number of concurrent users. The tests use tcpdump to collect the network traffic but only stores the summarized data. The summarized datasets will be included.

Krome, Cameron [Idaho National Laboratory (INL), I↗

Contribution to Open Molecules Dataset (coordcomplexsampling)

The Open Molecules Dataset is a project led by external collaborators, focused on the production of a wide diversity of molecular chemistries. This project will include high-throughput electronic structure calculations on metal coordination complexes, biomolecules, and electrolytes. Solvation effects, conformers, chemical reactivity, and spin/charge sampling will be pursued for elements on the periodic table up to, but not including the actinides. We aim to contribute code related to sampling for transition metal complexes, in particular, and expand software capabilities related to other thrusts desired.

Taylor, Michael G. [Los Alamos National Laboratory↗

SUNSet: The Software Understanding for National Security Dataset Repo

SAND2025-00511O SUNSet: The Software Understanding for National Security Dataset Repo serves as a repository for software understanding researchers to conduct systematic research in the field. It provides a platform for storing questions, answers, and scripts related to software programs, supporting research in software understanding. The repository is a nascent effort aimed at exploring the support needed by researchers and documenting how software understanding questions can be addressed using existing tools. The software consists of a simple database and front end interface for easy access and management of the stored information. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Amon, Tod [Sandia National Lab. (SNL-CA), Livermor↗

Spatiotemporally Registered In-Situ and Ex-Situ Datasets for Laser-based Blown Powder Directed Energy Deposition

This dataset is comprised of in situ sensing data collected during laser-based, blown powder directed energy deposition (DED) of Inconel 718 representing eight different printing conditions: (1) nominal, (2) +15% scan speed, (3) +12% laser power, (4) +42% powder feed rate, (5) +100% jerk limit, (6) +10% layer height, (7) +20% hatch spacing, (8) +20% carrier gas flow. All eight DED builds constructed an identical test coupon geometry consisting of geometric features representative of industrial print requirements (e.g., bulk deposition, thin walls, overhangs). In situ data consists of xyz-coordinates (100 Hz) and on-axis melt pool camera video (60 Hz), both of which have been temporally synchronized to spatially map the melt pool camera data. In addition, post-build X-ray computed tomography (XCT) data for each of the eight test geometries have been spatially registered to the recorded xyz-coordinates, allowing for comparisons between melt pool camera data and flaws identified in the XCT data.

additive manufacturing↗

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING↗

Dataset for manuscript "Equipartition and the temperature of maximum density of TIP4P/2005 water"

We simulate TIP4P/2005 water in the temperature range of 257 K to 318 K with time-steps 0.25, 0.50, 1.00, 2.00, and 4.00 fs. The density-temperature behavior obtained using 0.25 or 0.50 fs are in excellent agreement with each other but differ from those obtained using time-steps that have been shown earlier to lead to a breakdown of equipartition. The temperature of maximum density (TMD) is 277.15 K with time-step 0.25 or 0.50 fs, but is shifted to progressively lower values for longer time-steps, a trend that holds for different thermostat/barostat combinations. Enhancing the water-water dispersion interaction, as has been recommended for simulating disordered proteins in TIP4P/2005, degrades the description of the liquid-vapor phase envelope. We present a simple physically transparent reasoning to highlight the separation of the time-scales between translational and rotational motion. We also develop a metric, Chi, that we term the equipartition anomaly, to detect equipartition violations in simulations that include molecules that are treated as rigid objects. Calculating Chi is shown to be straightforward and sensitive to equipartition violations. A key takeaway from this study is that using sufficiently short time-steps (less than or equal to 0.5 fs) to preserve equipartition is essential for obtaining meaningful liquid water properties and for producing reliable simulation data, as correct-ensemble sampling is fundamental to ensure reproducibility across codes and simulation alogrithms. The included dataset provides the raw data used in the preparation of the graphs noted in the manuscript.

36 MATERIALS SCIENCE↗

Dataset for manuscript "Rotational Memory Function of SPC/E water"

Memory effect are essential for dynamics of condensed materials and are responsible for non-exponential relaxation of correlation functions of dynamic variables through the memory function entering the memory equation. Memory functions of dipole rotations for polar liquids have never been calculated. We present here calculations of memory functions and single-dipole rotations and of the overall system dipole moment for SPC/E water measured by dielectric spectroscopy. The memory functions for single-particle and collective dynamics turn out to be nearly identical. This result validates theories of dielectric spectroscopy in terms of single-particle time correlation function and the connection between the collective and single-particle relaxation times in terms of the Kirkwood factor. The dataset includes single particle and system dipole moments, including their time-dependence.

74 ATOMIC AND MOLECULAR PHYSICS↗

Gridded Vegetation and Land Unit Datasets over North America (1km x 1km) for E3SM Land Model, Version 2

Gridded 1 km land-surface-property dataset for Energy Exascale Earth System Model Land Model (ELM) simulations over North America. Thirty-metre NALCMS land cover is aggregated to the Daymet grid; vegetation classes are cross-walked to ELM plant functional types using mean-temperature-of-the-coldest-month (MTCO) rules. NetCDF includes PFT and land-unit fractions, counts, and MTCO.

54 ENVIRONMENTAL SCIENCES↗

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

IEEE39_IBL_Dataset

IEEE39 with IBL modified system EMT dataset.

Madurasinghe, Dulip [ORNL] (ORCID:0000000249316597↗

Dataset for Role of electron correlation on the adenine dimer interaction for non-equilibrium geometries: a benchmark Quantum Monte Carlo study

Datasets for the calculations reported in "Role of electron correlation on the adenine dimer interaction for non-equilibrium geometries: A benchmark Quantum Monte Carlo study" by L. Washburn, A. Sedova, P. R. C. Kent. J. Chem. Phys. (2026) 165 (5): 054118. https://doi.org/10.1063/5.0332651. Includes the molecular geometries, QMCPACK, PySCF, and ORCA inputs and outputs, analysis scripts and files needed to reproduce all the figures and tables.

59 BASIC BIOLOGICAL SCIENCES↗

National Dataset of EV Charging Stations With Estimates of Load and Vehicle Throughput

Current data from the AFDC provide locations and many details about EV charging stations, but not estimates of their peak loads or the number of vehicles they can accommodate. This dataset will augment the AFDC charging station locations with estimates of transmission load and vehicle throughput based on engineering specifications of the chargers, charging patterns based on vehicle types, battery capacities, and user behavior.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗