Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS↗

Climate Nowcasting

The climate is changing so rapidly that climatologies based on historical statistics cannot reliably capture the current risk of extreme weather events hazardous to society. Decision relevant projections of weather extreme probability over the next 10–15 years are needed to enable adaptation and resilience in the face of this evolving risk. Current weather forecasts/predictions and long-term climate projections for decades into the future are inadequate for providing this information to stakeholders that need it, targeting forecast horizons either too short or too far into the future. We argue that a new approach is needed: climate nowcasting. Climate nowcasting would focus on user-inspired extreme metrics over the next 10–15 year time frame, targeting specific impacts and locations down to a local scale, by engaging with stakeholders to understand their needs and provide information in a format relevant for decision making. Importantly, climate nowcasting will not consist of a single approach or data set, involving rather data fusion from different sources of information: simulations, observations and data driven methods, likely through different weights depending on the metric of interest. Predictions must be accompanied by serious engagement with stakeholders facing climate risks, and clearly present uncertainties and limitations of any prediction. Such a vision is very different from how typical climate or weather forecasts are applied today.

Gettelman, Andrew↗

Powered by dsgrid [Slides]

NREL's demand-side grid (dsgrid) toolkit harnesses decades of sector-specific energy modeling expertise to understand current and future U.S. electricity load for power systems analyses. The primary purpose of dsgrid is to create comprehensive electricity load data sets at high temporal, geographic, sectoral, and end-use resolution. These data sets enable detailed analyses of current patterns and future projections of end-use loads. This presentation will include NREL power grid researcher Elaine Hale.

24 POWER TRANSMISSION AND DISTRIBUTION↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Classification Strategy to Improve Streamflow Estimates in Diverse River Basins in the Colorado River Basin

Streamflow in the Colorado River Basin (CRB) is significantly altered by human activities including land use/cover alterations, reservoir operation, irrigation, and water exports. Climate is also highly varied across the CRB which contains snowpack-dominated watersheds and arid, precipitation-dominated basins. Recently, machine learning methods have improved the generalizability and accuracy of streamflow models. Previous successes with LSTM modeling have primarily focused on unimpacted basins, and few studies have included human impacted systems in either regional or single-basin modeling. We demonstrate that the diverse hydrological behavior of river basins in the CRB are too difficult to model with a single, regional model. We propose a method to delineate catchments into categories based on the level of predictability, hydrological characteristics, and the level of human influence. Lastly, we model streamflow in each category with climate and anthropogenic proxy data sets and use feature importance methods to assess whether model performance improves with additional relevant data. Overall, land use cover data at a low temporal resolution was not sufficient to capture the irregular patterns of reservoir releases, demonstrating the importance of having high-resolution reservoir release data sets at a global scale. On the other hand, the classification approach reduced the complexity of the data and has the potential to improve streamflow forecasts in human-altered regions.

54 ENVIRONMENTAL SCIENCES↗

HERO WEC V1.0 2024 - WEC-Sim Detailed Simulation Runs and Summary Data

This dataset includes results from simulations of NREL's hydraulic and electric reverse osmosis wave energy converter (HEREO WEC). Simulation runs include 135 wave cases that were based on the updated WEC-Sim model, which is linked below. The data represented in this repository is based on an updated WEC-Sim model using laboratory data to tune and refine the original WEC-Sim model for the V1.0 HERO WEC. The 135 wave cases represent waves with the following wave height and wave period ranges: - Significant Wave Height: 0.25 - 3.75m in 0.25m increments - Wave Period: 5 - 13 sec in 1 sec increments Each run was simulated using a Pierson-Moskowitz irregular wave spectrum with a 100 second ramp time, a total simulation time of 3,100 seconds, and a simulation time-step of 0.005s. A reference table has been included to map each multi condition run (MCR) case with each wave condition. Summary data set includes a spreadsheet and image files with matrices that are associated with data from simulation runs. All matrices cover the same significant wave height and wave periods from the simulation runs, in the same increments. The following matrices are included: - Power Abs: The average absorbed power from the WEC (calculated from anchor reaction force and heave velocity) - Power Hyd: The average hydraulic power output at pump (calculated from pump output flow and pressure) - Power - Hyd ROi: The average hydraulic power measured at the RO system inlet (calculated from RO system pressure and flow (pre-accumulator)) - Flow - Pump out: The average flowrate measured at the pump outlet - Flow - Perm: The average permeate (clean water) production - Flow - RO (pre): The average flowrate measured at the inlet of the RO system before the accumulators - Flow - RO (post): The average flowrate measured after the accumulator bank in the RO system - Pressure - RO: The average pressure measured at the inlet of the RO system This data set has been developed by the National Renewable Energy Laboratory, operated by the Alliance for Sustainable Energy, LLC, for the U.S. Department of Energy (DOE) under Contract No. DE-AC36-08GO28308. Funding provided by the U.S. Department of Energy Office of Energy Efficiency and Renewable Energy Water Power Technologies Office.

16 TIDAL AND WAVE POWER↗

Transfer Learning Trained LSTM Models for Household Load Profile Forecasting

Grid edge renewable energy resources, such as rooftop solar photovoltaics, closely interact with consumer load profiles. Therefore, forecasting future electricity demand, ideally at the individual household level, is indispensable. In this paper, we present a transfer learning enhanced household load profile forecasting method. First, we tune a long short-term memory forecasting model to perform day-ahead prediction of household electricity load profiles. Then we improve these individualized models using transfer learning, and we use k-means clustering to create optimal source data sets. We find average improvements of 4.38% (largest improvement of 10.71%) when the entire data set was used to train the source model and 2.45% (largest improvement of 11.57%) in the mean absolute error when households were first clustered and used to train separate source models for each cluster. We find that transfer learning with clustered data can effectively boost the forecasting performance of the LSTM models. We use realistic household power measurements for 148 real residential households in Austin, Texas.

deep learning↗

Applying Machine‐Learning Methods to Laser Acceleration of Protons: Lessons Learned From Synthetic Data

ABSTRACT In this study, we consider three different machine‐learning methods—a three‐hidden‐layer neural network, support vector regression, and Gaussian process regression—and compare how well they can learn from a synthetic data set for proton acceleration in the Target Normal Sheath Acceleration regime. The synthetic data set was generated from a previously published theoretical model by Fuchs et al. 2005 that we modified. Once trained, these machine‐learning methods can assist with efforts to maximize the peak proton energy, or with the more general problem of configuring the laser system to produce a proton energy spectrum with desired characteristics. In our study, we focus on both the accuracy of the machine‐learning methods and the performance on one GPU including memory consumption. Although it is arguably the least sophisticated machine‐learning model we considered, support vector regression performed very well in our tests.

Desai, Ronak↗

CACTI CSAPR2 Taranis Retrievals

Taranis is an end-to-end processing chain for radar data written in Python with C extension for computation performance. Features include: masking for quality control, specific differential phase (Kdp), attenuation correction for reflectivity factor (Z) and differential reflectivity (Zdr) in rain, and additional geophysical retrievals. Retrievals are mostly drawn from literature or open-source software when appropriate, and have been tested, tuned, and modified to work with one another cohesively rather than using isolated off-the-shelf algorithms. Incorporated algorithms include hydrometeor (echo) identification, rain water content, raindrop mass-weighted mean diameter (gamma size distribution assumption), and rainfall rate (QPE). Taranis data sets exist for CSAPR2 PPI, HSRHI, and sector RHI scans. Cartesian-gridded data sets were also produced as well as a near-surface rain rate retrieval. More details can be found in the README.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Verification, Validation, and Calibration Through a Causal Lens

While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.

97 MATHEMATICS AND COMPUTING↗

2024 Update of Comprehensive Review of Multi-arm Caliper Data for the Big Hill SPR Site

The Big Hill SPR site has a rich data set consisting of multi-arm caliper (MAC) logs collected from the cavern wells. This data set provides insight into the on-going casing deformation at the Big Hill site. This report summarizes the MAC surveys for each well and presents well longevity estimates where possible. Included in the report is an examination of the well twins for each cavern and a discussion on what may or may not be responsible for the different levels of deformation between some of the well twins. The report also takes a systematic view of the MAC data presenting spatial patterns of casing deformation and deformation orientation in an effort to better understand the underlying causes. The conclusions present a hypothesis suggesting the small-scale variations in casing deformation are attributable to similar scale variations in the character of the salt-caprock interface. These variations do not appear directly related to shear zones or faults. In addition, the deformation orientation shows no preferred directionality. This 2024 edition of this report represents an update to the original, 2023 edition. The updates primarily focus on the inclusion of MAC log data run since the December 2021 threshold date for the original report, but some new analyses are also included.

58 GEOSCIENCES↗

On the Abuse and Detection of Polyglot Files

A polyglot is a file that is valid in two or more formats. Polyglot files pose a problem for file-upload and generative AI web interfaces that rely on format identification to determine how to securely handle incoming files. In this work we found that existing file-format and embedded-file detection tools, even those developed specifically for polyglot files, fail to reliably detect polyglot files used in the wild. To address this issue, we studied the use of polyglot files by malicious actors in the wild, finding 30 polyglot samples and 15 attack chains that leveraged polyglot files. Using knowledge from our survey of polyglot usage in the wild---the first of its kind---we created a novel data set based on adversary techniques. We then trained a machine learning detection solution, PolyConv, using this data set. PolyConv achieves a precision-recall area-under-curve score of 0.999 with an F1 score of 99.20% for polyglot detection and 99.47% for file-format identification, significantly outperforming all other tools tested. We developed a content disarmament and reconstruction tool, ImSan, that successfully sanitized 100% of the tested image-based polyglots, which were the most common type found via the survey. Our work provides concrete tools and suggestions to enable defenders to better defend themselves against polyglot files, as well as directions for future work to create more robust file specifications and methods of disarmament.

Oesch, T [ORNL] (ORCID:0000000269091022)↗

Barge Site - Avian Radar System / Derived Data

This is a combined data set of 67,410 bird/bat tracks from an avian radar system deployed on a research barge (MERLIN True3D, DeTect, Panama City, Florida, USA) and concurrent wind measurements from two scanning lidars (WindCube v2.1, Vaisala, Vantaa, Finland, and Halo XR+, Halo Photonics, Lannion, France). The research barge (16.5 m x 61 m) was deployed as part of the Wind Forecast Improvement Project (WFIP-3) off the northeast coast of the United States south of Massachusetts (40.9 deg N, 70.79 deg W). This data set comprises 5 weeks of data between August 27th 2024 and September 27th 2024. Radar data were provided by DeTect and Lidar data were accessed through the Wind Data Hub (wfip3/barg.WINDPROF.z01.a0) The data have been filtered and sorted into two size groups ("big" and "small") based on a clustering approach. See Snortland, A., Clerc, J., Hein, C., & Cotter, E. (2025). Wind as Driver of Bird and Bat Abundance, Flight Direction, Altitude, and Speed on the North Atlantic Shelf. arXiv preprint arXiv:2511.14983 for complete details. Data are provided in 2 files: "Birds" and "Birds_hourly" Birds: This file contains information about each of the 67,410 flying animal tracks detected by the radar during the data collection period, including parameters measured by the radar and wind information interpolated from the lidar wind measurements. We note that the raw radar dataset contained 301,618 tracks; tracks in this processed dataset were filtered based on the requirements described in Snortland et al. (2025). Birds_hourly: This file contains timeseries of the number of tracks detected per hour over the course of the data collection period, including wind conditions and sun position for each hour. These data were used for generalized additive modeling in Snortland et al. (2025).

17 WIND ENERGY↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗