Search NASASearch

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

2012 California Household Travel Survey Supplement

# 2012 California Household Travel Survey Supplement The 2012 California Household Travel Survey Supplement focused on gathering specific travel information from residents for the development of next-generation, activity-based models. Called the "Augment Survey," it supplemented the [2010–2012 California Household Travel Survey](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-california-travel-survey). ## Data Collection Agency The Southern California Association of Governments (SCAG) hired Abt-SRBI, Inc. to conduct the survey. ## Methodology Travel data were collected from households via in-vehicle (625 vehicles) and wearable (244 participants) global positioning system (GPS) devices. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Study records include 473 households. ## Transportation Data The SCAG data set contains data from 473 households that participated in one or more areas of study. Of these, 141 completed the wearable GPS portion of the study and 332 completed the vehicle GPS portion. There was no overlap between households participating in the two study areas (wearable and vehicle GPS). For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/caltrans_scag_data_dictionary.pdf?sfvrsn=6ec36d7a_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Hourly gap-filled meteorological data from PIE LTER measurements (2004-2023) used as drivers to run ELM PFLOTRAN simulations

This dataset contains continuous gap-filled precipitation, solar radiation, photosynthetically active radiation (PAR), air temperature, relative humidity, wind speed, and barometric pressure data recorded primarily at the Marshview Farm weather station within the Plum Island Long Term Ecosystems Research (PIE LTER) in Newbury Massachusetts (MA) from 2004 to 2023. We compiled the data set from published annual data packages in 15min resolution available on DataOne. Gaps were filled using different statistical techniques or available observations from the vicinity, e.g. the US-PLo and the US-PHM Ameriflux sites, also located within the PIE LTER. Flags are included in this dataset to indicate the origin of each data point. Metadata files ELMPFLOTRAN_met_dd.csv and ELMPFLOTRAN_met_flmd.csv contain more information on site locations, gap filling protocols, data variables, flags, and QA/QC methods. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024).

54 ENVIRONMENTAL SCIENCES

Urban Parameters Arizona Urban Corridor 100m

132 Urban parameters based on building physical dimensions and location were generated for the cities in six Arizona Counties at 100m resolution using the NATURF model. To use the binary file with WRF, the binary file and the index file must be placed in their own directory in WRF_GEOG and accessed in the same way NUDAPT44 would be accessed.

Dumas, Melissa [ORNL] (ORCID:0000000233190846)

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude

NCAR-RAL Surface Hydrometeorological Observation Network Data for LASSO-CACTI Overview Paper

This data set contains the 15 minute resolution surface meteorology and soils data from the 15 NCAR/RAL weather stations that were operated around central Argentina during the RELAMPAGO (Remote sensing of Electrification, Lightning, And Meso-scale/micro-scale Processes with Adaptive Ground Observations) Extended Observing Period (EOP). Data providence, citation, and acknowledgement This ARM data set is a copy of v1.0 of the NCAR data set obtained in June 2024 from https://doi.org/10.26023/KW8Z-F2WX-H0Y. The citation for the original data source is: Gochis, D., et al. 2019. NCAR-RAL Surface Hydrometeorological Observation Network Data. Version 1.0. UCAR/NCAR - Earth Observing Laboratory. https://doi.org/10.26023/KW8Z-F2WX-H0Y Accessed June 2024. In addition to the citation reference and any other acknowledgements, please acknowledge NCAR/EOL in your publications with text such as: “Data provided by NCAR/EOL under the sponsorship of the National Science Foundation. https://data.eol.ucar.edu/”

air temperature

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY

A Climatology of Dust Deposition in the Upper Colorado River Basin for February-May 1980-2023

This data set is a long-term climatology of the average monthly total dust deposition, wet and dry, for the months of February-May 1980-2023 pulled from the MERRA-2 reanalysis data set over the Upper Colorado River Basin. This data set can be used to study the long-term spatiotemporal patterns of dust deposition, especially on snow. This data set is associated with the preprint article “A multi-decadal climatology of dust-on-snow from wet deposition in the Upper Colorado River Basin”.

dry_dust_deposition

DEPRECATED - State Policies and Programs for Community Solar (2024 Q3 Update)

This data set is no longer current – The most current data and all historical data sets can be found at https://data.nlr.gov/submissions/249 The purpose of this dataset is to summarize current community solar policies and low-income stipulations by state in the United States as of June 2024. The "State_Program" sheet summarizes the key policy details for each state. This list has been reviewed, but errors may exist, and the list may not be comprehensive. NREL invites input to update or add to the database. To submit updates, additions, or corrections please find contact information on the current data set page linked above.

14 SOLAR ENERGY

BLDAP Intro to Python/Data Science Curriculum v1

The Github repository contains the Jupyter notebooks for the intro to Python / Data Science course for Berkeley Lab Director's Apprenticeship Program (BLDAP). This course is designed for students with little to no experience in coding to learn skills in Python necessary for data science. Students utilize Jupyter notebooks throughout the course. The overall goal is for students to learn how to use Python to clean, analyze, and visualize large data sets in order to communicate effectively their conclusions about the data set. Students apply the skills they learned on actual data sets provided by researchers in Berkeley Lab.

Hales, Laurel [Lawrence Berkeley National Laborato

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)

Meteorological Services Annual Data Report for 2024

This document presents the meteorological data collected at Brookhaven National Laboratory (BNL) by Meteorological Services (Met Services) for the calendar year 2024. The purpose is to publicize the data sets available to emergency personnel, researchers and facility operations. Met services has been collecting data at BNL since 1949. Data from 1994 to the present is available in digital format. Data is presented in monthly plots of one-minute data. This allows the reader the ability to peruse the data for trends or anomalies that may be of interest to them. Full data sets are available to BNL personnel and to a limited degree outside researchers. The full data sets allow plotting the data on expanded time scales to obtain greater details (e.g., daily solar variability, inversions, etc.).

54 ENVIRONMENTAL SCIENCES