Search NASA⌕ Search

SEARCH · Search NASA

Results for “data publication and archiving”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

Spatially coordinated airborne data and complementary products for aerosol, gas, cloud, and meteorological studies: the NASA ACTIVATE dataset

The NASA Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE) produced a unique dataset for research into aerosol–cloud–meteorology interactions, with applications extending from process-based studies to multi-scale model intercomparison and improvement as well as to remote-sensing algorithm assessments and advancements. ACTIVATE used two NASA Langley Research Center aircraft, a HU-25 Falcon and King Air, to conduct systematic and spatially coordinated flights over the northwest Atlantic Ocean, resulting in 162 joint flights and 17 other single-aircraft flights between 2020 and 2022 across all seasons. Data cover 574 and 592 cumulative flights hours for the HU-25 Falcon and King Air, respectively. The HU-25 Falcon conducted profiling at different level legs below, in, and just above boundary layer clouds (< 3 km) and obtained in situ measurements of trace gases, aerosol particles, clouds, and atmospheric state parameters. Under cloud-free conditions, the HU-25 Falcon similarly conducted profiling at different level legs within and immediately above the boundary layer. The King Air (the high-flying aircraft) flew at approximately ~ 9 km and conducted remote sensing with a lidar and polarimeter while also launching dropsondes (785 in total). Collectively, simultaneous data from both aircraft help to characterize the same vertical column of the atmosphere. In addition to individual instrument files, data from the HU-25 Falcon aircraft are combined into “merge files” on the publicly available data archive that are created at different time resolutions of interest (e.g., 1, 5, 10, 15, 30, 60 s, or matching an individual data product's start and stop times). This paper describes the ACTIVATE flight strategy, instrument and complementary dataset products, data access and usage details, and data application notes. The data are publicly accessible through https://doi.org/10.5067/SUBORBITAL/ACTIVATE/DATA001 (ACTIVATE Science Team, 2020).

54 ENVIRONMENTAL SCIENCES↗

Why don't we share data and code? Perceived barriers and benefits to public archiving practices

The biological sciences community is increasingly recognizing the value of open, reproducible and transparent research practices for science and society at large. Despite this recognition, many researchers fail to share their data and code publicly. This pattern may arise from knowledge barriers about how to archive data and code, concerns about its reuse, and misaligned career incentives. Here, we define, categorize and discuss barriers to data and code sharing that are relevant to many research fields. We explore how real and perceived barriers might be overcome or reframed in the light of the benefits relative to costs. By elucidating these barriers and the contexts in which they arise, we can take steps to mitigate them and align our actions with the goals of open science, both as individual scientists and as a scientific community.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

The cosmic DANCe of Perseus

Context. Star-forming regions are excellent benchmarks for testing and validating theories of star formation and stellar evolution. The Perseus star-forming region, being one of the youngest (< 10 Myr), closest (280-320 pc), and most studied in the literature, is a fundamental benchmark. Aims. We aim to study the membership, phase-space structure, mass, and energy (kinetic plus potential) distribution of the Perseus star-forming region using public catalogues (Gaia, APOGEE, 2MASS, and Pan-STARRS). Methods. We used Bayesian methodologies that account for extinction to identify the Perseus physical groups in the phase-space, retrieve their candidate members, derive their properties (age, mass, 3D positions, 3D velocities, and energy), and attempt to reconstruct their origin. Results. We identify 1052 candidate members in seven physical groups (one of them new) with ages between 3 and 10 Myr, dynamical super-virial states, and large fractions of energetically unbounded stars. Their mass distributions are broadly compatible with that of Chabrier for masses ≳0.1 M ⊙ and do not show hints of over-abundance of low-mass stars in NGC 1333 with respect to IC 348. These groups’ ages, spatial structure, and kinematics are compatible with at least three generations of stars. Future work is still needed to clarify if the formation of the youngest was triggered by the oldest. Conclusions. The exquisite Gaia data complemented with public archives and mined with comprehensive Bayesian methodologies allow us to identify 31% more members than previous studies, discover a new physical group (Gorgophone: 7 Myr, 191 members, and 145 M ⊙ ), and confirm that the spatial, kinematic, and energy distributions of these groups support the hierarchical star formation scenario.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

SLOPE daily and 250 m gross primary productivity (GPP) for the CONUS, 2000-2019, Carbon Monitoring System (CMS)

The SatelLite Only Photosynthesis Estimation (SLOPE) GPP product is a daily, 250 m GPP dataset covering the Contiguous United States (CONUS) from 2000 to present with 1 day latency. Gross primary productivity (GPP) quantifies the amount of carbon dioxide (CO2) fixed by plants through photosynthesis. Although as a key quantity of terrestrial ecosystems, there is a lack of high-spatial-and-temporal-resolution, real-time, and observation-based GPP products. This product has been developed to address this critical gap, leveraging a number of MODIS land and atmosphere products. There are three distinct features of the SLOPE GPP production algorithm. (1) SLOPE couples machine learning models with MODIS atmosphere and land products to accurately estimate PAR. (2) SLOPE couples highly efficient and pragmatic gap-filling and filtering algorithms with surface reflectance acquired by both Terra and Aqua MODIS satellites to derive a soil-adjusted NIRv (SANIRv) dataset. (3) SLOPE couples a temporal pattern recognition approach with a long-term Crop Data Layer (CDL) product to predict dynamic C4 crop fraction. PAR, SANIRv and C4 fraction are used to drive a parsimonious model with only two slope parameters to estimate GPP along with a quantitative uncertainty on a per-pixel and daily basis. The slope GPP product has a R2 = 0.84 and a root-mean-square error (RMSE) of 1.65 gC m-2 d-1, evaluated against from 50 AmeriFlux eddy covariance sites (332 site-years). Archived data from 2000 to 2019 are publicly available at the NASA's Oak Ridge National Laboratory Distributed Active Archive Center (ORNL DAAC). Data from 2020 are available from the authors upon request. All data are projected in the standard MODIS Land Integerized Sinusoidal tile map projection. Each processing tile is in size of 4800 pixels by 4800 pixels, representing approximately 1200 km by 1200 km land region. In addition to the GPP product, SLOPE PAR, SANIRV, and C4 fraction, along with their uncertainties, are also released.

Jiang, Chongya↗

Bioenergy Feedstock Library Annual Summary Report

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at INL, is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use as well as samples generated from across U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2022 is to focus on the (1) publicly available analytical data and equipment tracked through the BFNUF, (2) physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities. This report highlights key statistics from FY22 and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS↗

Small-wedge synchrotron and serial XFEL datasets for Cysteinyl leukotriene GPCRs

Structural studies of challenging targets such as G protein-coupled receptors (GPCRs) have accelerated during the last several years due to the development of new approaches, including small-wedge and serial crystallography. Here, we describe the deposition of seven datasets consisting of X-ray diffraction images acquired from lipidic cubic phase (LCP) grown microcrystals of two human GPCRs, Cysteinyl leukotriene receptors 1 and 2 (CysLT 1 R and CysLT 2 R), in complex with various antagonists. Five datasets were collected using small-wedge synchrotron crystallography (SWSX) at the European Synchrotron Radiation Facility with multiple crystals under cryo-conditions. Two datasets were collected using X-ray free electron laser (XFEL) serial femtosecond crystallography (SFX) at the Linac Coherent Light Source, with microcrystals delivered at room temperature into the beam within LCP matrix by a viscous media microextrusion injector. All seven datasets have been deposited in the open-access databases Zenodo and CXIDB. Here, we describe sample preparation and annotate crystallization conditions for each partial and full datasets. We also document full processing pipelines and provide wrapper scripts for SWSX and SFX data processing. A Correction to this paper has been published: https://doi.org/10.1038/s41597-020-00759-w

97 MATHEMATICS AND COMPUTING↗

Benchmarking second and third-generation sequencing platforms for microbial metagenomics

Shotgun metagenomic sequencing is a common approach for studying the taxonomic diversity and metabolic potential of complex microbial communities. Current methods primarily use second generation short read sequencing, yet advances in third generation long read technologies provide opportunities to overcome some of the limitations of short read sequencing. Here, we compared seven platforms, encompassing second generation sequencers (Illumina HiSeq 300, MGI DNBSEQ-G400 and DNBSEQ-T7, ThermoFisher Ion GeneStudio S5 and Ion Proton P1) and third generation sequencers (Oxford Nanopore Technologies MinION R9 and Pacific Biosciences Sequel II). We constructed three uneven synthetic microbial communities composed of up to 87 genomic microbial strains DNAs per mock, spanning 29 bacterial and archaeal phyla, and representing the most complex and diverse synthetic communities used for sequencing technology comparisons. Our results demonstrate that third generation sequencing have advantages over second generation platforms in analyzing complex microbial communities, but require careful sequencing library preparation for optimal quantitative metagenomic analysis. Our sequencing data also provides a valuable resource for testing and benchmarking bioinformatics software for metagenomics.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Scale 3D Imaging for Machine Learning Property Upscaling: Mt. Simon Sandstone Case Study

Petrographic properties of principal target reservoirs for carbon sequestration, such as the Mt. Simon Sandstone, are relevant to broad interest groups. The Mt. Simon Sandstone is a deep, saline, regionally extensive Cambrian sandstone, overlain by low permeability sealing formations, making it one of the viable geologic carbon storage reservoirs in the Midwestern US. Its thickness (exceeding 2400 ft in some localities), depth, and lateral extent, combined with high porosity and permeability make it a high-priority target of multiple ongoing geologic carbon sequestration efforts in the United States of America. The National Energy Technology Laboratory in Morgantown, West Virginia, has been engaged in characterization efforts of the Mt. Simon for over a decade, with a strong focus on Computed Tomographic data acquisition. Data generated during this period has been hitherto not accessible to the public. This archival effort focused on preservation of historical CT data and associated metadata, and facilitating their accessibility, culminating with the publication of the entire dataset on NETL’s Energy Data eXchange (EDX) and the associated Gill et. al (2024) paper.

Gill, Magdalena K.↗

Transportation Secure Data Center: Frequently Asked Questions for Data Owners/Contributors

The Transportation Secure Data Center is a centralized repository for detailed transportation data from travel and transit surveys and studies conducted across the nation. It makes vital transportation data broadly available to users while preserving the privacy of survey participants. Hundreds of datasets from surveys and studies of household travel and transit passenger travel are archived in the TSDC, including surveys and studies conducted by state departments of transportation, metropolitan planning organizations, transit agencies, cities, and other public agencies. Detailed data from travel surveys and studies are extremely valuable for research purposes. However, the fine-grained information they contain could potentially be misused to identify individual travelers, so access to these data should only be granted with safeguards in place to protect participant privacy. The TSDC was created to address this challenge and to relieve public agencies from the burden of archiving their data and responding to data requests.

33 ADVANCED PROPULSION SYSTEMS↗

Bioenergy Feedstock Library Annual Summary Report 2024

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at Idaho National Laboratory (INL), is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use, as well as samples generated from U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) and U.S. Department of Agriculture-funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2024, similar to the 2023 Annual Summary Report , is to focus on the updates to: (1) publicly available analytical data and equipment tracked through the BFNUF, (2) significant increases in the physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities compared to the previous annual summary reports. This report highlights key statistics and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS↗

Municipality of Anchorage Household Travel Survey

The 2002 Household Travel Survey for the municipality of Anchorage, Alaska, collected demographic, socioeconomic, and travel information about households and persons (age five and older) on an assigned, 24-hour travel period. The main objective of the study, conducted by NuStats, was to improve the transportation system. 2,035 households were recruited to participate in the study. Of these, 1,293 completed travel logs depicting detailed data on driving habits, such as purpose of the trips, time of the data, and mode of transportation—spanning from April 1, 2002, to May 17, 2002. This dataset is part of the Metropolitan Travel Survey Archive, which includes travel surveys from numerous public agencies across the United States and is archived by the Transportation Secure Data Center to ensure their continued public availability.

1Hz data↗

EMDB—the Electron Microscopy Data Bank

Abstract The Electron Microscopy Data Bank (EMDB) is the global public archive of three-dimensional electron microscopy (3DEM) maps of biological specimens derived from transmission electron microscopy experiments. As of 2021, EMDB is managed by the Worldwide Protein Data Bank consortium (wwPDB; wwpdb.org) as a wwPDB Core Archive, and the EMDB team is a core member of the consortium. Today, EMDB houses over 30 000 entries with maps containing macromolecules, complexes, viruses, organelles and cells. Herein, we provide an overview of the rapidly growing EMDB archive, including its current holdings, recent updates, and future plans.

59 BASIC BIOLOGICAL SCIENCES↗

REMBI: Recommended Metadata for Biological Images—enabling reuse of microscopy data in biology

Bioimaging data have significant potential for reuse, but unlocking this potential requires systematic archiving of data and metadata in public databases. Here, we propose draft metadata guidelines to begin addressing the needs of diverse communities within light and electron microscopy. We hope this publication and the proposed Recommended Metadata for Biological Images (REMBI) will stimulate discussions about their implementation and future extension.

59 BASIC BIOLOGICAL SCIENCES↗

2001 Atlanta Household Travel Survey

The 2001 Atlanta Household Travel Survey collected demographic, socioeconomic, and travel information on work and non-work travel behavior for a 48-hour travel period. The study was conducted on behalf of the Atlanta Regional Commission, and it is an essential element in the transportation planning and modeling efforts for the 13-county Atlanta region. The main objective of the study was to produce data that could be used to develop and calibrate travel demand models for use in travel forecasting, land use planning, and air quality planning to improve the transportation system. The second component in the survey was the deployment of an electronic travel diary with person-based GPS and an accelerometer to collect health and activity information. Travel data includes trip generation, trip distribution, and modal choice. The survey recruited a total of 12,184 households to participate in the study. Of these, 8,069 households (66%) completed travel. The 8,069 households, when weighted, represent 21,323 persons, 14,449 vehicles, and 126,127 places visited from April 2001 through April 2002. This dataset is part of the Metropolitan Travel Survey Archive, which includes travel surveys from numerous public agencies across the United States and is archived by the Transportation Secure Data Center to ensure their continued public availability.

1Hz data↗