Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Patch-level chamber-based CH4 and CO2 flux measurements in freshwater and salt marshes of coastal Louisiana

This dataset contains chamber-based carbon dioxide and methane flux measurements from ecohydrological patches in two coastal wetlands in Louisiana (within the footprint of AmeriFlux sites US-LA2 and US-LA3). Measurements span four patches, two per wetland. At the freshwater marsh (US-LA2), patches were dominated by Sagittaria lancifolia or co-dominated by Sagittaria lancifolia and Typha latifolia. At the salt marsh site (US-LA3), patches were dominated by Spartina alterniflora or Juncus roemerianus. Flux measurements included (i) wetland surface-level measurements encompassing the entire soil-water-vegetation column and (ii) soil-water surface chamber measurements that excluded emergent vegetation. At US-LA3, open-water pools were common and were also measured as a patch type. With this data we aimed to assess differences in gas flux pathways within ecohydrological patches across different wetland types. Files can be read in any program or code typical of processing .csv file types. "PatchChambers_LA2LA3.csv" covers flux data from patch level chambers which includes bulk CO2/CH4 flux encompassing water, soil, and vegetation within each chamber. "Water_Soil_SurfaceChamber_LA2LA3.csv" has CO2/CH4 flux data from chambers measuring only the water-soil surface without the effects of vegetation. Metadata can be found in the similarly named files for each data sheet ("xxxx_md.csv").

54 ENVIRONMENTAL SCIENCES↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

Legacy Survey of Space and Time Data Preview 1: visit_detector_table dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the visit_detector_table dataset type. These are per-detector visit metadata. This release contains 1 dataset of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Legacy Survey of Space and Time Data Preview 1: CcdVisit searchable catalog

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of a searchable catalog named CcdVisit. This catalog contains per-detector visit metadata. This catalog contains 16,071 rows with 50 columns.

79 ASTRONOMY AND ASTROPHYSICS↗

MISIP: a data standard for the reuse and reproducibility of any stable isotope probing-derived nucleic acid sequence and experiment

DNA/RNA-stable isotope probing (SIP) is a powerful tool to link in situ microbial activity to sequencing data. Every SIP dataset captures distinct information about microbial community metabolism, process rates, and population dynamics, offering valuable insights for a wide range of research questions. Data reuse maximizes the information derived from the labor and resource-intensive SIP approaches. Yet, a review of publicly available SIP sequencing metadata showed that critical information necessary for reproducibility and reuse was often missing. Here, we outline the Minimum Information for any Stable Isotope Probing Sequence (MISIP) according to the Minimum Information for any (x) Sequence (MIxS) framework and include examples of MISIP reporting for common SIP experiments. Our objectives are to expand the capacity of MIxS to accommodate SIP-specific metadata and guide SIP users in metadata collection when planning and reporting an experiment. The MISIP standard requires 5 metadata fields—isotope, isotopolog, isotopolog label, labeling approach, and gradient position—and recommends several fields that represent best practices in acquiring and reporting SIP sequencing data (e.g., gradient density and nucleic acid amount). The standard is intended to be used in concert with other MIxS checklists to comprehensively describe the origin of sequence data, such as for marker genes (MISIP-MIMARKS) or metagenomes (MISIP-MIMS), in combination with metadata required by an environmental extension (e.g., soil). The adoption of the proposed data standard will improve the reuse of any sequence derived from a SIP experiment and, by extension, deepen understanding of in situ biogeochemical processes and microbial ecology.

Simpson, Abigayle↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

Common practices for quantifying methane emissions from plumes detected by remote sensing

This document provides a set of community-accepted practices for quantifying methane emissions based on plumes detected via spectroscopic remote sensing. Its primary goal is to promote consistency in the generation, validation, reporting, and quality assessment of methane emission estimates derived from remote sensing radiances. Developed by subject matter experts with deep experience across all stages of the measurement process, this guidance reflects a critical evaluation of current methodologies and highlights key practices needed to produce reliable, interoperable, and traceable products. The focus is specifically on methane emissions quantified from distinct plumes originating from localized sources, rather than diffuse emissions spread over large regions, which are beyond the scope of this work. This document is intended to serve both data producers and users. For producers, it offers a framework for aligning with field-recognized standards to ensure their outputs meet rigorous quality and transparency criteria. For users, it provides a reference to assess dataset fitness-for-purpose by highlighting essential metadata, assumptions, and methodological choices that underpin emission estimates. By fostering a shared understanding of best practices, this work aims to enhance comparability, confidence, and utility of remotely sensed methane emission products.

54 ENVIRONMENTAL SCIENCES↗

Neutron powder diffraction, Mossbauer Spectroscopy and Optical Spectroscopy to study magnetic and nuclear lattices in Fe-based oxychlorides Ca2FeO3Cl, Sr2FeO3Cl and Sr3Fe2O5Cl2

Data for Neutron Diffraction, Mossbauer Spectoscopy and Optical Spectroscopy are contained on all three samples (Ca2FeO3Cl, Sr2FeO3Cl, and Sr3Fe2O5Cl2). Mossbauer data are in the folder "MossbauerSpectroscopy". This contains details of files and examples to read the data. The Optical Spectroscopy are in the folder "AbsorptionData", this contains one file with explanatory headers. The neutron diffraction data are in the "NeutronDiffractionData" folder. The data was collected on the HB-2A powder diffractometer at HFIR. The autoreduced files are corrected using a vanadium standard and are in arbitrary intensity units. The RawData folder contains the uncorrected data with metadata for motor positions and temperature. All measurements were collected with a constant neuton wavelength of 2.41 Angstrom. An excel spreadsheet "NeutronDiffractionData_IPTS-29118_summary" contains details for each scan. Sr3Fe2O5Cl2 data were collected at only room temperature. Ca2FeO3Cl and Sr2FeO3Cl data were collected at 4K and room temperature. The autoreduced .dat fileformat is three columns corresponding to: Two-theta, Intensity and Intensity_error.

fe-based oxychlorides↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

Changuinola peat soil characteristics and gas emission raw data October 2019

This dataset comprises radiocarbon and geochemical measurements from peat and porewater samples collected across various depths at a site in Bocas del Toro, Panama. The study focuses on carbon cycling dynamics in tropical peatlands by examining carbon isotopic signatures (¹⁴C and ¹³C) and elemental compositions of bulk peat, dissolved organic carbon (DOC), carbon dioxide (CO₂), and methane (CH₄). Key parameters include radiocarbon ages and isotopic ratios (δ¹³C) of bulk peat, concentrations of carbon (%C) and nitrogen (%N), and radiocarbon content of porewater gases and dissolved organic carbon (DOC). The data provide insights into the vertical and spatial distribution of carbon sources and possible preservation and decomposition processes within tropical peat profiles, offering critical information for understanding carbon storage and greenhouse gas emissions in these ecosystems.This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) carbon isotopic signatures (¹⁴C and ¹³C); (5) concentrations of carbon (%C) and nitrogen (%N); (6) radiocarbon content of porewater carbon dioxide (CO₂), and methane (CH₄) ; (7) porewater DOC; (8) bulk peat sampling protocol; (9) porewater sampling protocol; (10) porewater gas collection methods; and (11) gas extraction methods. All files are in .csv format and can be opened with any software that supports this file types.

54 ENVIRONMENTAL SCIENCES↗

CHESS 2025: Location data for field observations and sampling

This dataset represents geolocation data associated with field observations and sampling from the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Location data were collected using Trimble DA2 Global Navigation Satellite System (GNSS) receivers with Trimble Catalyst 2 centimeter (cm) positioning service and the Environmental Systems Research Institute (Esri) Field Maps mobile app. Files in this data package include meadow site polygons, shrub site polygons, tree site polygons and stem point locations, and Leaf Area Index (LAI) plot polygons (.geojson). The geojson files can be opened with open-source GIS software (e.g, QGIS). A csv file is also provided with point coordinates for all locations. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Metabolomics (PB-DP5)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Culture samples were collected at 0, 0.5, 1, 2, 4, 6, and 8 hours for extracellular sucrose analysis. Circadian metabolomics data was acquired using a Agilent single quadrupole gas chromatography-mass spectrometer and processed using Agilent Mass Hunter for targeted sucrose quantification. Metabolomic analysis of PCC 7942 light-dark cycle cultures transitioned to constant light revealed distinct temporal patterns in sucrose production. Processed metabolomic datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed GC-MS results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

RefAHL: a curated quorum sensing reference linking diverse LuxI-type signal synthases with their acyl-homoserine lactone products

Some bacteria use acyl-homoserine lactone (AHL) signals in quorum sensing, a type of cell-cell communication. Here, we present “RefAHL,” an updated, curated collection of LuxI-type AHL synthases with their AHL products and associated metadata. RefAHL is publicly available as a community resource to help catalog LuxI-type diversity encoded in (meta) genomic data.

59 BASIC BIOLOGICAL SCIENCES↗

DOE BSSD Performance Management Metrics Report Q3

Microbiome data is complex, spanning information from microbial genomes within diverse communities, protein and metabolite readouts, and contextual information (metadata) captured from the environments from which these samples were collected. While the variety and scale of microbiome data generation has dramatically expanded over the past twenty years, infrastructure to support data management, sharing, and access has lagged. New ways to improve interoperability across existing resources and advancing community standards are necessary to support how researchers create, use, and reuse data. The National Microbiome Data Collaborative (NMDC) aims to advance a microbiome data sharing network through infrastructure, data standards, and community building.

54 ENVIRONMENTAL SCIENCES↗