Search NASA⌕ Search

SEARCH · Search NASA

Results for “data access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Human Liver Epithelial Cells (HuH7) Response to HCoV-229E Infection Epigenomics (ATAC-Seq) (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-299E) infection alters chromatin accessibility in infected cells. Sample data was obtained from mock-infected cells, UV-inactivated virus treated cells, and replication competent HCoV-229E infected immortalized human liver cells (HuH7) at 24 hours post infection. Samples were processed using ATAC-seq methods for reported bar coded libraries. Sample data was acquired using an Illumina Hi-Seq 2500 sequencer system and further processed for ATAC-Seq expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Advancing ocean monitoring and knowledge for societal benefit: the urgency to expand Argo to OneArgo by 2030

The ocean plays an essential role in regulating Earth’s climate, influencing weather conditions, providing sustenance for large populations, moderating anthropogenic climate change, encompassing massive biodiversity, and sustaining the global economy. Human activities are changing the oceans, stressing ocean health, threatening the critical services the ocean provides to society, with significant consequences for human well-being and safety, and economic prosperity. Effective and sustainable monitoring of the physical, biogeochemical state and ecosystem structure of the ocean, to enable climate adaptation, carbon management and sustainable marine resource management is urgently needed. The Argo program, a cornerstone of the Global Ocean Observing System (GOOS), has revolutionized ocean observation by providing real-time, freely accessible global temperature and salinity data of the upper 2,000m of the ocean (Core Argo) using cost-effective simple robotics. For the past 25 years, Argo data have underpinned many ocean, climate and weather forecasting services, playing a fundamental role in safeguarding goods and lives. Argo data have enabled clearer assessments of ocean warming, sea level change and underlying driving processes, as well as scientific breakthroughs while supporting public awareness and education. Building on Argo’s success, OneArgo aims to greatly expand Argo’s capabilities by 2030, expanding to full-ocean depth, collecting biogeochemical parameters, and observing the rapidly changing polar regions. Providing a synergistic subsurface and global extension to several key space-based Earth Observation missions and GOOS components, OneArgo will enable biogeochemical and ecosystem forecasting and new long-term climate predictions for which the deep ocean is a key component. Driving forward a revolution in our understanding of marine ecosystems and the poorly-measured polar and deep oceans, OneArgo will be instrumental to assess sea level change, ocean carbon fluxes, acidification and deoxygenation. Emerging OneArgo applications include new views of ocean mixing, ocean bathymetry and sediment transport, and ecosystem resilience assessment. Implementing OneArgo requires about $100 million annually, a significant increase compared to present Argo funding. OneArgo is a strategic and cost-effective investment which will provide decision-makers, in both government and industry, with the critical knowledge needed to navigate the present and future environmental challenges, and safeguard both the ocean and human wellbeing for generations to come.

ARGO↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

Strong Lensing by Galaxies

Strong gravitational lensing at the galaxy scale is a valuable tool for various applications in astrophysics and cosmology. Some of the primary uses of galaxy-scale lensing are to study elliptical galaxies’ mass structure and evolution, constrain the stellar initial mass function, and measure cosmological parameters. Since the discovery of the first galaxy-scale lens in the 1980s, this field has made significant advancements in data quality and modeling techniques. In this review, we describe the most common methods for modeling lensing observables, especially imaging data, as they are the most accessible and informative source of lensing observables. We then summarize the primary findings from the literature on the astrophysical and cosmological applications of galaxy-scale lenses. We also discuss the current limitations of the data and methodologies and provide an outlook on the expected improvements in both areas in the near future.

79 ASTRONOMY AND ASTROPHYSICS↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Grid Operator Analytics and Assessment Tools for Inverter- Based Resources Dominated Grid (GOAAT-IBR) Project Update

This presentation provides an update on the OPTIMA GOAAT project, with emphasis on the cloud-native data platform developed in-house to ingest, manage, and operationalize high-resolution power system data. Since our last NASPI presentation, accessible via OSTI ID #2671437, the project team advanced the design and deployment of a scalable architecture capable of handling both synchronized and non-synchronized streams, including PMU, point-on-wave (POW), COMTRADE, and SCADA data. These materials review the project status, recent progress, and key lessons learned. The core of the presentation examines the architecture and engineering of our cloud-native ingestion and data management platform. We then explain how pipelines were designed to collect, normalize, time-align, store, and serve heterogeneous data at scale. We will discuss design choices such as data models, streaming versus batch ingestion, storage tiers, and interoperability with analytics applications. Practical experiences with cloud-native technologies were shared during the event, including benefits, limitations, and integration challenges in a utility environment, along with methods used to improve performance, reduce latency, and optimize resource usage. The presentation also showcases user interface designs and visualization tools that convert raw measurements and analytics results into intuitive, actionable insights for operators and engineers. During the presentation examples were provided demonstrating how visualization, event views, and summarized analytics enhance situational awareness and support operational decision-making. These use cases illustrate how a well-designed data infrastructure can bridge the gap between high-volume measurements and practical grid operations.

Aminifar, Farrokh↗

Digitizing and Enhancing Accessibility of the Fusion Safety Archives

This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Thermal expansion of LaB 6 from 298 to 998 K

We propose LaB 6 as a temperature calibration standard for high-temperature (HT) X-ray diffractometry owing to its high temperature stability. Such HT applications require a reliable HT lattice parameter or, equivalently, peak position data, which have not been readily accessible to the diffraction community to date. As such, the thermal expansion behavior of NIST SRM 660a LaB6 was assessed in the temperature range 298–998 K using HT Bragg–Brentano parafocusing θ:θ X-ray diffractometry in conjunction with Rietveld analysis. Data were collected in the 2θ range 20–150° at a data collection rate of 0.5° θ min −1 in air and at 1 atm. The temperature was stepped in 50 K increments. The cubic unit-cell lattice parameter [a(T)] of LaB 6 in Å was found to vary as a(T) = 4.15678 (±0.00001) + Ξ(T − 298 K) + Ψ(T − 298 K) 2 , where Ξ = 2.4645 × 10 −5 (±4.8904 × 10 −8 ) Å K −1 and Ψ = 1.0325 × 10 −8 (±6.7376 × 10 −11 ) Å K −1 . The isobaric volume thermal expansion coefficient (TEC) was obtained as α V P = (5.9291 × 10 −6 ) + (4.9680 × 10 −9 )(T − 298 K) K −1 , from which the corresponding linear TEC was obtained as α L P = (1.9764 × 10 −6 ) + (1.6560 × 10 −9 )(T − 298 K) K −1 . The 3 × 3 matrix representations of the single-crystal isobaric linear TEC and the volume expansivity were obtained for the cubic crystal class to which LaB 6 belongs. Also, the temperature dependence of the lattice parameter data of this study was compared with past landmark studies on LaB 6 by Dutchak et al. [Inorg. Mater. (1972), 8, 1877–1880] and Aivazov et al. [Inorg. Mater. (1979), 15, 1015–1016].

LaB6 standard material↗

Equitable Energy Metrics for Integration into Building Performance Standard Tracking Platforms: Preprint

Building Performance Standards (BPS) are being adopted globally and in the United States of America, where 14 different states and jurisdictions have a policy in place and many others are under development (Department of Energy (DOE) 2023). Accurate and equitable data sources are essential to make informed decisions about focusing investment on upgrading buildings to meet jurisdictional goals. There have been multiple new tools developed related to Energy Equity and Environmental Justice (EEEJ) and the resulting datasets need to be integrated into large building port-folios for quick access and better scalability. Integrating EEEJ data in a user-friendly format can help decision makers more quickly assess impacts and analyze the multitude of potentially significant metrics for which there is not yet consensus. In the U.S. and Canada, many BPS ordinances rely primarily on ENERGY STAR Portfolio Manager (ESPM) to capture building characteristics and energy and water consumption data. These datasets can then be imported into city-specific building tracking tools like the Standard Energy Efficiency Data Platform (SEED). Crucially, BPS decision makers require an efficient means of identifying buildings in priority communities to effectively allocate resources and funding. This process must integrate seamlessly with existing jurisdictional toolsets for optimal utility. This paper will demonstrate, for the case of Washington D.C.'s (the District) data, a workflow that provides actionable data for building upgrade investment prioritization in disadvantaged communities.

BPS↗

Accelerated data-driven materials science with the Materials Project

The Materials Project was launched formally in 2011 to drive materials discovery forwards through high-throughput computation and open data. More than a decade later, the Materials Project has become an indispensable tool used by more than 600,000 materials researchers around the world. This Perspective describes how the Materials Project, as a data platform and a software ecosystem, has helped to shape research in data-driven materials science. We cover how sustainable software and computational methods have accelerated materials design while becoming more open source and collaborative in nature. Next, we present cases where the Materials Project was used to understand and discover functional materials. We then describe our efforts to meet the needs of an expanding user base, through technical infrastructure updates ranging from data architecture and cloud resources to interactive web applications. Finally, we discuss opportunities to better aid the research community, with the vision that more accessible and easy-to-understand materials data will result in democratized materials knowledge and an increasingly collaborative community.

Horton, Matthew K↗

Nuclear Data Management and Analysis System Plan

The United States Department of Energy Advanced Reactor Technologies Program was formed in Fiscal Year 2015 and encompasses the Next Generation Nuclear Plant Project and Very High Temperature Reactor (VHTR) Program as they were known previously. The VHTR Program was created to support design and licensing of the first VHTR nuclear plant. Data created for and used by the program must be qualified for use, stored in a readily accessible electronic form, categorized to assure the correct data are used, and controlled to prevent data corruption or inadvertent changes. The Nuclear Data Management and Analysis System was designed to support the data needs of the VHTR Program, at the time and now the Advanced Reactor Technologies Program. Since its inception, use of the Nuclear Data Management and Analysis System has expanded to support additional projects and programs with similar requirements for control, analysis, and availability of large data sets.

99 GENERAL AND MISCELLANEOUS↗

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Event generators for high-energy physics experiments

We provide an overview of the status of Monte-Carlo event generators for high-energy particle physics. Guided by the experimental needs and requirements, we highlight areas of active development, and opportunities for future improvements. Particular emphasis is given to physics models and algorithms that are employed across a variety of experiments. These common themes in event generator development lead to a more comprehensive understanding of physics at the highest energies and intensities, and allow models to be tested against a wealth of data that have been accumulated over the past decades. A cohesive approach to event generator development will allow these models to be further improved and systematic uncertainties to be reduced, directly contributing to future experimental success. Event generators are part of a much larger ecosystem of computational tools. They typically involve a number of unknown model parameters that must be tuned to experimental data, while maintaining the integrity of the underlying physics models. Making both these data, and the analyses with which they have been obtained accessible to future users is an essential aspect of open science and data preservation. It ensures the consistency of physics models across a variety of experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Videos, photos, and AI-derived grain size data associated with “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization” under review. This data package includes five data types: 1) raw photos and videos from drone survey and walking smartphone surveys; 2) images derived from raw videos; 3) manual labeling of reference scales; 4) metadata for all images and photo resolution derived from artificial intelligence (AI) models or manual labels, 5) grain size data obtained from AI models for all photos, 6) metadata and grain size data after quality control, 7) summaries of sample efficiency for all data, and 8) computational fluid dynamics (CFD) data used to support hydro-biogeochemical (HBGC) parameter estimation. Such data is used to 1) demonstrate significant improvements in accuracy, efficiency, and quality control for grain size data collection with the help of AI models, 2) study the spatial heterogeneity of grain size and observation reproducibility based on tens of thousands of data points generated by the AI models, and 3) evaluate the impacts of grain size heterogeneity on key HBGC parameters across sediment-to-reach and hourly-to-yearly scales. In particular, the data package contains 116 folders and 179696 files. The files include 41 videos in .mov format, 64047 photos in .jpg format, 13541 video-derived photos in .png format, 12747 segmentation mask data in .tif format, 12747 segmentation data in .json format, 24771 .csv files that with metadata and grain size for each individual photo as well as water depth and velocity data from CFD and observation, 51791 .txt files of raw AI predicted labels, and 11 flight record data in .srt format. The summary for all metadata and grain size statistics information is included in “Scales_V3_NG.csv” and “Statistics_V3_NG.csv”. The summary for data that pass data quality control (QC) level 0-2 is included in “QCStatistics_V3_NG.csv”. The QC level 0 represents photos whose photo resolution is positive, excluding photos that miss reference scale. The QC level 1 means reference scale circularity uncertainty is less than 5% for smartphone images while representing photo resolution is larger than 0.44 mm/pixel for drone images. The QC level 2 means excluding photos whose grain number is less than 100, a minimum number of grains recommended by classic literature. The summary for each video’s name, length, frame rates, survey area, grain number, survey efficiency, etc. can be found in “QCSummary_V3_NG.csv”. The summary for site name, GPS coordinates, and number of images at each site can be found in “SitesSummary_V3_*.csv” files. Overall computational efficiency summary is reported in Table 4 of accompanying manuscript. Additionally, the nitrate concentration data used in this work was downloaded from an existing dataset published on ESS-DIVE (Boat-Dragged Sensor Hanford Reach.csv; Conner A. et al., 2020). We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Port of Benton, and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the data were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate data collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

DeepLynx Ecosystem 2025

Poor data integration and governance continue to plague complex engineering projects, resulting in missed cost, schedule, and performance targets. Departments operate in isolated systems with manual data exchange, creating fragmented information that compounds errors and leads to significant delays and cost overruns. The DeepLynx ecosystem addresses these challenges through an open-source, modular data management platform that transforms fragmented project data into an integrated digital thread. Built on a federated microservice architecture, the ecosystem comprises seven specialized tools centered around DeepLynx Nexus, a unified data catalog with hierarchical organization and graph-based navigation capabilities. The ecosystem includes: DeepLynx Stream for real-time timeseries data ingestion from industrial sources; DeepLynx Ingest for governed data uploads with formal review workflows; DeepLynx Lattice for ontology-based entity and relationship extraction; DeepLynx Run for workflow orchestration and secure AI/ML compute; DeepLynx Visualize for 3D digital twin visualization; and DeepLynx Insight for AI-assisted document analysis with traceable, grounded responses. Deployable in cloud, on-premise, or hybrid environments using containerized Docker applications and Helm charts, the DeepLynx ecosystem provides flexible infrastructure that adapts to organizational requirements. By consolidating project data into a unified data lake with role-based access controls and OAuth2 authentication, DeepLynx enables digital thread and digital twin capabilities that improve decision-making, reduce risk, and support complex engineering workflows throughout the project lifecycle.

42 - ENGINEERING↗

Electricity Baseline 2022

The Electricity Baseline (2022) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2022" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569193. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; data inventory↗