Search NASASearch

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Community standards and future opportunities for synthetic communities in plant–microbiota research

Harnessing beneficial microorganisms is seen as a promising approach to enhance sustainable agriculture production. Synthetic communities (SynComs) are increasingly being used to study relevant microbial activities and interactions with the plant host. Yet, the lack of community standards limits the efficiency and progress in this important area of research. Here, to address this gap, we recommend three actions: (1) defining reference SynComs; (2) establishing community standards, protocols and benchmark data for constructing and using SynComs; and (3) creating an infrastructure for sharing strains and data. We also outline opportunities to develop SynCom research through technical advances, linking to field studies, and filling taxonomic blind spots to move towards fully representative SynComs.

59 BASIC BIOLOGICAL SCIENCES

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE

MRCI Subtask 2.3: Developing Industrial Partnerships and Regional Technical Collaboration Final Technical Summary Report

Under the objective of regional data collection and helping accelerate deployment, MRCI collaborated with industrial stakeholders in their project planning, characterization, and analysis. Some examples of these collaborations are given below. The data and information shared by the industrial collaborations added to the regional CCS framework development and were incorporated into the overall datasets, while addressing any proprietary data requirements. Three examples of collaborative partnerships with industry that have provided geologic characterization data relevant and beneficial to the MRCI program are discussed below, including: the UIC Class II Injection Facility in Eastern Ohio, the Core Energy CO2-EOR (enhanced oil recovery) operation in Otsego County Michigan, and the Marquis ethanol plant in Hennepin Illinois.

CCS,CCUS,MRCI,Midwest USA,Technical Challenges,inj

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM

Evaluating a Commercial Dynamic Line Rating Software with the National PMU Dataset

To accelerate the development of data-driven applications for power systems, the Department of Energy (DOE) supported the collection and curation of a synchrophasor dataset spanning two years of observations from transmission utilities across the US. This National PMU Dataset (NPDS) was anonymized and distributed to awardees of a DOE research grant under nondisclosure agreements (NDAs) but has also been retained at PNNL to enable further research. Agreements with data contributors prevent the data from being shared outside the organization. However, establishing a blind research validation methodology is envisioned to maximize the value proposition of the NPDS. In this validation strategy, researchers may share algorithms/software (potentially as executables to protect intellectual property) with PNNL, and PNNL will share feedback about the software’s performance on subsets of the NPDS. Such a blind methodology ensures that sensitive information about critical infrastructure remains protected, but the value of the NPDS can be extended to research beyond PNNL. Through iterative feedback, the algorithms may be tweaked to address real-world artifacts. As the NPDS data is temporally and geographically diverse, it may capture features absent in smaller datasets used during the development of the algorithm under test. This report presents lessons learned from applying the blind validation methodology to LineID™, a synchrophasor-based dynamic line rating software developed by Topolonet Corporation. Improvements made to the software through iterative feedback, limitations of the validation methodology, as well as how the limitations of the NPDS affected the evaluation process are discussed. Observations indicate that the proposed validation methodology can be valuable for evaluating other tools in the future.

97 MATHEMATICS AND COMPUTING

Physical, socio-psychological, and behavioural determinants of household energy consumption in the UK

Determining which attitudes and behaviours predict household energy consumption can help accelerate the low-carbon energy transition. Conventional approaches in this domain are limited, often relying on survey methods that produce data on individuals’ motivations and self-reported activities without pairing these with actual energy consumption records, which are particularly hard to collect for large, nationally representative samples. This challenge precludes the development of empirical evidence on which attitudes and behaviours influence patterns of energy consumption, thus limiting the extent to which these can inform energy interventions or conservation programs. This study demonstrates a novel methodology for estimating energy consumption in the absence of actual energy records by using a large, publicly available data set of energy consumption in the UK. We develop a predictive model using the Smart Energy Research Laboratory (SERL) data portal (with records from nearly 13,000 UK households) and then use this model to predict energy consumption (both electric and gas) for a sample of 1,000 UK householders for which we separately collect over 200 variables relating to climate change attitudes and practices. Our approach uses a set of over 50 independent variables that are shared between the data sets, allowing us to train a model on the SERL data and use it to analyse the relationship between energy consumption and the opinions, motivations, and daily practices of survey respondents. Results show that electricity consumption is influenced by a broader range of factors compared to gas. Household energy use is best explained by physical dwelling characteristics, socio-demographic variables, and certain behavioural and attitudinal measures. Notably, pro-environmental attitudes, frugality, and conscientiousness correlate with lower energy use, while income and consumerism are linked to higher consumption. We discuss how these findings can inform efforts to decarbonise home energy use in the UK.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION

Dataset 1: A National and City Dataset on Human Factors in Pooled Rideshare, 2021

Dataset 1: A National and City Dataset on Human Factors in Pooled Rideshare, 2021. Dataset Description: Pooled Rideshare Acceptance Survey - Phase 1 (2021, N = 5,385). This dataset captures responses from a nationally representative sample of 5,385 adults across the United States to understand public acceptance, preferences, and behavioral intentions related to pooled rideshare (PR) services. The primary objective of this research is to provide actionable insights to inform the design, deployment, and policy development of sustainable shared mobility systems. Data was collected via an online survey administered through a national panel provider. Participants ranged in age from 18 to 95 years, and representation from all U.S. regions. The survey instrument was designed to explore numerous dimensions related to PR adoption including demographic traits, current travel habits, rideshare familiarity, trust, safety, environmental attitudes, and user experience preferences. Both rideshare users and non-users were included, offering a diverse range of perspectives. - Phase_1_Final - The dataset includes survey items developed from literature reviews, and prior field studies. Each row represents an individual respondent, and each column corresponds to a variable such as willingness to use pooled rideshare, attitudes toward specific service features, and sociodemographic data. The data is available in both .CSV and .SAV formats. - Phase_1_Final_MapFile - The accompanying data dictionary explains all variable labels, response scales, and codes. An .XLSX format of the full survey instrument is also included to support interpretation and reuse of the dataset.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product

Distributed and Secure Spectrum Sharing for 5G and 6G Networks

Secure spectrum sharing or spectrum co-existence of multiple 5G networks and future 6G networks is a powerful enabler technology. The National Spectrum Strategy (NSS) published by the White House in November, 2023, and the subsequent NSS implementation plan led by the National Telecommunication and Information Administration (NTIA) is the driver of a national effort to enable co-existence of government incumbents and commercial networks in selected spectrum bands. Cellular networks such as 5G & 6G and non-cellular Wi-Fi 6E & 7 are the prominent wireless technologies considered for co-existence with incumbent wireless links. Security of the spectrum sharing solutions is a must to make this transformation of spectrum use possible, specially for mission critical communications. However, current spectrum sharing solutions rely on centralized data bases with inherent vulnerabilities. This paper focuses on secure spectrum sharing among multiple 5G networks using unlicensed and shared frequency bands. It presents an innovative AI/ML based distributed spectrum sharing approach that can be autonomously used by multiple networks. Each sharing network uses its own observation of the Radio Frequency (RF) environment, which consists of RF measurements reported from the 5G User Equipment (UE), to adjust the transmission power levels for secure co-existence. Data is presented to illustrate the superior performance of this solution compared to other spectrum sharing solutions where each network can utilize usage data of the other networks. Finally it discusses how this efficient spectrum sharing solution can evolve in the future for the 6G networks.

5G

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan

Large language models for batteries

Large Language Models (LLMs) are advanced artificial intelligence systems capable of solving diverse tasks using language, reasoning, and external tools. Despite their growing deployment in academia and industry, their potential remains underexplored in battery research. This review presents a comprehensive overview of existing and emerging applications of LLMs in batterie field, addressing two critical questions: What can LLMs offer to support battery-related tasks, and how to develop more effective models for this purpose. We begin by outlining the principles of LLMs and criteria for selecting appropriate models and tools for battery research and development. We then explore their roles in text-mining, data interpretation, and the development of intelligent battery systems. In parallel, we discuss technical challenges, such as data standardizing and sharing, model evaluation, and tool integration. Lastly, we propose future research directions with short-, medium-, and long-term goals and highlight more broad perspectives for connecting experts and cross-disciplinary collaborations.

SoC

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES

STREAMS guidelines: standards for technical reporting in environmental and host-associated microbiome studies

The interdisciplinary nature of microbiome research, coupled with the generation of complex multi-omics data, makes knowledge sharing challenging. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines provide a checklist for the reporting of study information, experimental design and analytical methods within a scientific manuscript on human microbiome research. Here, in this Consensus Statement, we present the standards for technical reporting in environmental and host-associated microbiome studies (STREAMS) guidelines. The guidelines expand on STORMS and include 67 items to support the reporting and review of environmental (for example, terrestrial, aquatic, atmospheric and engineered), synthetic and non-human host-associated microbiome studies in a standardized and machine-actionable manner. Based on input from 248 researchers spanning 28 countries, we provide detailed guidance, including comparisons with STORMS, and case studies that demonstrate the usage of the STREAMS guidelines. In conclusion, STREAMS, like STORMS, will be a living community resource updated by the Consortium with consensus-building input of the broader community.

59 BASIC BIOLOGICAL SCIENCES