Search NASASearch

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES

Summary Report of the FY24 DOE Contributions to the GIF VHTR CMVB

The Generation-IV Forum (GIF) Very-High-Temperature Reactor-Computational Methods Validation and Benchmark (VHTR-CMVB) initiative, involving organizations from Korea Atomic Energy Research Institute (KAERI) (South Korea), Institute of Nuclear and New Energy Technology of Tsinghua University (INET) (China), U.S. Department of Energy (DOE) (U.S.), Joint Research Centre (JRC) (Europe), and Japan Atomic Energy Agency (JAEA) (Japan), is dedicated to the verification and validation of tools for High-Temperature Gas-Cooled Reactors (HTGRs) analysis, using data shared by Computational Methods Validation and Benchmark (CMVB) signatories. For FY24, the US DOE CMVB has committed to several critical activities. Under WP1, led by the US, the integration of the High Temperature Gas Cooled Reactor - Pebble-Bed Module (HTR-PM) Phenomena Identification and Ranking Table (PIRT) into the comparative PIRT is progressing, with a new draft of the comparison tables issued earlier this year and currently being utilized by INET for their contribution. Neutronic validation efforts under WP3 include the preparation of the burnup analysis benchmark, preliminary calculations, and the development of reference models and results. In WP2, a validation exercise for hot gas mixing in the lower plenum of HTR-PM is in progress, using experimental data from INET (China) to validate modeling approaches. A model of the experimental facility has been developed using StarCCM+, with initial calculations slated for presentation at the GIF CMVB meeting this fall. Another WP2 activity focuses on validating numerical models for air-cooled Reactor Cavity Cooling System (RCCS) with experimental data from the Wisconsin Madison RCCS facility. A high-fidelity model, developed using NEK-RS, is currently being validated with available data from a low power forced convection test. These efforts are aimed at enhancing and confirming the accuracy of HTGR analysis tools, ensuring their alignment with experimental data and regulatory requirements.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Utilizing AI and Spatial Data to Identify & Rapidly Disseminate Energy Infrastructure Insights

GeoGov Summit Final Presentation entitled "Utilizing AI and Spatial Data to Identify & Rapidly Disseminate Energy Infrastructure Insights". Maintaining the integrity of energy infrastructure plays a critical role in ensuring energy security. Robust foundational AI models using data from federal, state, industry, and other sources can help address integrity risk management & mitigation issues as well as evaluate extended use strategies. Trusted foundational models can help with industry adoption and accelerate innovation by enhancing integrity predictions, reduce costs, and informing infrastructure build-out. Coordination, collaboration & data sharing to develop robust models to aid in: Optimizing operations; Minimizing costs; Ensuring energy security.

Advanced Infrastructure Integrity Model (AIIM)

Exponential Backoff and Its Security Implications for Safety-Critical OT Protocols over TCP/IP Networks

The convergence of Operational Technology (OT) and Information Technology (IT) networks has become increasingly prevalent with the growth of Industrial Internet of Things (IIoT) applications. This shift, while enabling enhanced automation, remote monitoring, and data sharing, also introduces new challenges related to communication latency and cybersecurity. Oftentimes, legacy OT protocols were adapted to the TCP/IP stack without an extensive review of the ramifications to their robustness, performance, or safety objectives. To further accommodate the IT/OT convergence, protocol gateways were introduced to facilitate the migration from serial protocols to TCP/IP protocol stacks within modern IT/OT infrastructure. However, they often introduce additional vulnerabilities by exposing traditionally isolated protocols to external threats. This study investigates the security and reliability implications of migrating serial protocols to TCP/IP stacks and the impact of protocol gateways, utilizing two widely used OT protocols: Modbus TCP and DNP3. Our protocol analysis finds a significant safety-critical vulnerability resulting from this migration, and our subsequent tests clearly demonstrate its presence and impact. A multi-tiered testbed, consisting of both physical and emulated components, is used to evaluate protocol performance and the effects of device-specific implementation flaws. Through this analysis of specifications and behaviors during communication interruptions, we identify critical differences in fault handling and the impact on time-sensitive data delivery. The findings highlight how reliance on lower-level IT protocols can undermine OT system resilience, and they inform the development of mitigation strategies to enhance the robustness of industrial communication networks.

DNP3

Position-Enhanced Gradient Attack (PEGA) on Medical Language Models

Federated Learning (FL) enables collaborative training of language models on sensitive clinical notes without sharing the data. However, this paradigm is vulnerable to gradient inversion attacks that can reconstruct private data from shared gradients. We find that state-of-the-art attacks are less effective in the medical domain, failing to overcome the unique challenges posed by its specialized vocabulary and unstructured format. To address this, we introduce the Position-Enhanced Gradient Attack (PEGA), a novel attack that makes gradients position-aware by optimizing token and position embeddings simultaneously. PEGA employs two key innovations: a periodic sorting of positional embeddings to resolve token order ambiguity and a late-stage embedding replacement strategy to correct hard-to-recover critical tokens. To evaluate the leakage of sensitive data more directly, we also propose the Unified PHI-Recall (UPHI), a new metric measuring the recovery of Protected Health Information. Experiments on the MIMIC-III dataset show that PEGA significantly outperforms leading attacks like TAG and LAMP, particularly in its ability to reconstruct identifiable patient information, exposing a more severe and nuanced privacy risk in federated medical NLP.

Xu, Nuo [University of Minnesota]

SAM-I-Am: Semantic boosting for zero-shot atomic-scale electron micrograph segmentation

Image segmentation is a critical enabler for tasks ranging from medical diagnostics to autonomous driving. However, the correct segmentation semantics — where are boundaries located? what segments are logically similar? — change depending on the domain, such that state-of-the-art foundation models can generate meaningless and incorrect results. Moreover, in certain domains, fine-tuning and retraining techniques are infeasible: obtaining labels is costly and time-consuming; domain images (micrographs) can be exponentially diverse; and data sharing (for third-party retraining) is restricted. To enable rapid adaptation of the best segmentation technology, we propose the concept of semantic boosting: given a zero-shot foundation model, guide its segmentation and adjust results to match domain expectations. Here, we apply semantic boosting to the Segment Anything Model (SAM) to obtain microstructure segmentation for transmission electron microscopy. Our booster, SAM-I-Am, serves as a post-processing engine that extracts geometric and textural features of various intermediate masks to perform mask removal and mask merging operations. We demonstrate a zero-shot performance increase of (absolute) +21.35%, +12.6%, +5.27% in mean IoU, and a -9.91%, -18.42%, -4.06% drop in mean false positive masks across images of three difficulty classes over vanilla SAM (ViT-L).

36 MATERIALS SCIENCE

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES

White paper on light sterile neutrino searches and related phenomenology

This white paper provides a comprehensive review of our present understanding of experimental neutrino anomalies that remain unresolved, charting the progress achieved over the last decade at the experimental and phenomenological level, and sets the stage for future programmatic prospects in addressing those anomalies. It is purposed to serve as a guiding and motivational "encyclopedic" reference, with emphasis on needs and options for future exploration that may lead to the ultimate resolution of the anomalies. We see the main experimental, analysis, and theory-driven thrusts that will be essential to achieving this goal being: 1) Cover all anomaly sectors -- given the unresolved nature of all four canonical anomalies, it is imperative to support all pillars of a diverse experimental portfolio, source, reactor, decay-at-rest, decay-in-flight, and other methods/sources, to provide complementary probes of and increased precision for new physics explanations; 2) Pursue diverse signatures -- it is imperative that experiments make design and analysis choices that maximize sensitivity to as broad an array of these potential new physics signatures as possible; 3) Deepen theoretical engagement -- priority in the theory community should be placed on development of standard and beyond standard models relevant to all four short-baseline anomalies and the development of tools for efficient tests of these models with existing and future experimental datasets; 4) Openly share data -- Fluid communication between the experimental and theory communities will be required, which implies that both experimental data releases and theoretical calculations should be publicly available; and 5) Apply robust analysis techniques -- Appropriate statistical treatment is crucial to assess the compatibility of data sets within the context of any given model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data

Abstract BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB’s positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset and plans for future growth and development.

Liu, Tiqing

Netload Range Cost Curves for Coordinated Transmission-Distribution Planning Under DER Growth Uncertainty

The increasing penetration of distributed energy resources (DERs) requires better coordination between transmission and distribution (T&D) planning to ensure system security and cost efficiency. However, misaligned planning horizons, computational burdens, and privacy concerns hinder effective coordination, leading to either underutilized resources caused by overinvestments or reliability risks due to underinvestment. To address this challenge, we introduce netload range cost curves (NRCCs), a novel approach for managing long-term DER growth uncertainty through T&D coordination, while preserving existing data-sharing and regulatory structures. NRCCs provide pairs of (i) peak substation netload guarantees and (ii) corresponding distribution upgrade options and costs, enabling their seamless integration into transmission planning workflows. To compute NRCCs efficiently, we develop a transmission-aware distribution network planning (TADNP), which is subsequently integrated to an iterative computation procedure. These NRCCs are then embedded into an NRCC-informed transmission planning model to enable resource-efficient coordination. We illustrate our proposed approach with a case study based on realistic distribution and transmission systems in the San Francisco Bay Area, California. Our results indicate the possibility of dramatic savings in transmission investments by incorporating the proposed NRCC-integrated T&D coordination framework.

Li, Yujia

Resilient Control of Networked Microgrids Using Vertical Federated Reinforcement Learning: Designs and Real-Time Test-Bed Validations

Improving system-level resiliency of networked microgrids against adversarial cyber-attacks is an important aspect in the current regime of increased inverter-based resources (IBRs). To achieve that, this paper contributes in designing a hierarchical control layer, in conjunction with the existing control layers, resilient to adversarial attack signals. Considering model complexities, unknown dynamical behaviors of IBRs, and privacy issues regarding data sharing in multi-party-owned microgrids, designing such a control layer is non-trivial. Here, to tackle these issues, a novel federated reinforcement learning (Fed-RL) method is proposed. To grasp the interconnected dynamics of networked microgrids, the paper develops Federated Soft Actor-Critic (FedSAC) algorithm following the vertical structure of implementing Fed-RL. Next, utilizing the OpenAI Gym interface, we built a custom set-up in GridLAB-D/HELICS co-simulation platform, named Resilient RL Co-simulation (ResRLCoSIM), to train the RL agents with IEEE 123-bus benchmark comprising 3 interconnected microgrids. Finally, the learned policies in the simulation are transferred to the real-time hardware-in-the-loop (HIL) test-bed developed using the high-fidelity Hypersim platform. Finally, experiments show that the simulator-trained RL controllers achieve desirable performance with the test-bed platform, validating the minimization of the sim-to-real gap.

24 POWER TRANSMISSION AND DISTRIBUTION

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop

Deep Design Data Portal (D3P) v0.01

The Deep Design Data Portal (D3P) tool was developed to demonstrate how readily accessible data sources, such as building energy model reports for design and baseline energy performance data for projects, can provide the data required for reporting to an industry initiative (AIA 2030 commitment), as well as more detailed data that makes the industry dataset more valuable to all stakeholders, enabling project level analysis and analysis of BEM industry trends. D3P provides an easier and less time-consuming way for firms to auto-extract data from this data source, compared to the current reporting workflows of the firms. The BEM reports are the first of several data sources that D3P could integrate. D3P also provides the ability for firms to review, compare, and evaluate the performance of their projects to not only their portfolio, but also to the larger anonymized industry dataset created each time a project is added to D3P. The intent of D3P is to become part of a data-sharing ecosystem to assist creating large anonymized industry datasets that are accessible to industry.

Regnier, Cynthia [Lawrence Berkeley National Labor

The ECP SICM project: Managing complex memory hierarchies for exascale applications

The Exascale Computing Project (ECP)’s Simplified Interface to Complex Memories (SICM) effort focuses on developing universal interfaces for discovering, managing, and sharing data across complex memory hierarchies. These facilitate the exploitation of emerging memory technologies and support precise control over their various trade-offs such as high-bandwidth versus low-latency, persistent versus ephemeral, high-capacity versus low-capacity, and near-CPU versus near-GPU. SICM comprises three interrelated components: a low-level interface, a high-level interface, and a persistent-heap interface. The low-level SICM interface is intended for system and run-time developers as well as expert application developers who prefer full control of the memory objects used within their application. The high-level SICM interface builds upon the low-level interface, employing application-level profiling and analysis to optimize data management for complex memory hierarchies. The persistent-heap interface provides applications with a persistent memory allocator that can allocate custom C++ data structures in both block-storage and byte-addressable persistent memories.

97 MATHEMATICS AND COMPUTING

Radio Frequency Spectrum Audit to Inventory Private Cellular Base Station Infrastructure

The ever-changing cellular communication landscape makes it difficult to identify, map, and localize cellular base stations. Localizing cellular base stations provides various advantages, including information security, cybersecurity, spectrum management, and interference detection. For example, the MITRE ATT&CK® (Adversarial Tactics, Techniques, and Common Knowledge architecture) [1] and Common Attack Pattern Enumeration and Classification [2] emphasize the importance of being able to minimize the cyber security threat presented by unregulated private cellular base stations (PCBS). The majority of published research looks at the malicious use of PCBSs and focuses on using data retrieved from user equipment (UE), data obtained from an application on the UE, or data shared between the UE and a mobile network to locate it. This innovative strategy, however, focuses on the passively discovered uniqueness of radio frequency (RF) transmissions from commercial cellular infrastructure received in a designated monitoring position (DMP).

42 ENGINEERING

Metadata Standards for the NSE: Core Fields

This standard presents a core set of metadata fields required for each managed digital object within the Nuclear Security Enterprise (NSE). Metadata standardization is a critical enabler for two primary objectives: 1) effectively sharing data, documents, and other digital objects between NSE sites; and 2) supporting digital engineering through the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a foundational field standard, recommending a core set of fields that should be uniformly required for all managed digital objects within the NSE.

99 GENERAL AND MISCELLANEOUS

Addressing Investment Barriers by Improving Documentation of Sustainable Biomass Resources (Workshop Report)

On May 8, 2025, Oak Ridge National Laboratory (ORNL), in collaboration with IEA Bioenergy Task 43 and the Biofuture Platform, convened an international workshop in Vancouver, Canada to improve the Global Biomass Resource Assessment. This effort addresses investment barriers in the global bioeconomy by improving the transparency, consistency, and usability of biomass supply data. The workshop gathered 38 participants from 11 countries, representing government agencies, academia, and industry. Participants reviewed the status of the biomass dataset, tested the data-sharing platform, and provided direct input on priorities for improvement.

09 BIOMASS FUELS

Theorems in Service of Sound Composition, Rapid Modeling and Scalable Analysis

This project extends the state of the art in formal verification modeling with modules and automatically checkable data-sharing patterns such that component modules can retain their assurance case when composed within a larger system. For users, smaller models make reasoning easier and help to ensure they accurately reflect text specifications. For automated methods, smaller models give exponential benefits for verification algorithm execution time.

97 MATHEMATICS AND COMPUTING