Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

EMP - Environmental Radiological Air Monitoring Plan: PNNL Operations in Washington

The Environmental Radiological Air Monitoring Plan (EMP) for Pacific Northwest National Laboratory (PNNL) describes systems/processes/practices related to radiological operations in Richland and Sequim, Washington, that are associated with environmental radiological air monitoring and surveillance activities. The activities described support the lab’s responsibility to maintain safe operations and minimize negative impacts to both onsite and offsite persons and environment. Dose assessments required by regulations and DOE Orders for the public and biota are described. PNNL conducts environmental air surveillance monitoring as part of the PNNL Site Radioactive Air Emissions License (RAEL)-005 for Richland Campus, issued in 2010 with its most recent renewal effective in January 2021. The radioactive air emissions license for the PNNL-Sequim campus (RAEL-014) was issued to the U.S. Department of Energy in 2012 with its most recent renewal effective in January 2023. The EMP is a compilation of the following four documents: - Environmental Radiological Air Monitoring Plan (this main document) (PNNL-20919) - Sampling and Analysis Plan (Attachment 1) (PNNL-20919-1) - Data Management Plan (Attachment 2) (PNNL-20919-2) - Dose Assessment Guidance (Attachment 3) (PNNL-20919-3).

40 CFR 61 Subpart H↗

Graphite Oxidation Rate Study on ET-10 and ETU-10 Grades - Task 4: QA Support and Testing for Structural Graphite Oxidation

INL performed targeted oxidation tests to measure oxidation rates for samples of ET-10 and ETU-10 graphite under CRADA No. 21CRA22 Mod. 3, Annex A, “Tritium Testing to Support Kairos Power Advanced Reactor Demonstration” (04/02/2024). All testing was conducted within INL’s Carbon Characterization Laboratory (CCL) using test standard ASTM D7542-21 "Standard Test Method for Air Oxidation of Carbon and Graphite in the Kinetic Regime" [ASTM International, 2021]. Kairos Power provided all test specimens through its graphite vendor Ibiden, Inc. to INL and ASTM specimen specified dimensions. Information within this report only provides the Arrhenius oxidation rate plots as a function of temperature for each graphite grade tested. The raw mass loss per time data will be provided on the Nuclear Data Management and Analysis System (NDMAS) portal located on the INL information system.

36 MATERIALS SCIENCE↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Just-In-Time Workflow Management for DUNE

The poster describes the justIN workflow management system funded by the UK for DUNE and now used for all DUNE centrally managed data processing and simulation

McNab, Andrew [CERN]↗

Multi-Entity Simulation with CoSim Toolbox

Co-simulation is an analysis technique for linking multiple software models during runtime by facilitating data exchange and simulation time synchronization. There are numerous challenges when constructing an effective co-simulation including simulation tool installation, data management, and writing new models in a manner compatible with the co-simulation framework of choice. CoSim Toolbox is an integration of multiple pieces of software designed to make assembling such a co-simulation in HELICS easier. This report summarizes the existing capabilities of CoSim Toolbox and outlines future development plans.

97 MATHEMATICS AND COMPUTING↗

Powered By CADET

The Capacity Expansion Decision Support for Distribution Networks (CADET) is a Python-based library and framework for creating electrical distribution system capacity planning tools for cost-effective, reliable power delivery. It enables the creation of modular, scalable, and extensible distribution capacity planning tools by providing a high-level optimization interface, parameter and options data managers, optimization constraint and objective libraries, generalized nomenclature, a system for tracking and modifying distribution network changes, optimization solution validation, and other capabilities. This webinar will describe 1) the motivation for creating CADET, 2) key designs, and 3) several use cases.

24 POWER TRANSMISSION AND DISTRIBUTION↗

From Raw to Curated Data: A Lakehouse Approach for Scientific Workflows

This report provides a technical overview of how to go from raw to curated data in three stages using a lakehouse approach. We focus on the application of open source tools in scientific use cases (while noting parallels to enterprise and commercial alternatives). Our goal is to provide scientific data managers and infrastructure providers with a common frame of reference for understanding and applying modern lakehouse technologies and approaches.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

District Geothermal Heating + Cooling Deployment in a CT Environmental Justice Community

The report marks the team’s completion of all required tasks and milestones. Work completed for Task 1 (Technical and Economic Feasibility Assessment & Procurement Drafting) included development of analysis and design model; completion of technical, economic, and environmental assessments; and technical outreach and coalition design. Components for Task 2 (Outreach & Community Engagement) involved broad outreach and community-engagement efforts (including stakeholder meetings and a webinar as well as development of a formal engagement plan) and development of a web page and a case study. For Task 3 (Workforce Transition, Development, & Training Plan), the team undertook a formal statewide geothermal workforce needs assessment, developed corresponding recommendations for both the state as a whole and the Wallingford project, and held several workshops. For Task 4 (Project Management & Data Sharing), the team drafted a data-sharing plan.

15 GEOTHERMAL ENERGY↗

Scientific Data Compression for Large Scale Computational Fluid Dynamics (CFD) Simulations

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and General Electric (GE) investigated methods for reducing the size of large computational fluid dynamics (CFD) simulation datasets using scientific data compression techniques. The work focused on adapting the MultiGrid Adaptive Reduction of Data (MGARD) compression framework and integrating it with high-performance I/O and visualization tools used in CFD workflows. MGARD uses hierarchical multilevel decomposition to enable error-controlled compression of floating-point scientific data while preserving quantities of interest. During the project, MGARD compression was integrated with the ADIOS I/O framework and visualization tools such as ParaView to enable efficient storage, transfer, and analysis of simulation data. The collaboration also explored approaches for improving compression performance for CFD data defined on unstructured meshes. Results demonstrate that scientific data compression can significantly reduce storage requirements and improve data management for large-scale CFD simulations.

97 MATHEMATICS AND COMPUTING↗

PRIMO – The Oil & Gas Well Plugging Optimizer

This work presents PRIMO’s main capabilities and introduces the PRIMO web application, an intuitive user interface that leverages our sophisticated mathematical optimization model to rigorously optimize P&A priorities and plugging campaign efficiency. The web app simplifies user interaction, provides a powerful data management framework, and supports a broad user base (e.g., state agencies, well owners/operators, and plugging companies) to use PRIMO for decision-making. Specifically, we provide a demonstration of how to input the information on candidate wells, plugging campaign budget, user-defined priority and efficiency criteria to PRIMO. A real-world case study that consists of 1411 oil and gas wells and impact and efficiency priorities (e.g., well age, well proximity to schools/hospitals, well accessibility, distance between wells in projects) is presented to showcase PRIMO’s core capabilities: (i) ranking a candidate well population based on priorities, (ii) recommending high-impact and high-efficiency P&A projects, and (iii) assigning impact and efficiency scores to projects allowing for rigorous quantitative comparison among them.

02 PETROLEUM↗

Unalakleet Microgrid Optimization for Tribal Community Resilience

The Unalakleet Microgrid Optimization Project aimed to strengthen the reliability and efficiency of the isolated electric power system that serves the Tribal community of Unalakleet, Alaska. The community relies entirely on a local wind-diesel microgrid, consisting of four 475 kW diesel generators and six 100 kW wind turbines, to provide electricity to approximately 745 residents and Tribal facilities. Because Unalakleet is not on a road system and is located nearly 400 miles from the nearest major power grid, maintaining a resilient and efficient local energy system is critical. The scope of this project included upgrading a portion of the transmission line between the wind farm and the power plant to increase voltage and reduce line losses, along with modernizing the Supervisory Control and Data Acquisition (SCADA) system to improve monitoring, control, and data management of the power system. These upgrades were designed to increase wind energy utilization, reduce diesel fuel consumption by tens of thousands of gallons annually, and improve overall grid stability. By allowing more of the community’s electricity to be supplied by local renewable wind resources, the project was designed to lower operating costs, reduce dependence on imported fuel, and strengthen the long-term resilience of the power system. These improvements represent an important step toward the community’s long-term energy vision of expanding renewable generation, incorporating energy storage, and eventually achieving “diesels-off” operation, where essential Tribal loads are powered primarily by local renewable resources.

17 WIND ENERGY↗

Evaluation of AI-Enabled Digital Documented Safety Analysis: A Case Study

Safety basis documentation development and review under U.S. Department of Energy (DOE) authorization have emerged as critical constraint throttling deployment of advanced nuclear reactors, with traditional processes demanding extraordinary resource investment that delays the delivery of these technologies. Traditional Documented Safety Analysis (DSA) processes rely on static documents with limited traceability [U.S. DOE]. The regulatory review and engagement processes are similarly constrained, often requiring significant effort and extensive manual verification. The scale of this challenge is exemplified by the U.S. Nuclear Regulatory Commission (NRC) review of the NuScale application, which required over 250,000 staff hours and the evaluation of approximately two million pages of documentation [Bergman 2021]. The volume and complexity of information within nuclear licensing applications or authorization reviews demands innovative approaches to document generation and data management.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗