Search NASA⌕ Search

SEARCH · Search NASA

Results for “research data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

Management and Storage of Scientific Data

Scientific discoveries rely heavily on efficient access, search, and management of massive data sets. Data management technologies have, for decades, provided foundational capabilities for scientific computing. Just as storage, input/output (I/O), and data management have been fundamental to simulation-based science for many years, so too are capable data-management technologies key to the success of today’s scientific workflows utilizing data intensive and machine learning (ML) techniques. The Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program has invested broadly in data-management research focused on high-performance computing (HPC) systems, from parallel file systems that store data to application software that makes these systems more productive. Still, advances in technology combined with growing diversity of supported science strongly motivate continued investment in this area. In January 2022, ASCR convened a workshop to identify priority research directions in the area of data management for high-performance and scientific computing. Attendees were challenged to identify promising approaches that would support the breadth of the DOE mission, including the explosion of artificial intelligence (AI) uses and the growing needs of experimental and observational science. Technological and science drivers were identified and considered as they relate to key aspects of data management such as interfaces, architectural design, and FAIR principles (Findable, Accessible, Interoperable, and Reusable). The thoughts of the workshop participants were distilled into a set of four priority research directions with the potential for high impact on DOE science. These research directions are summarized in the following pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Community-Driven Methods for Open and Reproducible Software Tools for Analyzing Datasets from Atom Probe Microscopy

Atom probe tomography, and related methods, probe the three-dimensional architecture of a material. The software tools that microscopists use, and how these tools are connected into workflows, makes a substantial contribution to the accuracy and precision of such a material characterization experiment. Typically, we adapt methods from other communities like mathematics, data science, computational geometry, artificial intelligence, or scientific computing. We also realize that improving on research data management is a challenge when it comes to align with the FAIR data stewardship principles. Faced with this global challenge, we are convinced that collaborating is useful. Here, we report the results and challenges with an inter-laboratory call for developing test cases for several types of atom probe software tools. The results support why defining detailed recipes of software workflows and sharing these recipes is necessary and rewarding: Open source tools and (meta)data exchange can help to make our day-to-day data processing tasks become more efficient, the training of new users and knowledge transfer become easier, and assist us with automated quantification of uncertainties to gain access to substantiated results.

36 MATERIALS SCIENCE↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

32 examples of LLM applications in materials science and chemistry: towards automation, assistants, agents, and accelerated scientific discovery

Abstract Large language models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 32 total projects developed during the second annual LLM hackathon for applications in materials science and chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.

Computer Science↗

DOE Data Days 2022 Report

The DOE Data Days (D3) workshop brings together data managers, developers, researchers, and program managers across the Department of Energy (DOE) and national laboratories to highlight data management successes, identify potential synergies and common problems, and establish channels for collaboration across the DOE data management community.

97 MATHEMATICS AND COMPUTING↗

DOE Data Days 2023 Report

The DOE Data Days (D3) workshop brings together data managers, developers, researchers, and program managers across the Department of Energy (DOE) and national laboratories to highlight data management successes, identify potential synergies and common problems, and establish channels for collaboration across the DOE data management community. The fourth D3 workshop was held on October 24th to 26th, 2023 held entirely in-person at Lawrence Livermore National Laboratory (LLNL). The workshop was organized by a multi laboratory committee in an effort to bring data management practitioners at the DOE laboratories together to share their work and results, facilitating knowledge transfers and best practices across project teams. Tools and platforms to support data management and analysis are rapidly evolving and provide enormous opportunities. This report summarizes the important discussions and recommendations from the different working sessions and contains the agenda, submitted abstracts, presentations with links to recorded presentations, breakout session summaries, list of registered attendees, and lessons learned for future organizing committee. The report will be distributed to the DOE, each participating institution’s programmatic stakeholders, and attendees. The dedicated D3 website will host presentations, agenda, and report that is accessible by all labs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

DOE Data Days 2025 Report

The DOE Data Days (D3) workshop brings together data managers, developers, researchers, and program managers across the Department of Energy (DOE) and its national laboratories to highlight data management successes, identify potential synergies and common problems, and establish channels for collaboration across the DOE data management community.

97 MATHEMATICS AND COMPUTING↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

FAIR for AI: An interdisciplinary and international community building perspective

A foundational set of findable, accessible, interoperable, and reusable (FAIR) principles were proposed in 2016 as prerequisites for proper data management and stewardship, with the goal of enabling the reusability of scholarly data. The principles were also meant to apply to other digital assets, at a high level, and over time, the FAIR guiding principles have been re-interpreted or extended to include the software, tools, algorithms, and workflows that produce data. FAIR principles are now being adapted in the context of AI models and datasets. Here, we present the perspectives, vision, and experiences of researchers from different countries, disciplines, and backgrounds who are leading the definition and adoption of FAIR principles in their communities of practice, and discuss outcomes that may result from pursuing and incentivizing FAIR AI research. The material for this report builds on the FAIR for AI Workshop held at Argonne National Laboratory on June 7, 2022.

97 MATHEMATICS AND COMPUTING↗

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES↗

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

A Lakehouse Architecture for the Management and Analysis of Heterogeneous Data for Biomedical Research and Mega-biobanks

Data Lakehouse is a new paradigm in data architectures that embodies and integrates already established concepts for the systematic management of disparate, large-scale data – a data lake for heterogeneous data management, use of open standards for high-performance querying, and systematic maintenance of the data "freshness". In addition to being a new concept, the data lakehouse is also still a conceptual construct. Many projects that use the lakehouse require maturing, empirical studies, and specific implementations. In this paper, we present our implementation of the data lakehouse concept in a biomedical research and health data analytics domain, and we discuss the implementation of some unique and novel features such as support for specialized access controls in support of HIPAA regulation and IRB protocols, and support for the FAIR standard.

Begoli, Edmon↗