Search NASASearch

SEARCH · Search NASA

Results for “Advanced Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Bringing Research to New Heights: How CASEI Integrates Data Curation, Discovery, and Education in Earth and Atmospheric Science

A challenging aspect of any project is finding all the relevant data and information needed to address the research objective. Searching for data and its contextual metadata can become overwhelming for both undergraduate and graduate students, potentially hindering their work and affecting the scientific discoveries that could be made in the long run. To ease this, the NASA Airborne Data Management Group (ADMG), part of the Interagency Implementation and Advanced Concepts Team (IMPACT), has developed the new Catalog of Archived Suborbital Earth science Investigations (CASEI). CASEI includes a web portal that users, be they professionals or students, can use to search, browse, discover, and locate relevant observations associated with NASA’s airborne and field campaigns. Users are able to query data in a variety of ways (via keywords, locations, timeframe, etc) from one online portal, minimizing the amount of time needed to search. CASEI also allows access to key contextual metadata and data from a wide array of Earth and Atmospheric Science topics such as aerosols and boundary layer processes, as well as ice and glacial properties or processes. Users are able to access the data via DOI links to data set landing pages. This presentation will demonstrate how CASEI can be used for classwork and student research. Teachers can provide CASEI to their students as a tool for their studies, or use it to find data themselves while constructing their curriculums. Additionally, users can leverage CASEI to learn about NASA’s Earth and Atmospheric Science research efforts and to find data relevant for assignments or other research projects. The metadata in CASEI has been carefully curated, and highlights important information about the campaigns and their data. Students can explore and learn about the scientific objectives of the campaigns, as well as descriptions of the campaign’s best research days. Having access to contextual metadata in an easy to understand way can help plant the seeds of new ideas in students at any point in their academic journey. From class projects to theses/dissertations and other research, CASEI is a valuable emerging tool for data discovery, giving access to all users and guiding researchers to NASA’s unique airborne data to answer the burning Earth Science questions of our time.

education

Status on Microstructural Studies of Nuclear Graphite

This report documents the completion of the Advanced Reactor Technologies (ART) Level 3 Milestone (M3AT-26OR0605054), “Provide status on microstructural studies of nuclear graphite,” due May 1, 2026. The contents of this report summarize recent progress in the development of the Nuclear Graphite Microstructure Library, including datasets that have been generated, curated, and submitted for publication, as part of the broader effort to establish and integrate the library within the NDMAS platform.

22 GENERAL STUDIES OF NUCLEAR REACTORS

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Hayabusa Recovery, Curation and Preliminary Sample Analysis: Lessons Learned from Recent Sample Return Mission

I describe lessons learned from my participation on the Hayabusa Mission, which returned regolith grains from asteroid Itokawa in 2010 [1], comparing this with the recently returned Stardust Spacecraft, which sampled the Jupiter Family comet Wild 2. Spacecraft Recovery Operations: The mission Science and Curation teams must actively participate in planning, testing and implementing spacecraft recovery operations. The crash of the Genesis spacecraft underscored the importance of thinking through multiple contingency scenarios and practicing field recovery for these potential circumstances. Having the contingency supplies on-hand was critical, and at least one full year of planning for Stardust and Hayabusa recovery operations was necessary. Care must be taken to coordinate recovery operations with local organizations and inform relevant government bodies well in advance. Recovery plans for both Stardust and Hayabusa had to be adjusted for unexpectedly wet landing site conditions. Documentation of every step of spacecraft recovery and deintegration was necessary, and collection and analysis of launch and landing site soils was critical. We found the operation of the Woomera Text Range (South Australia) to be excellent in the case of Hayabusa, and in many respects this site is superior to the Utah Test and Training Range (used for Stardust) in the USA. Recovery operations for all recovered spacecraft suffered from the lack of a hermetic seal for the samples. Mission engineers should be pushed to provide hermetic seals for returned samples. Sample Curation Issues: More than two full years were required to prepare curation facilities for Stardust and Hayabusa. Despite this seemingly adequate lead time, major changes to curation procedures were required once the actual state of the returned samples became apparent. Sample databases must be fully implemented before sample return for Stardust we did not adequately think through all of the possible sub sampling and analytical activities before settling on a database design - Hayabusa has done a better job of this. Also, analysis teams must not be permitted to devise their own sample naming schemes. The sample handling and storage facilities for Hayabusa are the finest that exist, and we are now modifying Stardust curation to take advantage of the Hayabusa facilities. Remote storage of a sample subset is desirable. Preliminary Examination (PE) of Samples: There must be some determination of the state and quantity of the returned samples, to provide a necessary guide to persons requesting samples and oversight committees tasked with sample curation oversight. Hayabusa s sample PE, which is called HASPET, was designed so that late additions to the analysis protocols were possible, as new analytical techniques became available. A small but representative number of recovered grains are being subjected to in-depth characterization. The bulk of the recovered samples are being left untouched, to limit contamination. The HASPET plan takes maximum advantage of the unique strengths of sample return missions

Zolensky, Michael E.

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY

SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

The rapid advancement of generative models has made the detection of AI-generated images a critical challenge for both research and society. Recent works have shown that most state-of-the-art fake image detection methods overfit to their training data and catastrophically fail when evaluated on curated hard test sets with strong distribution shifts. In this work, we argue that it is more principled to learn a tight decision boundary around the real image distribution and treat the fake category as a sink class. To this end, we propose SimLBR, a simple and efficient framework for fake image detection with Latent Blending Regularization (LBR). Our method significantly improves cross-generator generalization, achieving up to +24.85% accuracy and +69.62% recall on the challenging Chameleon benchmark. SimLBR is also highly efficient, training orders of magnitude faster than existing approaches. Furthermore, we emphasize the need for reliability-oriented evaluation in fake image detection, introducing risk-adjusted metrics and worst-case estimates to better assess model robustness. All the code and models are availabe at: https://github.com/mvrl/SimLBR

Dhakal, Aayush [Washington University, St. Louis]

Genesis Solar Wind Interstream, Coronal Hole and Coronal Mass Ejection Samples: Update on Availability and Condition

Recent refinement of analysis of ACE/SWICS data (Advanced Composition Explorer/Solar Wind Ion Composition Spectrometer) and of onboard data for Genesis Discovery Mission of 3 regimes of solar wind at Earth-Sun L1 make it an appropriate time to update the availability and condition of Genesis samples specifically collected in these three regimes and currently curated at Johnson Space Center. ACE/SWICS spacecraft data indicate that solar wind flow types emanating from the interstream regions, from coronal holes and from coronal mass ejections are elementally and isotopically fractionated in different ways from the solar photosphere, and that correction of solar wind values to photosphere values is non-trivial. Returned Genesis solar wind samples captured very different kinds of information about these three regimes than spacecraft data. Samples were collected from 11/30/2001 to 4/1/2004 on the declining phase of solar cycle 23. Meshik, et al is an example of precision attainable. Earlier high precision laboratory analyses of noble gases collected in the interstream, coronal hole and coronal mass ejection regimes speak to degree of fractionation in solar wind formation and models that laboratory data support. The current availability and condition of samples captured on collector plates during interstream slow solar wind, coronal hole high speed solar wind and coronal mass ejections are de-scribed here for potential users of these samples.

Allton, J. H.

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B

Spaceflight Environmental-Telemetry Data for Biological Science

There is a critical need for better access and visualization of spaceflight environmental telemetry and mission hardware data from sensors including relative humidity, carbon dioxide, oxygen, radiation, airflow, temperature, acceleration, and acoustics. Under the stewardship of the Ames Life Sciences Data Archive (ALSDA) and GeneLab, an effort is underway to consolidate, normalize and provide accessibility of archived mission environmental data and hardware information, with the purpose of providing important context to biological data. This effort is necessary to provide scientific context of its impact upon biological and biomedical data from spaceflight missions and experiments (genomic, metagenomic, gene expression, proteomic, metabolomic, physiological, phenomics, behavioral; tabular, imaging, video). Environmental spaceflight data is derived from dozens of sources, with various formats, and in the past year a pipeline is in development to collect, curate and present this data efficiently. In the upcoming year, a new Data Visualization Portal will utilize the standardized pipeline data to provide easy user access to compare parameters and environmental conditions between missions, locations, subjects, and durations. Environmental and hardware data enables broad accessibility and analytics, without the need for advanced data informatic expertise. Familiarity with the capabilities and limitations of a variety of existing hardware/tools is a strength that could be applied to creation of improved hardware for future ecosystems on the Moon and Mars. The intention is to make biological and environmental telemetry data maximally open-access and FAIR (findable, accessible, interoperable, reusable) for data mining-informatic approaches to support knowledge discovery necessary for low Earth orbit, cis-Lunar, Mars transit, and Mars surface missions.

Danielle K. Lopez

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors

Advances in X-ray Instruments to Support Mars Sample Return

The Mars 2020 Perseverance rover is currently collecting drill cores of ancient igneous and sedimentary rock in and around Jezero crater for potential transport to Earth. These samples from the martian surface will enable detailed mineralogical, geochemical, and petrological measurements to characterize ancient depositional and diagenetic environments, quantitatively age-date the samples, and identify the building blocks for life or evidence for life itself. Furthermore, these drill cores are especially precious because they may represent the most pristine samples from the martian surface and our best chance at identifying martian life, as future sample return missions may be conducted by humans that can introduce biological contaminants to the samples. Because of the importance of these samples, we must take great care in their handling, curation, and preliminary analyses so that they are preserved for scientific measurements for decades to come. In-situ measurements by Perseverance have identified minerals that further warrant special treatment of the returned samples. Hydrated sulfate carbonate, swelling clay minerals, and oxychlorine salts are extremely sensitive to changes in temperature and relative humidity. The structures of hydrated sulfates and oxychlorine minerals, in particular, readily change when exposed to different conditions, meaning the mineral assemblage of the as-returned samples may be lost if the samples aren’t handled properly. Characterizing the as-returned mineral assemblage, particularly of the salts, is essential for reconstructing past aqueous conditions and habitability. To characterize the as-returned mineral assemblage, the samples must be analyzed rapidly before phase changes occur and/or under controlled conditions (e.g., within a glove box). Significant recent advances in X-ray instrumentation for robotic exploration of the solar system have resulted in high-resolution miniaturized instruments that would provide mineralogical, geochemical, and petrological information on the returned martian samples without degradation of the mineral assemblage. Here, we describe a combined X-ray diffractometer/X-ray fluorescence spectrometer (XRD/XRF), an X-ray computed tomographic (XCT) instrument, and a scanned beam XRF mapping instrument that could be used in a glove box so that the martian samples remain under controlled conditions.

E. B. Rampe

A standardized workflow for kinetic metabolic model curation and dissemination

Kinetic metabolic models provide invaluable insights into cellular metabolism, supporting applications in synthetic biology, metabolic engineering, and systems biology. However, reproducibility and utility of these models hinge on clear and rigorous documentation, standardized annotation, and accessible visualization. This paper presents a workflow for building, annotating, visualizing, and sharing kinetic metabolic models. Our method integrates community standards and open-source tools to ensure reproducibility, interoperability, and user accessibility. This procedure enables researchers to produce reusable and well-documented kinetic models, advancing their role as powerful tools in metabolic research.

Cook, Margaret [Univ. of Washington, Seattle, WA (

Infusion of AI/ML Technology into Operational NASA Data Systems

NASA has been developing a variety of Artificial Intelligence / Machine Learning technologies related to Earth Observations. In most cases, the full value of such a technology is realized when it is infused into an operational system. NASA’s Earth Science Data Systems program has been formulating repeatable methods to execute technology infusion. These efforts include the Advancing Collaborative Connections for Earth System Science (ACCESS) program, a Technology Infusion Playbook, and an assemblage of working groups investigating methods for infusion collaboration, community development, and capacity building. ESDS has also been executing a pathfinder activity to infuse a machine-learning-driven recommender of science keywords for Earth Observation datasets, which is intended to be used for metadata curation in the Earth Observation System Data and Information System.

C Lynnes

NETL Coal Energy Atlas: A Collection of Coal/Energy Related Maps

The NETL Coal Energy Atlas contains a comprehensive collection of coal and energy-related maps and graphics curated by the National Energy Technology Laboratory (NETL) Systems Analysis group. It serves as a living document providing an overview of the U.S. coal and energy sectors. The volume is structurally organized into six key thematic areas. Ultimately, the atlas functions as a modular baseline for data integration, allowing researchers to drill down into specific regional locations or customize geographic base layers for advanced systems analysis.

bituminous coal

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin

Toward accelerating rare-earth metal extraction using equivariant neural networks

The separation of rare-earth metals, vital for numerous advanced technologies, is hampered by their similar chemical properties, making ligand discovery a significant challenge. Traditional experimental and quantum chemistry approaches for identifying effective ligands are often resource-intensive. We introduce a machine learning protocol based on an equivariant neural network, Allegro, for the rapid and accurate prediction of binding energies in rare-earth complexes. Key to this work is our newly curated dataset of rare-earth metal complexes—made publicly available to foster further research—systematically generated using the Architector program. This dataset distinctively features functionalized derivatives of proven rare-earth-chelating scaffolds, hydroxypyridinone (HOPO), catecholamide (CAM), and their thio-analogues, selected for their established efficacy in binding these elements. Trained on this valuable resource, our Allegro models demonstrate excellent performance, particularly when trained to directly predict DFT-level binding energies, yielding highly accurate results that closely correlate with theoretical calculations on a diverse test set. Furthermore, this strategy exhibited strong out-of-sample generalization, accurately predicting binding energies for an isomeric HOPO-derivative ligand not seen during training. By substantially reducing computational demands, this machine learning framework, alongside the provided dataset, represent powerful tools to accelerate the high-throughput screening and rational design of novel ligands for efficient rare-earth metal separation.

Gupta, Ankur K. [Lawrence Berkeley National Labora

Leaky ribosomal scanning enables tunable translation of bicistronic ORFs in green algae

Advances in sequencing technology have unveiled examples of nucleus-encoded polycistrons, once considered rare. Exclusively polycistronic transcripts are prevalent in green algae, although the mechanism by which multiple polypeptides are translated from a single transcript is unknown. Here, we used bioinformatic and in vivo mutational analyses to evaluate competing mechanistic models for translation of bicistronic mRNAs in green algae. High-confidence manually curated datasets of bicistronic loci from two divergent green algae, Chlamydomonas reinhardtii and Auxenochlorella protothecoides, revealed a preference for weak Kozak-like sequences for ORF 1 and an underrepresentation of potential initiation codons before the ORF 2 start codon, which are suitable conditions for leaky ribosome scanning to allow ORF 2 translation. We used mutational analysis in A. protothecoides to test the mechanism. In vivo manipulation of the ORF 1 Kozak-like sequence and start codon altered reporter expression at ORF 2, with a weaker Kozak-like sequence enhancing expression and a stronger one diminishing it. A synthetic bicistronic dual reporter demonstrated inversely adjustable activity of green fluorescent protein expressed from ORF 1 and luciferase from ORF 2, depending on the strength of the ORF 1 Kozak-like sequence. Our findings demonstrate that translation of multiple ORFs in green algal bicistronic transcripts is consistent with episodic leaky scanning of ORF 1 to allow translation at ORF 2. This work has implications for the potential functionality of upstream open reading frames (uORFs) found across eukaryotic genomes and for transgene expression in synthetic biology applications.

59 BASIC BIOLOGICAL SCIENCES

Real-time confinement regime detection in fusion plasmas with convolutional neural networks and high-bandwidth edge fluctuation measurements

Abstract A real-time detection of the plasma confinement regime can enable new advanced plasma control capabilities for both the access to and sustainment of enhanced confinement regimes in fusion devices. For example, a real-time indication of the confinement regime can facilitate transition to the high-performing wide-pedestal (WP) quiescent H-mode, or avoid unwanted transitions to lower confinement regimes that may induce plasma termination. To demonstrate real-time confinement regime detection, we use the 2D beam emission spectroscopy (BES) diagnostic system to capture localized density fluctuations of long wavelength turbulent modes in the edge region at a 1 MHz sampling rate. BES data from 330 discharges in either L-mode, H-mode, quiescent H (QH)-mode, or WP QH-mode were collected from the DIII-D tokamak and curated to develop a high-quality database to train a deep-learning classification model for real-time confinement detection. We utilize the 6×8 spatial configuration with a time window of 1024 µ s and recast the input to obtain spectral-like features via fast Fourier transform preprocessing. We employ a shallow 3D convolutional neural network for the multivariate time-series classification task and utilize a softmax in the final dense layer to retrieve a probability distribution over the different confinement regimes. Our model classifies the global confinement state on 44 unseen test discharges with an average F 1 score of 0.94, using only ∼1 ms snippets of BES data at a time. This activity demonstrates the feasibility for real-time data analysis of fluctuation diagnostics in future devices such as ITER, where the need for reliable and advanced plasma control is urgent.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY