Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Safety Culture at the World’s Premier Multi-User Spaceport

NASA’s Agency-wide Safety Culture is implemented at the Kennedy Space Center (KSC) using a unique strategy due an unparalleled approach in making human spaceflight history. Kennedy Space Center, the world’s premier multi-user spaceport, enables U.S. government and commercial space access, while providing the world a resource to allow the exploration of and the ability to work in space. A consistently healthy safety culture at KSC is imperative for mission success: desired achievements, protection of space flight hardware, and ultimately, the preservation of human life requires the support of a healthy safety culture. Emphasis on the NASA Agency-wide development of Safety Culture began with the conception of the NASA Agency Safety Culture Working Group. After the devastating loss of life and mission of the Space Shuttle Colombia, a broken safety culture was identified as an organizational cause by the Columbia Accident Investigation Board Report. Thus, the Agency Safety Culture Working Group was developed in 2009 to assess the status of the Agency’s Safety Culture, while addressing safety culture concerns at the NASA Center-level. A Five-factor model was developed to serve as the guiding principles for Safety Culture: 1) Reporting Culture, 2) Just Culture, 3) Flexible Culture, 4) Learning Culture, and 5) Engaged Culture. These five factors are included in the NASA Safety Culture logo, which was intentionally designed as a DNA double helix to prompt the permeation of safety into day-to-day work. KSC specifically implements the NASA Agency-wide Safety Culture principles in a tailored approach that is relevant to the diversity of work being performed. An emphasis is placed on the implementation to include safety at home, not exclusively at work. This emphasis is a KSC-specific element that has been intentionally added to promote a closed loop Safety Culture. To advertise the safety culture, various safety and health events are held throughout the calendar year, providing innovative speakers and engaging activities, while also promoting a wide range of curated safety initiatives. Development and continuous improvement of the KSC safety tracking database, allows for advanced tracking-to-closure, along with providing data sets used to identify areas of emphasis. Other safety initiatives rely solely on employee participation, such as the photo challenges; participants are encouraged to identify and capture themselves, coworkers, or family members participating in safe or healthy activities to share with others within the Center and at Agency levels. Fabrication of exclusive videos and graphics are utilized to advertise and inform employee of safety initiatives, upcoming safety events, and general dispersion of safety information. In addition, an anonymous Agency-wide Safety Culture Survey is advertised, administered, and analyzed at Kennedy Space Center, with the purpose of receiving basic feedback on Safety Culture perceptions to help prevent future incidents from occurring. Through these briefly identified means, and many other forms of employee engagement, Kennedy Space Center aims to maintain safety in the forefront, while creating an environment where everyone trusts that safety is a priority.

Larrin E. Moody↗

Artemis Curation: Preparing for Sample Return from the Lunar South Pole

Space Policy Directive-1 mandates that “the United States will lead the return of humans to the Moon for long-term exploration and utilization, followed by human missions to Mars and other destinations.” In addition, the Vice President stated that “It is the stated policy of this administration and the United States of America to return American astronauts to the Moon within the next five years,” that is, by 2024. These efforts, under the umbrella of the recently formed Artemis Program, include such historic goals as the flight of the first woman to the Moon and the exploration of the lunar south-polar region. Among the top priorities of the Artemis Program is the return of a suite of geologic samples, providing new and significant opportunities for progressing lunar science and human exploration. In particular, successful sample return is necessary for understanding the history of volatiles in the Solar System and the evolution of the Earth-Moon system, fully constraining the hazards of the lunar polar environment for astronauts, and providing the necessary data for constraining the abundance and distribution of resources for in-situ resource utilization (ISRU). Here we summarize the ef-forts of the Astromaterials Acquisition and Curation Office (hereafter referred to as the Curation Office) to ensure the success of Artemis sample return (per NASA Policy Directive (NPD) 7100.10E).

Mitchell, J. L.↗

MoonDB: Restoration and Synthesis of Lunar Petrological and Geochemical Data

About 2,200 samples were collected from the Moon during the Apollo missions, forming a unique and irreplaceable legacy of the Apollo program. These samples, obtained at tremendous cost and great risk, are the only samples that have ever been returned by astronauts from the surface of another planetary body. These lunar samples have been curated at NASA Johnson Space Center and made available to the global research community. Over more than 45 years, a vast body of petrological, geochemical, and geochronological studies of these samples have been amassed, which helped to expand our understanding of the history and evolution of the Moon, the Earth itself, and the history of our entire solar system. Unfortunately, data from these studies are dispersed in the literature, often only available in analog format in older publications, and/or lacking sample metadata and analytical metadata (e.g., information about analytical procedure and data quality), which greatly limits their usage for new scientific endeavors. Even worse is that much lunar data have never been published, simply because no forum existed at the time (e.g., electronic supplements). Thousands of valuable analyses remain inaccessible, often preserved only in personal records, and are in danger of being lost forever, when investigators retire or pass away. Making these data and metadata publicly accessible in a digital format would dramatically help guide current and future research and eliminate duplicated analyses of precious lunar samples.

Lehnert, Kerstin A.↗

Real-time confinement regime detection in fusion plasmas with convolutional neural networks and high-bandwidth edge fluctuation measurements

Abstract A real-time detection of the plasma confinement regime can enable new advanced plasma control capabilities for both the access to and sustainment of enhanced confinement regimes in fusion devices. For example, a real-time indication of the confinement regime can facilitate transition to the high-performing wide-pedestal (WP) quiescent H-mode, or avoid unwanted transitions to lower confinement regimes that may induce plasma termination. To demonstrate real-time confinement regime detection, we use the 2D beam emission spectroscopy (BES) diagnostic system to capture localized density fluctuations of long wavelength turbulent modes in the edge region at a 1 MHz sampling rate. BES data from 330 discharges in either L-mode, H-mode, quiescent H (QH)-mode, or WP QH-mode were collected from the DIII-D tokamak and curated to develop a high-quality database to train a deep-learning classification model for real-time confinement detection. We utilize the 6×8 spatial configuration with a time window of 1024 µ s and recast the input to obtain spectral-like features via fast Fourier transform preprocessing. We employ a shallow 3D convolutional neural network for the multivariate time-series classification task and utilize a softmax in the final dense layer to retrieve a probability distribution over the different confinement regimes. Our model classifies the global confinement state on 44 unseen test discharges with an average F 1 score of 0.94, using only ∼1 ms snippets of BES data at a time. This activity demonstrates the feasibility for real-time data analysis of fluctuation diagnostics in future devices such as ITER, where the need for reliable and advanced plasma control is urgent.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NASA Earthdata Knowledge Base

Prototype work for connecting together the main elements of Earth Observation knowledge and context in a way that is: machine-readable, human-usable and curatable. Using the latest in graph technologies and cloud managed services.

earth data↗

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Developing a Vision for Heliophysics Infrastructure: The LIKED Resource and the DIARieS Ecosystem

Heliophysics data and computational infrastracture are not equipped for 21st science, suffering from holes in the know-how to build better systems. Without a clear vision, efforts to improve the infrastructure have been incremental and incoherent. This poster presents both the vision and the technology required: an online LIbrary KnowledgE and Discovery (LIKED) resource for discovering and implementing knowledge, data, and infrastructure resources; and an online analysis ecosystem to simplify Discovery, Implementation, Analysis, Reproducibility, and Sharing (DIARieS) of scientific results and environments. The LIKED and DIARieS solutions adopt FAIR data principles and the best practices from the budding field of open science. The proposed new infrastructure components will close many of the current gaps in heliophysics’ infrastructure, such as the ability to search for data and knowledge by phenomenon across domains, and to find software and examples relevant to the desired data set (including model data). Further, these components will enable community members to more efficiently use the resources already present and improve upon the content via a community-curated and trusted library. Combining these solutions lowers the barriers to heliophysics resources for all, increasing the return on our investments. Finally, the structure behind these ideas are topic-agnostic, so they are fully extensible to other fields, leading to invaluable connections to other disciplines. Just as with the development and construction of a long-term satellite mission, we must work together as a community to build a vision of the infrastructure that will most benefit the community, and then collaborate to construct, assemble, and test all the necessary pieces individually and as a unit. Our purpose in presenting this work is to not only describe the proposed vision, but also to gather feedback from the community on this topic.

infrastructure↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

What if chondritic porous interplanetary dust particles are not the real McCoy

To select a target comet for a Comet Nucleus Sample Return Mission (CNSRM) it is necessary to have an experimental data base to evaluate the extent of diversity and similarity of comets. For example, the physical properties (e.g., low density) of chondritic porous (CP) interplanetary dust particles (IDPs) are believed to resemble these properties of cometary dust although it is yet to be demonstrated that the porous structure of CP IDPs is inherent to presolar dust particles stored in comet nuclei. Porous structures of IDPs could conceivably form during sublimation at the surface of active comet nuclei. Porous structures are also obtained during annealing of amorphous Mg-SiO smokes which initially forms porous aggregates of olivine + platey tridymite and which, upon continued annealing, react to fluffy enstatite aggregates. It is therefore uncertain that CP IDPs are entirely composed of unmetamorphosed presolar dust. Conceivably, new minerals and textures may form in situ in nuclei of active comets as a function of their individual thermal history. Unmetamorphosed comet dust is probably structurally amorphous. Thermal annealing of this dust can produce ultra fine-grained minerals and this ultrafine grain size of CP IDPs should be considered in assessments of aqueous alterations that could affect presolar dust in comet nuclei between 200 and 400 K. Devitrification and hydration may occur in situ in ice-dust mixtures and the mantle of active comet nuclei. Devitrification, or uncontrolled crystallization, of amorphous precursor dust can produce a range of chemical compositions of ultrafine-grained minerals and (non-equilibrium) mineral assemblages and textures in dust contained in comet nuclei as a function of period and trajectory of orbit and number of perihelion passages (not considering internal heating). Thus, experimental data on relevant processes and reaction rates between 200 and 400 K are needed in order to evaluate comet selection, penetration depth for sampling device and curation of samples for CNSRM.

Rietmeijer, Frans J. M.↗

RTN-124: Photometric Redshifts for the Vera C. Rubin Observatory Data Preview

We present the photometric redshifts (photo-z) inferred using algorithms implemented in the Redshift Assessment Infrastructure Layers (RAIL) for the NSF-DOE Vera C. Rubin Observatory Data Preview 2 (DP2). We produce a compilation of reference redshift catalog using spectroscopic, grism and many band photometric redshift dataset hosted on the LIneA Photo-z Server. We curate training and testing set for assessing the scientific and technical performance of Rubin photo-z. The algorithm applied to the object catalog are FlexZBoost, BPZ, kNN, GPz, DNF and TPz; with a combination of 6-band and 4-band photo-z depending on availability of u and y photometry. The redshift point estimates and uncertainty estimation in tabular format through the Large Survey DataBase (LSDB).

79 ASTRONOMY AND ASTROPHYSICS↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

The Science Discovery Engine: Connecting Heterogeneous Scientific Data and Information

Transformative science often occurs at the boundaries of different disciplines. Making interdisciplinary science data, software and documentation discoverable and accessible is essential to enabling transformative science. However, connecting this diverse and heterogeneous information is often a challenge due to several factors including the dispersed and sometimes isolated nature of data and the semantic differences between topical areas. NASA’s Science Discovery Engine (SDE) has developed several approaches to tackling these challenges. The SDE is a unified, insightful search experience that enables discovery of NASA’s open science data across five topical areas: astrophysics, biological and physical sciences, Earth science, heliophysics and planetary science. In this presentation, we will discuss our efforts to develop a systematic scientific curation workflow to integrate diverse content into a single search environment. We will also share lessons learned from our work to create a metadata crosswalk across the five disciplines.

Kaylin Bugbee↗

Underwater Target Detection Software Demonstration on the RivGen Turbine

This repository contains data and processing scripts necessary to train the object detection models utilized in the underwater target detection software demonstration on the RivGen turbine project and to produce performance metrics (precision, recall, mAP50, mAP50-95). - Contents - Data consist of "images" and "labels". Each image has an associated label, both share the same time string in its file name (e.g., 2024_05_25_09_01_57.98.jpg and 2024_05_25_09_01_57.98.txt). Time strings have the format %yyyy_%mm_%dd_%HH_%MM_%SS.%3f. Images and labels were curated from 2021 and 2024 smolt outmigration periods at the project site in Igiugig, AK. Images are monochrome 8-bit images of objects (smolt, debris, and other) passing through the field of view of the deployed cameras during various operational stages of the RivGen turbine. Labels are text files indicating the class and bounding polygon of each object in an image. The provided labels use the "YOLO" label format. - Requirements - Python3.8+ is required to install and run the train and validation script. The README.md provides instruction for installing the requirements from the requirements.py file. - Instructions - The "example_train.py" file ingests the provided data, trains a model, and produces model performance metrics at completion. NOTE: model performance metrics will vary from run to run as a consequence of the random selection of training and validation data.

16 TIDAL AND WAVE POWER↗

Extending the Reach of IGSN Beyond Earth: Implementing IGSN Registration to Link Nasa's Apollo Lunar Samples and Their Data

The rock and soil samples returned from the Apollo missions from 1969-72 have supported 46 years of research leading to advances in our understanding of the formation and evolution of the inner Solar System. NASA has been engaged in several initiatives that aim to restore, digitize, and make available to the public existing published and unpublished research data for the Apollo samples. One of these initiatives is a collaboration with IEDA (Interdisciplinary Earth Data Alliance) to develop MoonDB, a lunar geochemical database modeled after PetDB (Petrological Database of the Ocean Floor). In support of this initiative, NASA has adopted the use of IGSN (International Geo Sample Number) to generate persistent, unique identifiers for lunar samples that scientists can use when publishing research data. To facilitate the IGSN registration of the original 2,200 samples and over 120,000 subdivided samples, NASA has developed an application that retrieves sample metadata from the Lunar Curation Database and uses the SESAR API to automate the generation of IGSNs and registration of samples into SESAR (System for Earth Sample Registration). This presentation will describe the work done by NASA to map existing sample metadata to the IGSN metadata and integrate the IGSN registration process into the sample curation workflow, the lessons learned from this effort, and how this work can be extended in the future to help deal with the registration of large numbers of samples.

Todd, Nancy S.↗

Kaona: Deep Searching and Curating Aviation Safety Reporting Systems

Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given in augmenting existing processes of safety reporting systems. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety reporting systems activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using all public NASA's ASRS narratives and another using all public C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.

asrs↗

Kaona: Deep Searching and Curating Safety Reporting Systems

Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given in augmenting existing processes of safety reporting systems. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety reporting systems activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using all public NASA's ASRS narratives and another using all public C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.

asrs↗

Utilizing Ontology Structures To Curate the DOE-NETL Carbon Storage Open Database

The specialized ontology for the Carbon Storage Open Database will enable more rapid assignment of appropriate symbology standards for visualization improvements, optimize topical and spatial tagging within keywords, and improve flexibility for utilization in existing data repositories such as EDX. This effort also aims to establish a foundation for utilization of ontologies for organization of other data related to geologic carbon storage in the future.

Martin, Abigail↗

Analyzing Natural Language Context in Human-Machine Teaming using Supervised Machine Learning

Building a foundation for trustworthiness and trust verification in multi-asset teaming is the research challenge of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR). The Design Reference Mission (DRM) for ATTRACTOR is a search and rescue mission objective governed by a multi-member team consisting of human and machine operators. A crucial component to the effort is the communication between humans and autonomous agents throughout both planning and execution stages of the mission. Intuitive communication methods and modalities are posited as critical enablers for certifying trust and trustworthiness. This paper reports on the data collection and analysis conducted in support of the Human Informed Natural-language GANs Evaluation (HINGE)project to attain explainable and trusted communication between human-machine assets. Two identically curated image description datasets were acquired for HINGE, both consisting of two unique input modalities (typed vs. verbal) and retrieved in two distinct contexts (general vs. specific). The gathered datasets were assessed and compared using Parts-of-Speech (POS)features, sentence similarity metrics, and linguistic analysis. Then, the datasets were modeled and tested separately and in combination with one another using machine learning algorithms. The comparison and testing results reveal a superior dataset, by which a preferred context and input is understood, for generating image representations of missing persons using a Generative Adversarial Network (GAN).

Bryan A Barrows↗