Search NASASearch

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Open-Source Science-Driven Development of the Science Data System (SDS) for Earth System Observatory (ESO) Atmospheric Missions

The NASA Earth System Observatory (ESO) atmospheric missions will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The Science Data System (SDS) will deploy the adaptive processing system (APS) developed within the Cloud to manage the research and operational processing of ESO atmospheric mission orbital and suborbital sensors and curate these data for near real-time and collection reprocessing and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage and distribution. Further, the SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The SDS follows NASA’s commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the SDS system components will be developed with open-source concepts including components of APS itself as well as ESO atmospheric mission algorithms. This presentation describes the framework of the SDS and its integral part in facilitating OSS within the ESO atmospheric missions.

David M. Giles

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis

Laying the Foundations for FAIR-er Science: ISA and the LSDA Data Submission Process in NASA's Evolving Data Management Environment

The Life Sciences Data Archive (LSDA) archives data resulting from research on the effects of spaceflight on humans and the development of countermeasures to mitigate spaceflight hazards. Archivists work with researchers to ensure that unique and high value data products and their metadata are preserved and managed to support current and future research. Currently, LSDA is updating its procedures and data submission requirements in response to the evolving data preservation environment at NASA. LSDA is implementing best practices for research data management through the establishment of clear data submission guidelines, integration of the FAIR (Findability, Accessibility, Interoperability, Reusability) principles, and use of the ISA (Investigation, Study, Assay) research metadata framework for data discoverability and transparency into the data management processes. These changes directly impact LSDA’s requirements for research data submissions. The newly revised Research Data Submission Agreement (RDSA), formerly the Data Submission Agreement (DSA), introduces ISA-compatible metadata collection standards to LSDA’s process. Adherence to LSDA’s data submission guidelines enhances the FAIR-ness of the repository’s collections for future users. This presentation will discuss (1) how submission of research data and associated metadata are impacted by current data management policies, (2) benefits of the adoption of FAIR principles and the ISA metadata framework for retrospective studies utilizing existing LSDA datasets and historic data collections, and (3) the support LSDA will provide to researchers during this transition.

LSDA

Geolab in NASA's First Generation Pressurized Excursion Module: Operational Concepts

We are building a prototype laboratory for preliminary examination of geological samples to be integrated into a first generation Habitat Demonstration Unit-1/Pressurized Excursion Module (HDU1-PEM) in 2010. The laboratory GeoLab will be equipped with a glovebox for handling samples, and a suite of instruments for collecting preliminary data to help characterize those samples. The GeoLab and the HDU1-PEM will be tested for the first time as part of the 2010 Desert Research and Technology Studies (DRATS), NASAs annual field exercise designed to test analog mission technologies. The HDU1-PEM and GeoLab will participate in joint operations in northern Arizona with two Lunar Electric Rovers (LER) and the DRATS science team. Historically, science participation in DRATS exercises has supported the technology demonstrations with geological traverse activities that are consistent with preliminary concepts for lunar surface science Extravehicular Activities (EVAs). Next years HDU1-PEM demonstration is a starting point to guide the development of requirements for the Lunar Surface Systems Program and test initial operational concepts for an early lunar excursion habitat that would follow geological traverses along with the LER. For the GeoLab, these objectives are specifically applied to enable future geological surface science activities. The goal of our GeoLab is to enhance geological science returns with the infrastructure that supports preliminary examination, early analytical characterization of key samples, insight into special considerations for curation, and data for prioritization of lunar samples for return to Earth.

Evans, C. A.

Machine learning framework for predicting uranium enrichments from M400 CZT gamma spectra

A machine learning framework was developed for predicting uranium enrichments from M400 CZT gamma spectra. This framework leverages the availability of a large amount of measured M400 gamma spectra and uses a recently updated version of Gamma Detector Response and Analysis Software (GADRAS) for gamma spectrum analysis and generation. It also leverages the existing machine learning modules in Python for gamma spectrum data processing, curation, model training, benchmarking, and optimization of the deep machine learning models. The framework is used to develop a deep learning model to analyze gamma spectra from a set of U 3 O 8 samples with enrichments ranging from 0.31 to 93.17% and UF 6 cylinders with enrichments ranging from 0.2 to 4.95%, and the model performance is tested using a set of measured spectra and the respective declared enrichment values. Results show that the model can correctly classify 99.35% of the U 3 O 8 sample enrichments, and can predict the samples’ enrichments within an average absolute error of 0.099% (in percentage points of enrichment). For the UF 6 cylinders, the average absolute error was approximately 0.03%, with an accuracy of 98% in classifying discrete enrichment values of UF 6 samples. Finally, the results also show that the model has performed significantly better in terms of predicting enrichments in UF 6 cylinders based on measured gamma spectra than the GEM code, with a standard deviation (of the relative errors) of 2.23% (compared with the 11.51% value for the GEM code) based on results from a set of test data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE

Changes in Four Decades of Near‐CONUS Tropical Cyclones in an Ensemble of 12 km Thermodynamic Global Warming Simulations

We evaluate tropical cyclones (TCs) in a set of thermodynamic global warming (TGW) simulations over the continental United States (CONUS). A 12 km simulation forced by ERA5 provides a 40‐year historical (1980–2019) control. Four complimentary future scenarios are generated using thermodynamic deltas applied to lateral boundary, interior, and surface forcing. We curate a data set of 4,498 6‐hourly TC snapshots in the control and find a corresponding “twin” in each counterfactual, permitting a paired comparison. Warming results in an increase in mean dynamical TC intensity and moisture‐related quantities, with the latter being more pronounced. TC inner cores contract slightly but outer storm size remains unchanged. The frequency with which TCs become more intense is only moderately consistent, with snapshots having increased hazards ranging from 50% to 80% depending on warming level. The fractions of TCs undergoing rapid intensification and weakening both increase across all warming simulations, suggesting elevated short‐term intensity variability.

54 ENVIRONMENTAL SCIENCES

Edge AI-Enhanced Traffic Monitoring and Anomaly Detection Using Multimodal Large Language Models

This paper addresses the challenge of traffic monitoring and incident detection in remote areas, utilizing multimodal large language models (LLMs) deployed on edge AI devices. The key novelty of the LLM is to convert real-time video streams into descriptive texts, enabling low-bandwidth transmissions and reliable detection of anomalies and incidents in environments of intermittent connectivity. The model is developed based on fine-tuning open-source LLMs and extending it with multi-modal capabilities to analyze video frames. Our work also involves deploying this model on edge devices such as Nvidia IGX Orin and is planned to be tested in realistic environments in future work. The methodology includes data set curation, iterative model fine-tuning and compression, and hardware-based optimization. This approach aims to enhance traffic safety and response speed in remote areas, marking a significant advancement in the application of AI for traffic monitoring and safety management.

Peruski, Ryan [University of Tennessee, Knoxville

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

Decadal Seasonal Shifts of Precipitation and Temperature in TRMM and AIRS Data

We present results from an analysis of seasonal phase shifts in the global precipitation and surface temperatures. We use data from the TRMM (Tropical Rainfall Measuring Mission) Multi-satellite Precipitation Algorithm (TMPA), and the Atmospheric Infrared Sounder (AIRS) on Aqua satellite, all hosted at NASA Goddard Earth Science Data and Information Services Center (GES DISC). We explore the information content and data usability by first aggregating daily grids from the entire records of both missions to pentad (5-day) series which are then processed using Singular Value Decomposition approach. A strength of this approach is the normalized principal components that can then be easily converted from real to complex time series. Thus, we can separate the most informative, the seasonal, components and analyze unambiguously for potential seasonal phase drifts. TMPA and AIRS records represent correspondingly 20 and 15 years of data, which allows us to run simple “phase learning†from the first 5 years of records and use it as reference. The most recent 5 years are then phase-compared with the reference. We demonstrate that the seasonal phase of global precipitation and surface temperatures has been stable in the past two decades. However, a small global trend of delayed precipitation, and earlier arrival of surface temperatures seasons, are detectable at 95% confidence level. Larger phase shifts are detectable at regional level, in regions recognizable from the Eigen vectors to having strong seasonal patterns. For instance, in Central North America, including the North American Monsoon region, confident phase shifts of 1-2 days per decade are detected at 95% confidence level. While seemingly symbolic, these shifts are indicative of larger changes in the Earth Climate System. We thus also demonstrate a potential usability scenario of Earth Science Data Records curated at the NASA GES DISC in partnership with Earth Science Missions.

surface temperatures

Empowering Open Science with the Science Discovery Engine

: This presentation will describe the work to date in building the SDE as well as what the team has learned about the SMD ecosystem, information curation, and data governance. A short demonstration of the SDE will be presented, and an overview of near- and long-term goals for future development will be shared. Community feedback will be welcomed about the interface, content, and other features to help inform actions to maximize the SDE’s performance and usability. Whether users aim to discover Earth-like atmospheres on planets outside of our solar system or better understand the impacts of solar energy on our own planet, the Science Discovery Engine provides a means for scientists and all curious individuals to find content to further their understanding of science across all time and space scales.

Emily Foshee

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection

Cislunar Trajectory Design and Maneuver Autonomy for NASA's Moon to Mars Architecture

NASA’s Moon to Mars architecture is an ambitious roadmap of manned cislunar and deep space exploration. The extensive amount of orbital assets required will place a significant burden on ground-based resources, such as communication networks and operations facilities. Spacecraft autonomy is essential for maintaining a vast number of complex missions beyond Earth orbit. To achieve full autonomy, spacecraft must be able to employ methods of robust maneuver design without an explicit dependence on commands sent from the ground. This level of autonomy is needed not only for stationkeeping, but also for outbound transfers. To address the need of spacecraft maneuver design autonomy, this work investigates the use of neural networks (NNs) in a supervised learning environment. A supervised learning approach for NNs allows for a curated training data set, consisting exclusively of perturbations applied to a desired mission concept of operations (ConOps). The proposed approach allows humans on the ground to design a specific mission ConOps before flight, then employ NNs to fly the mission robustly and autonomously. This investigation numerically tests maneuver autonomy in four highly sensitive regions of flight: orbit raising, translunar injection burns, powered lunar flybys, and invariant manifold insertion burns. These straining cases are contextualized by testing them in a demonstration mission, targeting an Earth-Moon L3 orbit. The study first establishes feasibility by automating impulsive burn maneuvers. However, some guidance algorithms will need more intensive commands, such as inertial pointing and angular rates. To validate this method, NN maneuver autonomy is applied to a finite burn model of the demonstration mission. The use of sequential, mission specific maneuvers provide an appropriate testbed to demonstrate the robustness of a NN trained on feasible perturbed states. Moreover, these scenarios provide preliminary proof-of-concept for fully autonomous missions that execute maneuvers without dependence upon explicit command uplinks. As a result, the technological advancement proposed in this work may significantly ease the strain on ground-based mission operations. This would enable complex and autonomous mission execution in cislunar and deep space regimes, filling a technology gap required to support future manned missions.

NASA

Cognitive Performance in ISS Astronauts on 6-Month Low Earth Orbit Missions

Introduction: Current and future astronauts will endure prolonged exposure to spaceflight hazards and environmental stressors that could compromise cognitive functioning, yet cognitive performance in current missions to the International Space Station remains critically under-characterized. We systematically assessed cognitive performance across 10 cognitive domains in astronauts on 6-month missions to the ISS. Methods: Twenty-five professional astronauts were administered the Cognition Battery as part of National Aeronautics and Space Administration (NASA) Human Research Program Standard Measures Cross-Cutting Project. Cognitive performance data were collected at five mission phases: pre-flight, early flight, late flight, early post-flight, and late post-flight. We calculated speed and accuracy scores, corrected for practice effects, and derived z-scores to represent deviations in cognitive performance across mission phases from the sample’s mean baseline (i.e., pre-flight) performance. Linear mixed models with random subject intercepts and pairwise comparisons examined the relationships between mission phase and cognitive performance. Results: Cognitive performance was generally stable over time with some differences observed across mission phases for specific subtests. There was slowed performance observed in early flight on tasks of processing speed, visual working memory, and sustained attention. We observed a decrease in risk-taking propensity during late flight and post-flight mission phases. Beyond examining group differences, we inspected scores that represented a significant shift from the sample’s mean baseline score, revealing that 11.8% of all flight and post-flight scores were at or below 1.5 standard deviations below the sample’s baseline mean. Finally, exploratory analyses yielded no clear pattern of associations between cognitive performance and either sleep or ratings of alertness. Conclusions: There was no evidence for a systematic decline in cognitive performance for astronauts on a 6-month missions to the ISS. Some differences were observed for specific subtests at specific mission phases, suggesting that processing speed, visual working memory, sustained attention, and risk-taking propensity may be the cognitive domains most susceptible to change in Low Earth Orbit for high performing, professional astronauts. We provide descriptive statistics of pre-flight cognitive performance from 25 astronauts, the largest published preliminary normative database of its kind to date, to help identify significant performance decrements in future samples.

Data curation

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 350 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab Sequencing Lab. The GLDS contains rich metadata about each experiment and has integrated radiation dosimetry data from experiments flown on the Space Shuttle, International Space Station, and Free Flying spacecrafts. With the increasing amount and complexity of omics data being generated, GeneLab utilizes community-defined, common models for metadata and terminology so that omics data and results are discoverable and reliably reproducible. GeneLab uses the ISA-Tab specification and semantic model for organizing and representing omics metadata. In addition to metadata standards, data files must be open-source file or common exchange formats to ensure accessibility and usability by all users. To ease data ingestion and transfer, the web-based submission tool allows PIs a user-friendly user interface to curate, organize, and publish their space relevant omics data. In the more recent years, data curation and submission portal has incorporated the FAIR principles making data findable, accessible, interoperable, and reusable. To increase reusability of data, GeneLab has implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 200 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. To train the next generation of scientists, NASA offers training programs such as GeneLab 4 High School (GL4HS) and GeneLab 4 Universities. NLM Curation at a Scale Workshop 2022 | NASA GeneLab (GL4U) to teach students bioinformatics and computational biology methods to analyze omics data. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

GeneLab

NASA's Astromaterials Database: Enabling Research Through Increased Access to Sample Data, Metadata and Imagery

The Astromaterials Acquisition & Curation Office at NASA's Johnson Space Center (JSC) is the designated facility for curating all of NASA's extraterrestrial samples. Today, the suite of collections includes the lunar samples from the Apollo missions, cosmic dust particles falling into the Earth's atmosphere, meteorites collected in Antarctica, comet and interstellar dust particles from the Stardust mission, asteroid particles from Japan's Hayabusa mission, solar wind atoms collected during the Genesis mission, and space‐exposed hardware from several missions. To support planetary science research on these samples, JSC's Astromaterials Curation Office hosts NASA's Astromaterials Curation digital repository and data access portal [http://curator.jsc.nasa.gov/], providing descriptions of the missions and collections, and critical information about each individual sample. Our office is designing and implementing several informatics initiatives to better serve the planetary research community. First, we are re‐hosting the basic database framework by consolidating legacy databases for individual collections and providing a uniform access point for information (descriptions, imagery, classification) on all of our samples. Second, we continue to upgrade and host digital compendia that summarize and highlight published findings on the samples (e.g., lunar samples, meteorites from Mars). We host high resolution imagery of samples as it becomes available, including newly scanned images of historical prints from the Apollo missions. Finally we are creating plans to collect and provide new data, including 3D imagery, point cloud data, micro CT data, and external links to other data sets on selected samples. Together, these individual efforts will provide unprecedented digital access to NASA's Astromaterials, enabling preservation of the samples through more specific and targeted requests, and supporting new planetary science research and collaborations on the samples.

Evans, Cindy