Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

New Ways of Facilitating Improved Data Discovery and Access for NASA's Suborbital Earth Science Observations

NASA conducts field research in various Earth Science disciplines utilizing airborne and other non-satellite platforms to acquire in situ and remotely sensed observations indicative of physical processes across a range of scales. Field efforts are key in the development and validation of instruments and satellite algorithm refinements. The heterogeneous data, with a range of file formats, scales, and acquisition methods, support research in several science areas. NASA’s archive process assigns data products to discipline-oriented Distributed Active Archive Centers (DAACs) for stewardship. Over time, individual DAACs have developed tools for data browsing and serving disparate user bases. As science becomes more interdisciplinary, researchers need to incorporate observations from multiple campaigns, and multiple DAACs, into their work. Motivated in part by this shifting paradigm of needs, the Catalog of Archived Suborbital Earth Science Investigations (CASEI) was created. CASEI provides a single starting point to browse, search, and discover airborne and field data. Contextual metadata are organized and inter-linked allowing intuitive, integrated exploration across all NASA DAACs. Campaign science objectives, platform and instrument configurations, geographical details, geophysical concepts, and more are tracked in CASEI’s database, facilitating multi-parameter search, browse, and discovery of relevant data products. Researchers are able to directly access associated data products, via DOI links, regardless of the DAAC where they reside. Significant events, key time periods of high science interest within the longer-duration campaign effort, are also indicated and allow for a more efficient identification of critical data subsets. This presentation describes CASEI’s development, intensive metadata curation process, and demonstrates the web interface experience. Initial content metrics and plans for continued maintenance will also be discussed.

metadata↗

Active learning for the design of polycrystalline textures using conditional normalizing flows

Generative modeling has opened new avenues for solving previously intractable materials design problems. However, these new opportunities are accompanied by a drastic increase in the required amount of training data. This is in stark juxtaposition to the high expense and difficulty in curating such large materials datasets. In this work, we propose a novel framework for integrating generative models within an active learning loop. Further, this enables the training of generative models with datasets significantly smaller than what has previously been demonstrated, providing a direct route for their application in data constrained environments. The functionality of this framework is then demonstrated by addressing the challenge of designing polycrystalline textures associated with target anisotropic mechanical properties. The developed protocol exhibited a cost reduction between 14 to 18 times over a randomly sampled experimental design.

36 MATERIALS SCIENCE↗

New GES DISC Services Shortening the Path in Science Data Discovery

The Current GES DISC available services only allow user to select variables from a single dataset at a time and too many variables from a dataset are displayed, choice is hard. At American Geophysical Union (AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a new service: Datalist. A Datalist is a collection of predefined or user-defined data variables from one or more archived datasets. Our science support team curated predefined datalist and provided value to the user community. Imagine some novice user wants to study hurricane and typed in hurricane in the search box. The first item in the search result is GES DISC provided Hurricane Datalist. It contains scientists recommended variables from multiple datasets like TRMM, GPM, MERRA, etc. Datalist uses the same architecture as that of our new website, which also provides one-stop shopping for data, metadata, citation, documentation, visualization and other available services.We implemented Datalist with new GES DISC web architecture, one single web page that unified all user interfaces. From that webpage, users can find data by either type in keyword, or browse by category. It also provides user with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping.

Datalist↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

DeepBench: A simulation package for physical benchmarking data

We introduce **DeepBench**, a python library that generates simple simulated image data from first principles, such as basic geometric shapes and astronomical objects. These data are highly valuable for developing (calibration, testing, and benchmarking) statistical and machine learning models because they make it possible to connect the final data product to physically interpretable inputs. This software includes tools to curate and store the datasets to maximize reproducibility.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-Resolution Imaged-Based 3D Reconstruction Combined with X-Ray CT Data Enables Comprehensive Non-Destructive Documentation and Targeted Research of Astromaterials

Providing web-based data of complex and sensitive astromaterials (including meteorites and lunar samples) in novel formats enhances existing preliminary examination data on these samples and supports targeted sample requests and analyses. We have developed and tested a rigorous protocol for collecting highly detailed imagery of meteorites and complex lunar samples in non-contaminating environments. These data are reduced to create interactive 3D models of the samples. We intend to provide these data as they are acquired on NASA's Astromaterials Acquisition and Curation website at http://curator.jsc.nasa.gov/.

Blumenfeld, E. H.↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Evaluating a Commercial Dynamic Line Rating Software with the National PMU Dataset

To accelerate the development of data-driven applications for power systems, the Department of Energy (DOE) supported the collection and curation of a synchrophasor dataset spanning two years of observations from transmission utilities across the US. This National PMU Dataset (NPDS) was anonymized and distributed to awardees of a DOE research grant under nondisclosure agreements (NDAs) but has also been retained at PNNL to enable further research. Agreements with data contributors prevent the data from being shared outside the organization. However, establishing a blind research validation methodology is envisioned to maximize the value proposition of the NPDS. In this validation strategy, researchers may share algorithms/software (potentially as executables to protect intellectual property) with PNNL, and PNNL will share feedback about the software’s performance on subsets of the NPDS. Such a blind methodology ensures that sensitive information about critical infrastructure remains protected, but the value of the NPDS can be extended to research beyond PNNL. Through iterative feedback, the algorithms may be tweaked to address real-world artifacts. As the NPDS data is temporally and geographically diverse, it may capture features absent in smaller datasets used during the development of the algorithm under test. This report presents lessons learned from applying the blind validation methodology to LineID™, a synchrophasor-based dynamic line rating software developed by Topolonet Corporation. Improvements made to the software through iterative feedback, limitations of the validation methodology, as well as how the limitations of the NPDS affected the evaluation process are discussed. Observations indicate that the proposed validation methodology can be valuable for evaluating other tools in the future.

97 MATHEMATICS AND COMPUTING↗

Physics-aware adaptive checkpointing with shadow systems for nonlinear PDE simulations

Large-scale simulations of nonlinear partial differential equations (PDEs) that exhibit strongly transient behavior and pattern-forming dynamics produce enormous amounts of data, which, even with modern storage systems, cannot be stored for later curation. Current I/O strategies either write dense time series of snapshots, which is often prohibitive in I/O and storage, or store a few checkpoints that enable restart but incur expensive recomputation cost and provide no control over post-restart error growth, especially when lossy compression is used. Moreover, most, if not all, existing strategies take no account of the actual physical state of the system. Here, we present a simple physics-aware I/O framework in which a low-cost shadow system adaptively triggers lossy checkpoints when the shadow system deviates from the fine-scale simulation. The shadow system can be a coarsened replica of the fine-scale simulation that evolves concurrently. This means that checkpoints are taken based on the physical state of the system: fewer checkpoints are triggered when the system is quiescent while more are taken when the system undergoes a rapid change. This type of behavior is observed in many systems such as Brusselator and FitzHugh–Nagumo. We illustrate that our framework maintains stable restarts, keeps fine-scale restart errors bounded by shadow errors, and reconstructs the time history with significantly lower error and storage than interpolating fixed-interval snapshots, with low-cost shadow replay and modest online synchronization overhead.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

Integrating Multi-agency Data Products in a Cloud-based Platform for Streamlined Discovery, Visualization, and Use

Earth science data users almost always have an interest in utilizing geospatial data from multiple agencies. As computing capability and cloud-based infrastructures accelerate the pace at which scientific research can be done, there is a growing need to enable search, discovery, and use of multi-agency geospatial observations relevant for a common use case - without undergoing the search and discovery process in a less efficient, disparate path with each agency. NASA’s Earth Observing System Data and Information System (EOSDIS) and NOAA’s National Environmental Satellite, Data and Information Service (NESDIS) both support a wide range of Earth science disciplines’ research, operations, and applications activities. Presently, however, there are few examples of data discovery frameworks supporting an inquiry of both NASA’s and NOAA’s extensive archives of Earth observations that are equally suitable for a particular science scenario, regardless of the agency that “owns” the data. NASA and NOAA are collaborating on a data expedition platform for exploring fire weather using data products from both agencies. Users will be able to search, discover, and visualize NASA and NOAA products in one interface. Each agency will curate metadata for its respective datasets, providing for a rich search experience. The collaboration will pilot a shared search interface into these metadata datastores. Data products will be stored in the cloud in cloud-optimized format(s). These formats will allow for optimized data access and visualization to support the “data expedition”. Avenues for further development and application of this cloud-based, multi-agency data provisioning platform will also be discussed.

cloud-based technology↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Evolution in Metadata Quality: Common Metadata Repository's Role in NASA Curation Efforts

Metadata Quality is one of the chief drivers of discovery and use of NASA EOSDIS (Earth Observing System Data and Information System) data. Issues with metadata such as lack of completeness, inconsistency, and use of legacy terms directly hinder data use. As the central metadata repository for NASA Earth Science data, the Common Metadata Repository (CMR) has a responsibility to its users to ensure the quality of CMR search results. This poster covers how we use humanizers, a technique for dealing with the symptoms of metadata issues, as well as our plans for future metadata validation enhancements. The CMR currently indexes 35K collections and 300M granules.

metadata quality↗

Apollo Next Generation Sample Analysis (ANGSA): an Apollo Participating Scientist Program to Prepare the Lunar Sample Community for Artemis

As a first step in preparing for the return of samples from the Moon by the Artemis Program, NASA initiated the Apollo Next Generation Sample Analysis Program (ANGSA). ANGSA was designed to function as a low-cost sample return mission and involved the curation and analysis of samples previously returned by the Apollo 17 mission that remained unopened or stored under unique conditions for 50 years. These samples include the lower portion of a double drive tube previously sealed on the lunar surface, the upper portion of that drive tube that had remained unopened, and a variety of Apollo 17 samples that had remained stored at -27 °C for approximately 50 years. ANGSA constitutes the first preliminary examination phase of a lunar “sample return mission” in over 50 years. It also mimics that same phase of an Artemis surface exploration mission, its design included placing samples within the context of local and regional geology through new orbital observations collected since Apollo and additional new “boots-on-the-ground” observations, data synthesis, and interpretations provided by Apollo 17 astronaut Harrison Schmitt. ANGSA used new curation techniques to prepare, document, and allocate these new lunar samples, developed new tools to open and extract gases from their containers, and applied new analytical instrumentation previously unavailable during the Apollo Program to reveal new information about these samples. Most of the 90 scientists, engineers, and curators involved in this mission were not alive during the Apollo Program, and it had been 30 years since the last Apollo core sample was processed in the Apollo curation facility at NASA JSC. There are many firsts associated with ANGSA that have direct relevance to Artemis. ANGSA is the first to open a core sample previously sealed on the surface of the Moon, the first to extract and analyze lunar gases collected in situ, the first to examine a core that penetrated a lunar landslide deposit, and the first to process pristine Apollo samples in a glovebox at -20 °C. All the ANGSA activities have helped to prepare the Artemis generation for what is to come. The timing of this program, the composition of the team, and the preservation of unopened Apollo samples facilitated this generational handoff from Apollo to Artemis that sets up Artemis and the lunar sample science community for additional successes.

79 ASTRONOMY AND ASTROPHYSICS↗

Comprehensive Non-Destructive Conservation Documentation of Lunar Samples Using High-Resolution Image-Based 3D Reconstructions and X-Ray CT Data

Established contemporary conservation methods within the fields of Natural and Cultural Heritage encourage an interdisciplinary approach to preservation of heritage material (both tangible and intangible) that holds "Outstanding Universal Value" for our global community. NASA's lunar samples were acquired from the moon for the primary purpose of intensive scientific investigation. These samples, however, also invoke cultural significance, as evidenced by the millions of people per year that visit lunar displays in museums and heritage centers around the world. Being both scientifically and culturally significant, the lunar samples require a unique conservation approach. Government mandate dictates that NASA's Astromaterials Acquisition and Curation Office develop and maintain protocols for "documentation, preservation, preparation and distribution of samples for research, education and public outreach" for both current and future collections of astromaterials. Documentation, considered the first stage within the conservation methodology, has evolved many new techniques since curation protocols for the lunar samples were first implemented, and the development of new documentation strategies for current and future astromaterials is beneficial to keeping curation protocols up to date. We have developed and tested a comprehensive non-destructive documentation technique using high-resolution image-based 3D reconstruction and X-ray CT (XCT) data in order to create interactive 3D models of lunar samples that would ultimately be served to both researchers and the public. These data enhance preliminary scientific investigations including targeted sample requests, and also provide a new visual platform for the public to experience and interact with the lunar samples. We intend to serve these data as they are acquired on NASA's Astromaterials Acquisistion and Curation website at http://curator.jsc.nasa.gov/. Providing 3D interior and exterior documentation of astromaterial samples addresses the increasing demands for accessability to data and contemporary techniques for documentation, which can be realized for both current collections as well as future sample return missions.

Blumenfeld, E. H.↗

From natural language to control signals: a conceptual framework for semantic channel finding in complex experimental infrastructure

Modern experimental platforms such as particle accelerators, fusion devices, telescopes, and industrial process control systems expose tens to hundreds of thousands of control and diagnostic channels, accumulated over decades of hardware evolution. Operators and AI systems alike depend on informal expert knowledge, inconsistent naming conventions, and scattered documentation to locate the signals required for monitoring, troubleshooting, and automated control, creating a persistent bottleneck for reliability, scalability, and emerging language-model-driven interfaces. We formalize semantic channel finding, the task of mapping natural-language intent to concrete control-system signals, as a general problem in complex experimental infrastructure, and introduce a four-paradigm conceptual framework to guide architecture selection based on facility-specific data regimes. The paradigms span (i) direct in-context lookup over small, curated channel dictionaries, (ii) constrained hierarchical navigation through structured trees, (iii) interactive agent exploration using iterative reasoning and tool-based database queries, and (iv) ontology-grounded semantic search that decouples channel meaning from facility-specific naming conventions. We demonstrate the practical feasibility of each paradigm through proof-of-concept implementations at four operational facilities spanning two orders of magnitude in scale: from compact free-electron lasers to large synchrotron light sources, operating under diverse control-system architectures ranging from clean hierarchical naming schemes to legacy environments with decades of heterogeneous conventions. Where evaluated against expert-curated operational queries, these instantiations achieve 90%–97% accuracy, validating the framework’s applicability across real-world deployment scenarios. To accelerate adoption across the broader scientific and industrial control-system community, we release open-source, plug-and-play implementations of all three interactive paradigms-direct lookup, hierarchical navigation, and middle-layer exploration-within the Osprey framework, together with tools for channel database generation, interactive testing, and minimal-configuration deployment. This work establishes semantic channel finding as a foundational capability for human-centric and agentic AI interfaces at large-scale facilities, providing both a systematic framework for architecture design and practical resources to enable adoption without building custom infrastructure from scratch.

channel finding↗

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

With the end of Moore’s law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intelligent subsampling? To explore this, we develop SICKLE, a sparse intelligent curation framework for efficient learning, featuring a novel maximum entropy (MaxEnt) sampling approach, scalable training, and energy benchmarking. We compare MaxEnt with random and phase-space sampling on large direct numerical simulation (DNS) datasets of turbulence. Evaluating SICKLE at scale on Frontier, we show that subsampling as a preprocessing step can, in many cases, improve model accuracy and substantially lower energy consumption, with observed reductions of up to 38×.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗