Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Developing a Vision for Heliophysics Infrastructure: The LIKED Resource and the DIARieS Ecosystem

Heliophysics data and computational infrastracture are not equipped for 21st science, suffering from holes in the know-how to build better systems. Without a clear vision, efforts to improve the infrastructure have been incremental and incoherent. This poster presents both the vision and the technology required: an online LIbrary KnowledgE and Discovery (LIKED) resource for discovering and implementing knowledge, data, and infrastructure resources; and an online analysis ecosystem to simplify Discovery, Implementation, Analysis, Reproducibility, and Sharing (DIARieS) of scientific results and environments. The LIKED and DIARieS solutions adopt FAIR data principles and the best practices from the budding field of open science. The proposed new infrastructure components will close many of the current gaps in heliophysics’ infrastructure, such as the ability to search for data and knowledge by phenomenon across domains, and to find software and examples relevant to the desired data set (including model data). Further, these components will enable community members to more efficiently use the resources already present and improve upon the content via a community-curated and trusted library. Combining these solutions lowers the barriers to heliophysics resources for all, increasing the return on our investments. Finally, the structure behind these ideas are topic-agnostic, so they are fully extensible to other fields, leading to invaluable connections to other disciplines. Just as with the development and construction of a long-term satellite mission, we must work together as a community to build a vision of the infrastructure that will most benefit the community, and then collaborate to construct, assemble, and test all the necessary pieces individually and as a unit. Our purpose in presenting this work is to not only describe the proposed vision, but also to gather feedback from the community on this topic.

infrastructure↗

Propagation of Cosmic Rays: Nuclear Physics in Cosmic-ray Studies

The nuclei fraction in cosmic rays (CR) far exceeds the fraction of other CR species, such as antiprotons, electrons, and positrons. Thus the majority of information obtained from CR studies is based on interpretation of isotopic abundances using CR propagation models where the nuclear data and isotopic production cross sections in p- and alpha-induced reactions are the key elements. This paper presents an introduction to the astrophysics of CR and diffuse gamma-rays and dimsses some of the puzzles that have emerged recently due to more precise data and improved propagation models. Merging with cosmology and particle physics, astrophysics of CR has become a very dynamic field with a large potential of breakthrough and discoveries in the near fume. Exploiting the data collected by the CR experiments to the fullest requires accurate nuclear cross sections.

Moskalenko, Igor V.↗

Development of an Improved Spatial Metadata Simplification Algorithm

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata↗

Use of Spatial Metadata Simplification for TEMPO

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

What's new, Voyager: The discoveries continue

The twin Voyager spacecraft, launched nearly two decades ago, continue to operate and are now searching for the edge of our solar system, the heliopause. Voyager's giant-planet flybys of Jupiter, Saturn, Uranus, and Neptune have provided data that are likely to remain the definitive data set for the foreseeable future and have led to many ongoing discoveries. As the spacecraft move toward the heliopause, they are also providing data on the structure of the heliosphere. This article discusses the discoveries resulting from the flyby and heliosphere data that have been made within the past five years.

Miner, Ellis D.↗

Create your own science planning tool in 3 days with SOA

Scientific discovery and advancement of knowledge has been, and continues to be, the goal for space missions at Jet Propulsion Laboratory. Scientist must plan their observation/experiments to get the maximum data return in order to make those discoveries. However, each mission has different science objectives, a different spacecraft and different instrument payloads, as well as, different routes to different destinations with different spacecraft restrictions and characteristics. In the current reduced cost environment, manageable cost for mission planning software is a must. Science Opportunity Analyzer (SOA), a planning tool for scientists and mission planners, utilizes a simple approach to reduce cost and promote reusability.

Science Opportunity Analyzer (SOA).↗

Facilitating Science--The International Space Station Fluids and Combustion Facility

Scientists in many fields would like to perform experiments on the International Space Station (ISS) to take advantage of the unique environment of microgravity (which is the near-absence of gravity). The ISS will provide the opportunity for scientists to perform microgravity tests over much longer time periods than previously available on the space shuttle--months rather than hours or days--providing more data that could lead to new discoveries. Many of the experiments on ISS will be conducted through the use of new microgravity science facilities. A microgravity science facility is a complete system of on-orbit and ground (on-Earth) hardware, software, operations, and plans that have been optimized to perform sustained microgravity research in one or two scientific disciplines. The facility concept includes hardware that remains on-orbit (because of its general usefulness) and a small amount of unique hardware that is developed for each principal investigator. Such unique hardware customizes the facility to perform a given principal investigator's experiment effectively. Many facilities are planned for the ISS to accommodate scientists' needs. While the quality and quantity of scientific data are being improved, per-experiment costs will be lowered relative to other ways of performing such experiments. The NASA Lewis Research Center is developing a Fluids and Combustion Facility (FCF) to perform microgravity fluids and combustion experiments on the ISS. The FCF will be the lowest cost, most resource efficient approach to performing fluid physics and combustion science experiments on the ISS. Experiments performed in the FCF will be 3 to 9 times less expensive than similar Spacelab experiments. Moreover, use of key ISS resources, such as upmass, power, cooling, and astronaut crew time, will be cut by the same factor.

Winsa, Edward A.↗

Physics and chemistry from parsimonious representations: image analysis via invariant variational autoencoders

Electron, optical, and scanning probe microscopy methods are generating ever increasing volume of image data containing information on atomic and mesoscale structures and functionalities. This necessitates the development of the machine learning methods for discovery of physical and chemical phenomena from the data, such as manifestations of symmetry breaking phenomena in electron and scanning tunneling microscopy images, or variability of the nanoparticles. Variational autoencoders (VAEs) are emerging as a powerful paradigm for the unsupervised data analysis, allowing to disentangle the factors of variability and discover optimal parsimonious representation. Here, we summarize recent developments in VAEs, covering the basic principles and intuition behind the VAEs. The invariant VAEs are introduced as an approach to accommodate scale and translation invariances present in imaging data and separate known factors of variations from the ones to be discovered. We further describe the opportunities enabled by the control over VAE architecture, including conditional, semi-supervised, and joint VAEs. Several case studies of VAE applications for toy models and experimental datasets in Scanning Transmission Electron Microscopy are discussed, emphasizing the deep connection between VAE and basic physical principles. Python codes and datasets discussed in this article are available at https://github.com/saimani5/VAE-tutorials and can be used by researchers as an application guide when applying these to their own datasets.

36 MATERIALS SCIENCE↗

Analysis and Review of NASA Earth Science Metadata: How Automation Plays a Role

The Analysis and Review of the Common Metadata Repository (CMR ARC) Team reviews all EOSDIS metadata. The team’s objective is to achieve consistency, correctness, and completeness for all metadata records in the CMR, as well as improve the discoverability of NASA's Earth Science data within the CMR framework. This work is currently being completed at Marshall Space Flight Center. CMR makes a single discovery point possible for NASA's Earth Science data users. The CMR team, in collaboration with three other core metadata teams, contributes to the stewardship of NASA's Earth Science data through a process of continual curation and the ongoing development of the Unified Metadata Model (UMM). A key tool now used in the curation process, referred to as the NASA CMR Dashboard, is an online curation dashboard developed in collaboration with software development company, Element 84. This tool facilitates the review of Earth Science metadata records and subsequent stakeholder collaboration on the resolution of identified issues. A key capability of the new tool is a suite of automated compliance checks written in Python 3.6 that verify the integrity of various metadata elements across multiple standards.

Staton, Patrick↗

Data-driven design of electrolyte additives supporting high-performance 5 V LiNi 0.5 Mn 1.5 O 4 positive electrodes

LiNi 0.5 Mn 1.5 O 4 (LNMO) is a high-capacity spinel-structured material with an average lithiation/de-lithiation potential at ca. 4.6–4.7 V vs Li + /Li, far exceeding the stability limits of electrolytes. An efficient way to enable LNMO in lithium-ion batteries is to reformulate an electrolyte composition that stabilizes both graphitic (Gr) negative electrode with solid-electrolyte-interphase and LNMO with cathode-electrolyte-interphase. In this study, we select and test a diverse collection of 28 single and dual additives for the Gr||LNMO battery system. Subsequently, we train machine learning models on this dataset and employ the trained models to suggest 6 binary compositions out of 125, based on predicted final area-specific-impedance, impedance rise, and final specific-capacity. Such machine learning-generated new additives outperform the initial dataset. This finding not only underscores the efficacy of machine learning in identifying materials in a highly complicated application space but also showcases an accelerated material discovery workflow that directly integrates data-driven methods with battery testing experiments.

batteries↗

High entropy oxides prediction and discovery by the Mixed Enthalpy-Entropy Descriptor

The vast, high-dimensional composition space of high-entropy oxides (HEOs) offers exceptional opportunities for functional materials discovery, yet it also poses a fundamental challenge: the rational and efficient prediction of stable, synthesizable compositions and the corresponding structure–property relationships. Despite growing interest, the field still lacks broadly applicable, physically grounded descriptors capable of navigating various large chemical spaces. Here, we introduce a Mixed Enthalpy–Entropy Descriptor (MEED) that enables rapid, first-principles–based prediction of HEOs synthesizability across diverse chemistries. Using MEED, we perform high-throughput screening of two distinct HEO families: rocksalt oxides and perovskite oxides. The predicted top candidates in each family were experimentally validated. MEED reveals unifying thermodynamic and structural principles governing stability across both chemical compositions and polymorphs, providing mechanistic insight into the formation of high-entropy phases. This work significantly broadens the accessible chemical design space for HEOs and establishes a data-efficient framework for accelerating the discovery of next-generation functional materials.

Yu, Liping [University of Central Florida]↗

Discovering System Health Anomalies Using Data Mining Techniques

We present a data mining framework for the analysis and discovery of anomalies in high-dimensional time series of sensor measurements that would be found in an Integrated System Health Monitoring system. We specifically treat the problem of discovering anomalous features in the time series that may be indicative of a system anomaly, or in the case of a manned system, an anomaly due to the human. Identification of these anomalies is crucial to building stable, reusable, and cost-efficient systems. The framework consists of an analysis platform and new algorithms that can scale to thousands of sensor streams to discovers temporal anomalies. We discuss the mathematical framework that underlies the system and also describe in detail how this framework is general enough to encompass both discrete and continuous sensor measurements. We also describe a new set of data mining algorithms based on kernel methods and hidden Markov models that allow for the rapid assimilation, analysis, and discovery of system anomalies. We then describe the performance of the system on a real-world problem in the aircraft domain where we analyze the cockpit data from aircraft as well as data from the aircraft propulsion, control, and guidance systems. These data are discrete and continuous sensor measurements and are dealt with seamlessly in order to discover anomalous flights. We conclude with recommendations that describe the tradeoffs in building an integrated scalable platform for robust anomaly detection in ISHM applications.

Sriastava, Ashok, N.↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

Accelerating Space Life Sciences: Successes and Challenges of Biospecimen and Data Sharing

NASA's current human space flight research is directed towards enabling human space exploration beyond Low Earth Orbit (LEO). To that end, NASA Space Flight Payload Projects; Rodent Research, Cell Science, and Microbial Labs, flown on the International Space Station (ISS), benefit the global life sciences and commercial space communities. Verified data sets, science results, peer-reviewed publications, and returned biospecimens, collected and analyzed for flight and ground investigations, are all part of the knowledge base collected by NASA's Human Exploration and Operations Mission Directorate's Space Life and Physical Sciences Research and Applications (SLPSRA) Division, specifically the Human Research and Space Biology Programs. These data and biospecimens are made available through the public Life Sciences Data Archive (LSDA) website to promote basic discovery, pre-clinical and clinical science.The NASA Institutional Scientific Collection (ISC), stores flight and ground biospecimens from Space Shuttle and ISS programs. These specimens are curated and managed by the Ames Life Sciences Data Archive (ALSDA), an internal node of NASA's LSDA. The ISC stores over 30,000 specimens from experiments dating from 1984 to present. Currently available specimens include tissues from the circulatory, digestive, endocrine, excretory, integumentary, muscular, neurosensory, reproductive, respiratory and skeletal systems.NASA's biospecimen collection represents a unique and limited resource of unique spaceflight payload and ground control research subjects. These specimens are harvested according to well established SOPs that maintain their quality and integrity. Once the primary scientific objectives have been met, the remaining specimens are made available to provide secondary opportunities for complementary studies or new investigations to broaden research without large expenditures of time or resources. Website: https://lsda.jsc.nasa.gov/

Scott, Ryan T.↗

Large Amplitude Whistlers in the Magnetosphere Observed with Wind-Waves

We describe the results of a statistical survey of Wind-Waves data motivated by the recent STEREO/Waves discovery of large-amplitude whistlers in the inner magnetosphere. Although Wind was primarily intended to monitor the solar wind, the spacecraft spent 47 h inside 5 R(sub E) and 431 h inside 10 R(sub E) during the 8 years (1994-2002) that it orbited the Earth. Five episodes were found when whistlers had amplitudes comparable to those of Cattell et al. (2008), i.e., electric fields of 100 m V/m or greater. The whistlers usually occurred near the plasmapause. The observations are generally consistent with the whistlers observed by STEREO. In contrast with STEREO, Wind-Waves had a search coil, so magnetic measurements are available, enabling determination of the wave vector without a model. Eleven whistler events with useable magnetic measurements were found. The wave vectors of these are distributed around the magnetic field direction with angles from 4 to 48deg. Approximations to observed electron distribution functions show a Kennel-Petschek instability which, however, does not seem to produce the observed whistlers. One Wind episode was sampled at 120,000 samples/s, and these events showed a signature that is interpreted as trapping of electrons in the electrostatic potential of an oblique whistler. Similar waveforms are found in the STEREO data. In addition to the whistler waves, large amplitude, short duration solitary waves (up to 100 mV/m), presumed to be electron holes, occur in these passes, primarily on plasma sheet field lines mapping to the auroral zone.

Kellogg, P. J.↗

The Federation of Earth Science Information Partners ESIP

A broad-based, distributed community of science, data and information technology practitioners. With over 150 member organizations, the ESIP Federation brings together public, academic, commercial, and nongovernmental organizations to share knowledge, expertise, technology and best practices to improve opportunities for increasing access, discovery, integration and usability of Earth science data.

ESIP↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗