Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

The Federation of Earth Science Information Partners ESIP

A broad-based, distributed community of science, data and information technology practitioners. With over 150 member organizations, the ESIP Federation brings together public, academic, commercial, and nongovernmental organizations to share knowledge, expertise, technology and best practices to improve opportunities for increasing access, discovery, integration and usability of Earth science data.

ESIP↗

Data as a Key Resource in Catalysis: A Community Account

The deployment of artificial intelligence (AI) is transforming the scientific fields central to interdisciplinary catalysis research. By enabling more effective use of data, AI (including simpler machine learning and data science tools) holds great promise for accelerating discoveries. However, progress has so far been modest, largely due to the lack of standardized, machine-readable, and openly shared catalysis data. This perspective, accounting for community insights emerging at conferences, analyses the underlying reasons for these challenges and proposes solutions to a future whereFAIR data management becomes an integral part of research in catalysis. In the short-term, we deem that mandatory FAIR data depositing prior to scientific publications along with consensualized top-down guidelines on data sharing powered by ease-to-use tools can make the necessary step change happen to catalyse data as key resource in our community.

36 - MATERIALS SCIENCE↗

Giotto data analysis: Electron plasma and heavy ion composition measurements at Comet Halley

This investigation involved the analysis of electron plasma and heavy ion composition measurements made by the COPERNIC (COmplete Positive ion, Electron, and Ram Negative Ion measurements near Comet Halley) plasma experiment during the close fly-by of Halley by the European Space Agency's Giotto spacecraft. The experiment provided measurements of the full 3-dimensional distribution of 10 eV-30 keV electrons, and mass analysis of cold cometary ions from 10-210 amu. The analysis of the COPERNIC data has yielded some remarkable results, including: The discovery of negatively charged ions in the inner coma; the discovery of far heavier (mass is greater than 50 amu) ions than predicted, dominated by complex molecular ions made up of C, H, O, and N; the discovery of an adiabatic heating effect on electrons from the compression of the solar wind plasma; the identification of several organic and sulfur bearing ions; and the discovery of a new 'mystery region' where electrons are accelerated to high energies. These discoveries were in addition to the detailed analysis of 'expected' features at Comet Halley. Although this grant has expired, analysis continues on the data at a low (unfunded) level, and it is expected that more significant results will be obtained. A bibliography of the papers resulting from this research is attached, and a copy of each paper is included.

Lin, Robert P.↗

Space‐Time Causal Discovery in Earth System Science: A Local Stencil Learning Approach

Causal discovery tools enable scientists to infer meaningful relationships from observational data, spurring advances in fields as diverse as biology, economics, and climate science. Despite these successes, the application of causal discovery to space-time systems remains immensely challenging due to the high-dimensional nature of the data. For example, in climate sciences, modern observational temperature records over the past few decades regularly measure thousands of locations around the globe. To address these challenges, we introduce Causal Space-Time Stencil Learning (CaStLe), a novel meta-algorithm for discovering causal structures in complex space-time systems. CaStLe leverages regularities in local space-time dependencies to learn governing global dynamics. This local perspective eliminates spurious confounding and drastically reduces sample complexity, making space-time causal discovery practical and effective. For causal discovery, CaStLe flexibly accepts any appropriately adapted time series causal discovery algorithm to recover local causal structures. These advances enable causal discovery of geophysical phenomena that were previously unapproachable, including non-periodic, transient phenomena such as volcanic eruption plumes. Regularities in local space-time dependencies are transformed into informative spatial replicates, which actually improve CaStLe's performance when applied to ever-larger spatial grids. We successfully apply CaStLe to discover the atmospheric dynamics governing the climate response to the 1991 Mount Pinatubo volcanic eruption. We provide validation experiments to demonstrate the effectiveness of CaStLe over existing causal-discovery frameworks on a range of geophysics-inspired benchmarks while identifying the method's limitations and domains where its assumptions may not hold.

Nichol, J. Jake [Univ. of New Mexico, Albuquerque,↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Deep-learning atomistic semi-empirical pseudopotential model for nanomaterials

The semi-empirical pseudopotential method (SEPM) has been widely applied to provide computational insights into the electronic structure, photophysics, and charge carrier dynamics of nanoscale materials. We present “DeepPseudopot”, a machine-learned atomistic pseudopotential model that extends the SEPM framework by combining a flexible neural network representation of the local pseudopotential with parameterized non-local and spin-orbit coupling terms. Trained on bulk quasiparticle band structures and deformation potentials from GW calculations, the model captures many-body and relativistic effects with very high accuracy across diverse semiconducting materials, as illustrated for silicon and group III-V semiconductors. DeepPseudopot’s accuracy, efficiency, and transferability make it well-suited for data-driven in silico design and discovery of novel optoelectronic nanomaterials.

Lin, Kailai [University of California, Berkeley, C↗

Beyond the four core effects: revisiting thermoelectrics with a high-entropy design

Low-exergy waste heat, which constitutes the majority of industrial-scale thermal losses, remains largely unrecoverable with conventional technologies. Thermoelectrics offer a solid-state solution for converting this hard-to-access energy into electricity, making them attractive for decentralized power generation and sensor applications. High-entropy materials (HEMs) have gained traction as a strategy for better-performing thermoelectrics, but the mechanisms driving their benefits require further exploration. This article highlights key insights for heat and electronic transport in HEMs. For heat transport, we argue that reduced, and often ultralow, lattice thermal conductivity in HEMs—with respect to ordered counterparts—can be taken for granted, emerging naturally as a fifth core effect of high-entropy systems. While band convergence is often considered beneficial for electronic transport, its impact depends strongly on the electronic structure. We summarize the scenarios where it can be detrimental to thermoelectric performance. These insights motivate strategies that align seamlessly with advancements in artificial intelligence and data-driven approaches, helping accelerate the discovery of next-generation thermoelectric materials.

Oses, Corey [Johns Hopkins Univ., Baltimore, MD (U↗

Daily modulations and broadband strategy in axion searches: An application with the CAST-CAPP detector

It has been previously advocated that the presence of the daily and annual modulations of the axion flux on the Earth’s surface may dramatically change the strategy of the axion searches. The arguments were based on the so-called Axion Quark Nugget (AQN) dark matter model which was originally put forward to explain the similarity of the dark and visible cosmological matter densities Ω dark ∼ Ω visible . In this framework, the population of galactic axions with mass 10 − 6 eV ≲ m a ≲ 10 − 3 eV and velocity ⟨ v a ⟩ ∼ 10 − 3 c will be accompanied by axions with typical velocities ⟨ v a ⟩ ∼ 0.6 c emitted by AQNs. Furthermore, in this framework, it has also been argued that the AQN-induced axion daily modulation (in contrast with the conventional weakly interactive massive particle paradigm) could be as large as (10–20)%, representing the main motivation for the present investigation. We argue that the daily modulations along with the broadband detection strategy can be very useful tools for the discovery of such relativistic axions. The data from the CAST-CAPP detector have been used following such arguments. Unfortunately, due to the dependence of the amplifier chain on temperature-dependent gain drifts and other factors, we could not conclusively show the presence or absence of a dark sector-originated daily modulation. However, this proof of principle analysis procedure can serve as a reference for future studies. Published by the American Physical Society 2025

Caspers, F.↗

Collaborative: in situ visual analytics technologies for extreme scale combustion simulations

This project aims to drastically enhance the usability of in situ analysis and visualization for extreme-scale scientific simulations. Current exascale computing capabilities promise to offer greater predictive ability of simulations and to further push the frontiers of science and technology. However, to validate the simulation output at extreme scale, examine the modeled phenomena, and discover previously unknowns from the output data, the output must be reduced or transformed in situ as it is being generated during the simulation such that the amount of data to examine and store is kept to a minimum. Such in situ approaches allow us to process and analyze the data and any embedded geometry to an extent that would be prohibitively expensive, if not impossible, to perform as a post hoc task. While in situ processing has been demonstrated to be a feasible and promising approach, its full potential has not yet been leveraged. In this project, we have developed comprehensive enhancements to in situ technology based on probability distributions in data. Our research focuses on jointly developing new ways of interacting with massive statistical samples while creatively utilizing new state-of-the-art computational resources to push the boundaries of in situ exploration. Moreover, we have developed new time-dependent techniques to enable previously unattainable capabilities in areas such as intelligent simulation steering and precise feature identification. We have experimentally studied our design and implementation at NERSC and OLCF, and are able to leverage existing in situ infrastructures whenever possible. While the exemplar in this project is combustion, many other fields for which turbulent transport is important, e.g., fusion, climate, astrophysics among others, encounter similar issues as simulations scale up to the exascale. This project shows its potential to generate high impact on DOE missions since the resulting technology promises to improve scientists’ ability to rapidly and correctly interpret and tune extreme-scale simulations, leading to new scientific understanding and advancements.

97 MATHEMATICS AND COMPUTING↗

Wilkins: HPC in situ workflows made easy

In situ approaches can accelerate the pace of scientific discoveries by allowing scientists to perform data analysis at simulation time. Current in situ workflow systems, however, face challenges in handling the growing complexity and diverse computational requirements of scientific tasks. In this work, we present Wilkins, an in situ workflow system that is designed for ease-of-use while providing scalable and efficient execution of workflow tasks. Wilkins provides a flexible workflow description interface, employs a high-performance data transport layer based on HDF5, and supports tasks with disparate data rates by providing a flow control mechanism. Wilkins seamlessly couples scientific tasks that already use HDF5, without requiring task code modifications. We demonstrate the above features using both synthetic benchmarks and two science use cases in materials science and cosmology.

HPC↗

A research in support of NASA's space science

Instrumentation, the interpretation of data from space-borne instruments and the development of theoretical studies of the Earth's environment are reported. New circuitry was introduced to the existing ion drift meter to enable the detection of light ion velocities that are different from the major ion species. Significant progress was made in the tailoring of magnetic mass analysis to stratospheric ions where care must be taken to preserve the original species and to obtain good mass resolution at high mass numbers. Also a rugged and durable zoom imaging spectrometer was successfully tested and important modifications are being undertaken to allow larger scanning ranges for observation of weak airglow emissions from the Earth's atmosphere. Data interpretation efforts led to the discovery of a new class of plasma irregularities on the bottomside of the F-region. Studies of all the available plasma properties from satellite measurements in the high latitude ionosphere revealed regions of field aligned currents where it is reasonable to expect thermal electrons to be the dominant current carriers.

Hanson, W. B.↗

Some background about satellites

Four tables of planetary and satellite data are presented which list satellite discoveries, planetary parameters, satellite orbits, and satellite physical properties respectively. A scheme for classifying the satellites is provided and it is noted that most known moons fall into three general classes: regular satellites, collisional shards, and irregular satellites. Satellite processes are outlined with attention given to origins, dynamical and thermal evolution, surface processes, and composition and cratering. Background material is provided for each family of satellites.

Burns, Joseph A.↗

Introduction to the Asteroids II data base

This paper describes the Asteroids II data base, which is a compilation of asteroid data published, or in press, as of March 1988 with some updates in early 1989. The Asteroids II machine-readable data base includes asteroid names and discovery circumstances; proper elements and family identifications; asteroid light-curve parameters; asteroid pole determinations; taxonomic classes; and absolute magnitudes and slope parameters, UBV colors, albedos, and diameters.

Tedesco, Edward F.↗

Continued reduction and analysis of data from the Dynamics Explorer Plasma Wave Instrument

The plasma wave instrument on the Dynamics Explorer 1 spacecraft provided measurements of the electric and magnetic components of plasma waves in the Earth's magnetosphere. Four receiver systems processed signals from five antennas. Sixty-seven theses, scientific papers and reports were prepared from the data generated. Data processing activities and techniques used to analyze the data are described and highlights of discoveries made and research undertaken are tabulated.

Gurnett, Donald A.↗

Operations on Graphical Models with Plates

This paper explains how graphical models, for instance Bayesian or Markov networks, can be extended to model problems in data analysis and learning. This provides a unified framework that combines lessons learned from the artificial intelligence, statistical and connectionist communities. This also offers a set of principles for developing a software generator for data analysis, whereby a learning or discovery system can be compiled from specifications. Many of the popular learning algorithms can be compiled in this way from graphical specifications. While in a sense this paper is a multidisciplinary review of learning, the main contribution here is the presentation of the material within the unifying framework of graphical models, and the observation that, as a result, the process of developing learning algorithms can be partly automated.

Buntine, Wray L.↗

Feasibility and Definition of a Lunar Polar Volatiles Prospecting Mission

The recent Lunar Crater Observing and Sensing Satellite (LCROSS) mission has provided evidence for significant amounts of cold trapped volatiles in Cabeus crater near the Moon's south pole. Moreover, LRO/Diviner measurements of extremely cold lunar polar surface temperatures imply that volatiles can be stable outside or areas of strict permanent shadows. These discoveries suggest that orbital neutron spectrometer data point to extensive deposits at both lunar poles. The physical state, composition and distribution of these volatiles are key scientific issues that relate to source and emplacement mechanisms. These issues are also important for enabling lunar in situ resource utilization (ISRU). An assessment of the feasibility of cold-trapped volatile ISRU requires a priori information regarding the location, form, quantity, and potential for extraction of available resources. A robotic mission to a mostly shadowed but briefly .unlit location with suitable environmental conditions (e.g. short periods of oblique sunlight and subsurface cryogenic temperatures which permit volatile trapping) can help answer these scientific and exploration questions. Key parameters must be defined in order to identify suitable landing sites, plan surface operations, and achieve mission success. To address this need, we have conducted an initial study for a lunar polar volatile prospecting mission, assuming the use of a solar-powered robotic lander and rover. Here we present the mission concept, goals and objectives, and landing site selection analysis for a short-duration, landed, solar-powered mission to a potential hydrogen volatile-rich site.

Heldmann, Jennifer↗

Development of VBA Tool for Document Term Search

Employees throughout different agencies such as NASA, have identified that the search of determined terms/words through documents, consume substantial research time of such. These types of searches are substantially limited towards one word in a one document identification; forward one, these usual types of searches lack efficiency & optimization through research aspects of work. Consequently, this reflects in the decrease productivity during work hours etc. The application of VBA (Visual Basic for Applications) is the programming language of Excel, which was conducted for the development of optimized tool for document term search. The project enables the search of single & multiple word/term search through single format documents for paragraph data extraction.

Ssytems Development↗