Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Data Assimilation Technology Transfer from Planetary Exploration to Terrestrial Applications

Data assimilation can be a valuable tool for discovery and exploration. Planetary missions--principally polar orbiting spacecraft--are now capable of returning enough information to constrain global dynamical models. However, such missions operate under a number of severe constraints: there is no ground truth to calibrate and validate either the models or the data: and there are limited computational resources available, particularly onboard the spacecraft, for the near real-time processing that is required for aerobraking and other semi-autonomous maneuvers. A modified data assimilation approach, in which the prime analysis variable is the observed quantity (i.e., infrared radiances), has been found to be effective in deriving global meteorology from Mars Global Surveyor Thermal Emission Spectrometer data. It may also be useful in a number of terrestrial applications (e.g. aerosol assimilation).

Houben, Howard↗

A universal language for finding mass spectrometry data patterns

Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here, in this study, we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.

Damiani, Tito [Czech Academy of Sciences (CAS), Pr↗

Operation of the Planetary Plasma Interactions Node of the Planetary Data System

Five years ago NASA selected the Planetary Plasma Interactions (PPI) Node at UCLA to help the scientific community locate, access and preserve particles and fields data from planetary missions. We propose to continue to serve for 5 more years. During the first five years we have served the scientific community by providing them with high quality data products. We worked with missions and individual scientists to secure the highest quality data possible and to thoroughly document it. We validated the data, placed it on long lasting media and made sure it was properly archived for future use. So far we have prepared and archived over 10(exp 11) bytes of data from 26 instruments on 4 spacecraft. We have produced 106 CD-ROMs with peer reviewed data. In so doing, we have developed an efficient system to prepare and archive the data and thereby have been able to steadily increase the rate at which the data are produced. Although we produced a substantial archive during the initial five years, we have an even larger amount of work in progress. This includes preparing CD-ROM data sets with all of the Voyager, Pioneer and Ulysses data at Jupiter and Saturn. We will have the Jupiter data ready for the Galileo encounter in December, 1995. We are also completing the Pioneer Venus data restoration. The Galileo Venus archive and radio science data from Magellan will be prepared early in the next period. We are assisting the Small Bodies Node of PDS in the preparation of comet data and will be archiving the asteroid data from Galileo. We will be moving in several new directions as well. We will archive the PPI Node's first Earth based data with data from the International Jupiter Watch and Hubble data taken in support of Ulysses particles and field observations. We will work with the Cassini mission in archive planning efforts. For the inner planets we will begin an archive of Mars data starting with Phobos data and will support the US and Russian Mars missions in the late 1990's. We will restore the Mercury data from Mariner 10 and prepare the lunar data from Clementine in time for the lunar data analysis program in 1995. We will work with the Discovery mission teams to plan their archive and have already started with one, NEAR. Finally we will begin archiving our first heliospheric data from Voyager, Galileo, and Mars observers. We will continue to serve the science community by providing access to the data products. During the past 19 months we have filled nearly 6000 requests for on-line and CD-ROM data. The data delivered directly by the PPI Node has been - 5 x 10(exp 11) bytes. In addition to providing the data, we have provided users with software tools to manage and read the data which are computer, operating system and format independent. We have developed scalable systems so that the same software we use to manage and access the data for the entire PPI Node can be used by individual investigators to manage the data on a single CD-ROM, thereby greatly reducing the software development effort for both the PPI Node and users. We deliver this software with the disks. Recent technical advances have made it possible for us to serve a broader community than before. In the next five year period we plan to extend our outreach to the general public and in particular to increase our support for education. Since planetary plasma data are varied and require expertise in many areas the PPI Node will continue to be distributed. In addition to the primary node at UCLA, the PPI Node has three subnodes with an Outer Planets Subnode at the University of Iowa, an Inner Planets Subnode at UCLA, and a Radio Science Subnode at Stanford University. During the first two years of the renewal period there will be a Radio Astronomy Data Node at GSFC. These organizations will provide scientific expertise on the data, participate in node data selection activities and help with data restoration and mission activities.

Walker, Raymond J.↗

Structure of terrestrial impact craters from SIR-B radar data - Preliminary results

The structure in and around the Charlevoix and Deep Bay impact craters was analyzed by comparing SIR-B radar image lineament maps with maps derived from Landsat data and aerial photography. Relationships revealed include the discovery in the Deep Bay lineament distribution of a mode, of slightly different location and size to that determined by aerial photography, which correlates with the regional glacial trend. The Deep Bay regional glacial trend displays a much more prominent signature in the aerial photography than in the radar data, confirming the influence of radar look direction on the interpretation of lineament distribution and orientation. Effects from the impact event are noted in the radar images of both craters, with lineaments of one km or less appearing to be associated with the rim structure.

Fisher, P. C.↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Mars: Updating Geologic Mapping Approaches and the Formal Stratigraphic Scheme

At the Fourth Mars Conference in 1989, Tanaka reviewed the stratigraphy and geologic history of Mars that had emerged based on systematic geologic mapping of the planet s surface using Viking data. This review looked at the stratigraphic column for Mars and assessed the global geologic history in terms of impact, fluvial, periglacial, aeolian, volcanic, and tectonic processes. Many significant new studies using Mars Global Surveyor (MGS) and now Mars Odyssey (MO) data are showing some important new insights and discoveries that are altering and deepening previous understandings. If we were to illustrate the current state of the science, we might compare it to a loose-leaf notebook in which pages are rapidly being added, removed, and rewritten, with plenty of room remaining. Much of the flux is due to new data, of course, but also much can be attributed to the re-examination of basic assumptions and approaches and our ability to employ ever more powerful computer techniques. Here, we will attempt to review, based on our experience, the areas where the most change seems to be occurring, what prospects we face in the immediate future, and where caution needs to be exercised.

K L Tanaka↗

Equivariant, safe and sensitive — graph networks for new physics

This study introduces a novel Graph Neural Network (GNN) architecture that leverages infrared and collinear (IRC) safety and equivariance to enhance the analysis of collider data for Beyond the Standard Model (BSM) discoveries. By integrating equivariance in the rapidity-azimuth plane with IRC-safe principles, our model significantly reduces computational overhead while ensuring theoretical consistency in identifying BSM scenarios amidst Quantum Chromodynamics backgrounds. The proposed GNN architecture demonstrates superior performance in tagging semi-visible jets, highlighting its potential as a robust tool for advancing BSM search strategies at high-energy colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning-guided design, synthesis, and characterization of atomically dispersed electrocatalysts

The recent integration of machine learning into materials design has revolutionized the understanding of structure–property relationships and optimization of material properties beyond the trial-and-error paradigm. On one hand, machine learning has significantly accelerated the development of atomically dispersed metal-nitrogen-carbon (M-N-C) electrocatalysts, which traditionally heavily relied on heuristic approaches. On the other hand, the primary challenge of leveraging machine learning to expedite M-N-C materials discovery lies in the cost associated with data collection. Here, we review recent machine learning integration strategies for M-N-C catalyst development, including discussions on the typical algorithms such as symbolic regression and convolutional neural networks employed for the theoretical design, synthesis optimization via active learning, and advanced microscopy characterization. Subsequently, we provide our perspective on potential near-future directions for furthering machine learning-assisted development of new M-N-C catalysts and elucidating the complex physicochemical mechanisms governing the selectivity, activity, and durability in this class of materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Towards a self-driving trigger at the LHC: adaptive response in real time

Real-time data filtering and selection—or trigger—systems at high-throughput scientific facilities such as the experiments at the Large Hadron Collider must process extremely high-rate data streams under stringent bandwidth, latency, and storage constraints. Yet these systems are typically designed as static, hand-tuned menus of selection criteria grounded in prior knowledge and simulation. In this work, we further explore the concept of a self-driving trigger, an autonomous data-filtering framework that reallocates resources and adjusts thresholds dynamically in real-time to optimize signal efficiency, rate stability, and computational cost as instrumentation and environmental conditions evolve. We introduce a benchmark ecosystem to emulate realistic collider scenarios and demonstrate real-time optimization of a menu including canonical energy sum triggers as well as modern anomaly-detection algorithms that target non-standard event topologies using machine learning. Using simulated data streams and publicly available collision data from the Compact Muon Solenoid experiment, we demonstrate the capability to dynamically and automatically optimize trigger performance under specific cost objectives without manual retuning. Our adaptive strategy shifts trigger design from static menus with heuristic tuning to intelligent, automated, data-driven control, unlocking greater flexibility and discovery potential in future high-energy physics analyses.

Emami, Shaghayegh [Michigan U.] (ORCID:00090007589↗

Frontal systems during passage of the Martian north polar hood over the Viking Lander 2 site prior to the first 1977 dust storm

Analysis of a 12-sol period of wind speed, wind direction, temperature, pressure, and optical depth at the Viking Lander 2 site presents the first in situ evidence of high- and low-pressure systems, complete with fronts, on the surface of Mars. The discovery of these systems in the Lander data occurred while analyzing a period during which the north polar hood was advected over the site at midday, dramatically decreasing the surface illumination and surface-to-atmospheric heat flux. This obscuration immediately preceded a global dust storm in the southern hemisphere and low latitudes of the northern hemisphere. The direct effects of the dust storm, reached 48 N, the Lander 2 latitude, later more gradually than they reached the 22 deg N latitude of Lander 1. The system responsible for the polar hood passages is a major disturbance, and it appears that radiational damping is inadequate to stop strong frontal formation. The front analyzed is characteristic of a repetitive series of systems that pass roughly every 3.3 sols. These systems are similar to those predicted by theoretical analyses and by general circulation modeling of the Martian atmosphere and those observed in laboratory experiments.

Tillman, J. E.↗

Sources Sought for Innovative Scientific Instrumentation for Scientific Lunar Rovers

Lunar rovers should be designed as integrated scientific measurement systems that address scientific goals as their main objective. Scientific goals for lunar rovers are presented. Teleoperated robotic field geologists will allow the science team to make discoveries using a wide range of sensory data collected by electronic 'eyes' and sophisticated scientific instrumentation. rovers need to operate in geologically interesting terrain (rock outcrops) and to identify and closely examine interesting rock samples. Enough flight-ready instruments are available to fly on the first mission, but additional instrument development based on emerging technology is desirable. Various instruments that need to be developed for later missions are described.

Meyer, C.↗

High redshift quasars and high metallicities

A large-scale code called Cloudy was designed to simulate non-equilibrium plasmas and predict their spectra. The goal was to apply it to studies of galactic and extragalactic emission line objects in order to reliably deduce abundances and luminosities. Quasars are of particular interest because they are the most luminous objects in the universe and the highest redshift objects that can be observed spectroscopically, and their emission lines can reveal the composition of the interstellar medium (ISM) of the universe when it was well under a billion years old. The lines are produced by warm (approximately 10(sup 4)K) gas with moderate to low density (n less than or equal to 10(sup 12) cm(sup -3)). Cloudy has been extended to include approximately 10(sup 4) resonance lines from the 495 possible stages of ionization of the lightest 30 elements, an extension that required several steps. The charge transfer database was expanded to complete the needed reactions between hydrogen and the first four ions and fit all reactions with a common approximation. Radiative recombination rate coefficients were derived for recombination from all closed shells, where this process should dominate. Analytical fits to Opacity Project (OP) and other recent photoionization cross sections were produced. Finally, rescaled OP oscillator strengths were used to compile a complete set of data for 5971 resonance lines. The major discovery has been that high redshift quasars have very high metallicities and there is strong evidence that the quasar phenomenon is associated with the birth of massive elliptical galaxies.

Ferland, Gary J.↗

Hard X-Ray Emission of X-Ray Bursters

The main results from this investigation were serendipitous. The long observation approved for the study of the hard X-ray emission of X-ray bursters lead, instead, to one of the largest early samples of the behavior of fast quasi-periodic oscillations (QPOS) in an atoll sources. Our analysis of this data set lead to the several important discoveries including the existence of a robust correlation between QPO frequency and the flux of a soft blackbody component of the X-ray spectrum in the atoll source 4U 0614+091.

Kaaret, Phillip↗

In Interactive, Web-Based Approach to Metadata Authoring

NASA's Global Change Master Directory (GCMD) serves a growing number of users by assisting the scientific community in the discovery of and linkage to Earth science data sets and related services. The GCMD holds over 8000 data set descriptions in Directory Interchange Format (DIF) and 200 data service descriptions in Service Entry Resource Format (SERF), encompassing the disciplines of geology, hydrology, oceanography, meteorology, and ecology. Data descriptions also contain geographic coverage information, thus allowing researchers to discover data pertaining to a particular geographic location, as well as subject of interest. The GCMD strives to be the preeminent data locator for world-wide directory level metadata. In this vein, scientists and data providers must have access to intuitive and efficient metadata authoring tools. Existing GCMD tools are not currently attracting. widespread usage. With usage being the prime indicator of utility, it has become apparent that current tools must be improved. As a result, the GCMD has released a new suite of web-based authoring tools that enable a user to create new data and service entries, as well as modify existing data entries. With these tools, a more interactive approach to metadata authoring is taken, as they feature a visual "checklist" of data/service fields that automatically update when a field is completed. In this way, the user can quickly gauge which of the required and optional fields have not been populated. With the release of these tools, the Earth science community will be further assisted in efficiently creating quality data and services metadata. Keywords: metadata, Earth science, metadata authoring tools

Pollack, Janine↗

Search for Far-Side Deep Moonquakes

A truly unexpected finding of the Apollo missions, 1969-1972, was a discovery of deep moonquakes. Analysis of the data from the seismic network, which operated for eight years from 1969 through 1977, identified more than 100 discrete source regions at depths approximately half way to the center of the moon. Their distribution, however, was not uniform, as all but one of the source regions found were on the front hemisphere of the moon. Thus, a question remains whether the observed one-sided distribution of deep moonquake sources represents their true distribution or instead occurs because all seismic stations are on the near side of the moon. If it is the former, it means that the interior of the moon is truly asymmetric, structurally and dynamically; if it is the latter, it means that we simply did not identify most moonquakes on the far side and that their identification will be a great help in investigating the deep interior of the moon, including existence of a possible metallic core.

Nakamura, Y.↗

Metadata Authoring with Versatility and Extensibility

NASA's Global Change Master Directory (GCMD) assists the scientific community in the discovery of and linkage to Earth science data sets and related services. The GCMD holds over 13,800 data set descriptions in Directory Interchange Format (DIF) and 700 data service descriptions in Service Entry Resource Format (SERF), encompassing the disciplines of geology, hydrology, oceanography, meteorology, and ecology. Data descriptions also contain geographic coverage information and direct links to the data, thus allowing researchers to discover data pertaining to a geographic location of interest, then quickly acquire those data. The GCMD strives to be the preferred data locator for world-wide directory-level metadata. In this vein, scientists and data providers must have access to intuitive and efficient metadata authoring tools. Existing GCMD tools are attracting widespread usage; however, a need for tools that are portable, customizable and versatile still exists. With tool usage directly influencing metadata population, it has become apparent that new tools are needed to fill these voids. As a result, the GCMD has released a new authoring tool allowing for both web-based and stand-alone authoring of descriptions. Furthermore, this tool incorporates the ability to plug-and-play the metadata format of choice, offering users options of DIF, SERF, FGDC, ISO or any other defined standard. Allowing data holders to work with their preferred format, as well as an option of a stand-alone application or web-based environment, docBUlLDER will assist the scientific community in efficiently creating quality data and services metadata.

Pollack, Janine↗