Search NASA⌕ Search

SEARCH · Search NASA

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Basin & Range Investigation for Developing Geothermal Energy

Hidden geothermal systems represent a potentially prolific energy resource that could support critical U.S. public and government energy priorities. Basin and Range Investigations for Developing Geothermal Energy (BRIDGE) addressed some the challenges associated with hidden system exploration by prioritizing cost-effective exploration early on through strategic workflow and informed decision-making that mitigates early risk and shifts resources to later exploration stages (e.g., drilling). Sandia National Laboratories partnered with U.S. Navy Geothermal Office, Geologic Geothermal Group, and independent consultants, with additional collaboration with U.S. Geological Survey and private industry. The primary tool of the BRIDGE project was to deploy a regional-scale airborne electromagnetic method to investigate the shallow resistivity structure in areas with high prospectivity. This was followed up at several prospects by a multidisciplinary exploration approach, including additional geologic, geophysical and geochemical studies. A central tenet to the BRIDGE methodology is that zones of low resistivity frequently occur over geothermal systems in the Basin and Range, and when paired with other data constraints, imaging these zones can enable discovery of these systems. In addition to exploring greenfield areas (i.e., Grover Point), the BRIDGE project also flew HTEM resistivity surveys over known geothermal systems including those with established power plants (Don A. Campbell and Salt Wells) and prospects that are known to the literature but remain undeveloped, at least in part, due to a lack of understanding on the location of their producible reservoirs. BRIDGE produced a comprehensive set of data from prospects identified in the Nevada Play Fairway Analysis along with conceptual models for top ranking prospects, wherein all of the observations are used to inform an interpreted model of the system. These models present a range of possible system parameters such as temperature and size, and they are further informed by system analogues in the Basin and Range province and elsewhere. The results of this work leave space for further exploration that may now occur at prospects ‘down the list’ rather than distribution exploration resources evenly across all prospects.

15 GEOTHERMAL ENERGY↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Probabilistic data fusion and physics-informed machine learning: A new paradigm for modeling under uncertainty, and its application to accelerating the discovery of new materials

In this report we summarize the work conducted by PI Perdikaris and his group under this Early Career project DE–SC0019116 during the period of 09/01/2018 – 08/31/2023. The central aim of the work was to introduce a new paradigm for scientific data analysis that can seamlessly synthesize rigorous mathematical modeling with data of variable fidelity (e.g., measurements at multiple scales/resolutions or predictions of variable fidelity models) and multiple modalities (e.g., images, time–series, or scattered measurements). The setting we are interested in involves complex systems that are partially observed and whose dynamical behavior could be hard to model or totally unknown. The inherent uncertainty associated with this setting necessitates a departure from the classical deterministic realm of modeling and scientific computation, and, consequently, our main building blocks can no longer be crisp deterministic numbers and governing laws, but instead we must operate with probabilistic models.

97 MATHEMATICS AND COMPUTING↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗

MCP-eGridGPT (MCP-Enabled Chatbot with Electrical Power System Analysis and Interactive Visualization Tool) [SWR-25-126]

This software is an advanced chatbot system that integrates the Model Context Protocol (MCP) to provide intelligent electrical power system analysis and automated visualization generation. The system enables users to interact with complex electrical engineering tools through natural language, automatically analyzes power system data for voltage violations and grid health assessment, and generates professional interactive HTML dashboards and reports. Key features include dynamic tool discovery from MCP servers, multi-LLM provider support, intelligent data interpretation using large language models, automated chart generation, and a web-based interface for real-time analysis. The software bridges sophisticated electrical engineering analysis with user-friendly interfaces, making power system diagnostics accessible through conversational AI.

Choi, Seong [National Laboratory of the Rockies (N↗

A universal language for finding mass spectrometry data patterns

Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here, in this study, we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.

Damiani, Tito [Czech Academy of Sciences (CAS), Pr↗

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]↗

Equivariant, safe and sensitive — graph networks for new physics

This study introduces a novel Graph Neural Network (GNN) architecture that leverages infrared and collinear (IRC) safety and equivariance to enhance the analysis of collider data for Beyond the Standard Model (BSM) discoveries. By integrating equivariance in the rapidity-azimuth plane with IRC-safe principles, our model significantly reduces computational overhead while ensuring theoretical consistency in identifying BSM scenarios amidst Quantum Chromodynamics backgrounds. The proposed GNN architecture demonstrates superior performance in tagging semi-visible jets, highlighting its potential as a robust tool for advancing BSM search strategies at high-energy colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine learning-guided design, synthesis, and characterization of atomically dispersed electrocatalysts

The recent integration of machine learning into materials design has revolutionized the understanding of structure–property relationships and optimization of material properties beyond the trial-and-error paradigm. On one hand, machine learning has significantly accelerated the development of atomically dispersed metal-nitrogen-carbon (M-N-C) electrocatalysts, which traditionally heavily relied on heuristic approaches. On the other hand, the primary challenge of leveraging machine learning to expedite M-N-C materials discovery lies in the cost associated with data collection. Here, we review recent machine learning integration strategies for M-N-C catalyst development, including discussions on the typical algorithms such as symbolic regression and convolutional neural networks employed for the theoretical design, synthesis optimization via active learning, and advanced microscopy characterization. Subsequently, we provide our perspective on potential near-future directions for furthering machine learning-assisted development of new M-N-C catalysts and elucidating the complex physicochemical mechanisms governing the selectivity, activity, and durability in this class of materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters

Microbial metabolic gene clusters encode the biosynthesis or catabolism of metabolites that facilitate ecological specialization, mediate microbiome interactions and constitute a major source of medicines and crop protection agents. Here, we present BiG-SCAPE and BiG-SLiCE 2.0, next-generation methods that facilitate scalable, accurate and interactive gene cluster analyses. BiG-SCAPE 2.0 updates its classification, alignment methods, and visualizations, enabling more accurate analysis, up to 8x faster runtimes and halved memory requirements. BiG-SLiCE 2.0 updates its distance metric, pHMM database, and classification logic, resulting in increased sensitivity nearing that of BiG-SCAPE. Analysis of 260,630 biosynthetic gene clusters from publicly available genomes reveals that both tools generate concurring estimates of gene cluster diversity, thus providing significantly extended methodological support for recent evidence indicating that the vast majority of natural product diversity remains unexplored. Together, these updates will facilitate global genome mining efforts for natural product discovery and microbiome analyses scalable with current data sizes.

Draisma, Arjan [Wageningen University & Research (↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Towards a self-driving trigger at the LHC: adaptive response in real time

Real-time data filtering and selection—or trigger—systems at high-throughput scientific facilities such as the experiments at the Large Hadron Collider must process extremely high-rate data streams under stringent bandwidth, latency, and storage constraints. Yet these systems are typically designed as static, hand-tuned menus of selection criteria grounded in prior knowledge and simulation. In this work, we further explore the concept of a self-driving trigger, an autonomous data-filtering framework that reallocates resources and adjusts thresholds dynamically in real-time to optimize signal efficiency, rate stability, and computational cost as instrumentation and environmental conditions evolve. We introduce a benchmark ecosystem to emulate realistic collider scenarios and demonstrate real-time optimization of a menu including canonical energy sum triggers as well as modern anomaly-detection algorithms that target non-standard event topologies using machine learning. Using simulated data streams and publicly available collision data from the Compact Muon Solenoid experiment, we demonstrate the capability to dynamically and automatically optimize trigger performance under specific cost objectives without manual retuning. Our adaptive strategy shifts trigger design from static menus with heuristic tuning to intelligent, automated, data-driven control, unlocking greater flexibility and discovery potential in future high-energy physics analyses.

Emami, Shaghayegh [Michigan U.] (ORCID:00090007589↗

Visualizing metagenomic and metatranscriptomic data: A comprehensive review

The fields of Metagenomics and Metatranscriptomics involve the examination of complete nucleotide sequences, gene identification, and analysis of potential biological functions within diverse organisms or environmental samples. Despite the vast opportunities for discovery in metagenomics, the sheer volume and complexity of sequence data often present challenges in processing analysis and visualization. This article highlights the critical role of advanced visualization tools in enabling effective exploration, querying, and analysis of these complex datasets. Emphasizing the importance of accessibility, the article categorizes various visualizers based on their intended applications and highlights their utility in empowering bioinformaticians and non-bioinformaticians to interpret and derive insights from meta-omics data effectively.

59 BASIC BIOLOGICAL SCIENCES↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

An AI-Enabled Chat Bot for DuraMAT

The DuraMAT Data Hub has evolved to meet the ever expanding research demands by integrating a chatbot interface that enhances FAIR compliance and simplifies discovery across our growing archive of projects and datasets. With the Data Hubs increasing use by both consortium researchers and the international community, traditional navigation methods have become unwieldy. Building on last year's feasibility study, we now report the successful implementation of the chatbot system, detailing its architecture and demonstrating its ability to deliver a more intuitive and rewarding user experience.

14 SOLAR ENERGY↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗

Modular Autonomous Experimentation for Biological Applications (Full Report)

The Modular Autonomous Research System (MARS) was developed to address the pressing need for faster, more reliable, and more adaptable scientific discovery. Traditional experimentation is limited by manual labor, long cycle times, and fragmented data streams, which constrain the ability to explore complex chemical and materials design spaces. To overcome these limitations, we created an integrated, modular platform that combines laboratory robotics, diverse measurement instruments, and a central data infrastructure with artificial intelligence–driven decision-making. The system links liquid handling robots, robotic arms, and optical plate readers into a closed loop where experiments are executed automatically, data is analyzed in real time, and subsequent experimental conditions are adaptively chosen to maximize information gain. Over the course of the project, MARS was validated on two primary test cases—spectroscopic metal–ligand binding assays and peptide-directed mineralization—which highlighted the system’s ability to handle uncertainty and variability in experimental measurements. To further demonstrate modularity and extensibility, we also established additional testbeds in electrochemistry for catalyst discovery and electrolyte formulation for advanced batteries. The results show that MARS can reliably conduct autonomous campaigns with minimal human intervention, adapt to distinct scientific domains, and provide a scalable model for future self-driving laboratories. This work establishes new capabilities for modular, uncertainty-aware automation and directly supports the need for advanced, data-driven research platforms capable of accelerating discovery across a wide range of scientific and national security missions.

59 BASIC BIOLOGICAL SCIENCES↗

Targeted materials discovery using Bayesian algorithm execution

Rapid discovery and synthesis of future materials requires intelligent data acquisition strategies to navigate large design spaces. A popular strategy is Bayesian optimization, which aims to find candidates that maximize material properties; however, materials design often requires finding specific subsets of the design space which meet more complex or specialized goals. We present a framework that captures experimental goals through straightforward user-defined filtering algorithms. These algorithms are automatically translated into one of three intelligent, parameter-free, sequential data collection strategies (SwitchBAX, InfoBAX, and MeanBAX), bypassing the time-consuming and difficult process of task-specific acquisition function design. Our framework is tailored for typical discrete search spaces involving multiple measured physical properties and short time-horizon decision making. We demonstrate this approach on datasets for TiO 2 nanoparticle synthesis and magnetic materials characterization, and show that our methods are significantly more efficient than state-of-the-art approaches. Overall, our framework provides a practical solution for navigating the complexities of materials design, and helps lay groundwork for the accelerated development of advanced materials.

42 ENGINEERING↗