Search NASASearch

SEARCH · Search NASA

Results for “Data-driven discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data-Driven Discovery and Experimental Validation of Solvent Polarity Effects on Conjugated Polymer Solution-to-Film Assembly Pathways

Understanding how solvent properties influence the solution-to-film assembly of conjugated polymers remains a critical challenge due to the complex and intertwined nature of polymer–solvent interactions. In this study, we integrate a data-driven framework with experimental validation to identify key parameters influencing the assembly and performance of poly[2,5-(2-octyldodecyl)-3,6-diketopyrrolopyrrole-alt-5,5-(2,5-di(thien-2-yl)thieno[3,2-b]thiophene)] (DPP-DTT) in organic field-effect transistors (OFETs). A machine learning (ML) approach identified the normalized Reichardt polarity parameter (E T N ) as a significant descriptor correlated with DPP-DTT hole mobility (μ). Systematic DPP-DTT devices fabricated using solvents across a wide E T N range revealed that higher E T N solvents yield enhanced μ. To elucidate the structural origins of high μ, we conducted comprehensive analyses using UV–vis–NIR spectroscopy and grazing incidence wide angle X-ray scattering (GIWAXS) measurements. The results revealed that films processed from high E T N solvents exhibit reduced paracrystallinity. By analyzing the solution-state behavior using optical microscopy and solution WAXS, we revealed polymer solubility differences in the various solvents and associated distinct polymer assembly pathways, elucidating why the high E T N solvent produces long-range ordered films. Notably, the high E T N solvent shows a pronounced preference for liquid-crystal (LC)-mediated assembly, providing a mechanistic explanation for the enhanced structural order. Therefore, these results demonstrate that solvent polarity, as evaluated by E T N , serves as an important parameter that plays a significant role in the DPP-DTT assembly pathway and resultant solid-state morphology. This work provides a strategy for integrating data science with experiments to identify critical parameters associated with complex polymer systems and helps guide rational process design for high-performance organic electronics.

36 MATERIALS SCIENCE

SODAs: sparse optimization for the discovery of differential and algebraic equations

Differential-algebraic equations (DAEs) integrate ordinary differential equations (ODEs) with algebraic constraints, providing a fundamental framework for developing models of dynamical systems characterized by time-scale separation, conservation laws and physical constraints. While sparse optimization has revolutionized model development by allowing data-driven discovery of parsimonious models from a library of possible equations, existing approaches for dynamical systems assume DAEs can be reduced to ODEs by eliminating variables before model discovery. This assumption limits the applicability of such methods for DAE systems with unknown constraints and time scales. We introduce sparse optimization for differential-algebraic systems (SODAs), a data-driven method for the identification of DAEs in their explicit form. By discovering the algebraic and dynamic components sequentially without prior identification of the algebraic variables, this approach leads to a sequence of convex optimization problems. It has the advantage of discovering interpretable models that preserve the structure of the underlying physical system. To this end, SODAs improves since SODAs is singular numerical stability when handling high correlations between library terms, caused by near-perfect algebraic relationships, by iteratively refining the conditioning of the candidate library. We demonstrate the performance of our method on biological, mechanical and electrical systems, showcasing its robustness to noise in both simulated time series and real-time experimental data.

DAE

Data mining the missing ordered phases of Li/Na metal oxides

Data-driven discovery of Li-ion and Na-ion battery materials has been pioneered by generic materials data platforms such as the Materials Project. After decades of progress, it is timely to ask whether there remain underexplored compositional spaces. Here, in this work, we present a systematic data-mining effort to uncover missing ordered binary, ternary and quaternary Li/Na-containing metal oxides using high-throughput density functional theory (DFT). Building on 19,120 stable and metastable oxides entries from the Materials Project, we performed 13,245 additional calculations through isovalent substitutions of known ground states, experimentally reported compounds, and specific prototype structures. Our study identifies 36 new ground states within the GGA/GGA + U convex hull and 45 within the r 2 SCAN convex hull. Additionally, we identified 840 metastable compounds from GGA/GGA + U and 979 from r 2 SCAN that are absent in the present Materials Project databases. Moreover, we have tripled the metastable materials in compositional spaces with a molar ratio of cation/anion >1, highlighting the overlooked opportunities in this compositional space.

25 ENERGY STORAGE

Machine-Learning-Driven Discovery of Water Splitting BaFe 2 O 4 and Human-in-the-Loop Improvement via Al-Substitution for Increased Thermal Stability

Thermochemical hydrogen (TCH) production offers a promising method for converting thermal energy into hydrogen fuel through heat-driven redox cycles of metal oxides. Here, in this work a defect graph neural network (dGNN) was used to predict oxygen vacancy formation energies ΔH V O combined with Materials Project predictions of oxygen chemical potential stability to screen candidate oxides via high-throughput database analysis. BaFe 2 O 4 was identified as a promising material for experimental validation based on its predicted ΔH V O , oxygen chemical potential stability range, and potential for tunable substitutions to improve thermal properties. Experimental validation using thermogravimetric analysis (TGA), stagnation flow reactor (SFR), X-ray diffraction (XRD), and electron microscopy confirmed positive water-splitting behavior but also revealed limitations in thermal stability under aggressive reduction conditions. To address this, a human-in-the-loop modification strategy was employed introducing Al substitution in BaFe 2–x Al x O 4 ; this modification improves thermal stability, alters the crystal structure and enhances overall performance. These results demonstrate a combined computational and experimental workflow in which machine learning accelerates identification of promising candidates, while targeted experimental design enables optimization of functional performance. This approach advances the development of robust, cost-effective TCH materials and highlights the importance of integrating data-driven discovery with human-guided materials design in paving the way for scalable hydrogen production technologies.

organic

Quasars Acting as Strong Lenses Found in DESI DR1

Quasars acting as strong gravitational lenses offer a rare opportunity to probe the redshift evolution of scaling relations between supermassive black holes and their host galaxies, particularly the M$_{BH}$–M$_{host}$ relation. Using these powerful probes, the mass of the host galaxy can be precisely inferred from the Einstein radius θ$_{E}$. Using 812,118 quasars from DESI DR1 (0.03 ≤ z ≤ 1.8), we searched for quasars lensing higher-redshift galaxies by identifying background emission-line features in their spectra. To detect these rare systems, we trained a convolutional neural network (CNN) on mock lenses constructed from real DESI spectra of quasars and emission-line galaxies (ELGs), achieving a high classification performance (AUC = 0.99). We also trained a regression network to estimate the redshift of the background ELG. Applying this pipeline, we identified seven high-quality (Grade A) lens candidates, each exhibiting a strong [O II] doublet at a higher redshift than the foreground quasar; four candidates additionally show Hβ, [O III] λ4959, and [O III] λ5007 emission. These results significantly expand the sample of quasar lens candidates beyond the 12 identified and 3 confirmed in previous work and demonstrate the potential for scalable, data-driven discovery of quasars as strong lenses in upcoming spectroscopic surveys.

McArthur, Everett [Stanford U., Phys. Dept.; KIPAC

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS

Towards a self-driving trigger at the LHC: adaptive response in real time

Real-time data filtering and selection—or trigger—systems at high-throughput scientific facilities such as the experiments at the Large Hadron Collider must process extremely high-rate data streams under stringent bandwidth, latency, and storage constraints. Yet these systems are typically designed as static, hand-tuned menus of selection criteria grounded in prior knowledge and simulation. In this work, we further explore the concept of a self-driving trigger, an autonomous data-filtering framework that reallocates resources and adjusts thresholds dynamically in real-time to optimize signal efficiency, rate stability, and computational cost as instrumentation and environmental conditions evolve. We introduce a benchmark ecosystem to emulate realistic collider scenarios and demonstrate real-time optimization of a menu including canonical energy sum triggers as well as modern anomaly-detection algorithms that target non-standard event topologies using machine learning. Using simulated data streams and publicly available collision data from the Compact Muon Solenoid experiment, we demonstrate the capability to dynamically and automatically optimize trigger performance under specific cost objectives without manual retuning. Our adaptive strategy shifts trigger design from static menus with heuristic tuning to intelligent, automated, data-driven control, unlocking greater flexibility and discovery potential in future high-energy physics analyses.

Emami, Shaghayegh [Michigan U.] (ORCID:00090007589

Ultra-light antennas via charge programmed deposition additive manufacturing

Abstract The demand for lightweight antennas in 5 G/6 G communication, wearables, and aerospace applications is rapidly growing. However, standard manufacturing techniques are limited in structural complexity and easy integration of multiple material classes. Here we introduce charge programmed multi-material additive manufacturing platform, offering unparalleled flexibility in antenna design and the capability for rapid printing of intricate antenna structures that are unprecedented or necessitate a series of fabrication routes. Demonstrating its potential, we present a transmitarray antenna composed of an interconnected, multi-layered array of dielectric/conductive S-ring unit cells, reducing 94% mass of conventional antenna configurations. A fully printed circular polarized transmitarray system fed by a source and a Risley prism antenna system operating at 19 GHz both show close alignment between testing results and numerical simulations. This printing method establishes a universal platform, propelling discovery of new antenna designs and enabling data-driven design and optimizations where rapid production of antenna designs is crucial.

Science & Technology - Other Topics

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

36 MATERIALS SCIENCE

Automated and High-Throughput Phase Separation Control for Supramolecular Polymer Blends Enabled by Machine Learning

Supramolecular polymer blends (SPBs) offer tunable morphologies that dictate their macroscopic properties, yet their rational design is limited by the absence of predictive structure−morphology models. Here, we introduce a data-driven highthroughput workflow that integrates modular polymer synthesis, robotic formulation, automated morphology characterization, and machine learning (ML) for accelerated SPB discovery. Using a plug-and-play synthetic strategy, 33 hydrogen-bonding endfunctional homopolymers were prepared and orthogonally combined to generate 260 SPBs in 1 day. A fully automated atomic force microscopy (AFM) pipeline enabled systematic imaging, producing 2340 morphology data sets with minimal human intervention. Domain spacings were extracted through complementary imageprocessing methods and used to train ML models. A support vector regression (SVR) model accurately predicted target phase-separation sizes (50, 100, and 150 nm), which were experimentally validated. This work demonstrates the power of coupling high-throughput experimentation with ML to accelerate morphology discovery and provides one of the first large-scale experimental data sets for supramolecular polymer systems.

ML-guided polymer design

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE

Bifunctional Electrocatalysts with High-Entropy Alloys: Bridging Hydrogen Evolution and Oxygen Reduction

High-entropy alloys (HEAs) have emerged as a promising class of bifunctional electrocatalysts capable of simultaneously driving the hydrogen evolution reaction (HER) and the oxygen reduction reaction (ORR) with high activity and durability. Their near-equiatomic multicomponent compositions give rise to unique physicochemical characteristics, including lattice distortion, sluggish diffusion, high-entropy stabilization, and pronounced electronic heterogeneity, that collectively generate diverse and synergistic active sites inaccessible in conventional alloys. This review summarizes recent progress in HEA-based bifunctional electrocatalysis, with a focus on the fundamental mechanisms governing HER and ORR activity, stability, and selectivity. We discuss advances in synthesis strategies, ranging from confined growth and step-alloying to scalable continuous-flow methods, that enable precise control over composition, size, and surface structure. Complementary computational and data-driven approaches, including density functional theory, machine-learning-assisted screening, and descriptor development, are highlighted as essential tools for navigating the vast HEA design space and establishing structure−property relationships. Particular attention is paid to adsorption-energy distributions, multisite cooperativity, and environmental effects under realistic electrochemical conditions. Finally, we outline current challenges and future opportunities for integrating mechanistic understanding with AI-guided, closed-loop design frameworks to accelerate the discovery of next-generation HEA bifunctional electrocatalysts for sustainable energy conversion.

Alloys

Finding the perfect imperfection: Accelerated, computationally driven discovery and design of quantum defects

Optically addressable spin defects have emerged as the leading platforms for quantum sensing and communication in solid-state systems. While traditional efforts have concentrated on a focused set of well-studied defects, recent advances in high-throughput computational methods have shown promise for large-scale exploration of defects across diverse semiconductor hosts. By cataloging key properties of quantum defects in computational databases, high-throughput screening techniques can systematically suggest and design novel candidates. In this article, we highlight recent advances in data-driven quantum defect design aimed at addressing critical materials science challenges such as host materials selection, defect stability, and desirable electronic and optical properties. Here, we emphasize the importance of electronic-structure-guided searches across various materials and illustrate how high-throughput computations contribute to our understanding of design principles for quantum defects. Additionally, we outline ongoing challenges and emerging opportunities in this rapidly developing field.

Xiong, Yihuang [Dartmouth College, Hanover, NH (Un

Flavor as an Incomplete Structure: Conceptual Questions and the Role of DUNE

Flavor remains one of the most successful yet least understood structures of the Standard Model. The discovery of the Higgs boson completed the electroweak account of mass generation, but did not explain the origin of fermion families, mass hierarchies, or mixing patterns. In this sense, flavor can be regarded as an empirically successful but conceptually incomplete structure. Neutrinos occupy a particularly sensitive place within this problem: their masses are tiny, their mixing is large, and their mass-generation mechanism may differ from that of charged fermions. In this article, we discuss flavor as an open conceptual problem and argue that DUNE, as a phased program spanning precision oscillation measurements and sensitivity to BSM and dark-sector phenomena, provides a powerful framework for testing the self-consistency and possible limits of the present three-flavor description. In particular, the complementarity between the long-baseline program and the Phase I near-detector complex, together with the DUNE-PRISM strategy for controlling interaction-model systematics and enabling data-driven near-to-far predictions, makes DUNE especially well-suited to search for small, correlated departures from the minimal flavor framework.

Montanari, Claudio S. [Fermilab; INFN, Pavia] (ORC

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization

Detecting thermodynamic phase transition via explainable machine learning of photoemission spectroscopy

Identifying thermodynamic signatures of electronic phases, such as superconductivity, is challenging in low-dimensional materials due to strong fluctuations and low probing volume. Spectroscopic methods are often used to identify new bulk phases, but their main measurable quantity—electronic energy gaps—is no longer an effective order parameter in low-dimensional and fluctuating systems. Combining angle-resolved photoemission with a domain-adversarial neural network, we report a data-driven method to identify thermodynamic phase transitions solely based on single-particle spectra. We demonstrate 97.6% accuracy in cuprate superconductor Bi 2 Sr 2 CaCu 2 O 8+δ with strong superconducting fluctuations. This model notably compensates for the scarcity of experimental data by leveraging virtually inexhaustible simulated data. Further, its explainability reveals the crucial role of in-gap spectral weight in detecting phase fluctuations and thermodynamic transitions. Our work pinpoints the spectroscopic signatures of fluctuating orders and enables using spectroscopy for machine-learning-assisted material discovery for low-dimensional and strong coupling systems.

2D materials