Search NASASearch

SEARCH · Search NASA

Results for “active learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A 3D Citizen Science Video Game for NeMO-Net, the NASA Neural Multi-Modal Observation and Training Network for Global Coral Reef Assessment

NeMO-Net, the NASA neural multi-modal observation and training network for global coral reef assessment, is an open-source deep convolutional neural network aimed at accurately assessing the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. We present here the active learning component of the project, which consists of an interactive video game prototype for tablet and mobile devices where players are able to intuitively label morphology classifications over mm-scale 3D coral reef imagery. Active learning applications present a novel methodology for engaging the public while efficiently providing large-scale training and test data for increasingly complex and data-intensive machine learning algorithms. NeMO-Net trains players on domain-specific knowledge through interactive tutorials and periodically checks players' input against pre-classified coral imagery to gauge their accuracy and utilize in-game mechanics to provide personalized classification training. Players can rate the classifications of other players, unlock rewards and join a global community as they explore and classify coral reefs and other shallow marine environments.

Citizen Science

Progress in Normalizing Flows for 4d Gauge Theories

Normalizing flows have arisen as a tool to accelerate Monte Carlo sampling for lattice field theories. This work reviews recent progress in applying normalizing flows to 4-dimensional nonabelian gauge theories, focusing on two advancements: an architectural improvement referred to as learned active loops, and the application of correlated ensemble methods to QCD with N f = 2 dynamical fermions.

Abbott, Ryan [Massachusetts Institute of Technolog

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo

Learning Extended Finite State Machines

We present an active learning algorithm for inferring extended finite state machines (EFSM)s, combining data flow and control behavior. Key to our learning technique is a novel learning model based on so-called tree queries. The learning algorithm uses the tree queries to infer symbolic data constraints on parameters, e.g., sequence numbers, time stamps, identifiers, or even simple arithmetic. We describe sufficient conditions for the properties that the symbolic constraints provided by a tree query in general must have to be usable in our learning model. We have evaluated our algorithm in a black-box scenario, where tree queries are realized through (black-box) testing. Our case studies include connection establishment in TCP and a priority queue from the Java Class Library.

Register Automata

Benchtop Autonomous Electrochemical Characterization System for Combinatorial Thin-Film Solid Oxide Electrodes

The design of materials for electrochemical energy conversion is complicated by a vast search space of candidate materials and multifaceted property requirements: multicarrier conductivity, stability, and catalytic activity are all necessary but rarely intersect. Although self-driving laboratories are rapidly rising to address such material optimization problems, the required infrastructure for integrated, large-scale robotic facilities can be cost-prohibitive. Here we develop and evaluate a closed-loop measurement system for efficient screening of proton-conducting oxide electrodes for ceramic fuel cells and electrolyzers, building on top of an existing benchtop instrument and integrating techniques for rapid impedance measurement and automated analysis. This system exemplifies a “minimum viable” self-driving implementation that can deliver substantial benefits with relatively simple infrastructure. Combinatorial thin-film microelectrode libraries are characterized with a recently developed joint time-domain and frequency-domain impedance measurement technique, which provides an order-of-magnitude acceleration relative to conventional impedance spectroscopy. The distribution of relaxation times is extracted from impedance data and analyzed without human intervention. These results feed an active learning and Bayesian optimization process that learns to predict electrochemical impedance as a function of material composition, measurement temperature, oxygen partial pressure, and electrical bias, which further reduces the screening time by tenfold with optimized experimental sequences. We apply this system to Ba⁡(Co,Fe,Zr,Y)⁢O 3−𝛿 combinatorial libraries and evaluate its effectiveness for learning material property trends and optimizing expensive-to-evaluate properties such as activation energy. This offers insights into key methodological aspects of practical autonomous experimentation, including surrogate model validation, cost-aware acquisition functions, and high-throughput data interpretation. Our results demonstrate the efficacy of the system for rapidly gathering information, but also highlight real-world experimental challenges of thin-film degradation and numerical instability in surrogate models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

How Silica Surface Chemistry Modulates Interfacial Water: Insights from Machine Learning Molecular Dynamics

Controlling water structure and dynamics at silica interfaces are central to a wide range of technologies, including protective oxide layers for solar water splitting and nanoporous membranes. In this work, we develop a machine learning interatomic potential, trained via active learning, to achieve ab initio accuracy for water confined between hydroxylated silica surfaces over a range of silanol coverages and slit widths. We find that partially hydroxylated surfaces (50 and 75% OH) support stronger water−surface hydrogen bonding and more extended interfacial density profiles than fully hydroxylated (100% OH) surfaces, indicating that increasing OH coverage does not necessarily strengthen interfacial hydrogenbond networks. Translational diffusion decreases approximately linearly with slit width and OH coverage, whereas rotational dynamics respond nonlinearly. In particular, at the smallest slit width of 5 Å, 75% OH coverage produces an enhanced local tetrahedral ordered interfacial network that strongly suppresses reorientation, while 100% coverage yields a crowded, disordered interfacial layer that also hinders rotation. In contrast, the 50% OH coverage is sufficiently sparse that it does not markedly alter water structure or dynamics under confinement. These results show that coupled control of pore size and surface chemistry enables nonlinear tuning of interfacial water structure and transport, providing a design strategy for optimizing porous silica for either enhanced interfacial stability and controlled reactivity or rapid and selective transport.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Dataset, Code, and Models for Training Deep Learning Potentials for Low Temperature Plasma-Surface Interactions

This repository contains datasets, training scripts, and finished models, and test simulations used in the development of DeepREBO— a machine-learned interatomic potential trained to emulate the REBO2 empirical potential. The data was generated to study deep potential development for simulations of plasma-surface interactions. It uses an active learning framework, starting from a minimal dataset and iteratively expanding it. Included are those generated datasets, the trained models, and simulations used to evaluate the performance of the training process. This resource supports reproducibility and provides a reference framework for training deep potentials in plasma-surface interaction studies.

active learning

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic

Bayesian Statistics and Uncertainty Quantification for Safety Boundary Analysis in Complex Systems

The analysis of a safety-critical system often requires detailed knowledge of safe regions and their highdimensional non-linear boundaries. We present a statistical approach to iteratively detect and characterize the boundaries, which are provided as parameterized shape candidates. Using methods from uncertainty quantification and active learning, we incrementally construct a statistical model from only few simulation runs and obtain statistically sound estimates of the shape parameters for safety boundaries.

Active Learning

Machine‐Learning‐Driven Exploration of Surface Reconstructions of Reduced Rutile TiO 2

Abstract Titanium dioxide (TiO 2 ) is widely used as a catalyst support due to its stability, tunable electronic properties, and surface oxygen vacancies, which are crucial for catalytic processes such as the reverse water‐gas shift (RWGS) reaction. Reduced TiO 2 surfaces undergo complex surface reconstructions that endow unique properties but are computationally challenging to describe. In this study, we utilize machine‐learning interatomic potentials (MLIPs) integrated with an active‐learning workflow to efficiently explore reduced rutile TiO 2 surfaces. This approach enabled the prediction of a phase diagram as a function of oxygen chemical potential, revealing a variety of reconstructed phases, including a previously unreported subsurface shear plane structure. We further investigate the electronic properties of these surfaces and validate our results by comparing experimental and theoretical high‐resolution transmission electron microscopy (HRTEM). Our findings provide new insights into how extreme surface reductions influence the structural and electronic properties of TiO 2 , with potential implications for catalyst design.

Lee, Yonghyuk [Chemistry and Biochemistry Universi

An Atomistic Study of Reactivity in Solid-State Electrolyte Interphase Formation for Li/Li7P3S11

Lithium metal batteries offer superior volumetric and gravimetric specific capacities compared to those based on traditional graphite anodes. Although advancements in solid-state electrolytes address safety concerns, challenges remain, particularly regarding interphase formation in lithium metal anodes. This work presents a computational framework based on high-throughput first-principles density functional theory and machine-learning interatomic potentials (MLIPs) including automated iterative, active learning to enable robust computational exploration of interphase formation between lithium metal anodes and an inorganic solid-state electrolyte. As a demonstration, we apply the framework to a Li/Li7P3S11 interface and find that it accurately identifies the experimentally observed, thermodynamically stable interphase products as well as their overall spatial arrangement within a heterogeneous, amorphous layered structure, with Li2S domains of nanocrystallinity. Our simulations show two stages, a fast and slow diffusion reaction regime, that corroborate the relative phase formation rate of Li x P, Li2S, and Li3P. Using the Onsager transport theory, we capture time-dependent ionic diffusion within the reacting interface, including cross-correlation effects. We found that cross-correlation effects between Li-P and P-S ionic motion significantly influence P-ion diffusion, making it highly sensitive to the local environment and potentially leading to "kinetic trapping" of Li-P phases. The passivation of the interface is shown as the ionic fluxes all approach zero, effectively halting interphase growth.

Diffusion

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning

Efficient Calibration of Expensive Computational Models

Accounting for uncertainty when calibrating expensive computational models is a common challenge faced by scientists and engineers. Often Bayesian techniques are adopted to estimate a probability density function over the model parameters given noisy empirical data. The methods used to perform this type of probabilistic calibration are computationally prohibitive in that they require a large number of evaluations of the expensive model. In these cases, surrogate modeling -- that is, using a fast-to-evaluate, lower fidelity stand-in for the original computational model -- may be the only option to alleviate this computational burden. However, the upfront cost of generating training data to build a surrogate model can itself be expensive. As such, it is important to be judicious when selecting training points at which the full-fidelity model is evaluated. Here, an active learning approach is proposed that enables efficient selection of training points using approximate samples of the calibrated parameter probability density function. In this way, the training points can be concentrated in regions where the calibration algorithm requires high model accuracy.

active learning

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning

Machine‐Learning‐Driven Exploration of Surface Reconstructions of Reduced Rutile TiO 2

Titanium dioxide (TiO 2 ) is widely used as a catalyst support due to its stability, tunable electronic properties, and surface oxygen vacancies, which are crucial for catalytic processes such as the reverse water-gas shift (RWGS) reaction. Reduced TiO 2 surfaces undergo complex surface reconstructions that endow unique properties but are computationally challenging to describe. In this study, we utilize machine-learning interatomic potentials (MLIPs) integrated with an active-learning workflow to efficiently explore reduced rutile TiO 2 surfaces. This approach enabled the prediction of a phase diagram as a function of oxygen chemical potential, revealing a variety of reconstructed phases, including a previously unreported subsurface shear plane structure. We further investigate the electronic properties of these surfaces and validate our results by comparing experimental and theoretical high-resolution transmission electron microscopy (HRTEM). Our findings provide new insights into how extreme surface reductions influence the structural and electronic properties of TiO 2 , with potential implications for catalyst design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Active deep kernel learning of molecular properties from structural embeddings

As vast databases of chemical identities become increasingly available, the challenge shifts to how we effectively explore and leverage these resources to study molecular properties. This paper presents an active learning approach for molecular discovery using deep kernel learning (DKL), demonstrated on the QM9 dataset. DKL links structural embeddings directly to properties, creating organized latent spaces that prioritize relevant property information. By iteratively recalculating embedding vectors in alignment with target properties, DKL uncovers concentrated maxima representing key molecular properties and reveals unexplored regions with potential for innovation. This approach underscores DKL’s potential in advancing molecular research and discovery.

Artificial neural networks

WIP: Collaborating to Introduce Second-Life Battery Technology Research in an Informal STEM Learning Environment

This work in progress, innovative practice, paper outlines an interdisciplinary collaborative model utilized by a team of educational researchers, and engineers to design a series of community-centered, place-based learning activities suitable for a middle or high school informal STEM learning environment. Based on the potential applications and benefits of Second-Life Electric Vehicle Batteries, the activities connect cutting-edge STEM research with multidimensional real-world challenges within their communities related to climate change, renewable energy integration and storage, intelligent transportation systems, and optimized energy management, that empowers students to think critically about the role they can play in shaping a more resilient, environmentally sustainable community.

Whiteside, Hope

Optimization and stabilization of Fermilab Booster using hybrid Bayesian/RL framework

PIPII project will raise Fermilab Booster intensity and ramp rate. Beam losses will limit average power and are hard to simulate. Presently, Booster uses operator-guided empirical tuning. This task is challenging due to high dimensionality, multiple objectives, critical safety constraints, and drifts. We developed a synergistic suite of Bayesian optimization (BO) and reinforcement learning (RL) tools to optimize and stabilize beam losses. First, active learning was used to build a rough model. Data was collected parasitically using two novel safety constraint types – nonlinear input space restrictions (based on optics model), and uncertainty constraints (to stop bad steps/beam aborts). We then applied online multi-objective BO with scalarized objectives and fitting to improve/rebalance losses, increasing safety margins by 25%. Using BO model as a safety veto, we tried several on/off-policy RL agents for long term stabilization; SAC had best performance. We found that adding contextual (state) information further improved performance, eventually integrating key knobs like linac phase and temperature into the parameter space. Long term testing is ongoing to enable operational use.

Kuklev, Nikita [Fermilab]