Search NASA⌕ Search

SEARCH · Search NASA

Results for “knowledge discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

AI-NERD: Elucidation of relaxation dynamics beyond equilibrium through AI-informed X-ray photon correlation spectroscopy

Abstract Understanding and interpreting dynamics of functional materials in situ is a grand challenge in physics and materials science due to the difficulty of experimentally probing materials at varied length and time scales. X-ray photon correlation spectroscopy (XPCS) is uniquely well-suited for characterizing materials dynamics over wide-ranging time scales. However, spatial and temporal heterogeneity in material behavior can make interpretation of experimental XPCS data difficult. In this work, we have developed an unsupervised deep learning (DL) framework for automated classification of relaxation dynamics from experimental data without requiring any prior physical knowledge of the system. We demonstrate how this method can be used to accelerate exploration of large datasets to identify samples of interest, and we apply this approach to directly correlate microscopic dynamics with macroscopic properties of a model system. Importantly, this DL framework is material and process agnostic, marking a concrete step towards autonomous materials discovery.

36 MATERIALS SCIENCE↗

Mapping of flumioxazin tolerance in a snap bean diversity panel leads to the discovery of a master genomic region controlling multiple stress resistance genes

Effective weed management tools are crucial for maintaining the profitable production of snap bean (Phaseolus vulgaris L.). Preemergence herbicides help the crop to gain a size advantage over the weeds, but the few preemergence herbicides registered in snap bean have poor waterhemp (Amaranthus tuberculatus) control, a major pest in snap bean production. Waterhemp and other difficult-to-control weeds can be managed by flumioxazin, an herbicide that inhibits protoporphyrinogen oxidase (PPO). However, there is limited knowledge about crop tolerance to this herbicide. We aimed to quantify the degree of snap bean tolerance to flumioxazin and explore the underlying mechanisms. We investigated the genetic basis of herbicide tolerance using genome-wide association mapping approach utilizing field-collected data from a snap bean diversity panel, combined with gene expression data of cultivars with contrasting response. The response to a preemergence application of flumioxazin was measured by assessing plant population density and shoot biomass variables. Snap bean tolerance to flumioxazin is associated with a single genomic location in chromosome 02. Tolerance is influenced by several factors, including those that are indirectly affected by seed size/weight and those that directly impact the herbicide's metabolism and protect the cell from reactive oxygen species-induced damage. Transcriptional profiling and co-expression network analysis identified biological pathways likely involved in flumioxazin tolerance, including oxidoreductase processes and programmed cell death. Transcriptional regulation of genes involved in those processes is possibly orchestrated by a transcription factor located in the region identified in the GWAS analysis. Several entries belonging to the Romano class, including Bush Romano 350, Roma II, and Romano Purpiat presented high levels of tolerance in this study. The alleles identified in the diversity panel that condition snap bean tolerance to flumioxazin shed light on a novel mechanism of herbicide tolerance and can be used in crop improvement.

60 APPLIED LIFE SCIENCES↗

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗

Protein data bank: From two epidemics to the global pandemic to mRNA vaccines and Paxlovid

Structural biologists and the open-access Protein Data Bank (PDB) played decisive roles in combating the COVID-19 pandemic. Global biostructure data were turned into global knowledge, allowing scientists and engineers to understand the inner workings of coronaviruses and develop effective countermeasures. Two mRNA vaccines, initially designed with guidance from PDB structures of the SARS-CoV-1 and MERS-CoV spike proteins, prevented infections entirely or reduced the likelihood of morbidity and mortality for more than five billion individual recipients worldwide. Structure-guided drug discovery by Pfizer, Inc (facilitated by PDB structures), initiated in the 2000s in response to SARS-CoV-1 and resumed in 2020, yielded nirmatrelvir (the active ingredient of Paxlovid) -- a potent, orally-bioavailable inhibitor of the SARS-CoV-2 main protease. You've got to love the Protein Data Bank!

Burley, Stephen K.↗

Thin film combinatorial sputtering of TaTiHfZr refractory compositionally complex alloys for rapid materials discovery

Many applications from advanced nuclear reactors to aerospace and automotive industries require materials to operate in extreme environments. In search of new materials that can operate in these extremes, the present work explores this space whereby: (1) guided by atomistic and thermodynamic calculations we utilize thin film combinatorial synthesis to rapidly explore mechanical and thermal properties in a broad range of refractory compositionally complex alloys, and (2) observe transformation induced plasticity via oscillations in the thin film nanoindentation load depth curves that are attributed to, (3) a stress-induced HCP-to-BCC phase transformation in the resulting nanogranular microstructure, which to our knowledge has not been observed before in this alloy system; and finally (4) scale to bulk materials to compare the thin film results.

36 MATERIALS SCIENCE↗

Building workflows for an interactive human-in-the-loop automated experiment (hAE) in STEM-EELS

Exploring the structural, chemical, and physical properties of matter on the nano- and atomic scales has become possible with the recent advances in aberration-corrected electron energy-loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). However, the current paradigm of STEM-EELS relies on the classical rectangular grid sampling, in which all surface regions are assumed to be of equal a priori interest. However, this is typically not the case for real-world scenarios, where phenomena of interest are concentrated in a small number of spatial locations, such as interfaces, structural and topological defects, and multi-phase inclusions. One of the foundational problems is the discovery of nanometer- or atomic-scale structures having specific signatures in EELS spectra. Herein, we systematically explore the hyperparameters controlling deep kernel learning (DKL) discovery workflows for STEM-EELS and identify the role of the local structural descriptors and acquisition functions in experiment progression. In agreement with the actual experiment, we observe that for certain parameter combinations the experiment path can be trapped in the local minima. We demonstrate the approaches for monitoring the automated experiment in the real and feature space of the system and knowledge acquisition of the DKL model. Based on these, we construct intervention strategies defining the human-in-the-loop automated experiment (hAE). This approach can be further extended to other techniques including 4D STEM and other forms of spectroscopic imaging. The hAE library is available on Github at https://github.com/utkarshp1161/hAE/tree/main/hAE.

Pratiush, Utkarsh [Univ. of Tennessee, Knoxville, ↗

An Atom-Precise Approach to Damp First-Order Phase Transitions and Its Implications for Neuromorphic Signal Processing

Neuromorphic computing inspired by mammalian intelligence aims to emulate the nonlinear dynamics of biological neurons and synapses to achieve fast, low-energy, and highly efficient information processing. Brain-inspired computing relies on the design and discovery of materials exhibiting nonlinear current–voltage profiles, frequently underpinned by electronic state transitions, to achieve spiking neurons and dynamically tunable synapses. A signature challenge in the design of artificial neurons is controlling the steepness of first-order transitions in active elements, as abrupt transitions are at risk of driving unstable voltage and temperature oscillations, which result in catastrophic device failure. A critical knowledge gap is the lack of structure–function correlations mapping the composition and atomistic structure of crystalline solids to nonlinear dynamical response characteristics. Here, we address the key question of how modification of atomistic structure correlates with alteration of neuron-like functionality. Constructing oscillator circuits from millimeter-scale single crystals enables high-resolution atomic structure solutions, which we use to demonstrate that the selective positioning of Pb cations modifies charge ordering along a one-dimensional CuxV2O5 framework even at low insertion stoichiometries, thereby providing an atom-precise design parameter for damping first-order transitions. We use temperature-variant X-ray diffraction and X-ray spectroscopy to elucidate the suppression of Cu-ion shuttling based on the precise positioning of Pb ions in seven-coordinated tunnel interstitial sites as the mechanistic basis for transition broadening, thus bridging a critical gap between statistical mechanics and quantum chemical descriptions of phase transitions. Such mechanistic understanding thus paves the way to site-selective modification strategies for modulating the sharpness of first-order transitions, with an exemplary demonstration here in tuning neuronal signal processing.

Crystal structure↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Screening a knowledge‐based library of low molecular weight compounds against the proline biosynthetic enzyme 1‐pyrroline‐5‐carboxylate 1 ( PYCR1)

Abstract Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 1 (PYCR1) is the last enzyme of proline biosynthesis and catalyzes the NAD(P)H‐dependent reduction of Δ 1 ‐pyrroline‐5‐carboxylate toL‐proline. High PYCR1 gene expression is observed in many cancers and linked to poor patient outcomes and tumor aggressiveness. The knockdown of thePYCR1gene or the inhibition of PYCR1 enzyme has been shown to inhibit tumorigenesis in cancer cells and animal models of cancer, motivating inhibitor discovery. We screened a library of 71 low molecular weight compounds (average MW of 131 Da) against PYCR1 using an enzyme activity assay. Hit compounds were validated with X‐ray crystallography and kinetic assays to determine affinity parameters. The library was counter‐screened against human Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 3 and proline dehydrogenase (PRODH) to assess specificity/promiscuity. Twelve PYCR1 and one PRODH inhibitor crystal structures were determined. Three compounds inhibit PYCR1 with competitive inhibition parameter of 100 μM or lower. Among these, (S)‐tetrahydro‐2H‐pyran‐2‐carboxylic acid (70 μM) has higher affinity than the current best tool compoundN‐formyl‐l‐proline, is 30 times more specific for PYCR1 over human Δ 1 ‐pyrroline‐5‐carboxylate reductase isoform 3, and negligibly inhibits PRODH. Structure‐affinity relationships suggest that hydrogen bonding of the heteroatom of this compound is important for binding to PYCR1. The structures of PYCR1 and PRODH complexed with 1‐hydroxyethane‐1‐sulfonate demonstrate that the sulfonate group is a suitable replacement for the carboxylate anchor. This result suggests that the exploration of carboxylic acid isosteres may be a promising strategy for discovering new classes of PYCR1 and PRODH inhibitors. The structure of PYCR1 complexed withl‐pipecolate and NADH supports the hypothesis that PYCR1 has an alternative function in lysine metabolism.

Biochemistry & Molecular Biology↗

Cluster expansion by transfer learning for phase stability predictions

Recent progress towards universal machine-learned interatomic potentials holds considerable promise for materials discovery. Yet the accuracy of these potentials for predicting phase stability may still be limited. In contrast, cluster expansions provide accurate phase stability predictions but are computationally demanding to parameterize from first principles, especially for structures of low dimension or with a large number of components, such as interfaces or multimetal catalysts. We overcome this trade-off via transfer learning. Using Bayesian inference, we incorporate prior statistical knowledge from machine-learned and physics-based potentials, enabling us to sample the most informative configurations and to efficiently fit first-principles cluster expansions. Furthermore, this algorithm is tested on Pt:Ni, showing robust convergence of the mixing energies as a function of sample size with reduced statistical fluctuations.

36 MATERIALS SCIENCE↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

PAH101: A GW+BSE Dataset of 101 Polycyclic Aromatic Hydrocarbon (PAH) Molecular Crystals

Abstract The excited-state properties of molecular crystals are important for applications in organic electronic devices. TheGWapproximation and Bethe-Salpeter equation (GW+BSE) is the state-of-the-art method for calculating the excited-state properties of crystalline solids with periodic boundary conditions. We present the PAH101 dataset ofGW+BSE calculations for 101 molecular crystals of polycyclic aromatic hydrocarbons (PAHs) with up to ~500 atoms in the unit cell. To the best of our knowledge, this is the firstGW+BSE dataset for molecular crystals. The data records include theGWquasiparticle band structure, the fundamental band gap, the static dielectric constant, the first singlet exciton energy (optical gap), the first triplet exciton energy, the dielectric function, and optical absorption spectra for light polarized along the three lattice vectors. The dataset can be used to (i) discover materials with desired electronic/optical properties, (ii) identify correlations between DFT andGW+BSE quantities, and (iii) train machine learned models to help in materials discovery efforts.

Science & Technology - Other Topics↗

Detection of visible-wavelength aurora on Mars

Mars hosts various auroral processes despite the planet’s tenuous atmosphere and lack of a global magnetic field. To date, all aurora observations have been at ultraviolet wavelengths from orbit. We describe the discovery of green visible-wavelength aurora, originating from the atomic oxygen line at 557.7 nanometers, detected with the SuperCam and Mastcam-Z instruments on the Mars 2020 Perseverance rover. Near–real-time simulations of a Mars-directed coronal mass ejection (CME) provided sufficient lead-time to schedule an observation with the rover. The emission was observed 3 days after the CME eruption, suggesting that the aurora was induced by particles accelerated by the moving shock front. To our knowledge, detection of aurora from a planetary surface other than Earth has never been reported, nor has visible aurora been observed at Mars. This detection demonstrates that auroral forecasting at Mars is possible, and that during events with higher particle precipitation, or under less dusty atmospheric conditions, aurorae will be visible to future astronauts.

Science & Technology - Other Topics↗

Advancing International Integration and Strengthening Responsible Peaceful Uses of Nuclear Applications Through Specialized Curriculums

Nuclear technology has been pivotal in addressing some of the most pressing global challenges, ranging from energy production to combatting infectious diseases to agricultural security. Oak Ridge National Laboratory (ORNL), as a leader in nuclear research and development including in the field of radioisotopes, is well situated to share lessons learned and enhance global collaboration from years of discoveries in the field. Accordingly, the U.S. Department of Energy’s National Nuclear Security Administration (NNSA) sponsors specialized educational programs at ORNL, with support from the IAEA, focused on advancing peaceful nuclear applications and associated industries while upholding strong nuclear safety, security, and safeguards standards. The Joint U.S./IAEA International School on Peaceful Uses of Nuclear Applications, launched in 2024, is a cornerstone of this collaboration. The school provides an opportunity for early-career professionals from around the globe to gain practical knowledge and skills in utilizing nuclear technologies for peaceful purposes. This initiative demonstrates the United States’ commitment to its obligations under the Treaty on the Non-Proliferation of Nuclear Weapons (NPT), specifically Article IV, which calls on nuclear-weapon states to facilitate access to the peaceful uses of nuclear energy while guarding against the proliferation of nuclear weapons. In this paper, we explore the curriculum, objectives, and global impact of the program. Participants of the ICARST-2025 conference are invited to learn more about this program, contribute to its development, and explore opportunities for collaboration. This school exemplifies how strategic partnerships, and educational initiatives can drive the peaceful and beneficial use of nuclear technology worldwide

Raffo Caiado, Ana [ORNL] (ORCID:0009000239304805)↗

Physics-Guided Continual Learning for Predicting Emerging Aqueous Organic Redox Flow Battery Material Performance

Aqueous organic redox flow batteries (AORFBs) have gained popularity in renewable energy storage due to their low cost, environmental friendliness and scalability. The rapid discovery of aqueous soluble organic (ASO) redox-active materials necessitates efficient machine learning surrogates for predicting battery performance. The physics-guided continual learning (PGCL) method proposed in this study can incrementally learn data from new ASO electrolytes while addressing catastrophic forgetting issues in conventional machine learning. Using a AORFB database with a thousand potential materials generated by a 780 $\text{cm}^2$ interdigitated cell model, PGCL incorporates AORFB physics to optimize the continual learning task formation and training strategies to retain previously learned battery material knowledge. Finally, the trained PGCL demonstrates its capability in assessing emerging ASO materials within the established parameter space when evaluated with the dihydroxyphenazine isomers.

25 ENERGY STORAGE↗

The power of lanthanides: same composition, but different lanthanides leading to different interesting materials properties, from magnetocalorics to molecular magnets and phosphors

Commonly accepted design concepts for ionic liquids (ILs) state that the constituting ions must be large and carry low, well-dispersed charges. A series of ILs based of pentadeca charged ILs with pentanuclear linear {Ln 5 } units ([Ln 5 (C 2 H 5 -C 3 H 3 N 2 -CH 2 COO) 16 (H 2 O) 8 ](Tf 2 N) 15 (C 3 H 3 N 2 = imidazolium moiety, Tf 2 N = bis(trifluoromethanesulfonyl)amide) with Ln = Er, Ho, Tm) demonstrates that these criteria are not absolute. Highly charged ions can also support IL formation, provided they are sufficiently large. Expanding the series of these unconventional, record pentadeca charged with new lanthanide representatives, led to the discovery of additional unprecedented properties for ILs: The Gd compound exhibits a strong magnetocaloric effect (MCE) in the liquid state with a maximum magnetic entropy change of −ΔS M = −11 J⋅kg −1 ⋅K −1 at 2 K for Δμ 0 H = 7 T. Albeit the Dy representative shows slow magnetic relaxation, the relaxation times are not favorable for practical application as a molecular magnet. Lastly, for both the Gd and the Y compound, phosphorescence in the seconds time scale is observed, which is, to the best of our knowledge, the longest ever reported for an IL.

Ionic Liquids↗