Search NASA⌕ Search

SEARCH · Search NASA

Results for “structure prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

Likelihood-based interactive local docking into cryo-EM maps in ChimeraX

The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.

59 BASIC BIOLOGICAL SCIENCES↗

Topological grain boundary segregation transitions

Engineering the structure of grain boundaries (GBs) by solute segregation is a promising strategy to tailor the properties of polycrystalline materials. Solute segregation triggering phase transitions at GBs has been suggested theoretically to offer different pathways to design interfaces, but an understanding of their intrinsic atomistic nature is missing. Here, we combined atomic resolution electron microscopy and atomistic simulations to discover that iron segregation to GBs in titanium stabilizes icosahedral units (“cages”) that form robust building blocks of distinct GB phases. Owing to their five-fold symmetry, the iron cages cluster and assemble into hierarchical GB phases characterized by a different number and arrangement of the constituent icosahedral units. Our advanced GB structure prediction algorithms and atomistic simulations validate the stability of these observed phases and the high excess of iron at the GB that is accommodated by the phase transitions.

36 MATERIALS SCIENCE↗

SFold v0.1

This is a scientific software package to integrate Small Angle X-ray Scattering (SAXS) experimental data into OpenFold deep learning models to improve protein structure prediction.

Prince, Stephanie [Lawrence Berkeley National Labo↗

Rational Design of Lanmodulin Variants for Size-Based Selectivity of Individual Rare Earth Elements

Rare earth elements (REEs) are essential to modern technologies, yet their high physical and chemical similarity makes separation of individual REEs difficult and environmentally taxing. Metalloproteins offer a promising alternative for selective REE binding, as they tend to have high metal ion affinity and specificity. Lanmodulin (LanM), in particular, has arisen as a potential candidate for REE separation as it exhibits picomolar affinity for elements in the REE family. Prior work has shown that the single point mutation D9N can shift LanM’s preference away from lanthanides toward actinides, motivating efforts to tune selectivity of LanM through targeted mutagenesis. Here, we tested the hypothesis that introducing selective aspartic acid to glutamic acid substitutions in the metal coordinating EF hands of LanM would impose steric constraints that would drive LanM affinity away from larger ions, such as La3+, to smaller ions, such as Y3+. To test this hypothesis, a combination of computational and experimental approaches were employed to evaluate the signal mutations LanM D5E and LanM D3E and the double mutants LanM D1ED5E and LanM D3ED9E. Surprisingly, increasing the number of mutations within the metal center did not enhance affinity for smaller REEs, or decrease affinity for larger ions. Only the single point mutation LanM D5E weakened La3+ binding by one order of magnitude relative to LanM wild type (WT), and pairing it with a second mutation to produce LanM D1ED5E drove La3+ affinity to be stronger than that seen for LanM WT. The D3E mutation alone prevented proper expression and folding, but paring it with D9E to produce LanM D3ED9E rescued expression and yielded La3+ affinities comparable to LanM WT. All variants that expressed (LanM D5E, LanM D1ED5E, LanM D3ED9E) displayed Y3+ affinities comparable to LanM WT. Overall, these results highlight the tunability of LanM’s metal-binding environment but also expose current limitations in predicting structural responses to point mutations within a protein sequence. This work establishes a foundation that can be used for refining computational and experimental strategies to engineer metalloproteins with tailored REE selectivity.

Close, Emily [Pacific Northwest National Laborator↗

Prediction of α $IIb$ $β$ 3 integrin structures along its minimum free energy activation pathway

The adhesion protein integrin is a transmembrane heterodimer that plays a pivotal role in cellular processes such as cell signaling and cell migration. To execute its function, integrin undergoes extensive conformational changes from a bent-closed to an extended-open state. Resolving the structures across these changes remains a challenge with both experimental and computational methods, but it is crucial for understanding the activation mechanism of integrin. We address this challenge for the platelet integrin α IIb β 3 by employing finite temperature string method with structures of the images along the initial guess path generated by a multiscale data-driven framework. The full-length all-atom structures along the resulting minimum free energy path between the inactive bent-closed and active extended-open states of α IIb β 3 integrin are consistent with a variety of experimentally resolved structures. Changes in these predicted structures along the path show that the extension and separation of the α and β subunits from the bent-closed to the extended-open state require correlated movements between the subdomain pairs in α IIb β 3 . Furthermore, these results provide new insights into integrin activation mechanism, and the predicted structures have potential applications in guiding the design of integrin-targeting therapeutics.

Dasetty, Siva [University of Chicago, IL (United S↗

Random forest prediction of crystal structure from electron diffraction patterns incorporating multiple scattering

Diffraction is the most common method to solve for unknown or partially known crystal structures. However, it remains a challenge to determine the crystal structure of a new material that may have nanoscale size or heterogeneities. Here, in this study, we train an architecture of hierarchical random forest models capable of predicting the crystal system, space group, and lattice parameters from one or more unknown two-dimensional electron diffraction patterns. Our initial model correctly identifies the crystal system of a simulated electron diffraction pattern from a 20-nm-thick specimen of arbitrary orientation 67% of the time. We achieve a topline accuracy of 79% when aggregating predictions from ten patterns of the same material but different zone axes. The space group and lattice predictions range from 70% to 90% accuracy and median errors of 0.01-0.5Å, respectively, for cubic, hexagonal, trigonal, and tetragonal crystal systems while being less reliable on orthorhombic and monoclinic systems. We apply this architecture to a four-dimensional scanning transmission electron microscopy scan of gold nanoparticles, where it accurately predicts the crystal structure and lattice constants. These random forest models can be used to significantly accelerate the analysis of electron diffraction patterns, particularly in the case of unknown crystal structures. Additionally, due to the speed of inference, these models could be integrated into live transmission electron microscopy experiments, allowing real-Time labeling of a specimen.

36 MATERIALS SCIENCE↗

Deep Learning Prediction of Protein Complex Structures

Proteins interact to form protein complex to carry out biological functions such as catalytic chemical reaction. Therefore, it is important to develop computational methods to predict protein-protein interaction and the structures of protein complexes to study and enhance protein function. In this project, we successfully developed several deep learning methods to predict inter-protein contacts and the reinforcement learning and optimization methods to reconstruct protein complex structures from predicted inter-chain contacts. The methods were integrated with the MULTICOM protein complex structure prediction system and applied to predict the complex structures of biomass production-related proteins of green algae. During the two and a half years of research and development, all the specific milestones of the project were achieved successfully. 16 publications/manuscripts were produced. 10 software tools were developed. A patent application was submitted. Our MULTICOM predictors leveraging some tools developed in this project were ranked among the top predictors in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022.

59 BASIC BIOLOGICAL SCIENCES↗

Virtual Growth of SRF Materials: A Machine Learning Approach to Predict the Crystalline Structural Ordering in Nb Surface Oxides

Niobium's native surface oxide affects SRF cavity and superconducting qubit performance, motivating interest in controlling its crystalline structure. We combine a literature-derived machine-learning analysis with temperature-dependent XRD to study crystalline ordering in Nb2O5. Random Forest models, trained on 74 processing conditions from 17 papers and validated by leave-one-group-out cross-validation, predicted broad crystallinity outcomes well (balanced accuracy 0.809), but struggled with specific polymorph identity (0.577). Annealing temperature was the dominant predictor across all targets; oxygen partial pressure showed negligible importance, reflecting narrow literature coverage rather than physical irrelevance. Temperature-dependent XRD on anodized and H2O2-treated Niobium showed structural evolution consistent with the machine learning predictions. Our model and overall approach provide a data-driven framework for identifying and optimizing conditions that promote crystallization in initially amorphous oxides. This framework can guide the selection of growth and post-annealing conditions for Nb surfaces by narrowing the experimental parameter space, thereby reducing trial-and-error efforts in developing oxide structures relevant to SRF applications.

Tilkin, Anthony [Unlisted, US, IL; Fermilab]↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗

A Comparison of Electronic Structure Methods for Predicting the Hydrogenation Energies of Candidate Molecules for Hydrogen Storage

The development of novel energy materials and fuels is required to expand current available energy sources. Aiming to reach this goal, there is growing interest in using molecular hydrogen as an energy carrier due to its abundance and high energy density. Liquid organic hydrogen carriers (LOHCs) are a promising route to the large-scale storage and transport of hydrogen for use in the energy economy. The search for thermodynamically viable LOHC molecules for real world use has led to a set of constraints on the dehydrogenation enthalpy and the minimum gravimetric hydrogen capacity. These constraints allow one to formulate the search for an ideal LOHC candidate molecule as an optimization problem well suited to the strengths of machine learning and artificial intelligence computational approaches. A critical barrier to a large-scale, high-throughput screening of LOHC candidate molecules is the lack of reliable training data. Computational electronic structure methods including density functional theory, coupled cluster approximations, and diffusion Monte Carlo can be used to provide training data where experimental data are either unreliable or do not exist. In this work, we use these methods to calculate the dehydrogenation energies and enthalpies of candidate LOHC molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Random Forest Prediction of Crystal Structure from Electron Diffraction Patterns

Transmission electron microscopy (TEM) diffraction patterns are regularly used to determine the structure of crystalline materials. Electron diffraction is the most common method to solve for unknown or partially known crystal structures, as it provides direct and interpretable feedback on the orientation of crystal grains under the beam [1]. However, it remains a challenge to determine the crystal structure of a new material or even a new phase of an existing material. Analysis of such materials commonly requires manual exploration and comparison with simulated diffraction patterns. This is often a time consuming process with no obvious start point when many similar structures are possible, and this method cannot be used to determine crystal structure or orientation from structures not included in the diffraction libraries. Therefore, we have developed a machine learning model to determine the crystal structure of a material from its electron diffraction pattern.

36 MATERIALS SCIENCE↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

A brief review on strain engineering of ferroelectric K x Na 1- x NbO 3 epitaxial thin films: Insights from phase-field simulations

Strains play a pivotal role in determining the phase equilibrium, domain configuration, and functional properties of the low-dimensional ferroelectrics. There is growing interest in the strain engineering of ferroelectric K x Na 1- x NbO 3 (KNN) epitaxial thin films, which exhibit excellent physical properties and promise as eco-friendly alternatives to lead-based ferroelectrics for microdevice applications. Further, advances have been made in understanding the phase equilibria and transitions, domains and domain walls, and their relations to the physical properties of KNN epitaxial thin films using a combination of experiments and theoretical modeling, particularly phase-field simulations. Here, we review recent progress in these aspects and showcase the phase-field method for establishing strain phase diagrams, elucidating the domain and domain wall structures at equilibrium, and predicting the structure–property relationships in ferroelectric KNN thin films. We also discuss challenges and opportunities to further advance our understanding of KNN thin films and potentially unlock new functionalities by leveraging phase-field simulations.

36 MATERIALS SCIENCE↗

Divergent carbon use efficiency-growth rate tradeoff in popular biological growth models

Carbon use efficiency (CUE) is an important trait emerging from processes regulating biological growth. CUE can be computed either based on the growth of structural biomass or total biomass divided by substrate uptake rate. Nonequilibrium thermodynamics and observations suggest that, for an exponentially growing population of cells, structural biomass CUE should first increase, then peak, and finally decrease with specific growth rate; meanwhile, total biomass CUE increases asymptotically with specific growth rate. We compared predictions from six popular models that are often used for plant and microbial growth in existing ecosystem models. We found that, for an exponentially growing population of biological cells, (1) the source-driven Pirt and Compromise models predict that structural biomass CUE increase asymptotically with growth rate; (2) the apparent sink-driven modified Droop model predicts that structural biomass CUE decreases with growth rate; and (3) the sink-driven variable internal storage model and two dynamic energy budget models predict that structural biomass CUE first increases, then peaks, and finally decreases with growth rate. Moreover, the modified Droop model predicts that total biomass CUE is constant with growth rate, while all other five models predict that total biomass CUE increases with growth rate asymptotically. For non-exponential biological growth, we show that there is no static relationship between total biomass CUE or structural biomass CUE with respect to either growth rate or temperature. Therefore, we contend that biological growth models should explicitly represent interactions between substrate acquisition, substate transformation, and maintenance respiration to better capture observed CUE dynamics, and the sink-driven model should be preferred for general ecosystem biogeochemistry modeling.

Tang, Jinyun [Lawrence Berkeley National Laborator↗

Transferable predictions of energetic and structural properties for refractory solid solution alloys across chemical compositions

We present a data-efficient approach to train graph neural networks (GNNs) on density functional theory (DFT) data for accurate and transferable predictions of energetic and structural properties of refractory solid solution alloys in the niobium-tantalum-vanadium (Nb-Ta-V) chemical space. We start by training the GNN model only on DFT data that describes refractory binary alloys niobium-tantalum (Nb-Ta), niobium-vanadium (Nb-V), and tantalum-vanadium (Ta-V) to predict formation enthalpy and root mean squared displacement. Once trained, the GNN predictions are tested on DFT data describing refractory ternary alloys Nb-Ta-V. While, unsurprisingly, direct transferability from binary to ternary is not sufficiently accurate, augmenting the training with only 1% of the available ternary data (uniformly distributed across the entire range of chemical compositions) improves significantly the quality of the GNN predictions. For comparison, we assess the transferability in the opposite direction by training GNN models on ternary Nb-Ta-V data and making predictions on binaries Nb-Ta, Nb-V, and Ta-V, which exhibits notably higher predictive errors. The proposed methodology, which favors transferability from lower-component to higher-component alloys, offers an efficient path towards avoiding the curse of dimensionality incurred when collecting DFT data for discovery and design of multi-component disordered alloys.

Density functional theory calculations↗

Discovering type I cis-AT polyketides through computational mass spectrometry and genome mining with Seq2PKS

Type 1 polyketides are a major class of natural products used as antiviral, antibiotic, antifungal, antiparasitic, immunosuppressive, and antitumor drugs. Analysis of public microbial genomes leads to the discovery of over sixty thousand type 1 polyketide gene clusters. However, the molecular products of only about a hundred of these clusters are characterized, leaving most metabolites unknown. Characterizing polyketides relies on bioactivity-guided purification, which is expensive and time-consuming. To address this, we present Seq2PKS, a machine learning algorithm that predicts chemical structures derived from Type 1 polyketide synthases. Seq2PKS predicts numerous putative structures for each gene cluster to enhance accuracy. The correct structure is identified using a variable mass spectral database search. Benchmarks show that Seq2PKS outperforms existing methods. Applying Seq2PKS to Actinobacteria datasets, we discover biosynthetic gene clusters for monazomycin, oasomycin A, and 2-aminobenzamide-actiphenol.

60 APPLIED LIFE SCIENCES↗