Search NASA⌕ Search

SEARCH · Search NASA

Results for “prediction algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Complexity-calibrated benchmarks for machine learning reveal when prediction algorithms succeed and mislead

Abstract Recurrent neural networks are used to forecast time series in finance, climate, language, and from many other domains. Reservoir computers are a particularly easily trainable form of recurrent neural network. Recently, a “next-generation” reservoir computer was introduced in which the memory trace involves only a finite number of previous symbols. We explore the inherent limitations of finite-past memory traces in this intriguing proposal. A lower bound from Fano’s inequality shows that, on highly non-Markovian processes generated by large probabilistic state machines, next-generation reservoir computers with reasonably long memory traces have an error probability that is at least $$\sim 60\%$$ ∼ 60 % higher than the minimal attainable error probability in predicting the next observation. More generally, it appears that popular recurrent neural networks fall far short of optimally predicting such complex processes. These results highlight the need for a new generation of optimized recurrent neural network architectures. Alongside this finding, we present concentration-of-measure results for randomly-generated but complex processes. One conclusion is that large probabilistic state machines—specifically, large $$\epsilon$$ ϵ -machines—are key to generating challenging and structurally-unbiased stimuli for ground-truthing recurrent neural network architectures.

97 MATHEMATICS AND COMPUTING↗

An Advanced Machine Learning and Artificial Intelligence System for Demonstrating Radiation Regulatory Compliance in DOE Accelerator Facilities

In this Phase II proposal, Applied Research LLC (ARLLC), Thomas Jefferson National Accelerator Facility (Jefferson Lab), and Old Dominion University (ODU) propose the combination of domain knowledge (beam characteristics, fixed structural shielding, earthen burden (the soil and foliage added to the dome of the experimental halls as additional shielding), etc.), machine learning (ML) and/or artificial intelligence (AI) to correlate a variety of multi-modal onsite signals and the radiation fields seen in accessible areas of the accelerator site and the site boundary. The ML/AI will consider the complex influence of environmental parameters affecting the radon contribution of the measurements, focusing on actual data obtained from Jefferson Lab. In Phase I, the coded beam and location data were fed into a deep learning model to predict doses at several designated locations in Jefferson Lab’s facility. Moreover, a dense radiation map was generated using only a sparse collection of the samples in a facility. In Phase II, we will develop a software prototype containing a radiation prediction algorithm, dense radiation map algorithms, and background noise prediction algorithms, with actual data used to evaluate the prototype. This work will provide a framework for evaluation of radiation measurement results around the site based on learned responses. In addition, the proposed approach allows more granular mapping of radiation levels. Better understanding and communication of these levels is related to the overall approach in keeping doses to personnel ALARA.

43 PARTICLE ACCELERATORS↗

Adaptive Online Model Update Algorithm for Predictive Control in Networked Systems

In this article, we introduce an adaptive on-line model update algorithm designed for predictive control applications in networked systems, particularly focusing on power distribution systems. Unlike traditional methods that depend on historical data for offline model identification, our approach utilizes real-time data for continuous model updates. This method integrates seamlessly with existing online control and optimization algorithms and provides timely updates in response to real-time changes. This methodology offers significant advantages, including a reduction in the communication network bandwidth requirements by minimizing the data exchanged at each iteration and enabling the model to adapt after disturbances. Furthermore, our algorithm is tailored for non-linear convex models, enhancing its applicability to practical scenarios. The efficacy of the proposed method is validated through a numerical study, demonstrating improved control performance using a synthetic IEEE test case.

data-driven model predictive control↗

Improved machine learning algorithm for predicting ground state properties

Finding the ground state of a quantum many-body system is a fundamental problem in quantum physics. In this work, we give a classical machine learning (ML) algorithm for predicting ground state properties with an inductive bias encoding geometric locality. The proposed ML model can efficiently predict ground state properties of an n-qubit gapped local Hamiltonian after learning from only $\mathcal{O}$(log(n)) data about other Hamiltonians in the same quantum phase of matter. This improves substantially upon previous results that require $\mathcal{O}$(n c ) data for a large constant c. Furthermore, the training and prediction time of the proposed ML model scale as $\mathcal{O}$(n log n) in the number of qubits n. Numerical experiments on physical systems with up to 45 qubits confirm the favorable scaling in predicting ground state properties using a small training dataset.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

APACE: AlphaFold2 and advanced computing as a service for accelerated discovery in biophysics

The prediction of protein 3D structure from amino acid sequence is a computational grand challenge in biophysics and plays a key role in robust protein structure prediction algorithms, from drug discovery to genome interpretation. The advent of AI models, such as AlphaFold, is revolutionizing applications that depend on robust protein structure prediction algorithms. To maximize the impact, and ease the usability, of these AI tools we introduce APACE, AlphaFold2 and advanced computing as a service, a computational framework that effectively handles this AI model and its TB-size database to conduct accelerated protein structure prediction analyses in modern supercomputing environments. We deployed APACE in the Delta and Polaris supercomputers and quantified its performance for accurate protein structure predictions using four exemplar proteins: 6AWO, 6OAN, 7MEZ, and 6D6U. Using up to 300 ensembles, distributed across 200 NVIDIA A100 GPUs, we found that APACE is up to two orders of magnitude faster than off-the-self AlphaFold2 implementations, reducing time-to-solution from weeks to minutes. This computational approach may be readily linked with robotics laboratories to automate and accelerate scientific discovery.

97 MATHEMATICS AND COMPUTING↗

Optimization and Multimachine Learning Algorithms to Predict Nanometal Surface Area Transfer Parameters for Gold and Silver Nanoparticles

Interactions between gold metallic nanoparticles and molecular dyes have been well described by the nanometal surface energy transfer (NSET) mechanism. However, the expansion and testing of this model for nanoparticles of different metal composition is needed to develop a greater variety of nanosensors for medical and commercial applications. In this study, the NSET formula was slightly modified in the size-dependent dampening constant and skin depth terms to allow for modeling of different metals as well as testing the quenching effects created by variously sized gold, silver, copper, and platinum nanoparticles. Overall, the metal nanoparticles followed more closely the NSET prediction than for Förster resonance energy transfer, though scattering effects began to occur at 20 nm in the nanoparticle diameter. To further improve the NSET theoretical equation, an attempt was made to set a best-fit line of the NSET theoretical equation curve onto the Au and Ag data points. An exhaustive grid search optimizer was applied in the ranges for two variables, 0.1≤C≤2.0 and 0≤α≤4, representing the metal dampening constant and the orientation of donor to the metal surface, respectively. Three different grid searches, starting from coarse (entire range) to finer (narrower range), resulted in more than one million total calculations with values C=2.0 and α=0.0736. The results improved the calculation, but further analysis needed to be conducted in order to find any additional missing physics. With that motivation, two artificial intelligence/machine learning (AI/ML) algorithms, multilayer perception and least absolute shrinkage and selection operator regression, gave a correlation coefficient, R2, greater than 0.97, indicating that the small dataset was not overfitting and was method-independent. This analysis indicates that an investigation is warranted to focus on deeper physics informed machine learning for the NSET equations.

Demers, Steven M. E. (ORCID:0000000192213246)↗

Distributed Load Shedding Application Architecture and Bi-Level Predictive Estimator Algorithm

Increasing penetrations of distributed renewables are decreasing the effectiveness of traditional decentralized under-frequency load shedding (UFLS) schemes. As more distribution circuits begin to back-feed the transmission system, operation of traditional UFLS may exacerbate frequency instability. This paper presents the conceptual framework for a data-rich environment to coordinate UFLS across multiple distribution providers based on the laminar coordination framework in order to ensure optimal adaptive setting of UFLS relays. Communication and control are enabled through a distributed implementation of the IEC 61968-1 Common Information Model message bus structure. In addition to the proposed architecture, a novel adaptive UFLS scheme informed by a bi-level state estimator to create optimal relay setpoints is introduced. Initial simulation results are presented for the IEEE 14-bus test system on scenarios leading to mis-operation of traditional UFLS.

Anderson, Alexander A.↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Data Science Enabled Enabled Discovery of Superconductors (Final Progress Report)

This Final Technical Report describes efforts by 4 PIs at the University of Florida (Peter Hirschfeld, Richard Hennig, Greg Stewart and James Hamlin), over the period September 2019-August 2023, to use data science and machine learning techniques to discover new conventional superconductors. The PIs constructed a discovery loop with two theorists and two experimentalists to: develop algorithms to machine learn descriptors correlating strongly with the critical temperature Tc (PI's Peter Hirschfeld, UF Physics and Richard Hennig, UF Materials Science and En), synthesize and measure properties of promising materials, and feed back the knowledge gained into the prediction algorithm. This work was motivated by the theoretical prediction and experimental discovery of high-pressure, high-pressure hydride superconductors, and to find ways to recreate the high critical temperatures in these systems at ambient pressure. Highlights from the grant include: 1) a new equation for Tc in terms of moments of the electron-phonon spectral function, improving on the so-called Allen-Dynes equation (1975); 2) study of the metastable A15 superconductor Nb3Si, formed under explosive compression at ~1000GPa to determine the kinetic barrier to the ground state structure; 3) the development of ultra-fast machine-learned atomic potentials for molecular dynamics, and 4) the discovery of superconductivity at 19K in WB2 arising from metastable defect structures in the crystal.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

System and method for wave prediction

A method and system for prediction of wave properties include collecting time series data streams from one or more wave measurement devices and processing the data using a wave-prediction algorithm to identify the frequency components of the data and compute wave parameters. The wave-field is propagated in space and time to predict wave height, speed, and velocity at a target location. A sliding window approach is used to continuously update the prediction in real-time.

Previsic, Mirko↗

Inference of response functions with the help of machine-learning algorithms

Response functions are a key quantity to describe the near-equilibrium dynamics of strongly interacting many-body systems. Recent techniques that attempt to overcome the challenges of calculating these ab initio have employed expansions in terms of orthogonal polynomials. We employ a neural network prediction algorithm to reconstruct a response function 𝑆⁡(𝜔) defined over a range in frequencies 𝜔. Here, we represent the calculated response function as a truncated Chebyshev series whose coefficients can be optimized to reduce the representation error. We compare the quality of response functions obtained using coefficients calculated using a neural network (NN) algorithm with those computed using the Gaussian integral transform (GIT) method. In the regime where only a small number of terms in the Chebyshev series are retained, we find that the NN scheme outperforms the GIT method.

Kurkcuoglu, Doga Murat [Fermi National Accelerator↗

Exploring the fragmentation efficiency of proteins analyzed by MALDI-TOF-TOF tandem mass spectrometry using computational and statistical analyses

Matrix-assisted laser desorption/ionization time-of-flight-time-of-flight (MALDI-TOF-TOF) tandem mass spectrometry (MS/MS) is a rapid technique for identifying intact proteins from unfractionated mixtures by top-down proteomic analysis. MS/MS allows isolation of specific intact protein ions prior to fragmentation, allowing fragment ion attribution to a specific precursor ion. However, the fragmentation efficiency of mature, intact protein ions by MS/MS post-source decay (PSD) varies widely, and the biochemical and structural factors of the protein that contribute to it are poorly understood. With the advent of protein structure prediction algorithms such as Alphafold2, we have wider access to protein structures for which no crystal structure exists. In this work, we use a statistical approach to explore the properties of bacterial proteins that can affect their gas phase dissociation via PSD. We extract various protein properties from Alphafold2 predictions and analyze their effect on fragmentation efficiency. Our results show that the fragmentation efficiency from cleavage of the polypeptide backbone on the C-terminal side of glutamic acid (E) and asparagine (N) residues were nearly equal. In addition, we found that the rearrangement and cleavage on the C-terminal side of aspartic acid (D) residues that result from the aspartic acid effect (AAE) were higher than for E- and N-residues. From residue interaction network analysis, we identified several local centrality measures and discussed their implications regarding the AAE. We also confirmed the selective cleavage of the backbone at D-proline bonds in proteins and further extend it to N-proline bonds. Finally, we note an enhancement of the AAE mechanism when the residue on the C-terminal side of D-, E- and N-residues is glycine. To the best of our knowledge, this is the first report of this phenomenon. Our study demonstrates the value of using statistical analyses of protein sequences and their predicted structures to better understand the fragmentation of the intact protein ions in the gas phase.

59 BASIC BIOLOGICAL SCIENCES↗

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

Vibrio MARTX toxin processing and degradation of cellular Rab GTPases by the cytotoxic effector Makes Caterpillars Floppy

Vibrio vulnificus causes life-threatening wound and gastrointestinal infections, mediated primarily by the production of a Multifunctional-Autoprocessing Repeats-In-Toxin (MARTX) toxin. The most commonly present MARTX effector domain, the Makes Caterpillars Floppy-like (MCF) toxin, is a cysteine protease stimulated by host adenosine diphosphate (ADP) ribosylation factors (ARFs) to autoprocess. Here, we show processed MCF then binds and cleaves host Ra s-related proteins in b rain (Rab) guanosine triphosphatases within their C-terminal tails resulting in Rab degradation. We demonstrate MCF binds Rabs at the same interface occupied by ARFs. Moreover, we show MCF preferentially binds to ARF1 prior to autoprocessing and is active to cleave Rabs only subsequent to autoprocessing. We then use structure prediction algorithms to demonstrate that structural composition, rather than sequence, determines Rab target specificity. We further determine a crystal structure of aMCF as a swapped dimer, revealing an alternative conformation we suggest represents the open, activated state of MCF with reorganized active site residues. The cleavage of Rabs results in Rab1B dispersal within cells and loss of Rab1B density in the intestinal tissue of infected mice. Collectively, our work describes an extracellular bacterial mechanism whereby MCF is activated by ARFs and subsequently induces the degradation of another small host guanosine triphosphatase (GTPase), Rabs, to drive organelle damage, cell death, and promote pathogenesis of these rapidly fatal infections.

Science & Technology - Other Topics↗

Topological grain boundary segregation transitions

Engineering the structure of grain boundaries (GBs) by solute segregation is a promising strategy to tailor the properties of polycrystalline materials. Solute segregation triggering phase transitions at GBs has been suggested theoretically to offer different pathways to design interfaces, but an understanding of their intrinsic atomistic nature is missing. Here, we combined atomic resolution electron microscopy and atomistic simulations to discover that iron segregation to GBs in titanium stabilizes icosahedral units (“cages”) that form robust building blocks of distinct GB phases. Owing to their five-fold symmetry, the iron cages cluster and assemble into hierarchical GB phases characterized by a different number and arrangement of the constituent icosahedral units. Our advanced GB structure prediction algorithms and atomistic simulations validate the stability of these observed phases and the high excess of iron at the GB that is accommodated by the phase transitions.

36 MATERIALS SCIENCE↗

Component pose reconstruction using a single robotic total station for panelized building envelopes

The construction industry can benefit greatly from increased automation in construction processes in terms of precision, time, and labor costs. This is especially true for prefabricated construction for either new buildings or envelope retrofits. Here, prefabricated components are manufactured offsite according to design specifications and installed onsite. Accurate information on the component’s position and orientation (pose) is needed to achieve this. A tool named “Real-Time Evaluator” (RTE), designed to autonomously track prefabricated components as they are being installed and provide real-time installation instructions, is currently under development using a single robotic total station. This is challenging since at least three points are needed to determine the pose of an object in space. To achieve this goal with a single robotic total station, two key algorithms were developed: the “resection” algorithm aligns the digital twin with the physical twin regardless of the location of the total station, and the “transformation” algorithm gives instructions of translations and rotations to installers to achieve the desired installation pose. The algorithms were evaluated experimentally on a lab-scale demonstration. Results show that the resection algorithm achieved an average error of < 3.1 mm, while the transformation algorithm predicted the rotation angle along a single axis with an error of < 0.5◦.

Tang, Mengjia↗