Search NASA⌕ Search

SEARCH · Search NASA

Results for “encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Massive twistor worldline in electromagnetic fields

We study the (ambi-)twistor model for spinning particles interacting via electromagnetic field, as a toy model for studying classical dynamics of gravitating bodies including effects of both spins to all orders. We compute the momentum kick and spin kick up to one-loop order and show precisely how they are encoded in the classical eikonal. The all-orders-in-spin effects are encoded as a dynamical implementation of the Newman-Janis shift, and we find that the expansion in both spins can be resummed to simple expressions in special kinematic configurations, at least up to one-loop order. We confirm that the classical eikonal can be understood as the generator of canonical transformations that map the in-states of a scattering process to the out-states. We also remark that cut contributions for converting worldline propagators from time-symmetric to retarded amount to the iterated action of the leading eikonal at one-loop order.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Gravitational scattering and beyond from extreme mass ratio effective field theory

We explore a recently proposed effective field theory describing electromagnetically or gravitationally interacting massive particles in an expansion about their mass ratio, also known as the self-force (SF) expansion. By integrating out the deviation of the heavy particle about its inertial trajectory, we obtain an effective action whose only degrees of freedom are the lighter particle together with the photon or graviton, all propagating in a Coulomb or Schwarzschild background. The 0SF dynamics are described by the usual background field method, which at 1SF is supplemented by a “recoil operator” that encodes the wobble of the heavy particle, and similarly computable corrections appearing at 2SF and higher. Our formalism exploits the fact that the analytic expressions for classical backgrounds and particle trajectories encode dynamical information to all orders in the couplings, and from them we extract multiloop integrands for perturbative scattering. As a check, we study the two-loop classical scattering of scalar particles in electromagnetism and gravity, verifying known results. We then present new calculations for the two-loop classical scattering of dyons, and of particles interacting with an additional scalar or vector field coupling directly to the lighter particle but only gravitationally to the heavier particle.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

Boosting efficiency and reducing graph reliance: Basis adaptation integration in Bayesian multi-fidelity networks

The computational cost of high-fidelity numerical models makes outer-loop analysis, which requires repeated interrogation of the model such as uncertainty quantification, computationally demanding. Multi-fidelity methods, which construct a surrogate model using data from an ensemble of models of varying cost and accuracy, can substantially reduce the cost of outer-loop analysis. However, these methods can be difficult to apply when the model ensemble does not admit a clear hierarchy a priori and the correlations between models are low. Consequently, in this paper, we present a multi-fidelity method that leverages dimension reduction to enhance the correlation between models, thereby reducing the amount of data needed to train a surrogate from an unordered ensemble of models. Our method utilizes basis adaptation to build low-dimensional polynomial chaos expansions of each model and employs Multi-fidelity Networks to encode the relationships among models. We show that the resulting method exhibit two notable advantages over its counterpart: (1) enhanced accuracy (both reduced bias and variance); and (2) reduced dependency on the graph structure encoding relationships among models. We demonstrate the approach on an analytical test problem and a challenging finite element model for a spent nuclear fuel. Our method produces a surrogate model that is significantly more accurate than either a single-fidelity surrogate or a multi-fidelity surrogate constructed without basis adaptation.

42 ENGINEERING↗

An orphan gene BOOSTER enhances photosynthetic efficiency and plant productivity

Organelle-to-nucleus DNA transfer is an ongoing process playing an important role in the evolution of eukaryotic life. Here, genome-wide association studies (GWAS) of non-photochemical quenching parameters in 743 Populus trichocarpa accessions identified a nuclear-encoded genomic region associated with variation in photosynthesis under fluctuating light. The identified gene, BOOSTER (BSTR), comprises three exons, two with apparent endophytic origin and the third containing a large fragment of plastid-encoded Rubisco large subunit. Higher expression of BSTR facilitated anterograde signaling between nucleus and plastid, which corresponded to enhanced expression of Rubisco, increased photosynthesis, and up to 35% greater plant height and 88% biomass in poplar accessions under field conditions. Overexpression of BSTR in Populus tremula × P. alba achieved up to a 200% in plant height. Similarly, Arabidopsis plants heterologously expressing BSTR gained up to 200% in biomass and up to 50% increase in seed.

60 APPLIED LIFE SCIENCES↗

Multiscale simulation of spatially correlated microstructure via a latent space representation

When deformation gradients act on the scale of the microstructure of a part due to geometry and loading, spatial correlations and finite-size effects in simulation cells cannot be neglected. We propose a multiscale method that accounts for these effects using a variational autoencoder to encode the structure–property map of the stochastic volume elements making up the statistical description of the part. In this paradigm the autoencoder can be used to directly encode the microstructure or, alternatively, its latent space can be sampled to provide likely realizations. Furthermore, we demonstrate the method on three examples using the common additively manufactured material AlSi10Mg in: (a) a comparison with direct numerical simulation of the part microstructure, (b) a push forward of microstructural uncertainty to performance quantities of interest, and (c) a simulation of functional gradation of a part with stochastic microstructure.

Elastoplasticity↗

A conserved chaperone protein is required for the formation of a noncanonical type VI secretion system spike tip complex

Type VI secretion systems (T6SSs) are dynamic protein nanomachines found in Gram-negative bacteria that deliver toxic effector proteins into target cells in a contact-dependent manner. Prior to secretion, many T6SS effector proteins require chaperones and/or accessory proteins for proper loading onto the structural components of the T6SS apparatus. However, despite their established importance, the precise molecular function of several T6SS accessory protein families remains unclear. In this study, we set out to characterize the DUF2169 family of T6SS accessory proteins. Using gene co-occurrence analyses, we find that DUF2169-encoding genes strictly co-occur with genes encoding T6SS spike complexes formed by valine-glycine repeat protein G (VgrG) and DUF4150 domains. Although structurally similar to Pro-Ala-Ala-Arg (PAAR) domains, “PAAR-like” DUF4150 domains lack PAAR motifs and instead contain a conserved PIPY motif, leading us to designate them PIPY domains. Next, we present both genetic and biochemical evidence that PIPY domains require a cognate DUF2169 protein to form a functional T6SS spike complex with VgrG. This contrasts with canonical PAAR proteins, which bind VgrG on their own to form functional spike complexes. By solving the first crystal structure of a DUF2169 protein, we show that this T6SS accessory protein adopts a novel protein fold. Furthermore, biophysical and structural modeling data suggest that DUF2169 contains a dynamic loop that physically interacts with a hydrophobic patch on the surface of its cognate PIPY domain. Based on these findings, we propose a model whereby DUF2169 proteins function as molecular chaperones that maintain VgrG–PIPY spike complexes in a secretion-competent state prior to their export by the T6SS apparatus.

DUF2169↗

DOME: Directional medical embedding vectors from Electronic Health Records

Motivation: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. Methods: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. Results: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHRembedding.

60 APPLIED LIFE SCIENCES↗

GrainGNN: A dynamic graph neural network for predicting 3D grain microstructure

We propose GrainGNN, a surrogate model for the evolution of polycrystalline grain structure under rapid solidification conditions in metal additive manufacturing. High fidelity simulations of solidification microstructures are typically performed using multicomponent partial differential equations (PDEs) with moving interfaces. The inherent randomness of the PDE initial conditions (grain seeds) necessitates ensemble simulations to predict microstructure statistics, e.g., grain size, aspect ratio, and crystallographic orientation. Here, currently such ensemble simulations are prohibitively expensive and surrogates are necessary.In GrainGNN, we use a dynamic graph to represent interface motion and topological changes due to grain coarsening. We use a reduced representation of the microstructure using hand-crafted features; we combine pattern finding and altering graph algorithms with two neural networks, a classifier (for topological changes) and a regressor (for interface motion). Both networks have an encoder-decoder architecture; the encoder has a multi-layer transformer long-short-term-memory architecture; the decoder is a single layer perceptron.We evaluate GrainGNN by comparing it to high-fidelity phase field simulations for in-distribution and out-of-distribution grain configurations for solidification under laser power bed fusion conditions. GrainGNN results in 80%–90% pointwise accuracy; and nearly identical distributions of scalar quantities of interest (QoI) between phase field and GrainGNN simulations compared using Kolmogorov-Smirnov test. GrainGNN's inference speedup (PyTorch on single x86 CPU) over a high-fidelity phase field simulation (CUDA on a single NVIDIA A100 GPU) is 150×–2000× for 100-initial grain problem. Further, using GrainGNN, we model the formation of 11,600 grains in 220 seconds on a single CPU core.

36 MATERIALS SCIENCE↗

Latent space dynamics identification for interface tracking with application to shock-induced pore collapse

Capturing sharp, evolving interfaces remains a central challenge in reduced-order modeling, especially when data is limited and the system exhibits localized nonlinearities or discontinuities. Here, we propose LaSDI-IT (Latent Space Dynamics Identification for Interface Tracking), a data-driven framework that combines low-dimensional latent dynamics learning with explicit interface-aware encoding to enable accurate and efficient modeling of physical systems involving moving material boundaries. At the core of LaSDI-IT is a revised autoencoder architecture that jointly reconstructs the physical field and an indicator function representing material regions or phases, allowing the model to track complex interface evolution without requiring detailed physical models or mesh adaptation. The latent dynamics are learned through linear regression in the encoded space and generalized across parameter regimes using Gaussian process interpolation with greedy sampling. We demonstrate LaSDI-IT on the problem of shock-induced pore collapse in high explosives, a process characterized by sharp temperature gradients and dynamically deforming pore geometries. The method achieves relative prediction errors below 9% across the parameter space, accurately recovers key quantities of interest such as pore area and hot spot formation, and matches the performance of dense training with only half the data. This latent dynamics prediction was 10 6 times faster than the conventional high-fidelity simulation, proving its utility for multi-query applications. These results highlight LaSDI-IT as a general, data-efficient framework for modeling discontinuity-rich systems in computational physics, with potential applications in multiphase flows, fracture mechanics, and phase change problems.

Gaussian process↗

AI-assisted object condensation clustering for calorimeter shower reconstruction at CLAS12

Several nuclear physics studies using the CLAS12 detector rely on the accurate reconstruction of neutrons and photons from its forward angle calorimeter system. These studies often place restrictive cuts when measuring neutral particles due to an overabundance of false clusters created by the existing calorimeter reconstruction software. In this work, we present a new AI approach to clustering CLAS12 calorimeter hits based on the object condensation framework. The model learns a latent representation of the full detector topology using GravNet layers, serving as the positional encoding for an event’s calorimeter hits which are processed by a Transformer encoder. This unique structure allows the model to contextualize local and long range information, improving its performance. Evaluated on one million simulated $e^-$ $+$ $p$ collision events, our method significantly improves cluster trustworthiness: the fraction of reliable neutron clusters, increasing from 8.88% to 30.73%, and photon clusters, increasing from 51.07% to 64.73%. In conclusion, our study also marks the first application of AI clustering techniques for hodoscopic detectors, showing potential for usage in many other experiments.

Calorimeters↗

Synthesis and characterization of isotopically barcoded nickel, molybdenum, and tungsten taggants for intentional nuclear forensics

Intentional nuclear forensics is a concept wherein the deliberate addition of benign and persistent material signatures to nuclear material can be used to reduce the time between the discovery of material outside of regulatory control and determination of its original provenance. One concept within intentional nuclear forensics involves the use of perturbed stable isotopes to generate unique isotope ratio “barcodes” to encode information (e.g., production batch, location, etc.) and track material throughout the nuclear fuel cycle. Synthesis of taggant species of nickel (Ni), molybdenum (Mo), and tungsten (W) was undertaken via a double-spike mechanism, wherein two highly enriched isotopes of interest per elemental taggant were mixed to form an enriched “double-spike” which was subsequently isotopically diluted with bulk material having a natural isotopic composition. Two taggant species perturbing isotopic ratios, alpha (α) and beta (β), for each of Ni, Mo, and W were synthesized. Independent measurements of double spikes and alpha and beta taggant species agreed within uncertainty and are clearly resolvable from natural compositions. High-precision analyses were independently performed by MC-ICP-MS at two U.S. National Laboratories, with consensus values and uncertainties calculated for all samples. Observed isotopic perturbations in the final taggant species measured on the order of hundreds to thousands of permille (‰) with respect to natural for isotope ratios of interest (e.g., 60 Ni/ 58 Ni, 100 Mo/ 98 Mo, 186 W/ 183 W). Discrepancies between modeled and measured isotopic compositions were observed and are largely attributed to imprecise vendor assay values for starting materials. Using measured starting material compositions as inputs for the mixing model improved the level of agreement between predicted and measured α and β taggant isotope ratios. Overall, characterization of all taggant species demonstrates that this “barcode” concept could have viability for use in nuclear forensics. Finally, it is expected that for any two-isotope mixing array dozens of isotopic barcodes could be encoded into a material system and subsequently resolved utilizing modern mass spectrometric methods.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Two deeply conserved non-coding sequences control PLETHORA1/2 expression and coordinate embryo and root development

Conserved non-coding sequences (CNSs) are integral elements of transcriptional regulation. Transcriptional tuning of PLETHORA (PLT) genes that encode master regulators of plant development is vital for embryogenesis and meristematic function. However, how the expression of PLT genes is modulated through CNSs remains unclear. Through motif-based mining of upstream sequences in 120 angiosperm genomes, we identified 21 conserved and lineage-specific CNSs, two of which are unusually long, similar, and colinear within eudicots. Using Arabidopsis thaliana, we demonstrate that these two deeply conserved elements, which we named BOX1 and BOX2, control PLT1 and PLT2 expression. CRISPR mutants within these elements specifically reduced PLT expression levels, and reporter lines revealed that deletion of either or both BOXes altered and/or abrogated the PLT2 expression pattern in the root tip, affecting the ability to rescue the plt1 plt2 double mutant. We further show that the influence of these elements on expression patterns is already exerted during embryogenesis and functional in the context of the early embryo. Finally, we reveal the existence of a BOX-mediated autoregulatory feedback loop that, in large part, explains CNS influence on expression patterns. We thus uncover a transcriptional mechanism by which genes encoding master regulators of embryo and root meristem development are regulated.

PLETHORA↗

Adaptive laboratory evolution and metabolic engineering of Cupriavidus necator for improved catabolism of volatile fatty acids

Bioconversion of high-volume waste streams into value-added products will be an integral component of the growing bioeconomy. Volatile fatty acids (VFAs) (e.g., butyrate, valerate, and hexanoate) are an emerging and promising waste-derived feedstock for microbial carbon upcycling. Cupriavidus necator H16 is a favorable host for conversion of VFAs into various bioproducts due to its diverse carbon metabolism, ease of metabolic engineering, and use at industrial scales. Here, in this study, we report that a common strategy to improve product titers in C. necator, deletion of the polyhydroxybutyrate (PHB) biosynthetic operon, results in a significant growth defect on VFA substrates. Using adaptive laboratory evolution, we identify mutations to the regulator gene phaR, the two-component response regulator-histidine kinase pair encoded by H16_A1372/H16_A1373, and the tripartite transporter assembly encoded by H16_A2296-A2298 as causative for improved growth on VFA substrates. Deletion of phaR and H16_A1373 led to significantly reduced NADH abundance accompanied by large changes to expression of genes involved in carbon metabolism, balance of electron carriers, and oxidative stress tolerance that may be responsible for improved growth of these engineered strains. These results provide insight into the role of PHB biosynthesis in carbon and energy metabolism and highlight a key role for the regulator PhaR in global regulatory networks. By combining mutations, we generated platform strains with significant growth improvements on VFAs, which can enable improved conversion of waste-derived VFA substrates to target bioproducts.

09 BIOMASS FUELS↗

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterizing Ultrafast Intersystem Crossing Pathways in Molecular Pt Dimers Using Time-Resolved Wide-Angle X-ray Scattering

Vibronic coupling between transition metal charge transfer states is a potential mechanism for enhancing the intersystem crossing (ISC) rate. Vibronic coupling-driven ISC has been observed in Pt­(II) dimer complexes, where the trajectory across excited-state pathways is tuned by atomic displacements via Pt–Pt stretching vibrations. Time-resolved wide-angle X-ray scattering (TR-WAXS) was utilized to quantify the Pt–Pt contraction following metal–metal-to-ligand charge transfer (MMLCT) excitation in Pt dimers with different bridging ligands. Both complexes exhibit Pt–Pt bond formation with a decrease in Pt–Pt distance of ∼ 0.25 Å and coherent vibrational wavepackets (CVWPs) encoded in the Pt–Pt contraction of both dimers. However, the complexes exhibit different time-dependent evolution of their CVWPs. Analysis of interference patterns between different CVWPs is used to track the trajectory across the excited-state surfaces. Furthermore, this work demonstrates that the interference between CVWPs in ultrafast TR-WAXS encodes indirect information regarding electronic excited-states to reveal the Pt dimer bridge-dependent ISC mechanism.

Computational chemistry↗

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN↗

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES↗