Search NASA⌕ Search

SEARCH · Search NASA

Results for “encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

The phage nucleus synergizes with an anti-defense protein to resist bacterial immunity

Chimallivirus bacteriophages enclose their replicating genomes in a protein-based compartment termed the phage nucleus. While the phage nucleus segregates phage DNA from host immune proteins, it is not known if additional factors are required to protect against DNA-targeting host defenses. Here, we identify a chimallivirus-encoded DarG2-like antitoxin that localizes to the phage nucleus and provides protection against phage-targeting DarTG2 toxin-antitoxin systems. This protein, which we term AdfM (anti-darT factor macro), contains a macrodomain and removes DarT2-mediated ADP-ribose modifications from DNA. In the absence of AdfM, DarT2 modifies phage DNA and restricts chimallivirus replication despite being largely excluded from the phage nucleus. Increasing the nuclear concentration of DarT2 while decreasing the nuclear concentration of AdfM reduces phage replication. These results show that the phage nucleus is insufficient to completely protect the chimallivirus genome from host defenses; rather, it is one component of a multilayered counter-defense strategy.

CRISPR↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

Deep operator network surrogate for phase-field modeling of metal grain growth during solidification

A deep operator network (DeepONet) has been constructed that generates accurate representations of phase-field model simulations for evolving two dimensional metal grain morphology growing from melt. These representations serve as lower resolution, computationally efficient stand-ins for quick parameter space exploration of solutions to the the Allen-Cahn equations that dictate the phase-field model simulations. The experimental target for the phase-field model is a uranium casting system cooling a 434 g uranium charge from a maximum temperature of 1400° C at an average rate of 30° C / min , traversing the crystallographic phases of the pure metal. Experimental parameters inform the phase-field model, whose higher resolution computational model solutions are used to train the DeepONet in a given parameter space with the aim of developing a faster, more efficient method for predicting the solidifying metal's microstructure at different potential experimental values. The final DeepONet generates high accuracy, lower resolution predictions with cumulative relative approximation error over all timesteps of less than 0.5%, while ensuring solutions remain within physically feasible ranges. Further, these relative error values are comparable with other state-of-the-art DeepONet models for microstructure evolution, while significantly reducing the amount of training data required. Training a convolutional neural network simultaneously with the DeepONet, enforcing realistic values at the complex metal grain boundaries, and mathematically encoding boundary conditions into the structure of the DeepONet improved prediction accuracy and computational efficiency over a standard DeepONet model.

36 MATERIALS SCIENCE↗

Genetically pliable green algae for bioproduction of modified fatty acids, nutritional therapeutic oils, and biopharmaceuticals

Homologous recombination (HR) is an essential tool for complex metabolic engineering in yeast, but transgene integration into plant and green algal nuclear genomes predominantly occurs by non-homologous end-joining. Species of the closely related, oleaginous trebouxiophytes Auxenochlorella and Prototheca, are unusual among the green algae in that HR is the favored mechanism for DNA integration into the nuclear genome. This property enables locus-specific targeting of gene cassettes encoding multiple enzymes for manipulating existing biochemical pathways or introducing new functions. Genetic malleability, and regulatory approval for human consumption, coupled with robust fermentation performance at industrial scale, establishes Auxenochlorella and Prototheca as prime candidates for algal production of biochemicals and biomaterials. The examples presented here highlight strain improvement and engineering for synthesis of hydroxylated fatty acids for biomaterials, structured triglycerides resembling human milk fat for infant nutrition, very-long-chain mono- and polyunsaturated fatty acids with nutraceutical or therapeutic potential, and cannabinoids for pharmacological applications.

Moseley, Jeffrey L. [University of California, Ber↗

Geometry-aware framework for deep energy method: An application to structural mechanics with hyperelastic materials

Here, in this work, we introduce a novel physics-informed framework named the Geometry-Aware Deep Energy Method (GADEM) for solving structural mechanics problems on different geometries. As the weak form of the physical system equation (or the energy-based approach) has demonstrated clear advantages compared to the strong form for solving solid mechanics problems, GADEM employs the weak form and aims to infer the solution on multiple shapes of geometries. Integrating a geometry-aware framework into an energy-based method results in an effective physics-informed deep learning model in terms of accuracy and computational cost. Different ways to represent the geometric information and to encode the geometric latent vectors are investigated in this work. We introduce a loss function of GADEM which is minimized based on the potential energy of all considered geometries. An adaptive learning method is also employed for the sampling of collocation points to enhance the performance of GADEM. We present some applications of GADEM to solve solid mechanics problems, including a loading simulation of a toy tire involving contact mechanics and large deformation hyperelasticity. The numerical results of this work demonstrate the remarkable capability of GADEM to infer the solution on various and new shapes of geometries using only one trained model.

97 MATHEMATICS AND COMPUTING↗

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

The connection between the chromatic numbers of a hypergraph and its 1-intersection graph

A well known problem from an excellent book of Lovász states that any hypergraph with the property that no pair of hyperedges intersect in exactly one vertex can be properly 2-colored. Motivated by this as well as recent works of Keszegh and of Gyárfás et al. we study the 1-intersection graph of a hypergraph. The 1-intersection graph encodes those pairs of hyperedges in a hypergraph that intersect in exactly one vertex. We prove for k ϵ {2, 4} that all hypergraphs whose 1-intersection graph is k-partite can be properly k-colored.

1-intersection graph of hypergraphs↗

Quantum computing approach for building surface sunlit in urban-scale energy modeling

Solar shadow calculations are needed in building energy modeling and performance simulation of PV systems installed on roofs or facades of buildings. We present a quantum computing approach for calculation of building surface sunlit fractions by recasting solar visibility as a binary optimization problem solved by quantum annealing. Each triangulated surface centroid is encoded as a binary qubit indicating sunlit or shaded status. Geometric visibility constraints are derived from the Möller-Trumbore intersection algorithm and converted into a constrained quadratic binary model compatible with contemporary quantum annealers. The coefficients were embedded to D-Wave quantum computer. To demonstrate feasibility, we conducted a case study in San Francisco for a target building with 52 triangles and roughly 2700 nearby triangles within 50 m evaluated at representative winter and summer solar positions. The results demonstrated that quantum annealing can reliably calculate and distinguish sunlit from shaded surfaces. Quantum samples achieved average accuracy exceeding 92.4 %, with the aggregate surface-level agreement approaching 99.9 %. The outputs of quantum computers agreed closely with classical algorithms, indicating practical feasibility and promising scalability. Finally, the hourly sunlit fractions of building surfaces can be obtained for urban energy modelling. This is the first study to apply quantum computing to the solar shadow and building surface sunlit calculation. It introduces a new paradigm that differs fundamentally from traditional approaches.

Deng, Zhipeng↗

A Decomposition-Based Learn-To-Optimize Approach with Feasibility Layer Assistance for Sub-Hourly Unit Commitment

Sub-hourly unit commitment (UC) with 15-min intervals is gaining significant attention as a way to respond rapidly to the fluctuations in electricity supply and demand introduced by renewable resources. However, the increased temporal resolution and complex inter-temporal dependencies pose substantial computational challenges for traditional optimization methods. To this end, this paper explores a decomposition-based learn-to-optimize approach. Building on recent advances in machine learning, our method revisits the long- overlooked Lagrangian relaxation framework, which is a classical decomposition technique that enables tractable subproblem solving. These smaller subproblems are inherently well-suited for machine learning, as their reduced dimensionality and structural regularity allow predictive models to efficiently learn and generalize solution patterns. We thus propose a generic predictive model, which embeds Gated Recurrent Units (GRUs) and Attention in the encoder-decoder structure, and integrate a rule-based feasibility layer to capture temporal dependencies, reduce training effort, and improve feasibility w.r.t. unit-level constraints. Our method has been validated on the IEEE 118-bus system, demonstrating promising performance in solving sub-hourly UC problems efficiently and feasibly.

97 MATHEMATICS AND COMPUTING↗

Tropical intertidal microbiome response to the 2024 Marine Honour oil spill

Marine fuel oil (MFO) spills in tropical coastal environments are under-characterized despite increasing risk from maritime activities. Microbial and geochemical responses to the June 2024 Marine Honour MFO spill on Singapore's intertidal sediments were analyzed in real time over 185 days. Using metagenomics and hydrocarbon profiling, microbial community shifts and hydrocarbon degradation were quantified across visibly oiled (high-impact) and clean (low-impact) sites. Microbiomes at all sites adapted rapidly to the spill through increased diversity and abundance of genes encoding alkane and aromatic compound degradation, detoxification, and biosurfactant production. The dominant hydrocarbon-degrading bacteria differed markedly from those reported in other crude oil spills and in regions with different climates. Oil deposition intensity strongly influenced microbial succession and hydrocarbon-degrading gene profiles, and this reflected early toxicity constraints in heavily oiled areas. The persistence of hydrocarbon degradation genes beyond hydrocarbon detection in sediments suggested long-term functional priming may occur. The study provides novel genome-resolved insight into the microbial response to MFO pollution, advances understanding of marine environmental biodegradation, and provides urgently needed baseline data for oil spill response strategies in Southeast Asia and beyond.

Coastal pollution↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

A collision operator for describing dissipation in noncanonical phase space

The phase space of a noncanonical Hamiltonian system is partially inaccessible due to dynamical constraints (Casimir invariants) arising from the kernel of the Poisson tensor. When an ensemble of noncanonical Hamiltonian systems is allowed to interact, dissipative processes eventually break the phase space constraints, resulting in a thermodynamic equilibrium described by a Maxwell–Boltzmann distribution. However, the time scale required to reach Maxwell–Boltzmann statistics is often much longer than the time scale over which a given system achieves a state of thermal equilibrium. Examples include diffusion in rigid mechanical systems, as well as collisionless relaxation in magnetized plasmas and stellar systems, where the interval between binary Coulomb or gravitational collisions can be longer than the time scale over which stable structures are self-organized. Here, we focus on self-organizing phenomena over spacetime scales such that particle interactions respect the noncanonical Hamiltonian structure, but yet act to create a state of thermodynamic equilibrium. We derive a collision operator for general noncanonical Hamiltonian systems, applicable to fast, localized interactions. This collision operator depends on the interaction exchanged by colliding particles and on the Poisson tensor encoding the noncanonical phase space structure, is consistent with entropy growth and conservation of particle number and energy, preserves the interior Casimir invariants, reduces to the Landau collision operator in the limit of grazing binary Coulomb collisions in canonical phase space, and exhibits a metriplectic structure. We further show how thermodynamic equilibria depart from Maxwell–Boltzmann statistics due to the noncanonical phase space structure, and how self-organization and collisionless relaxation in magnetized plasmas and stellar systems can be described through the derived collision operator.

Boltzmann equation↗

Optimizing fluvial flood mitigation strategies: A multi-objective approach for cost-effective and socially-aware infrastructure feasibility analysis

Effective levee planning must balance capital cost, risk reduction, and community priorities. These objectives are rarely optimized together. This study presents a feasibility phase, simulationin-the-loop framework that couples terrain-based flood modeling with a socially aware multiobjective optimizer. Flood risk is measured as Expected Annual Exposed Population (EAEP), obtained by integrating exposure over Annual Exceedance Probability (AEP) nodes, mirroring the Hydrologic Engineering Center's Flood Damage Reduction Analysis (HEC-FDA) expected-annual formulation but with people rather than dollars. Exposure per scenario is computed by overlaying binary inundation masks with a population surface at the tract level. Distributional fairness is encoded through a Group Benefit Share (GBS) constraint that requires high-SVI tracts to receive at least a baseline share of annualized benefits. Capital cost is represented by a height-dependent unit-cost model suitable for screening. This study addresses the two-objective problem, minimize cost and expected annual exposure subject to the GBS constraint, using Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and leveraging Pareto front for feasibility phase decision making. Implemented with terrain-based flood modeling, GeoFlood, for rapid scenario evaluation, the framework is demonstrated in Southeast Texas. The results reveal clear trade-offs among cost, risk, and social benefits and identify non-dominated levee height configurations that satisfy the benefit-share floor. The contributions are a scalable decision support method that operationalizes expected annual population-based risk, embeds enforceable benefit-sharing guarantees, and uses lightweight simulation to explore large design spaces before higher fidelity design stages.

Flood mitigation↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

A cobalamin-dependent pathway of choline demethylation from the human gut acetogen Eubacterium limosum

Elevated serum levels of trimethylamine N-oxide (TMAO) are reported to promote the development of atherosclerosis. TMAO is produced by hepatic oxidation of trimethylamine (TMA) produced by the gut microbiome from dietary quaternary amines such as choline. Net TMA production in the gut depends on microbial enzymes that either produce or consume TMA and its precursors. Here we report the elucidation of a novel microbial pathway consuming choline without TMA production. The human gut acetogen Eubacterium limosum grows by demethylating choline to N-N-dimethylaminoethanol. Quantitative mass spectral analysis of the proteome revealed a multi-protein choline to tetrahydrofolate (THF) methyltransferase system present only in choline-grown cells. The components are encoded in a gene cluster on the genome and include MthB, an MttB superfamily member; MthC, homologous to methylotrophic cobalamin-binding proteins; MthA, homologous to cobalamin:THF methyltransferases; and MthK, a protein related to serine kinases. Together, MthB, MthC, and MthA methylate THF with phosphocholine, but not choline or other quaternary amines. MthB specifically methylates Co(I)-MthC with phosphocholine. MthK acts as a bifunctional choline kinase which can utilize ATP or the MthB demethylation product, N,N-dimethylaminoethanol phosphate, to phosphorylate choline. Together, MthK, MthB, MthC, and MthA are proposed to carry out the methylation of THF with choline. These results outline a THF methylation pathway in which choline is first activated with ATP to phosphocholine prior to demethylation to form N,N-dimethylaminoethanol phosphate. Furthermore, the latter can be recycled by MthK to form more phosphocholine without expending additional ATP, thus minimizing energy utilization during choline-dependent acetogenesis.

acetogenesis↗

Anti-symmetric barron functions and their approximation with sums of determinants

A fundamental problem in quantum physics is to encode functions that are completely anti-symmetric under permutations of identical particles. The architecture of neural network models for the electron wave function typically comprises an equivariant component followed by a summation of determinants. The recently introduced Generic Antisymmetric (GA) block is designed to enhance the expressivity of such neural wave functions, and it was found that the 2-layer GA block achieved more accurate energies than the corresponding single-determinant FermiNet architecure, suggesting its promise as a way to improve the expressivity of neural wave functions. In this paper we show how the function expressed by the 2-layer GA block can be decomposed into a sum of determinants. We formalize this result by defining the antisymmetric Barron space as a generalized version of the 2-layer GA block and providing an appromation theorem for this function class. This result can be viewed as a negative result showing that the 2-layer GA block is not more expressive than using multiple determinants.

Abrahamsen, Nilin↗

A score-based diffusion model approach for adaptive learning of stochastic partial differential equation solutions

In this paper, we propose a novel framework for adaptively learning the time-evolving solutions of stochastic partial differential equations (SPDEs) using score-based diffusion models within a recursive Bayesian inference setting. SPDEs play a central role in modeling complex physical systems under uncertainty, but their numerical solutions often suffer from model errors and reduced accuracy due to incomplete physical knowledge and environmental variability. To address these challenges, we encode the governing physics into the score function of a diffusion model using simulation data and incorporate observational information via a likelihood-based correction in a reverse-time stochastic differential equation. This enables adaptive learning through iterative refinement of the solution as new data becomes available. To improve computational efficiency in high-dimensional settings, we introduce the ensemble score filter, a training-free approximation of the score function designed for real-time inference. Numerical experiments on benchmark SPDEs demonstrate the accuracy and robustness of the proposed method under sparse and noisy observations.

97 MATHEMATICS AND COMPUTING↗