Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Genetic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Inducible flippase-mediated metabolic engineering of Rhodosporidium toruloides for enhanced 3-hydroxypropionic acid production from corn stover hydrolysate

Rhodosporidium toruloides has gained increasing interests as a promising non-model host organism to produce a wide range of bioproducts from lignocellulosic biomass. Increasing the bioproduct titers, rates, and yields remains a challenge, largely due to a lack of robust and well-characterized genetic tools in this host. Here we developed an inducible flippase (FLP) and flippase recognition target (FRT) system that enables genetic manipulations without the need for additional selection markers. Synthetic inducible promoters were established, enabling regulation of FLP expression and efficient antibiotic marker removal. Leveraging this system, we engineered a strain to optimize 3-hydroxypropionic acid (3HP) production. Over four rounds of iterative genomic editing to resolve pathway bottlenecks, we achieved a 3HP titer of 69.4 g/L in fed-batch fermentation - the highest level reported in yeast from lignocellulosic hydrolysates. The engineered high 3HP producing strain offers a robust platform for sustainable bio-based chemical production from lignocellulosic feedstocks.

3-hydroxypropionic acid↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Optimization of a Secondary Air Injector for a Rich-Quench-Lean (RQL) Ammonia Combustor Using Computational Fluid Dynamics

Ammonia has emerged as a promising medium for moving hydrogen around the globe due to its energy density, existing infrastructure for production, transportation, and storage, and multiple applications from fertilizer to power generation. However, one of the most significant challenges with ammonia combustion for power generation is the formation of nitrogen oxides (NOx) during combustion. Several combustion strategies have been developed to minimize the formation of NOx. One such strategy, known as rich-quench-lean (RQL), is a method that combusts NH3 in a fuel-rich environment, followed by a quick mix section and a lean burnout section. Rapid mixing of secondary air before lean burnout is thought to be important to minimize the formation of NOx. A genetic algorithm (GA) is used to parametrically vary the secondary air injector design, including the diameter, count, and angles over a constrained design space. The designs are evaluated using a non-reacting OpenFOAM model, with the objective function being the uniformity index of the secondary air (modeled as a scalar). Optimization campaigns show that traditional correlation-based designs might not be adequate. After running 414 OpenFOAM models, the optimal design is 32, 1 mm diameter, angled air injectors, increasing the uniformity index at the outlet of the quick mix zone by 36.8%. The most promising designs will be manufactured and tested in a small RQL burner setup, which is expected to lead to validation and insight into the practical usage of NH3 combustion for power generation.

ammonia combustion↗

Harnessing the power of gradient-based simulations for multi-objective optimization in particle accelerators

Abstract Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The underlying problem enforces strict constraints on both individual states and actions as well as cumulative (global) constraints on energy requirements of the beam. Using historical accelerator data, we develop a physics-based surrogate model which is differentiable and allows for back-propagation of gradients. The results are evaluated in the form of a Pareto-front with two objectives. We show that the DDRL outperforms MFRL, BO, and GA on high dimensional problems.

43 PARTICLE ACCELERATORS↗

Adsorption Hysteresis Under Control: Tuning Host–Guest Interactions via a Genetic Algorithm

Mesoporous adsorbent materials offer a large volumetric capacity; however, cyclic adsorption/desorption processes in these systems often suffer from hysteresis and may require a significant pressure swing to access this capacity. To mitigate hysteresis, a proposed strategy is to include nucleation sites on the walls of the mesoporous material to facilitate droplet and bubble formation, lowering the free energy barriers to the respective phase transitions. It is unclear, however, what combination of adsorbate− adsorbent interactions and spatial patterning would be beneficial for a given application, considering that improvements to some sorption properties may come at the expense of other attributes. To understand these interconnected observables, we examine two model systems, planar-slit and cylindrical pores with tunable interaction sites, using GPU-accelerated transition matrix Monte Carlo simulations. The simulations provide a free energy map of the pressure−adsorption space in a matter of minutes, which we use to track adsorption isotherm characteristics as a function of adsorbent properties. We then leverage the rapid acquisition of simulation data to construct a genetic algorithm to iteratively modify interaction sites of the slit-pore wall to minimize the hysteresis of this system without sacrificing uptake. We find that the adsorption branch of the isotherm is easily modulated via the average host−guest interaction strength, but desorption is only adjustable if there is a suitable bubble nucleation site. Within the context of a slit-pore system, we identify relative interaction strengths and patch sizes required to gain control over both branches of the hysteresis loop.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multi-trait multi-environment genomic prediction strategies for Miscanthus sacchariflorus

Genomic selection holds the potential to serve as a strategic tool to enhance the genetic gain of complex traits in Miscanthus breeding programs. The development of improved cultivars requires their assessment for various traits across diverse environments to ensure suitable overall performance. Hence, the multi-trait multi-environment (MTME) genomic prediction (GP) models offer an opportunity to improve selection accuracy. This study aims to evaluate the potential of five GP models: (1) three MTME models including genotype-by-trait-by-environment interaction (G×E×T) and (2) two single-trait multi-environment (STME) models (with and without G×E interaction). A Miscanthus sacchariflorus population comprising 336 genotypes evaluated in three environments and scored for four traits (biomass yield YDY, total culm number TCM, average internode length AIL, and culm node number CNN) was analyzed. The predictive ability of the models was evaluated considering three cross-validation schemes resembling realistic scenarios (CV1: predicting new genotypes, CVP: predicting missing traits in a given environment, and CV2: predicting partially observed genotypes). On average, in all cross-validation schemes compared to the STME the predictive ability of the MTME models was 10% to 70% higher for TCM and AIL. On the other hand, for YDY and CNN, both STME models performed similarly or slightly better (between 5 to 64%) than the MTME models in most environments. While the MTME models were not successful for all traits when compared to their STME counterparts, MTME models improved the prediction of the performance of genotypes that were untested across environments or lacked trait information in a specific environment. Overall, our study suggests that MTME GP models can be implemented in Miscanthus breeding programs to improve the predictive ability of the complex traits, shorten breeding cycles, and accelerate selection decisions.

genomic prediction (GP)↗

Spatial analysis of cell patterning to aid genetic and phenotypic understanding of grass stomatal density: A case study in maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

59 BASIC BIOLOGICAL SCIENCES↗

Comments on “Failure analysis of corroded hydrogen-blended natural gas pipelines based on finite element analysis and genetic algorithm-back propagation neural network” [262 (2025) 111174]

This is a brief commentary paper to highlight and discuss the determination of hydrogen concentration in pipeline steel, effect of hydrogen embrittlement (HE) on the mechanical properties of the material, burst strength of corroded pipelines using finite element analysis (FEA) simulations, and curve-fit models for assessing remaining strength of X80 corroded pipelines for transporting hydrogen blended natural gas. Recently, Xie et al. [1] proposed a methodology to quantify the impact of HE on material properties and numerically determined burst pressure of X80 corroded pipelines. However, their HE quantification overestimated the degradation of tensile strength for hydrogen blending ratios beyond the original data range, and their FEA results of burst pressure are nonconservative. This work thus recharacterized the hydrogen concentration in the steel pipeline and the effect of HE on tensile strength, and then redetermined burst pressures for a set of typical corrosion defect cases considered by Xie et al. [1] based on an experimentally validated FEA modelling method. With the new FEA results, two empirical corrosion models were proposed for X80 corroded pipelines for hydrogen service. At zero hydrogen blending ratio, the novel empirical models predict burst pressures to be consistent with the industry-accepted corrosion models. Furthermore, both the numerical simulation method and the novel corrosion models are significant contributions to the pipeline industry and the hydrogen community. Application of these results will enhance the safety, reliability, and integrity of natural gas pipelines when used to transport hydrogen.

Burst pressure prediction↗

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)↗

Climate change drives convergent evolution of root traits on Sky Island climate relicts

Roots are essential to the strategies plants use to survive in variable environments, yet we know little of how they vary within species. Experimental conditions demonstrate that intraspecific plant root traits respond strongly to variation in the environment; however, it is unclear when these responses can be characterized as evolution in response to selective pressures of climate change over many generations. Sky Islands are model, natural climate relict ecosystems to examine climate-change driven evolution. Utilizing a common garden with replicate genotypes of Populus angustifolia (Narrowleaf cottonwood) from six Sky Island (SI) populations and nine adjacent Mountain Chain (MC) populations across three genetic provenances, we hypothesized that SI root traits have diverged due to historical isolation in warmer, drier climates. When grown in common conditions, populations originating on SI’s showed convergent evolution across three distinct genetic provenances, which was characterized by 44.16% decreased total root length, 42.64% decreased average root volume, 43.31% decreased root surface area, and significantly less root trait variation, relative to adjacent mountain chains. Convergent evolution of root traits from trees originating on SI’s is correlated with changes in mean annual precipitation and potential evapotranspiration in the field over the past ~ 125 years. These results demonstrate a consistent pattern in root trait evolution at the landscape scale and the role of climate on the evolution of root traits in a genetic and geographic context relevant to climate change.

Convergent evolution↗

An interpretable model of pre-mRNA splicing for animal and plant genes

Pre-mRNA splicing is a fundamental step in gene expression, conserved across eukaryotes, in which the spliceosome recognizes motifs at the 3' and 5' splice sites (SSs), excises introns, and ligates exons. SS recognition and pairing is often influenced by protein splicing factors (SFs) that bind to splicing regulatory elements (SREs). Here, we describe SMsplice, a fully interpretable model of pre-mRNA splicing that combines models of core SS motifs, SREs, and exonic and intronic length preferences. We learn models that predict SS locations with 83 to 86% accuracy in fish, insects, and plants and about 70% in mammals. Learned SRE motifs include both known SF binding motifs and unfamiliar motifs, and both motif classes are supported by genetic analyses. Our comparisons across species highlight similarities between non-mammals, increased reliance on intronic SREs in plant splicing, and a greater reliance on SREs in mammalian splicing.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

Quantitative dissection of Agrobacterium T-DNA expression in single plant cells reveals density-dependent synergy and antagonism

Agrobacterium pathogenesis, which involves transferring T-DNA into plant cells, is the cornerstone of plant genetic engineering. As the applications that rely on Agrobacterium increase in sophistication, it becomes critical to achieve a quantitative and predictive understanding of T-DNA expression at the level of single plant cells. Here we examine if a classic Poisson model of interactions between pathogens and host cells holds true for Agrobacterium infecting Nicotiana benthamiana. Systematically challenging this model revealed antagonistic and synergistic density-dependent interactions between bacteria that do not require quorum sensing. Using various approaches, we studied the molecular basis of these interactions. To overcome the engineering constraints imposed by antagonism, we created a dual binary vector system termed ‘BiBi’, which can improve the efficiency of a reconstituted complex metabolic pathway in a predictive fashion. Our findings illustrate how combining theoretical models with quantitative experiments can reveal new principles of bacterial pathogenesis, impacting both fundamental and applied plant biology.

Alamos, Simon↗

Did the exposure of coacervate droplets to rain make them the first stable protocells?

Membraneless coacervate microdroplets have long been proposed as model protocells as they can grow, divide, and concentrate RNA by natural partitioning. However, the rapid exchange of RNA between these compartments, along with their rapid fusion, both within minutes, means that individual droplets would be unable to maintain their separate genetic identities. Hence, Darwinian evolution would not be possible, and the population would be vulnerable to collapse due to the rapid spread of parasitic RNAs. In this study, we show that distilled water, mimicking rain/freshwater, leads to the formation of electrostatic crosslinks on the interface of coacervate droplets that not only suppress droplet fusion indefinitely but also allow the spatiotemporal compartmentalization of RNA on a timescale of days depending on the length and structure of RNA. We suggest that these nonfusing membraneless droplets could potentially act as protocells with the capacity to evolve compartmentalized ribozymes in prebiotic environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing pressurized-water reactor equilibrium cycle using a novel loading pattern encoding and rule-based genetic crossover operators

This work presents an extended multi-batch approach applied in shuffling scheme optimization for equilibrium cycle for pressurized water reactors using Genetic Algorithms (GAs). A new ruled based GA crossover operator called Inherited Location and Batch (ILB) was introduced to enhance offsprings reproduction efficiency specialized for equilibrium cycle optimization problem. This approach was implemented within the Plant ReLoad Optimization (PRLO) framework and validated using a generic reactor model based on the AP1000 design, with core parameters calculated via the CASMO/SIMULATE software package. The ILB approach is then applied for both single and multi-objective problems in maximizing cycle length and core average exposure while minimizing the average enrichment of the 57 fresh fuel assemblies (FAs) per cycle. The optimal solutions are selected based on their dominance to the objectives from all feasible solutions. This research identified three optimal solutions satisfied safety constraints: The first solution minimizes feed enrichment costs with a cycle length of 338.8 days and core exposure of 25.39 MWd/MT; the second solution extends cycle length to 361.2 days, with the highest core exposure of 26.84 MWd/MT, using 3.75 wt% average fuel enrichment; the third solution balances both objectives with a cycle length of 349.6 days, core exposure of 25.82 MWd/MT with a slight enrichment increase compared to the first solution. Collectively, these findings underscore the efficiency and effectiveness of the proposed approach in achieving practical multi-objective optimal equilibrium cycle designs using GAs optimizer.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Network of epistatic interactions in an enzyme active site revealed by large-scale deep mutational scanning

Cooperative interactions between amino acids are critical for protein function. A genetic reflection of cooperativity is epistasis, which is when a change in the amino acid at one position changes the sequence requirements at another position. To assess epistasis within an enzyme active site, we utilized CTX-M β-lactamase as a model system. CTX-M hydrolyzes β-lactam antibiotics to provide antibiotic resistance, allowing a simple functional selection for rapid sorting of modified enzymes. We created all pairwise mutations across 17 active site positions in the β-lactamase enzyme and quantitated the function of variants against two β-lactam antibiotics using next-generation sequencing. Context-dependent sequence requirements were determined by comparing the antibiotic resistance function of double mutations across the CTX-M active site to their predicted function based on the constituent single mutations, revealing both positive epistasis (synergistic interactions) and negative epistasis (antagonistic interactions) between amino acid substitutions. The resulting trends demonstrate that positive epistasis is present throughout the active site, that epistasis between residues is mediated through substrate interactions, and that residues more tolerant to substitutions serve as generic compensators which are responsible for many cases of positive epistasis. Additionally, we show that a key catalytic residue (Glu166) is amenable to compensatory mutations, and we characterize one such double mutant (E166Y/N170G) that acts by an altered catalytic mechanism. These findings shed light on the unique biochemical factors that drive epistasis within an enzyme active site and will inform enzyme engineering efforts by bridging the gap between amino acid sequence and catalytic function.

59 BASIC BIOLOGICAL SCIENCES↗

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence↗