Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Genetic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

RLMolLM: Reinforcement Learning-Enhanced Language Model Framework for Inverse Molecular Design

Inverse molecular design faces significant challenges due to vast chemical space and complex property requirements. While language models show promise for molecular generation, they struggle with validity, multi-property optimization, and structural constraints. This work presents RLMolLM, a reinforcement learning framework combining Proximal Policy Optimization (PPO) with genetic algorithms to address these limitations. Our approach optimizes multiple user-specified properties including quantitative estimates of drug-likeness (QED), synthetic accessibility (SA), and ADMET (absorption, distribution, metabolism, excretion, and toxicity) endpoints without requiring complete model retraining, while maintaining capability for scaffold-constrained generation where specific substructures must be preserved. We outperform state-of-the-art methods for molecular optimization, achieving best QED scores across GDB13, Moses, and Zinc datasets with up to 31% improvement over previous methods while maintaining excellent validity, uniqueness, and novelty metrics. For simultaneous multi-property optimization, our framework achieves substantial improvements in ADMET properties including 4.5-fold reduction in hERG toxicity and enhanced Caco-2 permeability compared to Moses dataset. Under structural constraints, the framework significantly improves molecular validity while preserving scaffolds and effectively optimizing properties. In conclusion, this versatile solution advances pharmaceutical and materials molecular design through effective integration of reinforcement learning and genetic algorithms with multi-property optimization and scaffold preservation.

Genetic algorithms↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Optimizing fluvial flood mitigation strategies: A multi-objective approach for cost-effective and socially-aware infrastructure feasibility analysis

Effective levee planning must balance capital cost, risk reduction, and community priorities. These objectives are rarely optimized together. This study presents a feasibility phase, simulationin-the-loop framework that couples terrain-based flood modeling with a socially aware multiobjective optimizer. Flood risk is measured as Expected Annual Exposed Population (EAEP), obtained by integrating exposure over Annual Exceedance Probability (AEP) nodes, mirroring the Hydrologic Engineering Center's Flood Damage Reduction Analysis (HEC-FDA) expected-annual formulation but with people rather than dollars. Exposure per scenario is computed by overlaying binary inundation masks with a population surface at the tract level. Distributional fairness is encoded through a Group Benefit Share (GBS) constraint that requires high-SVI tracts to receive at least a baseline share of annualized benefits. Capital cost is represented by a height-dependent unit-cost model suitable for screening. This study addresses the two-objective problem, minimize cost and expected annual exposure subject to the GBS constraint, using Non-Dominated Sorting Genetic Algorithm II (NSGA-II) and leveraging Pareto front for feasibility phase decision making. Implemented with terrain-based flood modeling, GeoFlood, for rapid scenario evaluation, the framework is demonstrated in Southeast Texas. The results reveal clear trade-offs among cost, risk, and social benefits and identify non-dominated levee height configurations that satisfy the benefit-share floor. The contributions are a scalable decision support method that operationalizes expected annual population-based risk, embeds enforceable benefit-sharing guarantees, and uses lightweight simulation to explore large design spaces before higher fidelity design stages.

Flood mitigation↗

Stock-specific spatial overlap among seabird predators and Columbia River juvenile Chinook Salmon suggests a mechanism for predation during early marine residence

Abstract Objective Because predation is thought to be the primary source of natural mortality for juvenile salmon first entering the ocean, we sought to identify regions where, on average, stock-specific spatial overlap between the distribution of threatened and endangered juvenile Chinook Salmon Oncorhynchus tshawytscha and abundant fish-eating seabirds (common murres Uria aalge and sooty shearwaters Ardenna grisea) suggests the greatest potential for ocean predation risk to juvenile Chinook Salmon. Methods The relative abundance and spatial distribution of seabird predators and juvenile Chinook Salmon were quantified as part of long-term ecosystem surveys during May 2003–2012 and June 2003–2022. Genetic stock identification methods were used to assign individual fish to their respective stock groups. Stock-specific species distribution models then generated maps and indices of average annual spatial overlap between predators and prey within the survey area. Result There is unequivocal evidence for spatial overlap between common murres, sooty shearwaters, and five genetic groups of interior and lower Columbia River juvenile Chinook Salmon. We found strongly positive (≥0.70) spatial correlations between predator and prey densities in both May and June, although spatial overlap was, in general, greater during May. The region of highest spatial overlap occurred on the inner continental shelf between the Columbia River mouth (46.2°N) and Grays Harbor (47.0°N), a region at the beginning of the juvenile salmon migratory pathway that is strongly affected by freshwater outflow from the river. Conclusion Our findings support the idea that ocean avian predation during early marine residence has the potential to affect marine survival of juvenile Chinook Salmon and should be further investigated to better inform and implement ecological models and possible recovery actions for Chinook Salmon populations of the Columbia River basin.

Zamon, Jeannette E.↗

Novel Microbial Routes to Synthesize Industrially Significant Precursor Compounds

Ethylene is the most widely employed organic precursor compound in industry. The potential to impact ethylene formation via recently discovered microbial processes is tenable using plentiful CO2 feedstocks. The overall long-term objective of this project was to develop an industrially compatible microbial process to synthesize ethylene in high yields. The key objective of this project was to fully define and initially characterized a recently discovered and genetically regulated anaerobic pathway to produce high levels of ethylene called the Dihydroxyacetone Phosphate - Ethylene Pathway in phototrophic bacteria. This was addressed through the following specific aims: 1. Fully probe the catalytic potential of all enzymes of the DHAP ethylene pathway and determine the regulatory mechanism of DHAP-ethylene pathway gene expression. 2. Discover effective and active ethylene enzymes encoded in cultured and uncultured organisms from anoxic environments. 3.Model the thermodynamics and kinetics of ethylene synthetic pathways to guide engineering efforts in integrating best performing DHAP-ethylene pathway enzymes into model bacteria chassis for enhance ethylene yields. Through this project we discovered the initially missing genetic and enzyme component of the DHAP-ethylene pathway that directly synthesized ethylene and other important industrial compounds like methane and ethane from specific substrates. We uncovered and partially characterized a nitrogenase-like reductase that functions in DHAP-ethylene pathway specifically and in methionine synthesis in general. This nitrogenase-like system is called the Methylthio-Alkane Reductase (MAR) for its ability to cleave volatile organic sulfur compounds into methanethiol (CH3-SH) for methionine synthesis and a hydrocarbon byproduct. Key to the DHAP-ethylene pathway, MAR is the essential enzyme that cleaves 2-methylthioethanol (CH3-S-CH2-CH2-OH) into ethylene. Coordinately, we uncovered that the MAR genes and genes associated with conversion of methanethiol (CH3-SH) to methionine are under genetic control of a LysR Type Transcriptional Regulator called SalR, whose activity is dependent upon the amount of sulfate available to the cell. When sulfate as the preferred sulfur source for cell growth drops below 200 micromolar, SalR become active for expressing the MAR and methionine biosynthesis genes to enable the cell to grow from volatile organic sulfur compounds and make ethylene. Metabolic thermos-kinetic modeling revealed that these MAR reactions for ethylene and other hydrocarbon production are highly thermodynamically favorable and are one of the largest driving forces for ethylene production by the DHAP-ethylene pathway for high ethylene yields. Modeling also indicated that a key aldolase and to a lesser extent an isomerase of the DHAP-ethylene pathway for production of the ethylene precursor, 2-methylthioethanol, also would increase ethylene yields. Through metagenomic mining and gene synthesis by the JGI DNA synthesis program, over 500 aldolase and isomerase homologs were synthesized and screened. From this, variants were uncovered with substantially higher activity that increased ethylene yields 5-fold via the aldolase reaction and 1.5-fold via the isomerase reaction. Each of these elements that increase ethylene production were integrated together via plasmid under appropriate gene promoter elements in the phototrophic bacterium, Rhodospirillum rubrum, resulting in at least 3 orders of magnitude increase in ethylene yield from carbon dioxide feedstock.

10 SYNTHETIC FUELS↗

Optimized Gear Selection to Maximize Energy Savings in Electric Traction Drives for Medium and Heavy Duty Vehicles

Multi‑gear transmission systems are commonly used in electric traction drives for medium and heavy‑duty vehicles, while most passenger‑vehicle electric drivetrains rely on a single fixed ratio to reduce cost, weight, and complexity. Using multiple gear ratios can enable downsizing of the motor and inverter while still meeting performance requirements. Additionally, appropriately chosen ratios allow the motor to operate more frequently in high‑efficiency regions, improving overall energy usage and reducing operating costs over the drive cycle. This paper presents a systematic approach for selecting optimal gear ratios for electric drive systems. A neural‑network model is first developed to represent motor losses across the full torque–speed range using data generated from finite element analysis. This model enables fast, accurate evaluation of motor efficiency under varying operating conditions. A genetic‑algorithm‑based optimization framework is then applied to identify gear ratios that maximize energy cost savings over the drive cycle, with the resulting optimal ratios stored for real‑time implementation.

Gadiyar, Nishanth [ORNL] (ORCID:0000000348267524)↗

Spatial Replication Is Important for Developing Landscape Genetic Inferences for a Wetland Salamander

Habitat fragmentation is a pressing threat to wildlife populations, and maintenance of gene flow between populations is an essential goal of conservation. Resistance surfaces have emerged as an important tool for modelling connectivity and developing management strategies to mitigate effects of habitat fragmentation. However, recent studies have noted inconsistencies in the factors most strongly associated with connectivity across different landscapes. Thus, replication of genetic-based resistance surface optimisation across landscapes may be necessary for making robust conclusions about the influence of environmental variables. Accordingly, replication represents a substantive challenge and opportunity in the field of landscape genetics. In this study, we conducted replicated landscape genetic analyses across five landscapes in Tennessee and Kentucky for a threatened wetland amphibian, the four-toed salamander (Hemidactylium scutatum). We tested multiple hypotheses of how different landscape features that could directly affect small, desiccation-intolerant amphibians (e.g., canopy cover) influenced gene flow and assessed the appropriate scale at which to model different features. We found some concordance in the landscape features that influenced gene flow (e.g., a common importance of forest cover and topography), but also some differences—potentially owing to the difference in variability of predictors across landscapes. We also found discordance in the scale of effect of different features across landscapes. In conclusion, our work emphasises that flat areas of moist forest not bisected by roads may be important for H. scutatum conservation, and our replicated design allows us to identify relationships that would have been missed if only using one study site.

59 BASIC BIOLOGICAL SCIENCES↗

Hydrogen Production System Scaling Using a High-Fidelity Simulation-Optimization Framework

Proton exchange membrane (PEM) electrolyzers are widely used for hydrogen production, yet few validated, high-fidelity tools can reliably guide scale-up. Using measured performance from a 50-hour hardware-in-the-loop pilot test, a physics-based, plant-level model of a 1.25 MW PEM electrolyzer and its balance-of-plant (BoP) subsystems is developed and validated. The model couples electrochemistry and thermal/flow submodels and is calibrated against pilot test data via a genetic algorithm (GA) workflow. Validation yields a mean absolute percentage error (APE) of 0.43% for cell voltage and stack power. Two scale-out strategies are then benchmarked under a common 7-day wind-and-photovoltaic (PV) profile: (i) linear duplication of 1.25 MW blocks and (ii) shared-BoP architectures. Sharing BoP between stacks reduces BoP energy by 27% at 10 MW and 34% at 100 MW (vs. linear duplication) and improves system specific energy consumption (SEC) to 52.9 and 52.6 kWh/kg, respectively (from 54.0 kWh/kg with linear duplication). Partial-load studies (25-100% set-point) show that cumulative hydrogen production remains nearly constant down to 50% load because all cases use the same weekly renewable-energy input. Below 50%, the power cap limits how much energy can be used within 168 h, which reduces hydrogen output. The model further indicates that the practical operating optimum lies between 50% and 85% load, where efficiency gains begin to appear without significant loss in hydrogen output. Moreover, the efficiency gains at lower loads are offset by reduced production. The validated framework supports scenario-based engineering trade-off studies for large configurations (10-100 MW) and for operating policies under variable renewables.

08 HYDROGEN↗

Intra- and inter-subtype HIV diversity between 1994 and 2018 in southern Uganda: a longitudinal population-based study

There is limited data on human immunodeficiency virus (HIV) evolutionary trends in African populations. We evaluated changes in HIV viral diversity and genetic divergence in southern Uganda over a 24-year period spanning the introduction and scale-up of HIV prevention and treatment programs using HIV sequence and survey data from the Rakai Community Cohort Study, an open longitudinal population-based HIV surveillance cohort. Gag (p24) and env (gp41) HIV data were generated from people living with HIV (PLHIV) in 31 inland semi-urban trading and agrarian communities (1994–2018) and four hyperendemic Lake Victoria fishing communities (2011–2018) under continuous surveillance. HIV subtype was assigned using the Recombination Identification Program with phylogenetic confirmation. Inter-subtype diversity was evaluated using the Shannon diversity index, and intra-subtype diversity with the nucleotide diversity and pairwise TN93 genetic distance. Genetic divergence was measured using root-to-tip distance and pairwise TN93 genetic distance analyses. Demographic history of HIV was inferred using a coalescent-based Bayesian Skygrid model. Evolutionary dynamics were assessed among demographic and behavioral population subgroups, including by migration status. 9931 HIV sequences were available from 4999 PLHIV, including 3060 and 1939 persons residing in inland and fishing communities, respectively. In inland communities, subtype A1 viruses proportionately increased from 14.3% in 1995 to 25.9% in 2017 (P < .001), while those of subtype D declined from 73.2% in 1995 to 28.2% in 2017 (P < .001). The proportion of viruses classified as recombinants significantly increased by nearly four-fold from 12.2% in 1995 to 44.8% in 2017. Inter-subtype HIV diversity has generally increased. While intra-subtype p24 genetic diversity and divergence leveled off after 2014, intra-subtype gp41 diversity, effective population size, and divergence increased through 2017. Intra- and inter-subtype viral diversity increased across all demographic and behavioral population subgroups, including among individuals with no recent migration history or extra-community sexual partners. This study provides insights into population-level HIV evolutionary dynamics following the scale-up of HIV prevention and treatment programs. Continued molecular surveillance may provide a better understanding of the dynamics driving population HIV evolution and yield important insights for epidemic control and vaccine development.

60 APPLIED LIFE SCIENCES↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

Stress intensity factor models using mechanics-guided decomposition and symbolic regression

The finite element method can be used to compute accurate stress intensity factors (SIFs) for cracks with complex geometries and boundary conditions. In contrast, handbook solutions act as surrogate SIF models that provide significantly faster evaluation times. However, the development of conventional surrogate SIF models relies on manual development based on low-order parameterizations. This limits surrogate model accuracy and generalizability. Here, in this paper, we develop a framework for the automated development of mechanics-guided handbook SIF solutions by using interpretable machine learning via genetic programming for symbolic regression (GPSR). Formalizing the mechanics-based approach of Raju and Newman, SIF training data is decomposed into multiple subsets. This decomposition enables parallel GPSR model development of subfunctions, each of which accounts for specific geometrical corrections with respect to a known analytical model. Using this mechanics-based approach with GPSR allows for equations to be learned with improved accuracy and reduced complexity relative to the Raju Newman equations while maintaining the inherent interpretability of mathematical expressions. In this paper, we present equations that match the complexity of the Raju Newman equations while having reduced error, as well as equations with similar errors and reduced complexity.

42 ENGINEERING↗

Constitutive model development of aluminum alloy 1100 for elevated temperature forming process

Commercially pure aluminum alloy, AA1100, presents good electrical and thermal conductivity, high formability, and low cost. Those favorable characteristics have the potential to enable bipolar plates with improved economics and enhanced performance compared to current stainless steel bipolar plates for proton exchange membrane fuel cells. An accurate constitutive model is essential to develop and optimize processing parameters and effectively control the forming process. Here, the objective of this work is to develop a constitutive model of AA1100 that is able to simulate stress-strain relation, formed geometry, and predict the onset of fracture strain to avoid forming failure. Initially, a set of tensile tests at temperature between 300 and 500°C and strain rate between 0.005 and 1.0/s were conducted to examine the deformation behavior. Then, a set of damage-based unified visco-plastic constitutive equations is proposed and calibrated based on the results of stress-strain data. A genetic algorithm optimization method is applied to search for best fitting material constants in constitutive equations. The proposed model shows good predictability of both the stress-strain relation and fracture strain at low strain rate and high temperature conditions. The accuracy of proposed model is also evaluated statistically. A comparison of the proposed model with three popular models (Arrhenius-type mode, Johnson-Cook model and Zerilli-Armstrong model) was made. The proposed model shows the best experimental agreement with correlation coefficient of 0.96 in contrast to 0.25, 0.38 and 0.75 for the popular models, respectively. The proposed model can help to optimize the elevated temperature forming process and guide die design to enable optimal geometric features in the formed components.

08 HYDROGEN↗

The unique architecture of umbrella toxins permits a two-tiered molecular bet-hedging strategy for interbacterial antagonism

Bacteria exist in competitive and rapidly changing environments in which the nature of future threats cannot be easily predicted. Streptomyces coelicolor produces three antibacterial umbrella particles that harbor distinct polymorphic toxin domains and an overlapping set of six diversified lectins. Here, we show that the exquisite specificity of umbrella particles derives from lectin-mediated species-specific binding to previously undescribed hypervariable surface glycoconjugates. A cryo-electron microscopy (cryo-EM) structure of one such lectin in complex with its oligosaccharide substrate defines the molecular basis for targeting through the coordinated recognition of multiple glycan features. Biochemical and genetic studies of several target species, in conjunction with lectin-swapping experiments, support a model whereby S. coelicolor umbrella toxin diversification at the levels of lectin composition and toxin polymorphism represents a unique, two-tiered bet-hedging strategy. Bioinformatic analyses support this as a means by which the unusual architecture of umbrella toxins offers Streptomyces a generalizable strategy to antagonize an unpredictable array of competitors.

59 BASIC BIOLOGICAL SCIENCES↗

Nonlinear thermodynamic computing out of equilibrium

We present the design for a thermodynamic computer that can perform arbitrary nonlinear calculations in or out of equilibrium. Simple thermodynamic circuits, fluctuating degrees of freedom in contact with a thermal bath and confined by a quartic potential, display an activity that is a nonlinear function of their input. Such circuits can therefore be regarded as thermodynamic neurons, and can serve as the building blocks of networked structures that act as thermodynamic neural networks, universal function approximators whose operation is powered by thermal fluctuations. We simulate a digital model of a thermodynamic neural network, and show that its parameters can be adjusted by genetic algorithm to perform nonlinear calculations at specified observation times, regardless of whether the system has attained thermal equilibrium. This work expands the field of thermodynamic computing beyond the regime of thermal equilibrium, enabling fully nonlinear computations, analogous to those performed by classical neural networks, at specified observation times.

Whitelam, Stephen [Lawrence Berkeley National Labo↗

Using leaf and stomatal traits to predict biomass production and water use efficiency in Populus

Climate change is reshaping ecosystems, driving plants to adapt through leaf-trait plasticity that reflects strategies for growth and water use. Predicting biomass production and intrinsic water use efficiency (iWUE) remains challenging because of genetic, taxonomic, and environmental variability. Here, we used eastern cottonwood and Populus hybrids as a model system to test whether easily measurable leaf traits can serve as reliable predictors of performance, and whether adding stomatal and biochemical traits improves predictive power. Across two field sites in Mississippi, leaf mass per area (LMA), biomass production, iWUE, leaf area, and foliar nitrogen ( N %) differed significantly among taxa and sites, while other traits were conserved. Factorial analysis of mixed data (FAMD) revealed distinct clustering of taxa and sites, indicating coordinated variation among leaf and stomatal traits. Pairwise correlations highlighted fundamental trade-offs, with biomass positively related to LMA and petiole length but negatively associated with iWUE, N %, and carbon isotopic ratios (δ 13 C). Leaf temperature and leaf angle varied among taxa and were significantly correlated with LMA and petiole length, suggesting mechanisms of heat dissipation and leaf movability that link simple traits to gas exchange and productivity. Weighted multiple linear regression models explained 80%–91% of variation in biomass production and iWUE. Models using only LMA, petiole length, and stomatal metrics performed nearly as well as those incorporating N %, and δ 13 C, with complex traits adding approximately 10% explanatory power. These results demonstrate that simple morphological traits capture integrated functional trade-offs, while complex traits refine predictions. This tiered approach provides an efficient framework for selecting high-yielding, water-efficient genotypes of Populus and other hardwood species, offering practical pathways to enhance carbon uptake and iWUE under climate change.

biomass production↗

Developing a media formulation to sustain ex vivo chloroplast function

Chloroplasts are critical organelles in plants and algae responsible for accumulating biomass through photosynthetic carbon fixation and cellular maintenance through metabolism in the cell. Chloroplasts are increasingly appreciated for their role in biomanufacturing, as they can produce many useful molecules, and a deeper understanding of chloroplast regulation and function would provide more insight for the biotechnological applications of these organelles. However, traditional genetic approaches to manipulate chloroplasts are slow, and generation of transgenic organisms to study their function can take weeks to months, significantly delaying the pace of research. To develop chloroplasts themselves as a quicker and more defined platform, we isolated chloroplasts from the green algae, Chlamydomonas reinhardtii, and examined their photosynthetic function after extraction. Combined with a metabolic modeling approach using flux-balance analysis, we identified key metabolic reactions essential to chloroplast function and leveraged this information into reagents that can be used in a “chloroplast media” capable of maintaining chloroplast photosynthetic function over time ex vivo compared to buffer alone. We envision this could serve as a model platform to enable more rapid design-build-test-learn cycles to study and improve chloroplast function in combination with genetic modifications and potentially as a starting point for the bottom-up design of a synthetic organelle-containing cell.

Chlamydomonas reinhardtii↗

Directed Evolution of an Adenylation Domain Alters Substrate Specificity and Generates a New Catechol Siderophore in Escherichia coli

Nonribosomal peptide synthetases (NRPS) biosynthesize numerous natural products with therapeutic, agricultural, and industrial significance. Reliably altering substrate selection in these enzymes has been a longstanding goal, as this would enable the production of tailor-made peptides with desired activities. In this study, the NRPS EntF and the associated biosynthesis of the siderophore enterobactin (ENT) were used as a model system to interrogate substrate selection by an adenylation (A) domain. We employed a directed evolution pipeline that harnesses an in vivo genetic selection for siderophore production to alter A domain substrate selection. Surprisingly, this led to the formation of a new, physiologically active catechol siderophore in Escherichia coli. We characterized the enzyme variants in vitro and demonstrated transferability of our findings to the well-studied TycC and GrsB NRPSs. Furthermore, this work identifies critical binding pocket residues that allow for altered substrate selection in our model system and expands upon our understanding of iron acquisition in E. coli.

59 BASIC BIOLOGICAL SCIENCES↗

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS↗