Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Genetic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling plasticity-mediated void growth at the single crystal scale: A physics-informed machine learning approach

Modeling the evolution of voids during plastic flow as well as their effects on plastic dissipation is critical for both component manufacturing and lifetime estimation purposes. To this end, we propose a rate-dependent constitutive model to homogenize the effects of semi-randomly distributed voids on single crystal plasticity whilst capturing void interaction and plastic anisotropy. Here, this present work focuses on the case of face centered cubic crystals to introduce an anisotropic gauge function applicable within the crystal plasticity formalism. The approach combines analytical methods to describe the micromechanics of the system in combination with symbolic regression to capture analytically intractable mechanisms from data. The hybrid framework uses a physics-informed genetic programming-based symbolic regression algorithm to solve a multiform optimization problem simultaneously producing a new gauge function and a new strain rate equation. This is also a multi-objective optimization problem with many competing objectives. A new search and selection step is introduced to the genetic algorithm that promotes convergence toward a global solution that better satisfies all the objectives. Overall, the symbolic equations produced leverage data-driven methods to achieve greater accuracy than comparable alternatives on an analytically intractable problem while maintaining model transparency.

36 MATERIALS SCIENCE↗

Characterization of Caenorhabditis elegans sphingomyelin synthases through heterologous expression

Sphingomyelin (SM) is a major component of mammalian cell membranes and particularly abundant in the myelin sheath that surrounds nerve fibers. Its production is catalyzed by SM synthases SMS1 and SMS2, which interconvert phosphatidylcholine and ceramide to diacylglycerol and SM in the Golgi and at the plasma membrane, respectively. As the lipids participating in this reaction fulfill both structural and signaling functions, SMS enzymes have considerable potential to influence diverse important cellular processes. The nematode Caenorhabditis elegans is an attractive model for studying both animal development and human disease. The organism contains five SMS homologues but none of these have been characterized in any detail. Here, we carried out the first systematic analysis of SMS family members in C. elegans . Using heterologous expression systems, genetic ablation, metabolic labeling and lipidome analyses, we show that C. elegans harbors at least three distinct SM synthases and one ceramide phosphoethanolamine (CPE) synthase. Moreover, C. elegans SMS family members have partially overlapping but also unique sub-cellular distributions and together occupy all principal compartments of the secretory pathway. Our findings shed light on crucial aspects of sphingolipid metabolism in a valuable animal model and opens avenues for exploring the role of SM and its metabolic intermediates in organismal development.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Influence of particle size on NIR spectroscopic characterization of sorghum biomass for the biofuel industry

NIR spectroscopy is a rapid and accurate green technology for high-throughput biomass characterization, including sorghum (Sorghum bicolor), a promising energy crop for the biofuel industry. This study assessed the influence of particle size on NIR spectroscopic analysis (wavelength range: 867–2535 nm) of sorghum biomass composition. Grown under field conditions, a total of 113 types of genetically diverse sorghum accessions were dried, ground, and sieved (<250, 250–600, 600–850, and > 850 µm particle size) for developing partial least square regression (PLSR) prediction models for moisture, ash, extractive, glucan, xylan, acid-soluble lignin (ASL), acid-insoluble lignin (AIL), and total lignin (ASL + AIL). Overall, smaller particle sizes provided better model performance, while no single particle size provided the best performance for all the selected components. With only 9 selected bands and 4 latent variables (LVs), the best PLSR model was obtained for moisture with particle size of 600–850 µm with the square root of the coefficient of determination (R) of 0.85, the ratio of prediction to deviation (RPD) of 2.2, and the root mean square error (RMSE) of 0.46 % in external validation. Similar model performances were also obtained for ash, extractive, glucan, and xylan. This study showed that size reduction could effectively improve NIR spectroscopic analysis for lipid-producing sorghum biomass for the biofuel industry.

09 BIOMASS FUELS↗

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)↗

Data for Influence of Particle Size on NIR Spectroscopic Characterization of Sorghum Biomass for the Biofuel Industry

NIR spectroscopy is a rapid and accurate green technology for high-throughput biomass characterization, including sorghum ( Sorghum bicolor ), a promising energy crop for the biofuel industry. This study assessed the influence of particle size on NIR spectroscopic analysis (wavelength range: 867–2535 nm) of sorghum biomass composition. Grown under field conditions, a total of 113 types of genetically diverse sorghum accessions were dried, ground, and sieved (<250, 250–600, 600–850, and > 850 µm particle size) for developing partial least square regression (PLSR) prediction models for moisture, ash, extractive, glucan, xylan, acid-soluble lignin (ASL), acid-insoluble lignin (AIL), and total lignin (ASL + AIL). Overall, smaller particle sizes provided better model performance, while no single particle size provided the best performance for all the selected components. With only 9 selected bands and 4 latent variables (LVs), the best PLSR model was obtained for moisture with particle size of 600–850 µm with the square root of the coefficient of determination (R) of 0.85, the ratio of prediction to deviation (RPD) of 2.2, and the root mean square error (RMSE) of 0.46 % in external validation. Similar model performances were also obtained for ash, extractive, glucan, and xylan. This study showed that size reduction could effectively improve NIR spectroscopic analysis for lipid-producing sorghum biomass for the biofuel industry.

Biomass Analytics↗

Recent developments of oleaginous yeasts toward sustainable biomanufacturing

Oleaginous yeast are remarkably versatile organisms, distinguished by their natural capacities to accumulate high levels of neutral lipids and broad substrate range. With recent growing interests in engineering non-model organisms as superior biomanufacturing platforms, oleaginous yeasts have emerged as promising chassis for oleochemicals, terpenoids, organic acids, and other valuable products. Advancement in systems biology along with genetic tool development have significantly expanded our understanding of the metabolism in these species and enabled engineering efforts to produce biofuels and bioproducts from diverse feedstocks. This review examines the latest technical advances in oleaginous yeast research toward sustainable biomanufacturing. We cover recent developments in systems biology-enabled metabolism understanding, genetic tools, feedstock utilization, and strain engineering approaches for the production of various valuable chemicals.

59 BASIC BIOLOGICAL SCIENCES↗

NCAP: Noncanonical Amino Acid Parameterization Software for CHARMM Potentials

Noncanonical Amino Acids (NCAAs) provide numerous avenues for introduction of novel functionality to peptides and proteins. NCAAs can be incorporated through solid phase synthesis or genetic code expansion in conjugation with heterologous expression of the encoded protein modification. Due to the difficulty of synthesis, wide chemical space and lack of empirically resolved structures modeling the effects of NCAA mutation is critical for rational protein design. To evaluate the structural and functional perturbations NCAAs introduce we utilize molecular potentials that describe the forces in protein structure. Most potentials such as CHARMM are designed to model canonical residues but can be parameterized in include novel NCAAs. Here, in this work, we introduce NCAP a software package to generate CHARMM compatible parameters from quantum chemical calculation. Unlike currently available tools NCAP is designed to recognize NCAA structure and automatically bridge the gap between DFT calculations and potential parameters. For our software we discuss workflow, validation against canonical parameter sets and comparison to published NCAA-protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

From the bench to the reactor: engineered filamentous fungi for biochemical and biomaterial production

Filamentous fungi can convert a wide variety of naturally occurring chemical compounds, including organic biomass and waste streams, into a range of products. They have long been used for industrial organic acid production and food preparation. In this review, we will discuss production of products such as organic acids, lipids, small molecules, enzymes, materials, and foods, and highlight advances in metabolic and protein engineering, including CRISPR-Cas9-mediated strain improvements. We discuss to what extent these products are already being made on a commercial scale, as well as what is still required to make certain promising concepts industrially and commercially relevant. Despite significant progress, the systematic application of synthetic biology to filamentous fungi remains in its infancy, with many opportunities for discovery and innovation as new strains and genetic tools are developed. The integration of fungal biotechnology into circular and bio-based economies promises to address critical challenges in waste management, resource sustainability, and the development of new materials for terrestrial and extraterrestrial applications, but requires further developments in genetic engineering and process design.

09 BIOMASS FUELS↗

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi↗

Data for Spatial Analysis of Cell Patterning to Aid Genetic and Phenotypic Understanding of Grass Stomatal Density: A Case Study in Maize

Biological processes involve complex hierarchies where composite traits result from multiple component traits. However, holistically understanding of how sets of component traits interact to underpin genotype-to-phenotype relationships is generally lacking. Stomatal density (SD) is a tractable model system for exploring how high-throughput phenotyping (HTP) data could be exploited by a new spatial analysis approach to better understand a developmentally and functionally important trait. SD is a composite trait, resulting from various components related to cell identity and size, which are themselves governed by a series of spatio-developmental processes. Data from 192 recombinant inbred lines of maize [Zea mays (L.)] were analyzed by a new stomatal patterning phenotype (SPP) to (1) describe the average spatial probability distribution of the nearest neighboring stomata; (2) derive a core set of component traits related to cell size, cell packing, and positional probabilities; (3) build a structural equation model of component traits underlying SD; and (4) identify stomatal patterning quantitative trait loci (QTL). The core set of SPP-derived traits explained 74% of the variation in SD. Analyzing SPP component traits allowed some loci previously identified as generic SD QTL to be recognized as specific to lateral versus longitudinal elements of stomatal patterning. Therefore, this study highlights how novel insights can be gained by decomposing a composite trait (e.g., SD) into a set of component traits that were present in HTP data but not previously exploited.

AI/ML↗

Translocation mechanism of xeroderma pigmentosum group D protein on single-stranded DNA and genetic disease etiology

Abstract XPD is a key nucleotide excision repair (NER) protein whose function is vital for genome integrity. During NER, XPD serves as a 5′−3′ single-strand DNA translocase that enables lesion scanning and verification in genomic DNA. Yet, its translocation mechanism is incompletely understood. Here we use molecular simulations and chain-of-replicas path optimization methods to model the ATP-driven translocation mechanisms of XPD and its bacterial homolog DinG, revealing all on-path metastable intermediates and corresponding kinetic rates. We identify the XPD(DinG) global domain motions that modulate the strength of DNA association at the opposing ends of the DNA-binding groove. During the ATP hydrolysis cycle, alternating weak and strong interactions at two defined groove constrictions enable DNA reptation and forward displacement of the ATPase. Moreover, we show that DNA- or ATP-binding residues directly involved in translocation are hotspots for genetic disease mutations. Thus, our findings shed light on the etiology of XPD-associated genetic syndromes.

Paul, Tanmoy↗

Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning

Abstract mRNA degradation is a central process that affects all gene expression levels, though it remains challenging to predict the stability of a mRNA from its sequence, due to the many coupled interactions that control degradation rate. Here, we carried out massively parallel kinetic decay measurements on over 50,000 bacterial mRNAs, using a learn-by-design approach to develop and validate a predictive sequence-to-function model of mRNA stability. mRNAs were designed to systematically vary translation rates, secondary structures, sequence compositions, G-quadruplexes, i-motifs, and RppH activity, resulting in mRNA half-lives from about 20 seconds to 20 minutes. We combined biophysical models and machine learning to develop steady-state and kinetic decay models of mRNA stability with high accuracy and generalizability, utilizing transcription rate models to identify mRNA isoforms and translation rate models to calculate ribosome protection. Overall, the developed model quantifies the key interactions that collectively control mRNA stability in bacterial operons and predicts how changing mRNA sequence alters mRNA stability, which is important when studying and engineering bacterial genetic systems.

Cetnar, Daniel P.↗

A Hierarchical Optimization Method for Electric Vertical Takeoff and Landing Aircraft Network Design

Electric vertical takeoff and landing aircraft (eVTOLs) are expected to serve urban air mobility in a station-to-station configuration, which makes the optimal network design of eVTOL stations a critical question to explore. Existing approaches often face limitations, such as the inability to interact station locations with demand or difficulty in finding the optimal solution for large study regions. Here, this paper first proposes a mathematical model to generate optimal eVTOL station locations while considering associated potential eVTOL demand, and then proposes a heuristic algorithm, Hierarchical Optimization MEthod (HOME), to efficiently solve the model. With a case study of Southern California, HOME was compared to 1) directly solving the original integer linear programming-based network design problem, and 2) employing the widely used genetic algorithm. Results suggest that HOME can find optimal solutions with limited computational resources. The proposed framework powered by HOME provides a computationally efficient way to support urban air mobility planning.

97 MATHEMATICS AND COMPUTING↗

DVMDOSTEM v0.8.3: a terrestrial ecosystem model designed to represent arctic, boreal and permafrost ecosystem dynamics

The impacts of climate change on natural ecosystems are the result of complex physical and ecological processes operating and interacting at a variety of spatio-temporal scales, that can be represented in process-based ecosystem models. DVMDOSTEM is an advanced process-based terrestrial ecosystem model (TEM) designed to study ecosystem responses to climate changes and disturbances. It has a particular focus on permafrost regions (i.e. regions characterized by soils that stay partially frozen all year round for at least two consecutive years), encompassing boreal, arctic, and alpine landscapes. The model couples two previous versions of the Terrestrial Ecosystem Model (TEM) (McGuire et al., 1992): DVMTEM that includes a dynamic vegetation module (DVM) (E. S. Euskirchen et al., 2009), and DOSTEM that includes a dynamic organic soil module (DOS) (H. Genet et al., 2013; Yi et al., 2010). DVMDOSTEM simulates processes at yearly and monthly scales, with some physical processes operating at an even finer temporal resolution. Its versatility allows for site-specific to regional simulations, making it valuable for predicting shifts in permafrost, vegetation, and carbon (C) and nitrogen (N) dynamics. While DVMDOSTEM has been described in the methods sections of many manuscripts, this paper is the first stand alone description of DVMDOSTEM, independent of a particular scientific investigation.

Carman, Tobey B. [Univ. of Alaska, Fairbanks, AK (↗

Web-Based Tools for Data-Informed Remedy Optimization: Software Theory and User Guide

This report documents the development and application of two web-based decision-support tools for pump-and-treat (P&T) groundwater remediation systems: PTOLEMY (Pump-and-Treat Optimized Location Evaluation to Maximize Yields) and OPTIMA (Optimization for Pump-and-Treat Implementation, Management, & Assessment). These tools enhance remedy design and management by leveraging advanced computational methods – specifically deep learning and multi-objective optimization – within a user-friendly platform. By integrating data-driven models with established hydrogeological knowledge, PTOLEMY and OPTIMA enable more efficient evaluation of well placement and operational strategies, helping site managers balance multiple remediation objectives under complex conditions. Both tools are implemented as modules within the SOCRATES (Suite Of Comprehensive Rapid Analysis Tools for Environmental Sites) web platform, which provides data access, visualization, and analytics to support remedy optimization across sites in the U.S. Department of Energy Office of Environmental Management complex. PTOLEMY is a rapid screening module designed to identify promising locations for new extraction wells. It employs a multi-channel three-dimensional convolutional neural network (MC3D-CNN) trained on high-fidelity simulation data to predict the relative performance (in terms of contaminant mass recovery) of potential well sites. Through an interactive web interface, PTOLEMY visualizes the probability of high performance across a site, highlighting areas where an extraction well is likely to yield above-threshold contaminant removal over a multi-year period. PTOLEMY’s map-based displays and exportable results support transparent communication of screening analyses. By focusing attention on the most favorable candidate locations, the tool augments traditional engineering judgment and physics-based modeling, providing a data informed basis for subsequent detailed evaluations. OPTIMA is a multi objective optimization module designed to find wellfield layouts and operating schedules that meet various cleanup goals. It quickly evaluates thousands of candidate setups – combinations of well locations, timing, and rates – and returns a small set of best trade-off options for comparison. At its core, OPTIMA uses a U-Net-based surrogate model – a deep-learning emulator of a groundwater flow and transport simulator – to dramatically accelerate scenario evaluations. Coupling this fast surrogate with the NSGA-II (Non-dominated Sorting Genetic Algorithm II) evolutionary algorithm, OPTIMA explores a wide decision space of well locations and schedules to identify Pareto-optimal solutions that trade off key objectives (e.g., minimizing cleanup time, maximizing contaminant mass removal, and minimizing plume extent). The tool outputs a family of optimal configurations and visualizes their trade-offs (Pareto frontiers of cleanup metrics and maps of optimized well placements). Site managers can use these results to understand the range of viable strategies and to select candidate designs for more detailed verification. OPTIMA is currently under active development and not yet fully released; this guide provides early documentation to support planning and gather user feedback.

54 ENVIRONMENTAL SCIENCES↗

Design and Characterization of a Transcriptional Repression Toolkit for Plants

Regulation of gene expression is essential for all life. Tools to manipulate the gene expression level have therefore proven to be very valuable in efforts to engineer biological systems. However, there are few well-characterized genetic parts that reduce gene expression in plants, commonly known as transcriptional repressors. We characterized the repression activity of a library consisting of repression motifs from approximately 25% of the members of the largest known family of repressors. Combining sequence information with our trans-regulatory function data, we next generated a library of synthetic transcriptional repression motifs with function predicted in advance. After characterizing our synthetic library, we demonstrated not only that many of our synthetic constructs were functional as repressors but also that our advance predictions of repression strength were better than random guesses. Finally, we assessed the functionality of known transcriptional repression motifs from a wide range of eukaryotes. Our study represents the largest plant repressor motif library experimentally characterized to date, providing unique opportunities for tuning transcription in plants.

59 BASIC BIOLOGICAL SCIENCES↗

Learning model combining convolutional deep neural network with a self-attention mechanism for AC optimal power flow

Alternating current optimal power flow (OPF) analysis is critical for efficient and reliable operation of power systems. For large systems or repetitive computations, the traditional methods such as the direct and gradient methods, or non-traditional methods, such as the genetic algorithm and simulating annealing, are time-consuming and unsuitable for real-time computing. The work in this paper proposes a novel framework to obtain the optimal solution of power flow in real-time using a combination of convolutional neural networks and a self-attention mechanism. All parameters of the power networks are rearranged in an image-like shape of a multi-channel image where each channel is a two-dimensional matrix. The proposed approach is adaptive with every input size of power systems as well as frequent variations of network topologies without intervention to the framework core. The encompassment of all power system contexts in which all parameters of internal elements, generation costs, and topology information are included, contributes to the higher accuracy of inference compared to other current machine-learning-based OPF-solving methods. Besides, the proposed framework established on ubiquitous platforms is effortlessly integrated into current infrastructures of power systems, and the great efficiency along with the computation speed may serve as a critical point for practical implications, such as enabling faster decision-making during real-time operations, predicting system contingencies, and remedial actions based on an offline pre-trained model. Furthermore, this supervised learning process is applied to the dataset of four case studies of meshed power systems: the IEEE 5-bus system (IEEE-5), the IEEE 30-bus system (IEEE-30), the IEEE 39-bus system (IEEE-39), and the IEEE 57-bus system (IEEE-57) to prove the efficacy of the proposed method.

42 ENGINEERING↗