Search NASA⌕ Search

SEARCH · Search NASA

Results for “genetic algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Gene Expression Dynamics Inspector (GEDI): for integrative analysis of expression profiles

Genome-wide expression profiles contain global patterns that evade visual detection in current gene clustering analysis. Here, a Gene Expression Dynamics Inspector (GEDI) is described that uses self-organizing maps to translate high-dimensional expression profiles of time courses or sample classes into animated, coherent and robust mosaics images. GEDI facilitates identification of interesting patterns of molecular activity simultaneously across gene, time and sample space without prior assumption of any structure in the data, and then permits the user to retrieve genes of interest. Important changes in genome-wide activities may be quickly identified based on 'Gestalt' recognition and hence, GEDI may be especially useful for non-specialist end users, such as physicians. AVAILABILITY: GEDI v1.0 is written in Matlab, and binary Matlab.dll files which require Matlab to run can be downloaded for free by academic institutions at http://www.chip.org/~ge/gedihome.html Supplementary information: http://www.chip.org/~ge/gedihome.html.

Database Management Systems↗

Stacking sequence optimization of simply supported laminates with stability and strain constraints

An integer programming formulation for the design of symmetric and balanced rectangular composite laminates with simply supported boundary conditions subject to buckling and strain constraints is presented. The design variables that define the stacking sequence of the laminate are ply-identity zero-one integers. The buckling constraint is linear in terms of the ply-identity design variables, but strains are nonlinear functions of these variables. A linear approximation is developed for the strain constraints so that the problem can be solved by sequential linearization using the branch and bound algorithm. Examples of graphite-epoxy plates under biaxial compression are presented. Optimum stacking sequences obtained using the linear approximation are compared with global optimum designs obtained using a genetic search procedure.

Nagendra, S.↗

Steady-State ALPS for Real-Valued Problems

The two objectives of this paper are to describe a steady-state version of the Age-Layered Population Structure (ALPS) Evolutionary Algorithm (EA) and to compare it against other GAs on real-valued problems. Motivation for this work comes from our previous success in demonstrating that a generational version of ALPS greatly improves search performance on a Genetic Programming problem. In making steady-state ALPS some modifications were made to the method for calculating age and the method for moving individuals up layers. To demonstrate that ALPS works well on real-valued problems we compare it against CMA-ES and Differential Evolution (DE) on five challenging, real-valued functions and on one real-world problem. While CMA-ES and DE outperform ALPS on the two unimodal test functions, ALPS is much better on the three multimodal test problems and on the real-world problem. Further examination shows that, unlike the other GAs, ALPS maintains a genotypically diverse population throughout the entire search process. These findings strongly suggest that the ALPS paradigm is better able to avoid premature convergence then the other GAs.

Hornby, Gregory S.↗

Utilizing digitized occurrence records of Midwestern feral Cannabis sativa to develop ecological niche models

Hemp (Cannabis sativa L.) has historically played a vital role in agriculture across the globe. Feral and wild populations have served as genetic resources for breeding, conservation, and adaptation to changing environmental conditions. However, feral populations of Cannabis, specifically in the Midwestern United States, remain poorly understood. This study aims to characterize the abiotic tolerances of these populations, estimate suitable areas, identify regions at risk of abiotic suitability change, and highlight the utility of ecological niche models (ENMs) in germplasm conservation. The Maxent algorithm was used to construct a series of ENMs. Validation metrics and MOP (Mobility-oriented Parity) analysis were used to assess extrapolation risk and model performance. We also projected the final projected under current and future climate scenarios (2021–2040 and 2061–2080) to assess how abiotic suitability changes with time. Climate change scenarios indicated an expansion of suitable habitat, with priority areas for germplasm collection in Indiana, Illinois, Kansas, Missouri, and Nebraska. This study demonstrates the application of ENMs for characterizing feral Cannabis populations and highlights their value in germplasm conservation and breeding efforts. Populations of feral C. sativa in the Midwest are of high interest, and future research should focus on utilizing tools to aid the collection of materials for the characterization of genetic diversity and adaptation to a changing climate.

59 BASIC BIOLOGICAL SCIENCES↗

Numerical Treatment of Degenerate Diffusion Equations via Feller's Boundary Classification, and Applications

A numerical method is devised to solve a class of linear boundary-value problems for one-dimensional parabolic equations degenerate at the boundaries. Feller theory, which classifies the nature of the boundary points, is used to decide whether boundary conditions are needed to ensure uniqueness, and, if so, which ones they are. The algorithm is based on a suitable preconditioned implicit finite-difference scheme, grid, and treatment of the boundary data. Second-order accuracy, unconditional stability, and unconditional convergence of solutions of the finite-difference scheme to a constant as the time-step index tends to infinity are further properties of the method. Several examples, pertaining to financial mathematics, physics, and genetics, are presented for the purpose of illustration.

Cacio, Emanuela↗

Scaling features of noncoding DNA

We review evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene, and utilize this fact to build a Coding Sequence Finder Algorithm, which uses statistical ideas to locate the coding regions of an unknown DNA sequence. Finally, we describe briefly some recent work adapting to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function, and reporting that noncoding regions in eukaryotes display a larger redundancy than coding regions. Specifically, we consider the possibility that this result is solely a consequence of nucleotide concentration differences as first noted by Bonhoeffer and his collaborators. We find that cytosine-guanine (CG) concentration does have a strong "background" effect on redundancy. However, we find that for the purine-pyrimidine binary mapping rule, which is not affected by the difference in CG concentration, the Shannon redundancy for the set of analyzed sequences is larger for noncoding regions compared to coding regions.

Non-NASA Center↗

Quantifying leaf symptoms of sorghum charcoal rot in images of field‐grown plants using deep neural networks

Abstract Charcoal rot of sorghum (CRS) is a significant disease affecting sorghum crops, with limited genetic resistance available. The causative agent, Macrophomina phaseolina (Tassi) Goid, is a highly destructive fungal pathogen that targets over 500 plant species globally, including essential staple crops. Utilizing field image data for precise detection and quantification of CRS could greatly assist in the prompt identification and management of affected fields and thereby reduce yield losses. The objective of this work was to implement various machine learning algorithms to evaluate their ability to accurately detect and quantify CRS in red‐green‐blue images of sorghum plants exhibiting symptoms of infection. EfficientNet‐B3 and a fully convolutional network emerged as the top‐performing models for image classification and segmentation tasks, respectively. Among the classification models evaluated, EfficientNet‐B3 demonstrated superior performance, achieving an accuracy of 86.97%, a recall rate of 0.71, and an F1 score of 0.73. Of the segmentation models tested, FCN proved to be the most effective, exhibiting a validation accuracy of 97.76%, a recall rate of 0.68, and an F1 score of 0.66. As the size of the image patches increased, both models’ validation scores increased linearly, and their inference time decreased exponentially. This trend could be attributed to larger patches containing more information, improving model performance, and fewer patches reducing the computational load, thus decreasing inference time. The models, in addition to being immediately useful for breeders and growers of sorghum, advance the domain of automated plant phenotyping and may serve as a foundation for drone‐based or other automated field phenotyping efforts. Additionally, the models presented herein can be accessed through a web‐based application where users can easily analyze their own images.

Gonzalez, Emmanuel M.↗

System and method for embedding emotion in logic systems

A system, method, and computer readable-media for creating a stable synthetic neural system. The method includes training an intellectual choice-driven synthetic neural system (SNS), training an emotional rule-driven SNS by generating emotions from rules, incorporating the rule-driven SNS into the choice-driven SNS through an evolvable interface, and balancing the emotional SNS and the intellectual SNS to achieve stability in a nontrivial autonomous environment with a Stability Algorithm for Neural Entities (SANE). Generating emotions from rules can include coding the rules into the rule-driven SNS in a self-consistent way. Training the emotional rule-driven SNS can occur during a training stage in parallel with training the choice-driven SNS. The training stage can include a self assessment loop which measures performance characteristics of the rule-driven SNS against core genetic code. The method uses a stability threshold to measure stability of the incorporated rule-driven SNS and choice-driven SNS using SANE.

Curtis, Steven A.↗

ALPS: The Age-Layered Population Structure for Reducing the Problem of Premature Convergence

To reduce the problem of premature convergence we define a new attribute of an individual, its age, and propose the Age-Layered Population Structure (ALPS), in which age is used to restrict competition and breeding between members of the population. ALPS differs from a typical EA by segregating individuals into different age-layers by their age - a measure of how long the genetic material has been in the population - and by regularly replacing all individuals in the bottom layer with randomly generated ones. The introduction of new, randomly generated individuals at regular intervals results in an EA that is never completely converged and is always looking at new parts of the fitness landscape. By using age to restrict competition and breeding search is able to develop promising young individuals without them being dominated by older ones. We demonstrate the effectiveness of the ALPS algorithm on an antenna design problem in which evolution with ALPS produces antennas more than twice as good as does evolution with two other types of EAs. Further analysis shows that the ALPS model does allow the offspring of newly generated individuals to move the population out of mediocre local-optima to better parts of the fitness landscape.

Hornby, Gregory S.↗

Focused Metabolite Profiling for Dissecting Cellular and Molecular Processes of Living Organisms in Space Environments

Regulatory control in biological systems is exerted at all levels within the central dogma of biology. Metabolites are the end products of all cellular regulatory processes and reflect the ultimate outcome of potential changes suggested by genomics and proteomics caused by an environmental stimulus or genetic modification. Following on the heels of genomics, transcriptomics, and proteomics, metabolomics has become an inevitable part of complete-system biology because none of the lower "-omics" alone provide direct information about how changes in mRNA or protein are coupled to changes in biological function. The challenges are much greater than those encountered in genomics because of the greater number of metabolites and the greater diversity of their chemical structures and properties. To meet these challenges, much developmental work is needed, including (1) methodologies for unbiased extraction of metabolites and subsequent quantification, (2) algorithms for systematic identification of metabolites, (3) expertise and competency in handling a large amount of information (data set), and (4) integration of metabolomics with other "omics" and data mining (implication of the information). This article reviews the project accomplishments.

Source record↗

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

G F Bomarito↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

Uncertainty quantification↗

NASA Strategy to Safely Live and Work in the Space Radiation Environment

In space, astronauts are constantly bombarded with energetic particles. The goal of the National Aeronautics and Space Agency and the NASA Space Radiation Project is to ensure that astronauts can safely live and work in the space radiation environment. The space radiation environment poses both acute and chronic risks to crew health and safety, but unlike some other aspects of space travel, space radiation exposure has clinically relevant implications for the lifetime of the crew. Among the identified radiation risks are cancer, acute and late CNS damage, chronic and degenerative tissue decease, and acute radiation syndrome. The term "safely" means that risks are sufficiently understood such that acceptable limits on mission, post-mission and multi-mission consequences can be defined. The NASA Space Radiation Project strategy has several elements. The first element is to use a peer-reviewed research program to increase our mechanistic knowledge and genetic capabilities to develop tools for individual risk projection, thereby reducing our dependency on epidemiological data and population-based risk assessment. The second element is to use the NASA Space Radiation Laboratory to provide a ground-based facility to study the health effects/mechanisms of damage from space radiation exposure and the development and validation of biological models of risk, as well as methods for extrapolation to human risk. The third element is a risk modeling effort that integrates the results from research efforts into models of human risk to reduce uncertainties in predicting the identified radiation risks. To understand the biological basis for risk, we must also understand the physical aspects of the crew environment. Thus, the fourth element develops computer algorithms to predict radiation transport properties, evaluate integrated shielding technologies and provide design optimization recommendations for the design of human space systems. Understanding the risks and determining methods to mitigate the risks are keys to a successful radiation protection strategy.

Cucinotta, Francis A.↗

Nasa GeneLab Computomics Reveal Horizontal Gene Transfer on International Space Station Environmental Metagenomes

Prokaryotic lifeforms can be observed to demonstrate many keen adaptive advantages, perhaps facilitated by a nature simplistic relative to divergent domains of life. In particular, decompartmentalized gene expression facilitates adaptation by allowing free exchange of genetic material, albeit at the cost of increased susceptibility to genetic damage. Thus, these lifeforms must compensate by embracing diverse investment strategies in an attempt to “brute force” the evolvability equation through precipitous genesis, lean metabolic efficiency, and sheer population. This prokaryotic archetype also enables symbiotic relationships with secondary mobile genetic elements known as plasmids, which have been shown to drive evolution on rapid temporal scales through processes such as conjugation and transformation. This study attempts to decipher whether these mechanisms of horizontal gene transfer (HGT) are major factors in determining prokaryote fitness within a unique isolated environment, the International Space Station (ISS). The ISS Microbial Tracking (MT) project has generated a wealth of data concerning the successive reigns of microbial genera that appear to thrive amidst harsh conditions for life. Despite relatively higher doses of ionizing radiation as compared to Earth, complications associated with microgravity, and the anti-microbial mélange deployed, microbial life still persists in this environment. The NASA GeneLab serves as a data repository and analysis platform to enable researchers to access space flight factor related data. With the use of GeneLab’s modern computational suites (computomics), phylogenetic and functional genomic investigations of HGT events were conducted on the data generated from the MT-1 project. The putative data concerning the plasmid population (plasmidome) of the ISS was algorithmically derived and compared to those of habitats with similar environmental dynamics- such as living quarters and hospitals- to investigate whether these HGT elements may play crucial role(s) in shaping the microbiome of this closed habitat that serves as the only inhabited structure in space.

Bense, Nicholas↗

A frugal CRISPR kit for equitable and accessible education in gene editing and synthetic biology

Equitable and accessible education in life sciences, bioengineering, and synthetic biology is crucial for training the next generation of scientists, fostering transparency in public decision-making, and ensuring biotechnology can benefit a wide-ranging population. As a groundbreaking technology for genome engineering, CRISPR has transformed research and therapeutics. However, hands-on exposure to this technology in educational settings remains limited due to the extensive resources required for CRISPR experiments. Here, we develop CRISPRkit, an affordable kit designed for gene editing and regulation in high school education. CRISPRkit eliminates the need for specialized equipment, prioritizes biosafety, and utilizes cost-effective reagents. By integrating CRISPRi gene regulation, colorful chromoproteins, cell-free transcription-translation systems, smartphone-based quantification, and an in-house automated algorithm (CRISPectra), our kit offers an inexpensive (~$2) and user-friendly approach to performing and analyzing CRISPR experiments, without the need for a traditional laboratory setup. Experiments conducted by high school students in classroom settings highlight the kit’s utility for reliable CRISPRkit experiments. Furthermore, CRISPRkit provides a modular and expandable platform for genome engineering, and we demonstrate its applications for controlling fluorescent proteins and metabolic pathways such as melanin production. We envision CRISPRkit will facilitate biotechnology education for communities of diverse socioeconomic and geographic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗