Search NASA⌕ Search

SEARCH · Search NASA

Results for “Genetic Code”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

Validity of the Aluminum Equivalent Approximation in Space Radiation Shielding

The origin of the aluminum equivalent shield approximation in space radiation analysis can be traced back to its roots in the early years of the NASA space programs (Mercury, Gemini and Apollo) wherein the primary radiobiological concern was the intense sources of ionizing radiation causing short term effects which was thought to jeopardize the safety of the crew and hence the mission. Herein, it is shown that the aluminum equivalent shield approximation, although reasonably well suited for that time period and to the application for which it was developed, is of questionable usefulness to the radiobiological concerns of routine space operations of the 21 st century which will include long stays onboard the International Space Station (ISS) and perhaps the moon. This is especially true for a risk based protection system, as appears imminent for deep space exploration where the long-term effects of Galactic Cosmic Ray (GCR) exposure is of primary concern. The present analysis demonstrates that sufficiently large errors in the interior particle environment of a spacecraft result from the use of the aluminum equivalent approximation, and such approximations should be avoided in future astronaut risk estimates. In this study, the aluminum equivalent approximation is evaluated as a means for estimating the particle environment within a spacecraft structure induced by the GCR radiation field. For comparison, the two extremes of the GCR environment, the 1977 solar minimum and the 2001 solar maximum, are considered. These environments are coupled to the Langley Research Center (LaRC) deterministic ionized particle transport code High charge (Z) and Energy TRaNsport (HZETRN), which propagates the GCR spectra for elements with charges (Z) in the range I <= Z <= 28 (H -- Ni) and secondary neutrons through selected target materials. The coupling of the GCR extremes to HZETRN allows for the examination of the induced environment within the interior' of an idealized spacecraft as approximated by a spherical shell shield, and the effects of the aluminum equivalent approximation for a good polymeric shield material such as genetic polyethylene (PE). The shield thickness is represented by a 25 g/cm spherical shell. Although one could imagine the progression to greater thickness, the current range will be sufficient to evaluate the qualitative usefulness of the aluminum equivalent approximation. Upon establishing the inaccuracies of the aluminum equivalent approximation through numerical simulations of the GCR radiation field attenuation for PE and aluminum equivalent PE spherical shells, we Anther present results for a limited set of commercially available, hydrogen rich, multifunctional polymeric constituents to assess the effect of the aluminum equivalent approximation on their radiation attenuation response as compared to the generic PE.

Badavi, Francis F.↗

Conceptual Design of a Counter-Rotating Fan System for Distributed Boundary Layer Ingesting Propulsion

The present paper details the design of the counter rotating fans for a Turboelectric Distributed Propulsion (TeDP) system. Sixteen propulsors installed in mail-slot-shape nacelles are embedded on an aerodynamically optimized hybrid wing-body configuration. The hybrid-wing/body (HWB) configuration which was previously designed to satisfy the conditions of trim, longitudinally static stability and specific cargo space is employed as the baseline configuration in pursuing an optimal distributed propulsion system. A set of distributed propulsors is conceptually designed and the collective performance is evaluated against the target thrust mandated by the mission requirements. The concept of the distributed propulsion allows the fan pressure ratio to be around 1.27~1.32 for the target thrust. In addition, further splitting of the fan pressure ratio by using the counter-rotating fans for each slot realizes the target pressure ratio with low tip speed. In the distributed propulsion system, the nature of the flow conditions and/or the thickness of the ingested boundary layer may differ and result in different propulsive reaction of each individual propulsor. The optimization is, thus, approached from both the propulsion system and individual propulsor perspectives. An optimal distribution of the thrust and power output is determined by how the system utilizes each passage's propulsive characteristics and its interaction with the airframe. These system level analysis and optimization are conducted using an actuator disk model to account for the propulsion-airframe integration numerically. With respect to the propulsor level, aerodynamic shape optimizations of the fan blades are performed in a sequential multi-objective optimization process for various design objectives, such as mass flow rate condition, fan pressure ratio, efficiency and the exit flow angle of the fan stage by using a genetic algorithm, NSGA-II. The radial chord distribution, and meanline distribution of the rotors are designed on the circumferentially averaged axi-symmetric inlet profiles and tested on the six inlet profiles from six divided sectors to reckon flow distortion. The performances of the counter rotating fans are, thus, evaluated accordingly for obtaining distortion tolerant fan. The performance of the distributed propulsion system is evaluated by two CFD tools, i.e., a multi-stage turbo-machinery CFD code and one propulsion-airframe integration flow solver coupled with a body-force model. The optimized boundary layer ingestion propulsion system of 16 distributed slots not only reaches the system target thrust, but also delivers a close to 20% fuel saving benefit against its counterpart 12 distributed clean inlet propulsion system.

Boundary-Layer-Ingestion Propulsion↗

Multi-Objective Optimization of Uranium Target Assembly–3: A Comparison of Genetic and Traditional Methods

Commonly produced as a byproduct of uranium fission, 99 Mo is a key medical isotope that is in high demand in the United States. An international goal is to switch from medical isotope production technologies that require highly enriched uranium to medical isotope production technologies that require only low-enriched uranium. Niowave Inc. is contributing to this goal by developing an accelerator-driven subcritical assembly called the Uranium Target Assembly (UTA). This work compares the performance of Dakota’s Multi-Objective Genetic Algorithm (MOGA) against traditional sensitivity analysis in the neutronic optimization of the UTA-3 system. The design objectives are k-eigenvalue (k eff ) and natural uranium fission power, which are directly correlated with the amount of 99 Mo produced. Dakota:MOGA did not perform as well as human engineering ingenuity in optimization studies with high numbers of input parameters, such as fuel rod type selection and fuel rod placement. However, Dakota:MOGA did outperform traditional sensitivity analysis in optimization studies with fewer than 20 parameters and revealed the degree to which each parameter influences the optimal design space for k eff and natural uranium fission power (to a lesser extent). As the design model became more complex in the final stage of design, the computational resources required to calculate the design objective values in the Monte Carlo N-Particle transport code from selected input parameter combinations limited Dakota:MOGA’s performance, and, unfortunately, human intervention was required to discern the optimal design space. In conclusion, future work will attempt to reduce computational resource constraints by incorporating areduced-order neutronics model into the optimization cycle.

Accelerator-driven systems↗

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew↗

The mGA1.0: A common LISP implementation of a messy genetic algorithm

Genetic algorithms (GAs) are finding increased application in difficult search, optimization, and machine learning problems in science and engineering. Increasing demands are being placed on algorithm performance, and the remaining challenges of genetic algorithm theory and practice are becoming increasingly unavoidable. Perhaps the most difficult of these challenges is the so-called linkage problem. Messy GAs were created to overcome the linkage problem of simple genetic algorithms by combining variable-length strings, gene expression, messy operators, and a nonhomogeneous phasing of evolutionary processing. Results on a number of difficult deceptive test functions are encouraging with the mGA always finding global optima in a polynomial number of function evaluations. Theoretical and empirical studies are continuing, and a first version of a messy GA is ready for testing by others. A Common LISP implementation called mGA1.0 is documented and related to the basic principles and operators developed by Goldberg et. al. (1989, 1990). Although the code was prepared with care, it is not a general-purpose code, only a research version. Important data structures and global variations are described. Thereafter brief function descriptions are given, and sample input data are presented together with sample program output. A source listing with comments is also included.

Goldberg, David E.↗

Rapid evolution of cis-regulatory sequences via local point mutations

Although the evolution of protein-coding sequences within genomes is well understood, the same cannot be said of the cis-regulatory regions that control transcription. Yet, changes in gene expression are likely to constitute an important component of phenotypic evolution. We simulated the evolution of new transcription factor binding sites via local point mutations. The results indicate that new binding sites appear and become fixed within populations on microevolutionary timescales under an assumption of neutral evolution. Even combinations of two new binding sites evolve very quickly. We predict that local point mutations continually generate considerable genetic variation that is capable of altering gene expression.

Non-NASA Center↗

Identification of a small tetraheme cytochrome c and a flavocytochrome c as two of the principal soluble cytochromes c in Shewanella oneidensis strain MR1

Two abundant, low-redox-potential cytochromes c were purified from the facultative anaerobe Shewanella oneidensis strain MR1 grown anaerobically with fumarate. The small cytochrome was completely sequenced, and the genes coding for both proteins were cloned and sequenced. The small cytochrome c contains 91 residues and four heme binding sites. It is most similar to the cytochromes c from Shewanella frigidimarina (formerly Shewanella putrefaciens) NCIMB400 and the unclassified bacterial strain H1R (64 and 55% identity, respectively). The amount of the small tetraheme cytochrome is regulated by anaerobiosis, but not by fumarate. The larger of the two low-potential cytochromes contains tetraheme and flavin domains and is regulated by anaerobiosis and by fumarate and thus most nearly corresponds to the flavocytochrome c-fumarate reductase previously characterized from S. frigidimarina to which it is 59% identical. However, the genetic context of the cytochrome genes is not the same for the two Shewanella species, and they are not located in multicistronic operons. The small cytochrome c and the cytochrome domain of the flavocytochrome c are also homologous, showing 34% identity. Structural comparison shows that the Shewanella tetraheme cytochromes are not related to the Desulfovibrio cytochromes c(3) but define a new folding motif for small multiheme cytochromes c.

Cytochrome c Group/chemistry/genetics/metabolism↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

Exon disruptive variants in Populus trichocarpa associated with wood properties exhibit distinct gene expression patterns

Abstract Forest trees may harbor naturally occurring exon disruptive variants (DVs) in their gene sequences, which potentially impact important ecological and economic phenotypic traits. However, the abundance and molecular regulation of these variants remain largely unexplored. Here, 24,420 DVs were identified by screening 1014Populus trichocarpafull genomes. The identified DVs were predominantly heterozygous with allelic frequencies below 5% (only 26% of DVs had frequencies greater than 5%). Using common garden‐grown trees, DVs were assessed for gene expression variation in the developing xylem, revealing that their gene expression can be significantly altered, particularly for homozygous DVs (in the range of 27%–38% of cases depending on the studied common garden). DVs were further investigated for their correlations with 13 wood quality traits, revealing that, among the 148 discovered DV associations, 15 correlated with more than one wood property and six genes had more than one DV in their coding sequences associated with wood traits. Approximately one‐third of DVs correlated with wood property variation also showed significant gene expression variation, confirming their non‐spurious impact. These findings offer potential avenues for targeted introduction of homozygous mutations using tree biotechnology, and while the exact mechanisms by which DVs may directly influence wood formation remain to be unraveled, this study lays the groundwork for further investigation.

Genetics & Heredity↗

Extension of an Object-Oriented Optimization Tool: User's Reference Manual

The National Aeronautics and Space Administration Armstrong Flight Research Center has developed a cost-effective and flexible object-oriented optimization (O (sup 3)) tool that leverages existing tools and practices and allows easy integration and adoption of new state-of-the-art software. This object-oriented framework can integrate the analysis codes for multiple disciplines, as opposed to relying on one code to perform analysis for all disciplines. Optimization can thus take place within each discipline module, or in a loop between the O (sup 3) tool and the discipline modules, or both. Six different sample mathematical problems are presented to demonstrate the performance of the O (sup 3) tool. Instructions for preparing input data for the O (sup 3) tool are detailed in this user's manual.

Multidisciplinary design optimization↗

Software Helps Retrieve Information Relevant to the User

The Adaptive Indexing and Retrieval Agent (ARNIE) is a code library, designed to be used by an application program, that assists human users in retrieving desired information in a hypertext setting. Using ARNIE, the program implements a computational model for interactively learning what information each human user considers relevant in context. The model, called a "relevance network," incrementally adapts retrieved information to users individual profiles on the basis of feedback from the users regarding specific queries. The model also generalizes such knowledge for subsequent derivation of relevant references for similar queries and profiles, thereby, assisting users in filtering information by relevance. ARNIE thus enables users to categorize and share information of interest in various contexts. ARNIE encodes the relevance and structure of information in a neural network dynamically configured with a genetic algorithm. ARNIE maintains an internal database, wherein it saves associations, and from which it returns associated items in response to a query. A C++ compiler for a platform on which ARNIE will be utilized is necessary for creating the ARNIE library but is not necessary for the execution of the software.

Mathe, Natalie↗

NASA Tech Briefs, June 2004

Topics covered include: COTS MEMS Flow-Measurement Probes; Measurement of an Evaporating Drop on a Reflective Substrate; Airplane Ice Detector Based on a Microwave Transmission Line; Microwave/Sonic Apparatus Measures Flow and Density in Pipe; Reducing Errors by Use of Redundancy in Gravity Measurements; Membrane-Based Water Evaporator for a Space Suit; Compact Microscope Imaging System with Intelligent Controls; Chirped-Superlattice, Blocked-Intersubband QWIP; Charge-Dissipative Electrical Cables; Deep-Sea Video Cameras Without Pressure Housings; RFID and Memory Devices Fabricated Integrally on Substrates; Analyzing Dynamics of Cooperating Spacecraft; Spacecraft Attitude Maneuver Planning Using Genetic Algorithms; Forensic Analysis of Compromised Computers; Document Concurrence System; Managing an Archive of Images; MPT Prediction of Aircraft-Engine Fan Noise; Improving Control of Two Motor Controllers; Electro-deionization Using Micro-separated Bipolar Membranes; Safer Electrolytes for Lithium-Ion Cells; Rotating Reverse-Osmosis for Water Purification; Making Precise Resonators for Mesoscale Vibratory Gyroscopes; Robotic End Effectors for Hard-Rock Climbing; Improved Nutation Damper for a Spin-Stabilized Spacecraft; Exhaust Nozzle for a Multitube Detonative Combustion Engine; Arc-Second Pointer for Balloon-Borne Astronomical Instrument; Compact, Automated Centrifugal Slide-Staining System; Two-Armed, Mobile, Sensate Research Robot; Compensating for Effects of Humidity on Electronic Noses; Brush/Fin Thermal Interfaces; Multispectral Scanner for Monitoring Plants; Coding for Communication Channels with Dead-Time Constraints; System for Better Spacing of Airplanes En Route; Algorithm for Training a Recurrent Multilayer Perceptron; Orbiter Interface Unit and Early Communication System; White-Light Nulling Interferometers for Detecting Planets; and Development of Methodology for Programming Autonomous Agents.

Source record↗

Genetic learning in rule-based and neural systems

The design of neural networks and fuzzy systems can involve complex, nonlinear, and ill-conditioned optimization problems. Often, traditional optimization schemes are inadequate or inapplicable for such tasks. Genetic Algorithms (GA's) are a class of optimization procedures whose mechanics are based on those of natural genetics. Mathematical arguments show how GAs bring substantial computational leverage to search problems, without requiring the mathematical characteristics often necessary for traditional optimization schemes (e.g., modality, continuity, availability of derivative information, etc.). GA's have proven effective in a variety of search tasks that arise in neural networks and fuzzy systems. This presentation begins by introducing the mechanism and theoretical underpinnings of GA's. GA's are then related to a class of rule-based machine learning systems called learning classifier systems (LCS's). An LCS implements a low-level production-system that uses a GA as its primary rule discovery mechanism. This presentation illustrates how, despite its rule-based framework, an LCS can be thought of as a competitive neural network. Neural network simulator code for an LCS is presented. In this context, the GA is doing more than optimizing and objective function. It is searching for an ecology of hidden nodes with limited connectivity. The GA attempts to evolve this ecology such that effective neural network performance results. The GA is particularly well adapted to this task, given its naturally-inspired basis. The LCS/neural network analogy extends itself to other, more traditional neural networks. Conclusions to the presentation discuss the implications of using GA's in ecological search problems that arise in neural and fuzzy systems.

Smith, Robert E.↗

Introducing GPU Acceleration into the Python-Based Simulations of Chemistry Framework

We introduce the first version of GPU4P Y SCF, a module that provides GPU acceleration of methods in P Y SCF. As a core functionality, this provides a GPU implementation of two-electron repulsion integrals (ERIs) for contracted basis sets comprising up to g functions using the Rys quadrature. As an illustration of how this can accelerate a quantum chemistry workflow, we describe how to use the ERIs efficiently in the integral-direct Hartree–Fock build and nuclear gradient construction. Benchmark calculations show a significant speedup of 2 orders of magnitude with respect to the multithreaded CPU Hartree–Fock code of P Y SCF and the performance comparable to other open-source GPU-accelerated quantum chemical packages, including GAMESS and QUICK, on a single NVIDIA A100 GPU.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Strategy to Safely Live and Work in the Space Radiation Environment

The goal of the National Aeronautics and Space Agency and the Space Radiation Project is to ensure that astronauts can safely live and work in the space radiation environment. The space radiation environment poses both acute and chronic risks to crew health and safety, but unlike some other aspects of space travel, space radiation exposure has clinically relevant implications for the lifetime of the crew. The term safely means that risks are sufficiently understood such that acceptable limits on mission, post-mission and multi-mission consequences (for example, excess lifetime fatal cancer risk) can be defined. The Space Radiation Project strategy has several elements. The first element is to use a peer-reviewed research program to increase our mechanistic knowledge and genetic capabilities to develop tools for individual risk projection, thereby reducing our dependency on epidemiological data and population-based risk assessment. The second element is to use the NASA Space Radiation Laboratory to provide a ground-based facility to study the understanding of health effects/mechanisms of damage from space radiation exposure and the development and validation of biological models of risk, as well as methods for extrapolation to human risk. The third element is a risk modeling effort that integrates the results from research efforts into models of human risk to reduce uncertainties in predicting risk of carcinogenesis, central nervous system damage, degenerative tissue disease, and acute radiation effects. To understand the biological basis for risk, we must also understand the physical aspects of the crew environment. Thus the fourth element develops computer codes to predict radiation transport properties, evaluate integrated shielding technologies and provide design optimization recommendations for the design of human space systems. Understanding the risks and determining methods to mitigate the risks are keys to a successful radiation protection strategy.

Corbin, Barbara J.↗

Statistical and linguistic features of DNA sequences

We present evidence supporting the idea that the DNA sequence in genes containing noncoding regions is correlated, and that the correlation is remarkably long range--indeed, base pairs thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationary" feature of the sequence of base pairs by applying a new algorithm called Detrended Fluctuation Analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and noncoding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to all eukaryotic DNA sequences (33 301 coding and 29 453 noncoding) in the entire GenBank database. We describe a simple model to account for the presence of long-range power-law correlations which is based upon a generalization of the classic Levy walk. Finally, we describe briefly some recent work showing that the noncoding sequences have certain statistical features in common with natural languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts, and the Shannon approach to quantifying the "redundancy" of a linguistic text in terms of a measurable entropy function. We suggest that noncoding regions in plants and invertebrates may display a smaller entropy and larger redundancy than coding regions, further supporting the possibility that noncoding regions of DNA may carry biological information.

Non-NASA Center↗

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much richer learning medium for LLMs, compared to policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across six tasks, GEPA outperforms GRPO by 6% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% (e.g., +12% accuracy on AIME-2025), and demonstrates promising results as an inference-time search strategy for code optimization. We release our code at https://github.com/gepa-ai/gepa.

97 MATHEMATICS AND COMPUTING↗