Search NASA⌕ Search

SEARCH · Search NASA

Results for “Maintainability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Design, performance, and field operation of twisting heliostats for 3,000-sun concentration and high temperature

The idea is to combine the solar energy reflected from a field of many heliostats to obtain a single focus of high concentration and high power. We exploit a new type of “twisting” heliostat” in which as the mount is moved to track the sun through the day, the reflector shape is twisted to maintain a focused image of the sun’s disc on the target, through the day over a wide range of angles of incidence. By combining the light reflected by many twisting heliostats into a single focus, we are not only able to accomplish but also maintain very high concentration and temperature through the day, with higher efficiency than has previously been possible with conventional heliostats having a fixed shape. A circular field of twisting heliostats is used to power a single intense focus atop a central tower. The light from all the heliostats is relayed to an upward facing receiver at the focus via a central Cassegrain secondary reflector located above the receiver. In a specific design targeting 1 MW of power, a 100 m diameter field arrays 431 twisting heliostats, each with a 7 m2 reflector. The secondary reflector ,7.2 m in diameter, is located 24 m above the heliostat field, bringing sunlight with annual average power of 1 MW to a circular focus 0.8 m in diameter, at a concentration averaging 3,000 suns.

14 SOLAR ENERGY↗

Microfluidic droplets with amended culture media cultivate a greater diversity of soil microorganisms

ABSTRACT Uncultivated but abundant soil microorganisms have untapped potential for producing broad ranges of natural products, as well as for bioremediation. However, cultivating soil microorganisms while maintaining a broad microorganism diversity to enable phenotyping and functional analysis of as diverse individual isolates as possible remains challenging. In this study, we developed and tested the ability of several culture media formulations that contain defined soil metabolites or soil extracts to maintain microorganism diversity during culture. We also assessed their performance in microfluidic droplet cultivation where single-soil microorganism isolates were encapsulated and cultivated in picoliter-volume water-in-oil emulsion droplets to enable clonal growth needed for downstream functional analyses. Our results show that droplet cultivation with media supplemented by soil extract or soil metabolites enables the recovery of soil microorganisms with higher diversity (up to 1.5-fold higher richness) compared to bulk cultivation methods. Importantly, 1.7-fold more of less abundant (<1%) phyla and 11-fold more of unique genera were recovered, demonstrating the utility of this method for interrogating highly diverse soil microorganisms for broad ranges of applications. IMPORTANCE Although soil microorganisms hold a significant value in bioproduction and bioremediation, only a small fraction—less than 1%—can be cultured under specific media and cultivation conditions. This indicates that there are ample opportunities in harvesting the diverse environmental microorganisms if isolating and recovering these uncultured microorganisms are possible. This paper presents a new cultivation technique composed of isolating single-soil microorganism cell from anin situsoil microorganism community in microfluidic droplets and conducting in-droplet cultivation in media supplemented by soil extract or soil metabolites. This method enables the recovery of a broader diversity of the original microorganism community, laying the groundwork for a high-throughput phenotyping of these diverse microorganisms from their natural habitats.

Biotechnology & Applied Microbiology↗

Engineered IL-18 variants with half-life extension and improved stability for cancer immunotherapy

Background The pro-inflammatory cytokine, interleukin-18 (IL-18), plays an instrumental role in bolstering anti-tumor immunity. However, the therapeutic application of IL-18 has been limited due to its susceptibility to neutralization by IL-18 binding protein (IL-18BP), short in vivo half-life, and unfavorable physicochemical properties. Methods In order to overcome the poor drug-like properties of IL-18, we installed an artificial disulfide bond, removed the native, unpaired cysteines, and fused the stabilized cytokine to an IgG Fc domain. The stability, potency, pharmacokinetic and pharmacodynamic properties as well as efficacy of disulfide-stabilized IL-18 Fc-fusion (dsIL-18-Fc) were assessed via in vitro and in vivo studies. Results The stability and mammalian host cell production yields of dsIL-18-Fc were improved, compared to the wild-type (WT) cytokine, while maintaining its biological potency and interactions with IL-18 receptor α (IL-18Rα) and IL-18BP. Recombinant fusion of the cytokine to an IgG Fc domain provided extended half-life. Notably, despite maintaining sensitivity to IL-18BP, dsIL-18-Fc was effective at activating both T and natural killer (NK) cells, and elicited a strong anti-tumor response, either as a single agent, or in conjunction with anti-programmed cell death-ligand 1 (anti-PD-L1) therapy. Conclusions We engineered IL-18 for reinforced stability, extended half-life, and improved manufacturability. The therapeutic benefit of dsIL-18-Fc, coupled with a more favorable manufacturability profile and enhanced drug-like properties, underscores the potential utility of this engineered cytokine in cancer immunotherapy.

Immunology↗

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks↗

Stability-preserving Lossy Compression for Large-scale Partial Differential Equations

Checkpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications.

Gong, Qian [ORNL] (ORCID:0000000235704142)↗

ObstacleSense: Low-Power Neuromorphic Vision for Corridor Obstacle Awareness in Low-Level ADAS

The automotive industry’s pursuit of Level 5 autonomy is constrained by substantial perception-compute power requirements, often reaching 1, 000 + watts in full autonomy stacks. Reducing this energy burden requires rethinking perception not only at the high-end autonomy level, but also at the foundational Advanced Driver Assistance Systems (ADAS) level where low-power, safety-critical sensing can have broad impact. Neuromorphic vision provides a promising starting point: HD Dynamic Vision Sensors (DVS) can operate below 100 mW at the sensor level by reporting only asynchronous brightness changes. However, low-power sensing alone is insufficient if downstream perception reintroduces dense, energy-intensive computation. In particular, many event-driven object-detection pipelines still rely on CNN backbones, while purely spiking alternatives often trade away accuracy or ignore deployment constraints. We introduce ObstacleSense, a highly compact, CNN-free hybrid ANN–SNN framework for Level 0–1 forward-corridor obstacle awareness. Instead of performing full-scene object detection with a convolutional feature backbone, ObstacleSense targets the safety-critical question of whether the ego corridor is occupied and how far the nearest obstacle is. The architecture combines polarity-conditioned event encoding, lightweight temporal spiking dynamics, axial spatial mixing, and coarse-to-fine range estimation within a regular fixed-grid compute pattern. This design avoids the dense CNN backbone commonly used in event-based detection while maintaining a small state footprint suitable for eventual small-FPGA deployment. Before hardware mapping, we evaluate the software implementation using a model-side power proxy derived from MACs, weight and activation traffic, and spiking state updates under shared FP16 assumptions. On simulated CARLA event corpora, the deployment-oriented model achieves 0.9464 objectness F1, 0.9978 grid-level mAP, and 0.8987 m distance Mean Absolute Error at an estimated 1.92 mW proxy cost, while maintaining performance on unseen generalization test sequences.

Johnson-Scott, Zac [ORNL]↗

A Highly Effective Polysulfide-Trapping Approach for the Development of High Energy Density, Scalable Lithium-Sulfur Batteries

Lithium-sulfur (Li-S) batteries are identified as one of the most promising next-generation battery technologies owing to their high theoretical specific energy, sustainability, and affordability. However, the commercialization of Li-S batteries has been hindered by severe technical challenges, including the lithium polysulfide (PS) dissolution/shuttling effect, a major cause of fast capacity degradation over cycling. We demonstrated that, for the first time, nanolayer polymer coated high surface area porous carbons (NPCs) were coated directly on sulfur electrodes (NPC-S), which led to a high specific capacity of ∼1,600 mAh g −1 approaching the theoretical specific capacity limit in the NPC-S based Li-S batteries. The NPC-S based Li-S batteries maintained their large initial specific capacity gain compared with the Baseline-S based Li-S batteries (control) over extended cycles. A follow-on study indicated that the NPC-S approach is a necessary and critical step to boost the near-theoretical specific capacity while being stabilized over long cycles with a synergistic strategy. Our experimental and computational results suggest that NPC coated on sulfur electrodes provides not only an effective and strong PS-trapping power but also an increased redox reaction kinetics for sulfur ↔ PS’s conversions during battery charge and discharge, rendering the realization of near-theoretical discharge specific capacity in the NPC-S based Li-S batteries. The findings presented in this study may inspire a new, simple, low-cost, and commercially scalable approach, without adding any appreciable dead weight or volume to the batteries, in the effort to tackle the technical challenges facing SOA Li-S batteries.

25 ENERGY STORAGE↗

Modeling Electrodeposition in 3D Porous Architectures for Solid-State Li-Metal Batteries

Li-metal storage in three-dimensional (3D) electrodes is considered a potential dendrite-mitigation strategy. The large surface area and high porosity of these electrodes result in reduced local Li-plating current densities. The porous topology provides a scaffold for Li-deposition and stripping, maintaining both mechanical integrity and Li accessibility. The goal of this study is to understand how characteristics, such as geometry and material properties, affect the current distribution and deposition pattern. To this end, we developed a computational method to track material growth driven by electrodeposition within a complex geometry. This method ensures that the finite-element discretization remains conforming to the moving boundary while preserving an adequate mesh quality, and thus maintains solution accuracy. Using this new computational tool, we analyze the conditions under which porous anode architectures effectively expand the surface area of the charge-transfer interface, and self-regulate current density and dendrite growth.

3D electrode architectures↗

MIP-4 is Induced by Bleomycin and Stimulates Cell Migration Partially via Nir-1 Receptor

Background. CC-chemokine ligand 18 also known as MIP-4 is a chemokine with roles in inflammation and immune responses. It has been shown that MIP-4 is involved in the development of several diseases including lung fibrosis and cancer. How exactly MIP-4 is regulated and exerts its role in lung fibrosis remains unclear. Therefore, in the present study, we examined how MIP-4 is regulated and whether it acts via its potential receptor Nir-1. Materials and Methods. A549 cells were grown and maintained in DMEM : F12 (1 : 1) and supplemented with 10% FBS and 1000 U of penicillin/streptomycin and maintained as recommended by the manufacturer (ATCC). Cell migration and invasion, immunohistochemistry (IHC), Western blot, qPCR, and siRNA Nir-1 were used to determine MIP-4 regulation and its role in cell migration. Results. Cell migration was increased following stimulation of cells with recombinant (r) MIP-4 and bleomycin (BLM), whereas quenching rMIP-4 with its antibody (Ab) or addition of the Ab to BLM or H 2 O 2 diminished rMIP-4-induced cell migration. Along with cell migration, rMIP-4, BLM, and H 2 O 2 induced the formation of actin filaments dynamic structures whereas costimulation with MIP-4 Ab limited BLM- and H 2 O 2 -induced effects. MIP-4 mRNA and protein were increased by BLM and H 2 O 2 , and the addition of its Ab significantly reduced treatments effect. Experiments with siRNA investigating whether Nir-1 is a potential MIR-4 receptor indicated that the inhibition of Nir-1 decreased cell migration/invasion but did not totally inhibit rMIP-4-induced cell migration. Conclusion. Therefore, our data indicate that MIP-4 is regulated by BLM and H 2 O 2 and costimulation with its Ab limits the effects on MIP-4 and that the Nir-1 receptor partially mediates MIP-4’s effects on increased cell migration. These data also evidenced that MIP-4 is regulated by fibrotic and oxidative stimuli and that quenching MIP-4 with its Ab or therapeutically targeting the Nir-1 receptor may partially limit MIP-4 effects under fibrotic or oxidative stimulation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Chlamydomonas cells transition through distinct Fe nutrition stages within 48 h of transfer to Fe-free medium

Low iron (Fe) bioavailability can limit the biosynthesis of Fe-containing proteins, which are especially abundant in photosynthetic organisms, thus negatively affecting global primary productivity. Understanding cellular coping mechanisms under Fe limitation is therefore of great interest. We surveyed the temporal responses of Chlamydomonas (Chlamydomonas reinhardtii) cells transitioning from an Fe-rich to an Fe-free medium to document their short- and long-term adjustments. While slower growth, chlorosis and lower photosynthetic parameters are evident only after one or more days in Fe-free medium, the abundance of some transcripts, such as those for genes encoding transporters and enzymes involved in Fe assimilation, change within minutes, before changes in intracellular Fe content are noticeable, suggestive of a sensitive mechanism for sensing Fe. Promoter reporter constructs indicate a transcriptional component to this immediate primary response. With acetate provided as a source of reduced carbon, transcripts encoding respiratory components are maintained relative to transcripts encoding components of photosynthesis and tetrapyrrole biosynthesis, indicating metabolic prioritization of respiration over photosynthesis. In contrast to the loss of chlorophyll, carotenoid content is maintained under Fe limitation despite a decrease in the transcripts for carotenoid biosynthesis genes, indicating carotenoid stability. These changes occur more slowly, only after the intracellular Fe quota responds, indicating a phased response in Chlamydomonas, involving both primary and secondary responses during acclimation to poor Fe nutrition. Overall design: Sampling of Chlamydomonas CC-4532 cells cultivated photoheterotrophically (TAP) under Fe-starvation condition (0 uM Fe-EDTA). Samples were collected at multiple timepoints from biological duplicate cultures after washing in TAP medium lacking Fe. Two time courses were collected. A short time course with t=0 (pre-wash), 0 (post-wash), 5, 10, 15, 30, 60, 120, and 240 min. A long time course with t= 0, 0.5, 1, 2, 4, 8, 12, 24 and 48 hours. Please note that, for long time course, the GSE44611/PRJNA190650 samples were re-used/re-analyzed together with the short time course data: GSM1087792 C.reinhardtii_Fe_Long_0_hours SRX245324 SAMN01924672 GSM1087793 C.reinhardtii_Fe_Long_0.5_hours SRX245325 SAMN01924673 GSM1087794 C.reinhardtii_Fe_Long_1_hours SRX245326 SAMN01924674 GSM1087795 C.reinhardtii_Fe_Long_2_hours SRX245327 SAMN01924675 GSM1087796 C.reinhardtii_Fe_Long_4_hours SRX245328 SAMN01924676 GSM1087797 C.reinhardtii_Fe_Long_8_hours SRX245329 SAMN01924677 GSM1087798 C.reinhardtii_Fe_Long_12_hours SRX245330 SAMN01924678 GSM1087799 C.reinhardtii_Fe_Long_24_hours SRX245331 SAMN01924679 GSM1087800 C.reinhardtii_Fe_Long_48_hours SRX245332 SAMN01924680

Source record↗

Flexible Resource Scheduler for FAST-DERMS (FRS-FASTDERMS) v0.9

The Flexible Resource Scheduler is a hierarchical controller that manages the distributed energy resources in a distribution substation or distribution feeder to provide a firm commitment of power flow at the substation or feeder head to be scheduled in transmission-level markets as an aggregated demand resource. It is the reference controller for the FAST-DERMS Architecture, developed in tandem with the architecture under the DOE FAST-DERMS project. It is comprised of a day-ahead stochastic optimization, which schedules substation power flow and reserves, a intra-hour MPC, which generates dispatch base points for DER, and a real-time PID controller maintaining that dispatches DER to maintain the substation power around the base points. The repository also includes a representative aggregator controller, and all of the necessary components to run a simulation using PNNL's GridAPPS-D software with the controller.

MacDonald, Jason [Lawrence Berkeley National Labor↗

HP-FLEX MPC v0.1.0

HP-FLEX MPC is control software developed by Lawrence Berkeley National Laboratory with support from the California Energy Commission (CEC) through EPIC-19-301. HP-FLEX aims to provide load flexibility for heat pumps (HPs) in response to dynamic grid signals (including Time-of-Use, Dynamic Pricing, and Critical Peak Pricing) while maintaining thermostat temperatures within user-specified bounds. The software includes a system-identification module, which models the dynamics of the building envelope with thermostat data, and a control module based on a model predictive controller (MPC) to make optimal decisions. HP-FLEX receives forecasts of outdoor air temperature, solar irradiation, and internal gain (if available), as well as trajectories of energy price, temperature lower and upper bounds over a prediction horizon. It then optimizes heating and cooling capacities to minimize energy cost and peak power (with a user-defined weight on peak power) over the prediction horizon, while maintaining room air temperature within the temperature constraints, and outputs the optimal thermostat setpoints.

Kim, Donghun↗

Divertor Plasma Detachment Control Neural Network

DivControlNN is a state-of-the-art software tool that leverages advanced machine learning techniques to predict and control divertor plasma behavior in fusion reactors. Plasma, a highly energetic and electrically charged gas, requires meticulous management to protect reactor components and maintain optimal energy production. Conventional simulation methods, although extremely detailed, typically demand extensive computational time-making them unsuitable for real-time control scenarios. DivControlNN addresses this challenge by learning from tens of thousands of high-fidelity simulations, thereby creating a rapid surrogate model that can deliver near-instantaneous predictions. At the core of its functionality is a sophisticated technique known as latent space mapping, which condenses complex, high-dimensional plasma data into a compact, lower-dimensional representation. This streamlined representation enables the system to quickly forecast essential plasma properties and determine the precise conditions required for effective detachment. Detachment is a crucial process in which the plasma is cooled before reaching the divertor plates, thereby reducing heat loads and mitigating material erosion. In recent experiments conducted on the KSTAR tokamak in South Korea, DivControlNN successfully guided the detachment process without any fine-tuning-even when applied to a new tungsten divertor configuration. By achieving a computational speed-up of over one hundred million times compared to traditional simulation methods while maintaining low prediction errors, DivControlNN stands to significantly enhance real-time control and diagnostic capabilities in future fusion reactors. This breakthrough paves the way for safer, more reliable reactor operation and represents a major advancement toward realizing fusion energy as a practical, sustainable, and clean power source.

Xu, Xueqiao [Lawrence Livermore National Laborator↗

SpectraCodec: A Hilbert curve-based method for encoding metadata in mass spectra for machine learning applications (SpectraCodec) v1

Machine learning approaches to mass spectrometry (MS) data analysis require structured metadata for optimal performance. However, current MS file formats necessitate external metadata sources, creating integration challenges that impede analytical workflows. Here, we present a novel approach for encoding metadata directly within mzML files using one-hot encoding of ASCII characters mapped via Hilbert space-filling curves. This strategy embeds metadata in the first spectrum's m/z-intensity space, ensuring persistence with the primary data, eliminating the need for external metadata files, and maintaining compatibility with existing MS software. We demonstrate that the Hilbert curve mapping efficiently utilizes the two-dimensional spectral space while maintaining robust data recovery. This method offers a practical solution for machine learning applications in mass spectrometry by ensuring metadata and spectral data remain unified through all stages of analysis.

Bowen, Benjamin [Lawrence Berkeley National Labora↗

AIF for Vis (Active Inference for simulating human interpretation of data visualization) [SWR-26-084]

AIF for Vis contains the Active Inference models and analysis scripts used to study a simple visualization-interpretation task: estimating the average value of two bars in a bar chart. The work is a proof of concept for translating hypothesized cognitive strategies into executable, inspectable process models. We implement two idealized strategies inspired by dual-process accounts of visualization-aided decision making: *Fast model: a compressed, heuristic strategy that estimates the visual midpoint of the two bars and maintains a single belief over their average. *Slow model: a sequential, analytic strategy that estimates the two bar heights separately and maintains them in working memory before computing an average. Both models use a common Active-Inference-inspired framework for sequential perception, belief updating, action selection, and reporting. Their different internal representations produce distinct predicted vulnerabilities: *the Fast model is more susceptible to tick-salience bias; *the Slow model is more susceptible to working-memory decay. The repository includes the model implementations, scripts used for the experiments reported in the paper, precomputed trial-level results, and plotting scripts.

Goldwyn, Harrison [National Laboratory of the Rock↗

Nonlinear Ensemble Filtering with Diffusion Models: Application to the Surface Quasigeostrophic Dynamics

The intersection between classical data assimilation methods and novel machine learning techniques has attracted significant interest in recent years. Here, we explore another promising solution in which diffusion models are used to formulate a robust nonlinear ensemble filter for sequential data assimilation. Unlike standard machine learning methods, the proposed ensemble score filter (EnSF) is completely training free and can efficiently generate a set of analysis ensemble members. Here, in this study, we apply the EnSF to a surface quasigeostrophic model and compare its performance against the popular local ensemble transform Kalman filter (LETKF), which makes Gaussian assumptions in the analysis step. Numerical tests demonstrate that EnSF maintains stable performance in the absence of localization and for a variety of experimental settings. We find that while LETKF maintains optimal performance in the case of linear observations of the entire state and a perfect model, EnSF shows improvements over LETKF when nonlinear observations are assimilated and the system is subject to unexpected model errors. A spectral decomposition of the analysis results in this nonlinear observation regime shows that the largest improvements over LETKF occur at large scales (small wavenumbers), where LETKF lacks sufficient ensemble spread. Overall, this initial application of EnSF to a geophysical model of intermediate complexity motivates further development of the algorithm for more realistic problems.

Artificial intelligence↗

ORCHA: A performance portability system for extreme heterogeneity

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Lee, Youngjun↗

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST↗