Search NASA⌕ Search

SEARCH · Search NASA

Results for “prioritization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Electrolyte and Cutoff Potential Effects on Cycle Life of Li4Ti5O12/LiNi0.9Mn0.1O2 Batteries for Behind-the-Meter Storage Applications

Behind-the-Meter Storage (BTMS) is a stationary battery energy storage system that is connected to the electrical distribution system on the customer's side of the utility's service meter. BTMS systems are used to store electrical energy from the grid as well as inconstant, renewable energy, such as local solar and wind generation. A successful BTMS system will allow the customer to pair their energy generation and storage to optimize electrical consumption from the grid, improving reliability and minimizing cost. For BTMS applications, batteries must be designed and optimized with different set of criteria from other leading segments of the Li-ion battery market, like transportation, due the system being stationary and proximal to the residential or commercial building it's benefitting. BTMS applications prioritize safety, cost (low/no-critical materials), reliability (20-year calendar life), and durability (10,000 cycle life), while having the ability to (minimally) compromise energy density and rate capability. Lithium titanate (Li4Ti5O12-, LTO) is a promising anode candidate for BTMS applications due to its high safety and capacity retention, while maintaining a reasonable 160 mAhg-1 reversable capacity and composition of relatively abundant materials. (1) Specifically, LTO has a high working voltage which helps to prevent Li dendrite formation, improving safety. Furthermore, LTO also has negligible lithiation-based volume change, leading to less mechanical pulverization, or loss of active material, upon cycling. For the cathode, materials with little or no Co are of high interest due to the high cost and low abundance of Co. LiMn2O4 (LMO) has been paired with LTO for BTMS applications in the past due to its safety, low cost (abundancy), and reasonably high operating voltage. (2-4) However, the low capacity of LMO limits energy density and specific energy. While not the highest priority for BTMS applications, increasing energy density will enable deployment in space constrained BTMS applications and decrease total cost. LiNi0.9Mn0.1O2 (LN-MO) is a recently developed material with promise due to its high operating voltage and relatively low price. (5) However, Ni-rich layered oxides, including LNMO, tend to struggle with capacity retention during high-voltage cycling due to mechanical pulverization, irreversible phase transitions, and unstable solid-electrolyte interphase. The study presented here focuses on building an understanding of how electrolyte solvent and varied cutoff potentials will impact the cycle life of LTO/LN-MO cells. Specifically, a comparison is provided between ethylene carbonate (EC), ethyl methyl carbonate (EMC), fluoroethylene carbonate (FEC), and Gen2 electrolyte solvents with 1M Lithium hexafluorophosphate (LiPF6) salt, cycling to two upper termination potentials, 2.6V and 2.7V. Electrochemical testing and diagnostics (e.g., differential capacity analysis, area specific impedance, constant voltage hold, and rate capability) and post-mortem characterization will be used to understand the aging behavior and failure mechanisms of the 8 cell combinations (four electrolytes and two voltage cutoffs). Cells with FEC electrolyte showed a lower initial capacity compared to cells with Gen2, EMC, and EC cycling at both voltages; however, the cells with FEC showed consistent trends in capacity retention with 2.6V and 2.7V termination potentials, while the cells with the other electrolytes showed much higher rates of capacity loss when cycling to the higher voltage. These results indicate that FEC may play a role in improving durability of high-voltage, Ni-rich electrode systems for use in high-cycle applications, such as BTMS.

electrolyte↗

Chlamydomonas cells transition through distinct Fe nutrition stages within 48 h of transfer to Fe-free medium

Low iron (Fe) bioavailability can limit the biosynthesis of Fe-containing proteins, which are especially abundant in photosynthetic organisms, thus negatively affecting global primary productivity. Understanding cellular coping mechanisms under Fe limitation is therefore of great interest. We surveyed the temporal responses of Chlamydomonas (Chlamydomonas reinhardtii) cells transitioning from an Fe-rich to an Fe-free medium to document their short- and long-term adjustments. While slower growth, chlorosis and lower photosynthetic parameters are evident only after one or more days in Fe-free medium, the abundance of some transcripts, such as those for genes encoding transporters and enzymes involved in Fe assimilation, change within minutes, before changes in intracellular Fe content are noticeable, suggestive of a sensitive mechanism for sensing Fe. Promoter reporter constructs indicate a transcriptional component to this immediate primary response. With acetate provided as a source of reduced carbon, transcripts encoding respiratory components are maintained relative to transcripts encoding components of photosynthesis and tetrapyrrole biosynthesis, indicating metabolic prioritization of respiration over photosynthesis. In contrast to the loss of chlorophyll, carotenoid content is maintained under Fe limitation despite a decrease in the transcripts for carotenoid biosynthesis genes, indicating carotenoid stability. These changes occur more slowly, only after the intracellular Fe quota responds, indicating a phased response in Chlamydomonas, involving both primary and secondary responses during acclimation to poor Fe nutrition. Overall design: Sampling of Chlamydomonas CC-4532 cells cultivated photoheterotrophically (TAP) under Fe-starvation condition (0 uM Fe-EDTA). Samples were collected at multiple timepoints from biological duplicate cultures after washing in TAP medium lacking Fe. Two time courses were collected. A short time course with t=0 (pre-wash), 0 (post-wash), 5, 10, 15, 30, 60, 120, and 240 min. A long time course with t= 0, 0.5, 1, 2, 4, 8, 12, 24 and 48 hours. Please note that, for long time course, the GSE44611/PRJNA190650 samples were re-used/re-analyzed together with the short time course data: GSM1087792 C.reinhardtii_Fe_Long_0_hours SRX245324 SAMN01924672 GSM1087793 C.reinhardtii_Fe_Long_0.5_hours SRX245325 SAMN01924673 GSM1087794 C.reinhardtii_Fe_Long_1_hours SRX245326 SAMN01924674 GSM1087795 C.reinhardtii_Fe_Long_2_hours SRX245327 SAMN01924675 GSM1087796 C.reinhardtii_Fe_Long_4_hours SRX245328 SAMN01924676 GSM1087797 C.reinhardtii_Fe_Long_8_hours SRX245329 SAMN01924677 GSM1087798 C.reinhardtii_Fe_Long_12_hours SRX245330 SAMN01924678 GSM1087799 C.reinhardtii_Fe_Long_24_hours SRX245331 SAMN01924679 GSM1087800 C.reinhardtii_Fe_Long_48_hours SRX245332 SAMN01924680

Source record↗

Cyber-informed Engineering Microgrid Analysis Tool

The Cyber-Informed Engineering Microgrid Analysis Tool (CIEMAT) leverages the Department of Energy’s Cyber-Informed Engineering to prompt engineering designers and operators through an analysis of the critical functions to be supported by a microgrid installations, the criticality of those functions, the impacts of denial, disruption or misuse of those functions on the microgrid and dependent functions, and the mitigations which could best prevent impacts to those functions resulting from cyber attack. Through use of this tool, microgrid designers and operators can quickly identify appropriate engineering mitigations to limit impacts from cyber attack and functions where engineering and operational staff can prioritize and guide the application of cybersecurity protections to best support the resiliency of the system.

Wright, VirginiaL [Idaho National Laboratory (INL)↗

Cyber-informed Engineering Battery Analysis Tool

The Cyber-Informed Engineering Battery Analysis Tool (CIEBAT) leverages the Department of Energy’s Cyber-Informed Engineering to prompt engineering designers and operators through an analysis of the critical functions to be supported by a BESS installation, the criticality of those functions, the impacts of denial, disruption or misuse of those functions within the BESS system, and the mitigations which could best prevent impacts to those functions resulting from cyber attack. Through use of this tool, BESS designers and operators can quickly identify appropriate engineering mitigation opportunities to limit impacts from cyber attack and functions where engineering and operational staff can prioritize and guide the application of cybersecurity protections to best support the resiliency of the system.

Lampe, BenjaminR [Idaho National Laboratory (INL),↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Host Onboarding Tool (HObT) v1.0.0

The Host OnBoarding Tool (Hobt) is a publicly accessible, web-based software designed to organize and share information about microbial hosts under development at the Agile BioFoundry (ABF). It streamlines the assessment, tracking, and sharing of information related to microbial host development and provides a centralized platform where users can rapidly evaluate hosts' readiness for various bio processes. HObT leverages the Tier System, a standardized host development framework that organizes and assesses microbial hosts based on their readiness for biomanufacturing. Each tier outlines key targets—including genetic tools, growth conditions, omics data, and predictive models—needed to transform new or emerging microbes into established production platforms. By applying clear criteria for advancement, the Tier System helps users quickly evaluate each organism's current development status, identify gaps in available knowledge or tools, and prioritize future strain improvement efforts. Through its user-friendly interface, HObT encourages contributions of new data and insights from researchers, fostering collaboration and accelerating host development. By providing structured guidance for microbial strain advancement, HObT and the Tier System support more systematic, rapid, and cost-effective development of non-traditional microbial hosts, ultimately enhancing the efficiency and impact of biomanufacturing research and applications.

Plahar, Hector [Lawrence Berkeley National Laborat↗

Cyber Informed Engineering Cie Analysis Tool

Main Benefits: • Collaborate on assessment via the web and access and share assessments on your mobile device. • Helps you maximize your cybersecurity investment and resources • Saves you significant time and money by eliminating the requirement to research each government and industry standard in order to understand your cybersecurity posture • Contains easy to follow, step by step instructions to guide you through the process of identifying the cybersecurity posture of your organization • Provides a place to begin with cybersecurity improvement and a way to prioritize your tasks and budgets. • Covers all major cyber relevant topic areas for a comprehensive assessment of your organization’s cybersecurity posture. • Dives deep into the details of each topic area. • Contributes to the organization's risk management and decision-making process • Highlights vulnerabilities and gaps in your organization's IT and control systems. • Raises awareness and facilitates discussion on cybersecurity within your organization • Educates the controls system community on cyber security.

Hansen, Barry [Idaho National Laboratory (INL), Id↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

ClusterWeave

ClusterWeave is a workflow for biosynthetic target discovery and prioritization. It assembles annotation, BGC detection, BiG-SCAPE family context, shortlist generation, and clinker-ready panel staging into one reproducible workflow.

Martin, StantonL↗

Neural Networks to Find the Optimal Forcing for Offsetting the Anthropogenic Climate Change Effects

Abstract Of great relevance to climate engineering is the systematic relationship between the radiative forcing to the climate system and the response of the system, a relationship often represented by the linear response function (LRF) of the system. However, estimating the LRF often becomes an ill-posed inverse problem due to high-dimensionality and nonunique relationships between the forcing and response. Recent advances in machine learning make it possible to address the ill-posed inverse problem through regularization and sparse system fitting. Here, we develop a convolutional neural network (CNN) for regularized inversion. The CNN is trained using the surface temperature responses from a set of Green’s function perturbation experiments as imagery input data together with data sample densification. The resulting CNN model can infer the forcing pattern responsible for the temperature response from out-of-sample forcing scenarios. This promising proof of concept suggests a possible strategy for estimating the optimal forcing to negate certain undesirable effects of climate change. The limited success of this effort underscores the challenges of solving an inverse problem for a climate system with inherent nonlinearity. Significance Statement Predicting the climate response for a given climate forcing is a direct problem, while inferring the forcing for a given desired climate response is often an inverse, ill-posed, problem, posing a new challenge to the climate community. This study makes the first attempt to infer the radiative forcing for a given target pattern of global surface temperature response using a deep learning approach. The resulting deeply trained convolutional neural network inversion model shows promise in capturing the forcing pattern corresponding to a given surface temperature response, with a significant implication on the design of an optimal solar radiation management strategy for curbing global warming. This study also highlights the technical challenges that future research should prioritize in seeking feasible solutions to the inverse climate problem.

Ren, Huiying↗

Substitution or Shared Utilization? Intrahousehold Vehicle Use in Mixed-Powertrain Households

While previous research has focused heavily on understanding the factors deriving alternative fuel vehicle adoption rates, there remains a significant gap in understanding how households distribute mileage across different powertrains. This study utilizes data from the 2022 Next Generation National Household Travel Survey to investigate vehicle miles traveled within a sample of 150 plug-in electric vehicle (PEV)-owning households (in which at least one battery electric vehicle is present), characterizing how different powertrains are integrated into daily mobility. Leveraging a Seemingly Unrelated Regression (SUR) framework the study jointly models the utilization of PEVs, hybrid electric vehicles (HEV), and internal combustion engine vehicles (ICEVs) while accounting for household-level substitution effects. The results provide evidence of an asymmetric substitution effect. In households with mixed-powertrain configurations, the ICEV captures a substantially higher share of household miles (compared with the PEV), acting as a utility sponge. Conversely, the model identifies specific socioeconomic and geographic cohorts that prioritize PEV as the primary household workhorse, indicating a systematic sorting effect. Although the sample size limits broader generalizability, these findings suggest that PEVs are used for frequent, specific routine-intensive roles, whereas the ICEV remains a specialized utility vehicle. These insights highlight distinct intrahousehold vehicle use behaviors that are often obscured by aggregate fleetwide statistics.

25 ENERGY STORAGE↗

Mitigative Strategies for Recovering From Large Language Model Trust Violations

In this study, we investigated strategies to address trust issues arising from errors in large language models (LLMs). The study examined the impact of confidence scores, system capability explanations, and user feedback on trust restoration post-error. 68 participants viewed the responses of an LLM to 20 general trivia questions, with an error introduced on the third trial. Each participant was presented with one mitigation strategy. Participants rated their overall trust in the model and the reliability of the answer. Results showed an immediate drop in trust after the error; however, there were no differences across the three strategies in trust recovery. All conditions had a logarithmic trend in trust recovery following error. Differences in overall trust were predicted by perceived reliability of the answer, suggesting that participants were evaluating results critically and using that to inform their trust in the model. Qualitative data supported this finding; participants expressed lasting distrust despite the LLM’s later accuracy. Results showcase the need to prioritize accuracy in LLM deployment, because early errors may irrevocably damage user trust calibration and later adoption.

97 MATHEMATICS AND COMPUTING↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of mechanisms driving heterogeneous void growth in ductile aluminum

Void growth plays a central role in ductile fracture, yet the specific mechanisms that control this remain obscure. Classical models, such as those proposed by Rice and Tracey in 1969, are able to capture average rates of void growth, but cannot capture the heterogeneity of individual void growth. Building on recent work, the present study employs laboratory-based diffraction contrast tomography and in-situ x-ray computed tomography to investigate the effect of grain structure and other microstructural factors on void growth in an Al-2219 alloy. Crystal plasticity finite element (CP-FE) modeling is used alongside experimental data to evaluate the contributions of local mechanical states, grain orientation, grain size, and neighboring microstructural features. No strong linear relationships are found with any of the considered descriptors and void growth rate. Potential complex nonlinear relationships are explored with the use of a random forest regression model, which identifies initial void volume, void aspect ratio, local normal stress state, local shear stress state, and local equivalent plastic strain (EQPS) as features that most improve void growth rate predictions. The combination of these analyses suggests that these features should be prioritized to improve models of void growth.

Diffraction contrast tomography (DCT)↗

Approach for energy efficient building design during early phase of design process

Energy consumption in the building sector is about 40% of total energy consumed globally and is trending upwards, along with its contribution to greenhouse gas (GHG) emissions. Given the adverse impacts of GHG emissions, it is crucial to integrate energy efficiency into building designs. The most significant opportunities for enhancing energy performance are present during the initial phases of building design, when there is less impact of other design constraints. Various tools exist for simulating different design options and providing feedback in terms of energy consumption and comfort parameters. These simulation outputs must then be analyzed to derive design solutions. This paper presents an innovative approach that utilizes user input parameters, processes them through cloud computing, and outputs easily understandable strategies for energy-efficient building design. The methodology employs Asynchronous Distributed Task Queues (DTQ) - a more scalable and reliable alternative to conventional speedup techniques-for conducting parametric energy simulations in the cloud. The goal of this approach is to assist design teams in identifying, visualizing, and prioritizing energy-saving design strategies from a range of possible solutions for each project. Furthermore, a tool ‘eDOT’ has been developed utilizing the discussed methodology. Unlike existing tools, eDOT leverages artificial intelligence to dynamically generate and provide design strategies during the early phases of design process. By simplifying the simulation process, eDOT enables design teams to make informed, data-driven decisions without needing to interpret complex simulation outputs. A case study simulated for two locations is provided in this paper to demonstrate the effectiveness of eDOT, further underscoring its practical impact on energy-efficient building design.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Enabling Real-Time Communication in Multi-Agent Systems: A Graph Neural Network Based Approach

Global connectivity enables effective coordination in Multi-Agent Systems (MAS). Solving these connection problems under hardware constraints is an NP-hard non-Euclidean Degree Constrained Minimum Spanning Tree (DCMST) problem. Prior MAS controllers coordinate team movement for task completion and collision avoidance; some considering Line-of-Sight (LOS) maintenance but prioritizing flexibility over guarantees. Evolutionary Algorithms (EA) have been shown to find good solutions for DCMST, but their performance degrades with larger populations required to support a large MAS. We present a method based on edge graph attention networks, trained offline to reduce online computation times. Empirical comparisons with greedy polynomial-time solvers and EA show that our method leverages latent graph information to consistently find constraint-satisfying solutions in less time.

connectivity maintenance↗

Data for "Land-based Resources for Engineered Carbon Dioxide Removal in the United States Exceed the Expected Needs"

Gigatonne-scale atmospheric carbon dioxide removal (CDR), alongside deep emission cuts, is critical to stabilizing the climate. However, some of the most scalable CDR technologies are also the most land intensive. Here, we examine whether adequate land resources exist in the contiguous United States to meet CDR targets when prioritizing grid emissions reduction, food production, and the protection of sensitive ecosystems. We focus on biomass carbon removal and storage (BiCRS) and direct air capture and storage (DACS) and show that suitable lands exceed the expected needs: 37.6 million hectares of land are available for BiCRS, resulting in 0.26 GtCO2 of CDR/year, and 34 million hectares are suitable for wind- and solar-powered DACS, resulting in 4.8 GtCO2 of CDR/year if facilities are co-located with geologic CO2 storage. We identify biomass and energy supply hotspots to meet CDR targets while ensuring land protection and minimizing land competition.

carbon↗

Learning epistatic polygenic phenotypes with Boolean interactions

Detecting epistatic drivers of human phenotypes is a considerable challenge. Traditional approaches use regression to sequentially test multiplicative interaction terms involving pairs of genetic variants. For higher-order interactions and genome-wide large-scale data, this strategy is computationally intractable. Moreover, multiplicative terms used in regression modeling may not capture the form of biological interactions. Building on the Predictability, Computability, Stability (PCS) framework, we introduce the epiTree pipeline to extract higher-order interactions from genomic data using tree-based models. The epiTree pipeline first selects a set of variants derived from tissue-specific estimates of gene expression. Next, it uses iterative random forests (iRF) to search training data for candidate Boolean interactions (pairwise and higher-order). We derive significance tests for interactions, based on a stabilized likelihood ratio test, by simulating Boolean tree-structured null (no epistasis) and alternative (epistasis) distributions on hold-out test data. Finally, our pipeline computes PCS epistasis p-values that probabilisticly quantify improvement in prediction accuracy via bootstrap sampling on the test set. We validate the epiTree pipeline in two case studies using data from the UK Biobank: predicting red hair and multiple sclerosis (MS). In the case of predicting red hair, epiTree recovers known epistatic interactions surrounding MC1R and novel interactions, representing non-linearities not captured by logistic regression models. In the case of predicting MS, a more complex phenotype than red hair, epiTree rankings prioritize novel interactions surrounding HLA-DRB1 , a variant previously associated with MS in several populations. Taken together, these results highlight the potential for epiTree rankings to help reduce the design space for follow up experiments.

59 BASIC BIOLOGICAL SCIENCES↗