Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

AEPF: Attention-Enabled Point Fusion for 3D Object Detection

Current state-of-the-art (SOTA) LiDAR-only detectors perform well for 3D object detection tasks, but point cloud data are typically sparse and lacks semantic information. Detailed semantic information obtained from camera images can be added with existing LiDAR-based detectors to create a robust 3D detection pipeline. With two different data types, a major challenge in developing multi-modal sensor fusion networks is to achieve effective data fusion while managing computational resources. With separate 2D and 3D feature extraction backbones, feature fusion can become more challenging as these modes generate different gradients, leading to gradient conflicts and suboptimal convergence during network optimization. To this end, we propose a 3D object detection method, Attention-Enabled Point Fusion (AEPF). AEPF uses images and voxelized point cloud data as inputs and estimates the 3D bounding boxes of object locations as outputs. An attention mechanism is introduced to an existing feature fusion strategy to improve 3D detection accuracy and two variants are proposed. These two variants, AEPF-Small and AEPF-Large, address different needs. AEPF-Small, with a lightweight attention module and fewer parameters, offers fast inference. AEPF-Large, with a more complex attention module and increased parameters, provides higher accuracy than baseline models. Experimental results on the KITTI validation set show that AEPF-Small maintains SOTA 3D detection accuracy while inferencing at higher speeds. AEPF-Large achieves mean average precision scores of 91.13, 79.06, and 76.15 for the car class’s easy, medium, and hard targets, respectively, in the KITTI validation set. Results from ablation experiments are also presented to support the choice of model architecture.

Chemistry↗

Power Analysis of an ePump Applied to the Linear Functions of an Agricultural Planter

Like many other industries, the agricultural industry has recently ex-perienced pressure to reduce vehicle emissions while improving productivity. Electric actuation is perceived as a viable solution to replace or augment hydrau-lic and mechanical actuation. However, electrification presents challenges with regards to linear functions, where hydraulic actuators have advantages in terms of compactness, tolerance to contamination and resistance to shocks. A combined electro-hydraulic actuation architecture can leverage the benefits of both electric and hydraulic actuation, while reducing the drawbacks of both approaches. This work investigates the potential of a centralized electric driven pump (ePump) system powering the pressure-controlled linear functions of an agricul-tural planter, with the goal of improving the operating point of the main supply pump. In this application, the rotary functions are hydraulically actuated, alt-hough such a solution could be applied also to electric rotary actuation. This is accomplished by setting the ePump to boost the pressure supplied by the tractor to the level required by the linear functions. An accumulator is used to stabilize the flow requirements of the linear functions, and control the pressure supplied to the actuators. Two control schemes are proposed for the regulation of the ac-cumulator pressure, one favoring an efficient operating point for the ePump, the other favoring stable steady state operation. A simulation model of the baseline system and proposed system is developed and validated using experimental data from a full-scale machine. Using the vali-dated simulation, both control architectures are then evaluated for improvement in system power consumption and dynamic requirements on the ePump to assess their effectiveness. Both systems demonstrate significant improvement in power consumption over the baseline system, with the best solution improving effi-ciency by 64%.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An explainable variational autoencoder model for three-dimensional acoustic emission source localization in hollow cylindrical structures

We introduce an explainable variational autoencoder for three-dimensional (3D) localization of acoustic emission sources in hollow cylindrical structures, with an unsupervised approach. This research capitalizes on multi-arrival waveforms generated by helical path propagation in cylindrical geometries to enable efficient two-receiver localization. By integrating the modal characteristics of Lamb modes under multi-path conditions, we demonstrate that two sets of time-of-arrival differences and peak amplitudes extracted from one receiver can serve as effective localization features. This initial approach identifies four potential source locations, highlighting the feasibility of two-receiver source localization using traditional feature extraction methods. However, direct extraction can be challenging when mode overlaps occur, complicating the localization process. To address this, our work proposes a novel waveform-based method. This method leverages the consistent dispersion characteristics within isotropic materials, where each unique combination of mode arrival times and peak amplitudes constructs a distinct waveform. This distinctiveness overcomes the ambiguities associated with mode overlaps, significantly enhancing the method’s precision and robustness. Our approach adopts a data-driven strategy for waveform-based localization using variational autoencoder (VAE). VAE discerns waveform patterns for localization, while also addressing data uncertainties. The VAE’s encoder and decoder networks capture the localization process and the source’s influence on waveform generation, respectively, guiding latent variables to segregate waveforms by source in the latent space. The design of the learning process focuses on specific localization characteristics to enhance result explainability. Localization predictions are generated by projecting test waveforms, not included in the training set, onto a trained latent space. The prediction is determined using a nearest-neighbor approach based on the closest latent representation of a source. Validation with pencil-lead-break tests on a metallic pipe confirmed our method’s effectiveness, achieving an averaged 3D localization accuracy of 0.84.

Lee, Guan-Wei↗

NMF-Based Anomaly Detection in CMS 2D Tracking Occupancy Histograms

The CMS experiment relies on Data Quality Monitoring (DQM) to ensure that recorded collision data are suitable for physics analysis. During LHC Run 3, each run contains many lumisections and tracking monitoring elements, making offline inspection challenging, especially for localized detector effects that may appear only for short periods of time. This poster presents an unsupervised machine-learning approach to identify anomalous lumisections in CMS tracking occupancy histograms using Non-Negative Matrix Factorization (NMF). The workflow uses offline CMS DQMIO tracking histograms retrieved with the CMS DIALS API and organized as two-dimensional occupancy maps for each lumisection. After selecting stable lumisections, the occupancy maps are normalized and arranged into a non-negative data matrix. The NMF model learns a compact set of basis patterns describing normal tracking occupancy. Each lumisection is then reconstructed from these learned components, and the reconstruction error is used as an anomaly score. Large residuals indicate occupancy patterns that deviate from normal detector behavior and are flagged for further inspection. This NMF-based approach provides a fast and interpretable way to flag lumisections whose tracking occupancy patterns differ from normal detector behavior. Preliminary studies show sensitivity to known tracking anomalies, and ongoing work is focused on validating the method across additional Run 3 Pixel and Strip detector issues.

Rodríguez Ramos, Iliomar [Puerto Rico U., Mayaguez↗

Maximizing Efficiency and Quality: Leveraging Automated Testing for Laboratory Commissioning

The traditional commissioning process uses sampling to select equipment for functional acceptance testing when large quantities of equipment are present. Although this approach is generally effective in identifying wide-spread issues, it has several shortcomings: it fails to evaluate equipment not included in the sample, provides only a one-time validation of equipment operation, and the standard documentation is a simple checklist of pass/fail questions. During the construction and commissioning process of the new Research and Innovation Laboratory (RAIL) in Golden, CO, the National Renewable Energy Laboratory team engaged Group14 Engineering to implement a Connected Commissioning process using fault detection and diagnostic software for automated functional acceptance testing. This presentation highlights the advantages offered by automated functional testing in this critical laboratory setting: (1) sampling 100% of BAS-connected equipment during functional testing, (2) testing results backed by data beyond the traditional pass/fail checklist, and (3) an automated test process that can be regularly executed by the building management team for ongoing commissioning throughout the life of the building. The presentation will also cover technical challenges associated with Connected Commissioning and the important conversations with key stakeholders that need to occur well before functional acceptance testing in order to successfully implement the automated testing processes.

automated testing↗

Location-Specific Microstructure Characterization Within AM Bench 2022 Nickel Alloy 718 3D Builds

Abstract The Additive Manufacturing Benchmark Test Series (AM Bench) is a broad effort to produce rigorous measurement datasets for validating AM computer simulations across the range of processing, structure, and properties, for many additive manufacturing (AM) build methods and material classes. Here, the microstructures of nickel alloy 718 AM Bench 2022 test artifacts produced using laser-based powder bed fusion (PBF-LB), in both as-built and fully heat-treated conditions, are examined. Cross sections are primarily characterized using large area scanning electron microscopy (SEM) electron backscatter diffraction (EBSD) and example analyses of the crystallographic textures are described. These data are part of a large set of in situ and ex situ measurements from both three-dimensional builds and laser tracks on bare plates. All the measurement data are available online with download links at www.nist.gov/ambench .

Levine, L. E. (ORCID:0000000334484229)↗

Inclusive cross section measurements in final states with and without protons for charged-current ν μ -Ar scattering in MicroBooNE

A detailed understanding of inclusive muon neutrino charged-current interactions on argon is crucial to the study of neutrino oscillations in current and future experiments using liquid argon time projection chambers. To that end, we report a comprehensive set of differential cross section measurements for this channel that simultaneously probe the leptonic and hadronic systems by dividing the channel into final states with and without protons. Measurements of the proton kinematics and proton multiplicity of the final state are also presented. For these measurements, we utilize data collected with the MicroBooNE detector from 6.4 × 10 20 protons on target from the Fermilab booster neutrino beam at a mean neutrino energy of approximately 0.8 GeV. We present in detail the cross section extraction procedure, including the unfolding, and model validation that uses data to model comparisons and the conditional constraint formalism to detect mismodeling that may introduce biases to extracted cross sections that are larger than their uncertainties. The validation exposes insufficiencies in the overall model, motivating the inclusion of an additional data-driven reweighting systematic to ensure the accuracy of the unfolding. The extracted results are compared to a number of event generators and their performance is discussed with a focus on the regions of phase space that indicate the greatest need for modeling improvements. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FORGE STRESS annual report

The project's goal is to combine high-fidelity numerical models and true-triaxial block fracturing tests at high temperatures to understand the relationship between in situ stress, thermal effects, wellbore orientations and hydraulic fracture patterns. The numerical models are calibrated against field data, such as well pressures and microseismic data, and employed to estimate the in situ stress at the FORGE site. Laboratory experiments investigate the complex physics driving hydraulic fracture nucleation in EGS, enhancing understanding of the role of parameters like temperature, well orientation, and stress. Additionally, they are employed to validate some of the numerical tools used in the project. The project will have a significant impact by: (1) improving the characterization of the in-situ stress field at FORGE; (2) demonstrating the use of high-fidelity modeling tools for EGS; (3) providing a unique set of high-temperature hydraulic fracturing results to identify key components for in-situ stress estimation, validate current theories, and propose new ones; (4) offering a validated set of numerical tools within an open-source simulation framework, GEOS, that will be available to any future user.

15 GEOTHERMAL ENERGY↗

Excited-state uncertainties in lattice-QCD calculations of multi-hadron systems

Excited-state effects lead to hard-to-quantify systematic uncertainties in lattice quantum chromodynamics (LQCD) spectroscopy calculations when computationally accessible imaginary times are smaller than inverse excitation gaps, as often arises for multi-hadron systems with signal-to-noise problems. Lanczos residual bounds address this by providing two-sided constraints on energies that do not require assumptions beyond Hermiticity, but often give very conservative systematic uncertainty estimates. Here, a more-constraining set of gap bounds is introduced for hadron spectroscopy. These bounds provide tighter constraints whose validity requires an explicit assumption about an energy gap. Exactly solvable lattice field theory correlators are used to test the utility of residual and gap bounds at finite and infinite statistics. Two-sided bounds and other analysis methods are then applied to a high-statistics LQCD calculation of nucleon-nucleon scattering at $m_π\sim 800$ MeV. Generalized eigenvalue problem (GEVP) and Lanczos energy estimators are compatible when applied to the same correlator data, but analyses including different interpolating operators show statistically significant inconsistencies. However, two-sided bounds from all operators are consistent. Under the assumption that the number of energy levels below $NΔ$ and $ΔΔ$ thresholds is the same as for non-interacting nucleons, gap bounds are sufficient to constrain nucleon-nucleon scattering amplitudes at phenomenologically relevant precision. Lanczos methods further reveal that energy-eigenstate estimates from previously studied asymmetric correlators have not converged over accessible imaginary times. Nevertheless, data-driven examples demonstrate why assumptions are required to draw conclusions about the natures of two-nucleon ground states at these masses.

Detmold, William [MIT, Cambridge, CTP]↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Meta-Analysis of Advanced Nuclear Reactor Cost Estimations

Supporting Data can be downloaded at: https://gain.inl.gov/content/uploads/4/2024/06/INL-RPT-24-77048-R1.xlsx Nuclear energy is a critical cornerstone of the current United States clean energy supply and may play a larger role in the future in support of a transition to a net-zero economy. The current fleet of nuclear reactors predominantly consists of large light-water reactors (LWRs), while many of the reactor designs under consideration are smaller and/or different technologies. Because these new designs have not yet been built, there is a high degree of uncertainty associated with their cost. This complicates energy-planning efforts because cost projections are not always standardized, consistent, and centralized in an easily accessible location. To help support energy planning in the US, this report provides advanced nuclear cost ranges using a transparent methodology along with other relevant information that can be used to help support decision making and energy planning. The purpose of this work was to conduct a methodical process for cost evaluation using only public information that was vetted with the end-goal to provide reference cost projections for nuclear energy. To provide a solid basis for these values, the approach and assumptions are explicitly laid out throughout the report allowing any user of the data to challenge or reconsider them. Because future US nuclear-reactor costs are still unknown due to little recent observed data, the report opted to compile a comprehensive list of bottom-up estimates and evaluate averages/trends within the data to identify reference ranges. This was deemed preferable to opining on the robustness or validity of one cost estimation versus another. To that end, the work evaluated thousands of lines of cost subaccounts from several bottom-up cost estimates. A wide variety of different reactor types captured in the data are of various sizes and technologies. Some of these reactors will be representative of advanced reactors under development while others will not. Thus, the results here are dependent on the data that are available and the accuracy of the estimates that are used. Each bottom-up estimate was reviewed to determine whether it was complete. Incomplete data sets were corrected to ensure an adequate basis of cross-comparison. The report is not without limitations and should be interpreted as an initial step to develop cost ranges for nuclear technology. Ultimately, future work can build upon the methodology with refined cost estimates to reduce uncertainty. US-based overnight capital cost (OCC) estimates were compiled from extensive data sets into ranges for both large and small reactor sizes for 2030. To project the cost declines over time, learning rates were sampled from literature sources. No SMRs were previously built; hence, learning rates based on bottom-up approaches (e.g., by quantifying the impact stemming from fabrication of different components, modular work, site construction, commissioning) were prioritized. For larger reactors, actual learning rates from deployments were used to project future costs (adjusted to account for standardization or lack thereof between designs). Other costs included are fixed and variable operations and maintenance costs. The final variables were capacity factors and ramp rates to support energy planning.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Control Room of the Future Testbed Workshop – After-Action Report

The U.S. Department of Energy’s Office of Electricity is supporting a one-year, multi-laboratory effort to define the needs and requirements for a Control Room of the Future testbed, or CROFT. The effort responds to increasing grid complexity driven by large new loads, dynamic generation resources, and the growing adoption of advanced technologies and tools, including artificial intelligence (AI) and machine learning (ML). To support safe, secure, and effective grid modernization, CROFT will focus on how emerging technologies and tools can be rigorously evaluated in realistic operational settings, with attention to human-machine interaction, cognitive load, and workforce readiness. The project team includes Argonne National Laboratory, Idaho National Laboratory, National Laboratory of the Rockies, and Pacific Northwest National Laboratory. As part of the scoping effort, the team conducted two industry-focused workshops: one at DTECH on February 5, 2026, informed by prior industry interviews, and a second on May 4, 2026, adjacent to IEEE T&D. These engagements brought together utilities, vendors, consultants, national laboratories, academia, and government stakeholders to identify and prioritize use cases, barriers, validation needs, data-sharing constraints, and near- and longer-term requirements. This feedback will directly inform CROFT’s architecture and research focus areas, ensuring the testbed is grounded in real-world operational needs and designed to evaluate emerging technologies and tools in realistic control-room environments.

artificial intelligence↗

Statistical Downscaling of Climate Models for Solar Resource Assessment

This study presents the development of statistical models to efficiently downscale future projections of solar irradiance for solar energy applications. A climate data set simulated from a Regional Climate Model (RCM) obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) is selected as input to the statistical models to create high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). Our approach builds statistical downscaling models that (1) regrid RCM data (0.22 degree and daily spatiotemporal resolution), (2) correct bias of GHI projections, (3) downscale the future GHI project from daily-scale to hourly-scale, and (4) spatially downscale to generate GHI at 8-km resolution. To calibrate and validate the statistical models, we adapt and use the National Solar Radiation Database (NSRDB). Preliminary results show that the statistical downscaling approach downscales future projections of GHI under two climate scenarios (RCP4.5 and RCP8.5) with a nBIAS of 3%, nMAE of 34% and nRMSE of 46% estimated against NSRDB for the contiguous United State. This presentation will summarize the implemented methodology and validation results as well as future extension of this research.

climate data↗

Predicting river turbidity in Pine Island Bayou using machine learning techniques coupled with variational mode decomposition

Elevated turbidity levels pose significant public health risks by facilitating the transport of harmful pollutants, including metals, organic compounds, and pathogenic microorganisms into the surface water. These conditions create serious challenges for public recreational water use and drinking water treatment, leading to economic losses and health risks. This study utilizes water monitoring data in Pine Island Bayou, Texas, and develops a Sequence-to-Sequence (S2S) model to predict turbidity using Attention-based Gated Recurrent Units with Encoder-Decoder (AT-GRU-ED) and Long Short-Term Memory (LSTM), coupled with Variational Mode Decomposition (VMD). Compared to the model without VMD, the model demonstrates satisfactory 72-hour turbidity prediction performance, achieving MAEs of 2.60 and 3.29 NTU (reductions of 53% and 58%), RMSEs of 21.08 and 31.49 NTU (reductions of 82% and 80%), and R² values of 0.96 and 0.84 on the validation and test sets, respectively. Feature importance analysis reveals that water temperature is the dominant factor influencing seasonal turbidity patterns, while real-time hourly rainfall significantly contributes to short-term variability. Turbidity typically peaks within 48 hours after rainfall events due to lagged effects from surface runoff and upstream flow. Findings suggest suspending recreational water use and water supply pumping for three days after heavy rainfall can benefit public health and improve water treatment processes. Discharges above 100 m3/s are found to accelerate sediment dilution and transport, reducing turbidity levels more quickly after the peak. In conclusion, the proposed model demonstrates reliable 72-hour turbidity prediction, supporting decision-making for water treatment plant operations and providing early warning for public recreational water use.

Deep learning↗

Measured Reaction Rate Ratios of 235 U, 238 U, 237 Np, 239 Pu Samples in the PFUNS 235 U Prompt Fission Neutron Spectrum Criticality Experiment

The Prompt Fission Uranium Neutron Spectrum experiment, an experiment to reduce uncertainties in the high energy tail of the 235 U prompt fission spectrum, achieved success by performing two separate irradiations measuring approximately 40 different IRDFF-II reactions total using over 20 different foil materials at the National Criticality Experiments Research Center in February 2024. The criticality experiment utilized a set of highly enriched uranium hemispherical shells of increasing diameters with a large void in the center where the samples were located. The focus of this work is the first of two PFUNS irradiations focused on irradiating two of each fission foils, one bare and one cadmium covered, along with metal activation foils containing reaction products with short half-lives, such as indium, iron, and gold along with nickel fluence monitors. This work focuses on presenting the initial reaction rate ratio results of the fission foils from the aforementioned first irradiation to assist in nuclear data validation of those species in a nearly pure 235 U prompt fission neutron spectrum and compares to the Lady Godiva and Flattop-25 historic experiments at the Los Alamos Critical Experiments Facility. Future work will combine fission foil and metallic activation foil reaction rate ratio results from both the first and second higher power irradiation and reaction rates for a final spectral adjustment.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Demonstration of Optimal Benchmark Selection Website and Validation of the q c Coverage Metric Using HEU-SOL-THERM-013-003 Experiment

In the work documented in this interim report, the experiment selection toolkit web site was demonstrated and q C coverage metric methodology was validated for IEU-MET-FAST-002-001, MIX-COMP-THERM 004-004, and HEU-SOL-THERM-013-003 experiments. 𝑞 𝐶 is an information-theoretic measure based on mutual information that quantifies the ability of candidate benchmark experiments to reduce the bias and uncertainty of a target criticality safety application. The metric and an accompanying open-source Python toolkit with a web-based interface were tested against a benchmark set of 425 experiments drawn from the International Criticality Safety Benchmark Evaluation Project Handbook. The interface is hosted at https://edim.covdef.com. It accepts sensitivity data files produced by the TSUNAMI-IP module of the SCALE code system and supports both (i) deterministic analysis using the ENDF/B-VII.0 covariance library and (ii) stochastic analysis based on user-supplied keff samples. Demonstrations on representative applications across a range of material composition, spectrum, and form show that q C -guided benchmark selection achieves greater uncertainty reduction with fewer experiments and yields more stable posterior bias and uncertainty estimates than traditional similarity coefficient ( c k )–based selection, while also capturing valuable low-ck experiments that one-to-one metrics overlook.

Abdel-khalik, Hany S. [Indiana Univ.-Purdue Univ. ↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗