Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Validating an Air Traffic Management Concept of Operation Using Statistical Modeling

Validating a concept of operation for a complex, safety-critical system (like the National Airspace System) is challenging because of the high dimensionality of the controllable parameters and the infinite number of states of the system. In this paper, we use statistical modeling techniques to explore the behavior of a conflict detection and resolution algorithm designed for the terminal airspace. These techniques predict the robustness of the system simulation to both nominal and off-nominal behaviors within the overall airspace. They also can be used to evaluate the output of the simulation against recorded airspace data. Additionally, the techniques carry with them a mathematical value of the worth of each prediction-a statistical uncertainty for any robustness estimate. Uncertainty Quantification (UQ) is the process of quantitative characterization and ultimately a reduction of uncertainties in complex systems. UQ is important for understanding the influence of uncertainties on the behavior of a system and therefore is valuable for design, analysis, and verification and validation. In this paper, we apply advanced statistical modeling methodologies and techniques on an advanced air traffic management system, namely the Terminal Tactical Separation Assured Flight Environment (T-TSAFE). We show initial results for a parameter analysis and safety boundary (envelope) detection in the high-dimensional parameter space. For our boundary analysis, we developed a new sequential approach based upon the design of computer experiments, allowing us to incorporate knowledge from domain experts into our modeling and to determine the most likely boundary shapes and its parameters. We carried out the analysis on system parameters and describe an initial approach that will allow us to include time-series inputs, such as the radar track data, into the analysis

Statistical emulation↗

Synergistic Use of Hyperspectral UV-Visible OMI and Broadband Meteorological Imager MODIS Data for a Merged Aerosol Product

The retrieval of optimal aerosol datasets by the synergistic use of hyperspectral ultraviolet(UV)–visible and broadband meteorological imager (MI) techniques was investigated. The Aura Ozone Monitoring Instrument (OMI) Level 1B (L1B) was used as a proxy for hyperspectral UV–visible instrument data to which the Geostationary Environment Monitoring Spectrometer (GEMS) aerosol algorithm was applied. Moderate-Resolution Imaging Spectroradiometer (MODIS) L1B and dark target aerosol Level 2 (L2) data were used with a broadband MI to take advantage of the consistent time gap between the MODIS and the OMI. First, the use of cloud mask information from the MI infrared (IR) channel was tested for synergy. High-spatial-resolution and IR channels of the MI helped mask cirrus and sub-pixel cloud contamination of GEMS aerosol, as clearly seen in aerosol optical depth (AOD) validation with Aerosol Robotic Network (AERONET) data. Second, dust aerosols were distinguished in the GEMS aerosol-type classification algorithm by calculating the total dust confidence index (TDCI) from MODIS L1B IR channels. Statistical analysis indicates that the Probability of Correct Detection (POCD) between the forward and inversion aerosol dust models (DS) was increased from 72% to 94% by use of the TDCI for GEMS aerosol-type classification, and updated aerosol types were then applied to the GEMS algorithm. Use of the TDCI for DS type classification in the GEMS retrieval procedure gave improved single-scattering albedo (SSA) values for absorbing fine pollution particles (BC) and DS aerosols. Aerosol layer height (ALH) retrieved from GEMS was compared with Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP) data, which provides high-resolution vertical aerosol profile information. The CALIOP ALH was calculated from total attenuated backscatter data at 1064 nm, which is identical to the definition of GEMS ALH. Application of the TDCI value reduced the median bias of GEMS ALH data slightly. The GEMS ALH bias approximates zero, especially for GEMS AOD values of>~0.4 and GEMS SSA values of<~0.95.Finally, the AOD products from the GEMS algorithm and MI were used in aerosol merging with the maximum-likelihood estimation method, based on a weighting factor derived from the standard deviation of the original AOD products. With the advantage of the UV–visible channel in retrieving aerosol properties over bright surfaces, the combined AOD products demonstrated better spatial data availability than the original AOD products, with comparable accuracy. Furthermore, pixel-level error analysis of GEMS AOD data indicates improvement through MI synergy.

aerosol↗

Comparison of soil dielectric mixing models for Soil Moisture Retrieval using SMAP Brightness Temperature over croplands in India

The accurate estimation of soil moisture (SM) using microwave remote sensing depends mostly on careful selection of retrieval parameters among which the soil dielectric mixing model is the important one. These models are often categorized into empirical, semi-empirical or volumetric based on their methodologies and input data requirements. To study in detail, the comparative performance of four dielectric mixing models -- Wang & Schmugge model, Hallikainen model, Dobson model and Mironov model were used with Soil Moisture Active Passive (SMAP) L-band brightness temperature and Single Channel Algorithm for SM retrieval over agricultural landscapes in India. The highest performance statistics combination in terms of Root Mean Square Error (RMSE), correlation coefficient (R^(2)) and percentage bias (PBIAS) against the concurrent in-situ SM measurements were calculated at the selected validation sites. The overall results indicate that the best performance was given by the Mironov model (RMSE = 0.07 cu. m/cu. m), followed by Wang & Schmugge model (RMSE = 0.08 cu. m3/cu. m), Hallikainen model (RMSE = 0.09 cu. m/cu. m), Dobson model (RMSE = 0.10 cu. m/cu. m) and original SMAP radiometer SM (RMSE = 0.12 cu. m/cu. m). Findings of this study provides important insights into application and performance of dielectric mixing models in mapping surface SM variations. This study also underlines the pivotal role of local conditions for SM retrieval which should be carefully included in the algorithms.

Swati Suman↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Measurements of Rainfall Rate, Drop Size Distribution, and Variability at Middle and Higher Latitudes: Application to the Combined DPR-GMI Algorithm

The Global Precipitation Measurement mission is a major U.S.–Japan joint mission to understand the physics of the Earth’s global precipitation as a key component of its weather, climate, and hydrological systems. The core satellite carries a dual-precipitation radar and an advanced microwave imager which provide measurements to retrieve the drop size distribution (DSD) and rain rates using a Combined Radar-Radiometer Algorithm (CORRA). Our objective is to validate key assumptions and parameterizations in CORRA and enable improved estimation of precipitation products, especially in the middle-to-higher latitudes in both hemispheres. The DSD parameters and statistical relationships between DSD parameters and radar measurements are a central part of the rainfall retrieval algorithm, which is complicated by regimes where DSD measurements are abysmally sparse (over the open ocean). In view of this, we have assembled optical disdrometer datasets gathered by research vessels, ground stations, and aircrafts to simulate radar observables and validate the scattering lookup tables used in CORRA. The joint use of all DSD datasets spans a large range of drop concentrations and characteristic drop diameters. The scaling normalization of DSDs defines an intercept parameter N(W), which normalizes the concentrations, and a scaling diameter D(m), which compresses or stretches the diameter coordinate axis. A major finding of this study is that a single relationship between N(W) and D(m), on average, unifies all datasets included, from stratocumulus to heavier rainfall regimes. A comparison with the N(W)–D(m) relation used as a constraint in versions 6 and 7 of CORRA highlights the scope for improvement of rainfall retrievals for small drops (D(m) < 1 mm) and large drops (D(m) > 2 mm). The normalized specific attenuation–reflectivity relationships used in the combined algorithm are also found to match well the equivalent relationships derived using DSDs from the three datasets, suggesting that the currently assumed lookup tables are not a major source of uncertainty in the combined algorithm rainfall estimates.

Viswanathan Bringi↗

CMB-PAInT: An inpainting tool for the cosmic microwave background

Abstract The presence of astrophysical emissions in microwave observations forces us to perform component separation to extract the Cosmic Microwave Background (CMB) signal. However, even in the most optimistic cases, there are still strongly contaminated regions, such as the Galactic plane or those with emission from extragalactic point sources, which require the use of a mask. Since many CMB analyses, especially the ones working in harmonic space, need the whole sky map, it is crucial to develop a reliable inpainting algorithm that replaces the values of the excluded pixels by others statistically compatible with the rest of the sky. This is especially important when working withQandUsky maps in order to obtainE- andB-mode maps which are free fromE-to-Bleakage. In this work we study a method based on Gaussian Constrained Realizations (GCR), that can deal with both intensity and polarization. Several tests have been performed to asses the validation of the method, including the study of the one-dimensional probability distribution function (1-PDF),E- andB-mode map reconstruction, and power spectra estimation. We have considered two scenarios for the input simulation: one case with only CMB signal and a second one including also Planck PR4 semi-realistic noise. Even if we are limited to low resolution maps, N side = 64 ifT,QandUare considered, we believe that this is a useful approach to be applied to future missions such as LiteBIRD, where the target are the largest scales.

Astronomy & Astrophysics↗

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv↗

Incorporating spatial context into statistical classification of multidimensional image data

Compound decision theory is employed to develop a general statistical model for classifying image data using spatial context. The classification algorithm developed from this model exploits the tendency of certain ground-cover classes to occur more frequently in some spatial contexts than in others. A key input to this contextural classifier is a quantitative characterization of this tendency: the context function. Several methods for estimating the context function are explored, and two complementary methods are recommended. The contextural classifier is shown to produce substantial improvements in classification accuracy compared to the accuracy produced by a non-contextural uniform-priors maximum likelihood classifier when these methods of estimating the context function are used. An approximate algorithm, which cuts computational requirements by over one-half, is presented. The search for an optimal implementation is furthered by an exploration of the relative merits of using spectral classes or information classes for classification and/or context function estimation.

Bauer, M. E.↗

Edge detection for synthetic aperture radar and other noisy images

The development is examined of a new edge detector which is shown to perform adequately in the non-Gaussian multiplicative noise environment which characterizes radar images. This edge detector operates over larger local neighborhood and is less susceptible to noise than previous edge detectors and is therefore more suitable for radar. In addition, a radar image noise model is employed for the design of this new operator. This edge detector is unique in that it is assumed that every local area belongs to either the class of local areas not containing edges or to the class of local areas containing edges. Each pixel's local neighborhood is then assigned to one of these two classes using a statistical hypothesis (a likelihood ratio) test. It is demonstrated that this algorithm is useful for detecting edges in radar images.

Frost, V. S.↗

Multisensor Arrays for Greater Reliability and Accuracy

Arrays of multiple, nominally identical sensors with sensor-output-processing electronic hardware and software are being developed in order to obtain accuracy, reliability, and lifetime greater than those of single sensors. The conceptual basis of this development lies in the statistical behavior of multiple sensors and a multisensor-array (MSA) algorithm that exploits that behavior. In addition, advances in microelectromechanical systems (MEMS) and integrated circuits are exploited. A typical sensor unit according to this concept includes multiple MEMS sensors and sensor-readout circuitry fabricated together on a single chip and packaged compactly with a microprocessor that performs several functions, including execution of the MSA algorithm. In the MSA algorithm, the readings from all the sensors in an array at a given instant of time are compared and the reliability of each sensor is quantified. This comparison of readings and quantification of reliabilities involves the calculation of the ratio between every sensor reading and every other sensor reading, plus calculation of the sum of all such ratios. Then one output reading for the given instant of time is computed as a weighted average of the readings of all the sensors. In this computation, the weight for each sensor is the aforementioned value used to quantify its reliability. In an optional variant of the MSA algorithm that can be implemented easily, a running sum of the reliability value for each sensor at previous time steps as well as at the present time step is used as the weight of the sensor in calculating the weighted average at the present time step. In this variant, the weight of a sensor that continually fails gradually decreases, so that eventually, its influence over the output reading becomes minimal: In effect, the sensor system "learns" which sensors to trust and which not to trust. The MSA algorithm incorporates a criterion for deciding whether there remain enough sensor readings that approximate each other sufficiently closely to constitute a majority for the purpose of quantifying reliability. This criterion is, simply, that if there do not exist at least three sensors having weights greater than a prescribed minimum acceptable value, then the array as a whole is deemed to have failed.

Immer, Christopher↗

Maintaining Atmospheric Mass and Water Balance Within Reanalysis

This report describes the modifications implemented into the Goddard Earth Observing System Version-5 (GEOS-5) Atmospheric Data Assimilation System (ADAS) to maintain global conservation of dry atmospheric mass as well as to preserve the model balance of globally integrated precipitation and surface evaporation during reanalysis. Section 1 begins with a review of these global quantities from four current reanalysis efforts. Section 2 introduces the modifications necessary to preserve these constraints within the atmospheric general circulation model (AGCM), the Gridpoint Statistical Interpolation (GSI) analysis procedure, and the Incremental Analysis Update (IAU) algorithm. Section 3 presents experiments quantifying the impact of the new procedure. Section 4 shows preliminary results from its use within the GMAO MERRA-2 Reanalysis project. Section 5 concludes with a summary.

IAU↗

Advances in Landslide Hazard Forecasting: Evaluation of Global and Regional Modeling Approach

A prototype global satellite-based landslide hazard algorithm has been developed to identify areas that exhibit a high potential for landslide activity by combining a calculation of landslide susceptibility with satellite-derived rainfall estimates. A recent evaluation of this algorithm framework found that while this tool represents an important first step in larger-scale landslide forecasting efforts, it requires several modifications before it can be fully realized as an operational tool. The evaluation finds that the landslide forecasting may be more feasible at a regional scale. This study draws upon a prior work's recommendations to develop a new approach for considering landslide susceptibility and forecasting at the regional scale. This case study uses a database of landslides triggered by Hurricane Mitch in 1998 over four countries in Central America: Guatemala, Honduras, EI Salvador and Nicaragua. A regional susceptibility map is calculated from satellite and surface datasets using a statistical methodology. The susceptibility map is tested with a regional rainfall intensity-duration triggering relationship and results are compared to global algorithm framework for the Hurricane Mitch event. The statistical results suggest that this regional investigation provides one plausible way to approach some of the data and resolution issues identified in the global assessment, providing more realistic landslide forecasts for this case study. Evaluation of landslide hazards for this extreme event helps to identify several potential improvements of the algorithm framework, but also highlights several remaining challenges for the algorithm assessment, transferability and performance accuracy. Evaluation challenges include representation errors from comparing susceptibility maps of different spatial resolutions, biases in event-based landslide inventory data, and limited nonlandslide event data for more comprehensive evaluation. Additional factors that may improve algorithm performance accuracy include incorporating additional triggering factors such as tectonic activity, anthropogenic impacts and soil moisture into the algorithm calculation. Despite these limitations, the methodology presented in this regional evaluation is both straightforward to calculate and easy to interpret, making results transferable between regions and allowing findings to be placed within an inter-comparison framework. The regional algorithm scenario represents an important step in advancing regional and global-scale landslide hazard assessment and forecasting.

Kirschbaum, Dalia B.↗

Virtual refrigerant charge sensing algorithm for residential CO₂ heat pumps

Natural refrigerants are increasingly adopted in next-generation heat pump systems, among which CO₂ heat pumps have attracted significant attention. However, due to their high operating pressures, the leakage risk is higher, resulting in undercharge conditions and degraded heat pump performance. Thus, developing an accurate refrigerant charge level detection technique is necessary to guarantee safe and efficient operation. Although virtual refrigerant charge (VRC) level calculation algorithms for CO₂ heat pumps exist, they typically rely on empirically selected features without a systematic selection framework, leading to multicollinearity and potential overfitting, which limit their prediction accuracy and generalizability. To address these issues, this study proposes a VRC algorithm framework with a systematic feature selection method that identifies physically meaningful and statistically significant features, and is applied using a residential CO₂ heat pump as a case study. The method is extended from previous work on conventional refrigerants to account for charge behavior in CO₂ gas coolers. The selected features include gas cooler outlet density, evaporator pressure, and superheat temperature. The results demonstrate that the proposed feature selection method significantly improves prediction accuracy compared to existing VRC approaches. A relatively small training dataset (∼30 samples) is sufficient for feature identification and model development. The developed algorithm achieves less than 3% prediction error under both undercharge and overcharge conditions, representing reductions of 46.7% and 35.3% compared to two recent reference VRC algorithms for transcritical CO₂ heat pumps reported in the literature. The proposed algorithm and feature selection method enhance leakage detection capability, facilitate the deployment of CO₂ heat pump systems, and contribute to reduced energy waste and maintenance costs.

Guo, Fangzhou [Lawrence Berkeley National Laborato↗

Statistical versus nonstatistical temperature inversion methods

Vertical temperature profiles are derived from radiation measurements by inverting the integral equation of radiative transfer. Because of the nonuniqueness of the solution, the particular temperature profile obtained depends on the numerical inversion technique used and the type of auxiliary information incorporated in the solution. The choice of an inversion algorithm depends on many factors; including the speed and size of computer, the availability of representative statistics, and the accuracy of initial data. Results are presented for a numerical study comparing two contrasting inversion methods: the statistical-matrix inversion method and the nonstatistical-iterative method. These were found to be the most applicable to the problem of determining atmospheric temperature profiles. Tradeoffs between the two methods are discussed.

Smith, W. L.↗

New graph-neural-network flavor tagger for Belle II and measurement of sin 2⁢𝜙 1 in 𝐵 0 → 𝐽/𝜓⁢𝐾$^0_ S$ decays

We present GFlaT, a new algorithm that uses a graph-neural-network to determine the flavor of neutral 𝐵 mesons produced in ϒ⁡(4⁢𝑆) decays. It improves previous algorithms by using the information from all charged final-state particles and the relations between them. We evaluate its performance using 𝐵 decays to flavor-specific hadronic final states reconstructed in a 362 fb −1 sample of electron-positron collisions collected at the ϒ⁡(4⁢𝑆) resonance with the Belle II detector at the SuperKEKB collider. We achieve an effective tagging efficiency of (37.40 ± 0.43 ± 0.36%), where the first uncertainty is statistical and the second systematic, which is 18% better than the previous Belle II algorithm. Demonstrating the algorithm, we use 𝐵 0 →𝐽/𝜓⁢𝐾$^0_ S$ decays to measure the mixing-induced and direct 𝐶⁢𝑃 violation parameters, 𝑆 = (0.724 ± 0.035 ± 0.009) and 𝐶 = (−0.035 ± 0.026 ± 0.029).

CP violation↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction↗

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗