Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Tropopause Laminar Cirrus and Its Role in the Lower Stratosphere Total Water Budget

Laminar cirrus are thin, extensive, isolated layers of ice clouds frequently observed in the tropical tropopause layer. Widespread laminar cirrus significantly affects tropical tropopause layer total water and thermal budget. In this study, we extract laminar cirrus from the Cloud‐Aerosol Lidar with Orthogonal Polarization Level 1 attenuated total backscatter images for January 2009, in order to characterize statistical properties of laminar cirrus cloud length, base, thickness, optical depth, and layer partial ice water path. These characteristics are used to develop an algorithm identifying laminar cirrus automatically from the Cloud‐Aerosol Lidar with Orthogonal Polarization Level 2 layer product for 2008–2017. The nearly 10‐year records reveal that tropopause laminar cirrus occurrence (30–40% of total cirrus) is strongly anticorrelated with the tropopause temperatures in that colder tropopause in frequent (super)saturation during boreal winter favors in situ formation of clouds. Interannually, anomalously warmer troposphere temperature (ΔT), easterly shear of the quasi‐biennial oscillation, and stronger upwelling branch of the Brewer‐Dobson circulation enhance laminar cirrus formation via cooling of the tropopause. The tropopause laminar cirrus carries ~0.05 mg/m(exp 3) (~0.5 ppmv) of ice water content during boreal winter and <0.01 mg/m(exp 3) during summer, which is anticorrelated with the seasonal variations of water vapor (H2O) observed by the Microwave Lime Sounder, indicating a temperature‐regulated partition between vapor and ice. Interannually, in cirrus‐rich region 1 ppmv decrease in H2O corresponds to 0.2–0.3 ppmv increase in ice water content. Frequently situated in (super)saturated air, tropopause laminar cirrus are likely to survive multiple lifecycles of the sublimation‐deposition processes, and may contribute up to 10% to the total water budget in the lower stratosphere. Satellites constantly observe thin, isolated, extensive layer of cirrus around the tropopause. The so‐called “laminar” cirrus occurrence and their ice amount are strongly regulated by temperature, such that colder temperatures favor more frequent (super)saturation, which results in more frequent laminar cirrus with more ice amount and therefore less water vapor. In this study we analyze laminar cirrus and water vapor from the Cloud‐Aerosol Lidar with Orthogonal Polarization and Microwave Lime Sounder observations, and hypothesize that laminar cirrus could act as an important transient water storage and contribute to the total water budget in the lower stratosphere.

Tao Wang↗

Statistical Analysis of Aquarius Radiometer Radio Frequency Interference (RFI) and Implications on RFI Detection and Mitigation

The Aquarius/SAC-D mission operated between August 2011 and June 2015 with the main goal of providing global estimates of sea surface salinity (SSS). It comprised both active and passive microwave sensors operating at L-band to observe the same surface area almost simultaneously. Measurements from both instruments underwent subsequent filtering to mitigate the effect of Radio Frequency Interference (RFI). This report describes the analysis of statistics of RFI in samples acquired by the Aquarius radiometers, and its results could be used to improve the performance of the interference detection algorithm.

De Matthaeis, Paolo↗

An Empirical State Error Covariance Matrix Orbit Determination Example

State estimation techniques serve effectively to provide mean state estimates. However, the state error covariance matrices provided as part of these techniques suffer from some degree of lack of confidence in their ability to adequately describe the uncertainty in the estimated states. A specific problem with the traditional form of state error covariance matrices is that they represent only a mapping of the assumed observation error characteristics into the state space. Any errors that arise from other sources (environment modeling, precision, etc.) are not directly represented in a traditional, theoretical state error covariance matrix. First, consider that an actual observation contains only measurement error and that an estimated observation contains all other errors, known and unknown. Then it follows that a measurement residual (the difference between expected and observed measurements) contains all errors for that measurement. Therefore, a direct and appropriate inclusion of the actual measurement residuals in the state error covariance matrix of the estimate will result in an empirical state error covariance matrix. This empirical state error covariance matrix will fully include all of the errors in the state estimate. The empirical error covariance matrix is determined from a literal reinterpretation of the equations involved in the weighted least squares estimation algorithm. It is a formally correct, empirical state error covariance matrix obtained through use of the average form of the weighted measurement residual variance performance index rather than the usual total weighted residual form. Based on its formulation, this matrix will contain the total uncertainty in the state estimate, regardless as to the source of the uncertainty and whether the source is anticipated or not. It is expected that the empirical error covariance matrix will give a better, statistical representation of the state error in poorly modeled systems or when sensor performance is suspect. In its most straight forward form, the technique only requires supplemental calculations to be added to existing batch estimation algorithms. In the current problem being studied a truth model making use of gravity with spherical, J2 and J4 terms plus a standard exponential type atmosphere with simple diurnal and random walk components is used. The ability of the empirical state error covariance matrix to account for errors is investigated under four scenarios during orbit estimation. These scenarios are: exact modeling under known measurement errors, exact modeling under corrupted measurement errors, inexact modeling under known measurement errors, and inexact modeling under corrupted measurement errors. For this problem a simple analog of a distributed space surveillance network is used. The sensors in this network make only range measurements and with simple normally distributed measurement errors. The sensors are assumed to have full horizon to horizon viewing at any azimuth. For definiteness, an orbit at the approximate altitude and inclination of the International Space Station is used for the study. The comparison analyses of the data involve only total vectors. No investigation of specific orbital elements is undertaken. The total vector analyses will look at the chisquare values of the error in the difference between the estimated state and the true modeled state using both the empirical and theoretical error covariance matrices for each of scenario.

Frisbee, Joseph H., Jr.↗

Mimas: Preliminary Evidence For Amorphous Water Ice from VIMS

We have conducted a statistical clustering analysis (1,2) on a mosaic of VIMS data cubes obtained on February 13, 2010, for Saturn s satellite Mimas. Seven VIMS cubes were geometrically projected and re-sampled to a common spatial resolution. The clustering technique consists of a partitioning algorithm coupled to a criterion that prevents sub-optimal solutions and tests for the influence of random noise in the measurements. The clustering technique is agnostic about the meaning of the clusters, and scientific interpretation requires their a posteriori evaluation. The preliminary results yielded five clusters, demonstrating that spectral variability across Mimas surface is statistically significant. The ratios of the means calculated for each of the clusters show structure within the 1.6- micron water ice band, as well as the shape and the central wavelength of the strong ice band at 2 micron, that map spatially in patterns apparently related to the topography of Mimas, in particular certain regions in and around Herschel crater. The mean spectra of the five clusters, show similarities with laboratory spectra of amorphous and crystalline H2O ice (3) that are suggestive of the presence of an amorphous ice component in certain regions of Mimas, notably on the central peak of Herschel, on the crater floor, and in faults surrounding the crater. This may represent a mixture of both ice phases, or perhaps a layer of amorphous ice on a base of crystalline ice. Another possible occurrence of amorphous ice appears southwest of Herschel, close to the south pole.

Cruikshank, Dale P.↗

Fast Solution in Sparse LDA for Binary Classification

An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.

Moghaddam, Baback↗

The Collection 6 'dark-target' MODIS Aerosol Products

Aerosol retrieval algorithms are applied to Moderate resolution Imaging Spectroradiometer (MODIS) sensors on both Terra and Aqua, creating two streams of decade-plus aerosol information. Products of aerosol optical depth (AOD) and aerosol size are used for many applications, but the primary concern is that these global products are comprehensive and consistent enough for use in climate studies. One of our major customers is the international modeling comparison study known as AEROCOM, which relies on the MODIS data as a benchmark. In order to keep up with the needs of AEROCOM and other MODIS data users, while utilizing new science and tools, we have improved the algorithms and products. The code, and the associated products, will be known as Collection 6 (C6). While not a major overhaul from the previous Collection 5 (C5) version, there are enough changes that there are significant impacts to the products and their interpretation. In its entirety, the C6 algorithm is comprised of three sub-algorithms for retrieving aerosol properties over different surfaces: These include the dark-target DT algorithms to retrieve over (1) ocean and (2) vegetated-dark-soiled land, plus the (3) Deep Blue (DB) algorithm, originally developed to retrieve over desert-arid land. Focusing on the two DT algorithms, we have updated assumptions for central wavelengths, Rayleigh optical depths and gas (H2O, O3, CO2, etc.) absorption corrections, while relaxing the solar zenith angle limit (up to 84) to increase pole-ward coverage. For DT-land, we have updated the cloud mask to allow heavy smoke retrievals, fine-tuned the assignments for aerosol type as function of season location, corrected bugs in the Quality Assurance (QA) logic, and added diagnostic parameters such as topographic altitude. For DT-ocean, improvements include a revised cloud mask for thin-cirrus detection, inclusion of wind speed dependence in the retrieval, updates to logic of QA Confidence flag (QAC) assignment, and additions of important diagnostic information. At the same time as we have introduced algorithm changes, we have also accounted for upstream changes including: new instrument calibration, revised land-sea masking, and changed cloud masking. Upstream changes also impact the coverage and global statistics of the retrieved AOD. Although our responsibility is to the DT code and products, we have also added a product that merges DT and DB product over semi-arid land surfaces to provide a more gap-free dataset, primarily for visualization purposes. Preliminary validation shows that compared to surface-based sunphotometer data, the C6, Level 2 (along swath) DT-products compare at least as well as those from C5. C6 will include new diagnostic information about clouds in the aerosol field, including an aerosol cloud mask at 500 m resolution, and calculations of the distance to the nearest cloud from clear pixels. Finally, we have revised the strategy for aggregating and averaging the Level 2 (swath) data to become Level 3 (gridded) data. All together, the changes to the DT algorithms will result in reduced global AOD (by 0.02) over ocean and increased AOD (by 0.02) over land, along with changes in spatial coverage. Changes in calibration will have more impact to Terras time series, especially over land. This will result in a significant reduction in artificial differences in the Terra and Aqua datasets, and will stabilize the MODIS data as a target for AEROCOM studie

Aerosol retrieval algorithms↗

Canadian and Alaskan Wildfire Smoke Particle Properties, Their Evolution and Controlling Factors, From Satellite Observations

The optical and chemical properties of biomass burning (BB) smoke particles greatly affect the impact that wildfires have on climate and air quality. Previous work has demonstrated some links between smoke properties and factors such as fuel type and meteorology. However, the factors controlling BB particle speciation at emission are not adequately understood nor are the factors driving particle aging during atmospheric transport. As such, modeling wildfire smoke impacts on climate and air quality remains challenging. The potential to provide robust, statistical characterizations of BB particles based on ecosystem type and ambient environmental conditions with remote sensing data is investigated here. Space-based Multi-angle Imaging SpectroRadiometer (MISR) observations, combined with the MISR Research Aerosol (RA) algorithm and the MISR Interactive Explorer (MINX) tool, are used to retrieve smoke plume aerosol optical depth (AOD) and to provide constraints on plume vertical extent; smoke age; and particle size, shape, light-absorption properties, and absorption spectral dependence. These tools are applied to numerous wildfire plumes in Canada and Alaska, across a range of conditions, to create a regional inventory of BB particle-type temporal and spatial distribution. We then statistically compare these results with satellite measurements of fire radiative power (FRP) and land cover characteristics, as well as short-term climate, meteorological, and drought information from the Modern-Era Retrospective analysis for Research and Applications (MERRA-2) reanalysis and the North American Drought Monitor. We find statistically significant differences in the retrieved smoke properties based on land cover type, with fires in forests producing the thickest plumes containing the largest, brightest particles and fires in savannas and grasslands exhibiting the opposite. Additionally, the inferred dominant aging mechanisms and the timescales over which they occur vary systematically between land types. This work demonstrates the potential of remote sensing to constrain BB particle properties and the mechanisms governing their evolution over entire ecosystems. It also begins to realize this potential, as a means of improving regional and global climate and air quality modeling in a rapidly changing world.

Katherine T. Junghenn Noyes↗

Evaluating Retrieval Algorithm Climate Stability: Estimating 3D Optical Thickness Bias Distributions by Cloud Type

Detecting climate trends on large spatiotemporal scales requires accurate, stable measurements and stable retrieval algorithms. We strive to estimate how time-variant retrieval algorithm biases may impact trend detection. Here we focus on the 3D cloud optical thickness (τc) bias, which is among the largest in passive cloud retrieval algorithms. If this bias is time dependent, a possibility with potential decadal changes in cloud morphology, it may obscure genuine trends in τc. Although previous studies have evaluated the cloud- and sun-view geometry-dependent 3D τc bias on small spatial scales, before our current study none have evaluated the stability of this well-known bias on climate-relevant large spatiotemporal scales. These studies must estimate large scale distributions of the 3D τc bias by cloud type and estimate how cloud type amount may change between two climate states. We employ a novel approach to estimate large scale distributions of 3D τc using a proxy of the bias that quantifies the departure of clouds from satisfying the 1D radiative transfer assumption used in passive τc retrievals. This existing globally-distributed proxy is an angular consistency metric that was developed using fused Moderate-Resolution Imaging Spectroradiometer (MODIS) and Multi-angle Imaging Spectroradiometer (MISR) measurements. Calculating the 3D τc bias and the proxy, for known cloud fields enables us to establish statistical relationships between these two quantities, which can be used to calculate large-scale distributions of the 3D τc bias. This approach limits the number of 3D radiative transfer simulations required to only those needed to estimate a statistical relationship between the 3D τc bias for known cloud fields and a proxy of the bias. It is likely that future studies will be needed to evaluate retrieval algorithm bias stability for other geophysical variables as the community develops climate data records from satellite observations and their retrievals. This must be done in addition to monitoring and correcting measurement errors and uncertainties and understanding their impact on retrieved essential climate variables.

Yolanda Shea↗

Interpolation of a surface from sets of discrete height data of different statistical characteristics

This paper presents and analyzes a method for the interpolation of a unique surface from two sets of independent digital height data of differing statistical characteristics. This method is based on linear prediction and thus relies on the concepts of auto- and cross-covariance functions. The linear prediction algorithm for two sets of digital height measurements is first derived and then evaluated using the method of moving averages and bilinear interpolation for comparison. It is found that the overall root mean square interpolation errors of linear prediction are similar to those from moving averages and bilinear interpolation. This accuracy performance, together with the well known potential for controlled filtering of measuring errors and good-behavior in areas of poor control, makes linear prediction a versatile and general method for interpolating a unique surface from two sets of digital height data, with applications in photogrammetric mapping, remote sensing, and other fields.

Leberl, F.↗

Galileo maneuver analysis

In the maneuver analysis of the Galileo spacecraft, analytic models have been developed to assess the performance of an interplanetary dual spin spacecraft. These models take into account all the important effects of dual spin and flexible body dynamics to determine the spacecraft capability to achieve precise velocity changes for a variety of maneuver modes, as dictated by the requirements and as are tested and verified by computer simulation. Proportional velocity change magnitude accuracies as small as 0.34%, proportional velocity change pointing accuracies as little as 10 milliradians and fixed velocity change accuracies as precise as 0.015 m/sec are indicative of the stringency of these requirements. Error sources considered in the statistical analysis include probabilistic uncertainties due to wobble, plume impingement, nutation, thruster and accelerometer misalignments and radial offsets, gyro drift, burn timing, mass properties and algorithm errors. With its twelve thrusters, the versatility of the spacecraft to maneuver among the Galilean moons for eleven encounters after delivering a probe into the Jovian atmosphere provides a new level of challenge in the area of maneuver analysis.

Longuski, J. M.↗

The stochastic evolution of asteroidal regoliths and the origin of brecciated and gas-rich meteorites

A model is constructed which views regolith evolution on asteroids as a stochastic process. Average values are shown to be poor descriptors of regolith depth. The utility of the average depth is not significantly increased by avoiding large craters or thick ejecta deposits, a procedure adopted in previous regolith studies. The statistical uncertainty associated with regolith depth severely limits the power of regolith models in predicting parent-body size for brecciated meteorites. A Monte Carlo algorithm was used to simulate the random walks and corresponding charged-particle irradiation histories of grains in regoliths. On rocky asteroids, only about 20 percent of the grains was exposed to solar cosmic ray ions. Results based on present-day conditions in the asteroid belt agree well with irradiation features observed in gas-rich meteorites. An origin during epochs of early solar system evolution is not required.

Housen, K. R.↗

General multiyear aggregation technology: Methodology and software documentation

A general methodology is presented for estimating a stratum's at-harvest crop acreage proportion for a given crop year (target year) from the crop's estimated acreage proportion for sample segments from within the stratum. Sample segments from crop years other than the target year are (usually) required for use in conjunction with those from the target year. In addition, the stratum's (identifiable) crop acreage proportion may be estimated for times other than at-harvest in some situations. A by-product of the procedure is a methodology for estimating the change in the stratum's at-harvest crop acreage proportion from crop year to crop year. An implementation of the proposed procedure as a statistical analysis system routine using the system's matrix language module, PROC MATRIX, is described and documented. Three examples illustrating use of the methodology and algorithm are provided.

Baker, T. C.↗

Bias correction for rainrate retrievals from satellite passive microwave sensors

Rainrates retrieved from past and present satellite-borne microwave sensors are affected by a fundamental remote sensing problem. Sensor fields-of-view are typically large enough to encompass substantial rainrate variability, whereas the retrieval algorithms, based on radiative transfer calculations, show a non-linear relationship between rainrate and microwave brightness temperature. Retrieved rainrates are systematically too low. A statistical model of the bias problem shows that bias correction factors depend on the probability distribution of instantaneous rainrate and on the average thickness of the rain layer.

Short, David A.↗

Real-Time Assimilation of Goes-Derived Products into A Mesoscale Model and It's Impact on Short-Term (06-36hr) Forecasts from 17 October 1998 through the Present

As the parameterizations of surface energy budgets in regional models have become more complete physically, models have the potential to be much more realistic in simulations of coupling between surface radiation, hydrology, and surface energy transfer. Realizing the importance of properly specifying the surface energy budget, many institutions are using land-surface models to represent the lower boundary forcing associated with biophysical processes and soil hydrology. However, the added degrees of freedom due to inclusion of such land-surface schemes require the specification of additional parameters within the model system such as vegetative resistances, green vegetation fraction, leaf area index, soil physical and hydraulic characteristics, stream flow, runoff, and the vertical distribution of soil moisture. Spatial heterogeneity of these parameters makes correct specification problematic since measurements are not routinely available. A technique has been developed for assimilating GOES-IR skin temperature tendencies, solar insolation, and surface albedo into the surface energy budget equation of a mesoscale model so that the simulated rate of temperature change closely agrees with the satellite observations. The technique has been successfully employed in a number of mesoscale models in case-study mode. We have taken the next step and developed a study to determine if assimilating these types of data into mesoscale models in real-time can improve short-term (648h) forecasts of temperature, relative humidity, and QPF on a daily basis over relatively large regions. Therefore, an operational modeling/assimilation system has been developed at the GHCC during the past summer that allows us to produce simulations out to 48 hours in a timely manor. The PSU/NCAR MM5 is used in a nested configuration with a 25 km grid covering the southeastern third of the US. The model has been on-line since 1 July 1998 and forecast products are posted on our web site. The satellite algorithms that generate data to be assimilated came on-line 17 October 1998. Quantitative assessment of the forecast quality is performed via traditional verification statistics. In addition, invaluable qualitative information is obtained through close collaboration with several NWSFO's who are using the MM5 products in real-time on a daily basis. The assimilation technique has been applied in an off-line mode since 17 October. Results based on bulk statistical verification of surface meteorology over the entire Southeastern US show that assimilating the GOES-derived land surface tendencies and solar radiation results in a significant reduction of the shelter air temperature and RH bias on a daily basis. In fact, the assimilation technique has produced improved temperature and RH forecasts for 97% of the 100 simulations performed to date. Work is currently underway to determine the sensitivity of the assimilation procedure to the availability of satellite data, length of assimilation period, model initialization, and synoptic-scale meteorological conditions. In addition, results from a detailed energy budget analysis using the Early Eta, our operational MM5, and the assimilation runs will help us to better understand the satellite assimilation the land-surface energy budge. Research during the spring-summer of 1999 will focus on the impact of the assimilation technique during the warm season where it is hypothesized that it can have a positive impact on QPF during conditions of weak synoptic-scale forcing.

Lapenta, William M.↗

Structural Health Monitoring of Composite Plates Under Ambient and Cryogenic Conditions

Methods for structural health monitoring are now being assessed, especially in high-performance, extreme environment, safety-critical applications. One such application is for composite cryogenic fuel tanks. The work presented here attempts to characterize and investigate the feasibility of using imbedded piezoelectric sensors to detect cracks and delaminations under cryogenic and ambient conditions. Different types of excitation and response signals and different sensors are employed in composite plate samples to aid in determining an optimal algorithm, sensor placement strategy, and type of imbedded sensor to use. Variations of frequency and high frequency chirps of the sensors are employed and compared. Statistical and analytic techniques are then used to determine which method is most desirable for a specific type of damage and operating environment. These results are furthermore compared with previous work using externally mounted sensors. More work is needed to accurately account for changes in temperature seen in these environments and be statistically significant. Sensor development and placement strategy are other areas of further work to make structural health monitoring more robust. Results from this and other work might then be incorporated into a larger composite structure to validate and assess its structural health. This could prove to be important in the development and qualification of any 2nd generation reusable launch vehicle using composites as a structural element.

Engberg, Robert C.↗

NASA Global Precipitation Mission Ground Validation Implementation

The Global Precipitation Mission (GPM; core-satellite launch 2013) will provide Ka/Ku-band dual-frequency precipitation radar (DPR) and accompanying passive microwave radiometer-diagnosed precipitation estimates over a latitude range of 65 N to 65 S. The extended latitudinal domain of GPM coverage combined with requirements to detect (and in the case of liquid, estimate) liquid and frozen precipitation rates for values ranging from several hundred to just a few tenths of a millimeter per hour present new challenges to the development of physically-based satellite precipitation retrieval algorithms. On regional scales select national and international resources such as existing calibrated radar and rain gauge networks can provide basic datasets that enable direct statistical validation of GPM core-satellite reflectivitys and core/constellation rain rate measurements. Near-term planned field campaign involvements include Finland/Baltic Sea (fall 2010; joint CloudSat,GPM, and European study of precipitation in low-altitude melting layers and snowfall in the vicinity of the Helsinki testbed), central Oklahoma (spring 2011; joint with DOE ARM- precipitation retrievals over a mid-latitude continental land surface), and the Great Lakes region (winter 2011-12, snowfall retrieval).

Petersen, Walter A.↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Enhancing segmentation fairness through curriculum learning and progressive loss: a centralized and federated perspective on radiograph analysis

Bias in medical image segmentation can lead to unequal performance across demographic subgroups, raising concerns about fairness and reliability in clinical AI systems. While deep learning models have achieved high segmentation accuracy, ensuring equitable performance across race and gender remains a significant challenge, particularly in privacy-sensitive healthcare environments. This study investigates fairness-aware medical image segmentation for hip and knee radiographs using deep learning models evaluated in both centralized and Federated Learning (FL) settings. We introduce Curriculum Learning (CL) strategies and Progressive Loss (PL) functions to regulate sample difficulty during training. In addition, we propose two novel fairness-oriented federated learning algorithms, Federated Intersection over Union (FedIoU) and Federated Intersection over Union with Outlier Analysis (FedIoUoutlier). Experiments are conducted using multiple segmentation backbones and simulated multi-site data partitions derived from the Osteoarthritis Initiative dataset. Model performance is evaluated using Intersection over Union (IoU), IoU standard deviation, Skewed Error Ratio (SER), and Min-Max Disparity across race and gender subgroups. Statistical significance was verified using paired t-tests to compare per-sample IoU performance against baseline configurations. Across both hip and knee segmentation tasks, curriculum learning and progressive loss strategies consistently improved segmentation accuracy and reduced demographic performance disparities in centralized training. In federated settings, fairness-aware aggregation further enhanced performance. Notably, FedIoUoutlier combined with balanced curriculum learning and tiered progressive loss achieved the highest mean IoU while yielding the lowest SER and Min-Max Disparity, indicating improved fairness without sacrificing accuracy. In several configurations, federated models matched or exceeded the performance of optimized centralized models, with statistically significant improvements in per-sample IoU over baseline configurations. The results demonstrate that structured training strategies and fairness-aware federated aggregation can jointly improve accuracy, stability, and demographic fairness in medical image segmentation. By integrating curriculum learning, progressive loss, and novel FL algorithms, this work provides a practical pathway toward equitable and privacy-preserving AI systems for medical imaging.

97 MATHEMATICS AND COMPUTING↗