Search NASASearch

SEARCH · Search NASA

Results for “data skew”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment

Resource-Adaptive Federated Text Generation with Differential Privacy

In cross-silo federated learning (FL), sensitive text datasets remain confined to local organizations due to privacy regulations, making repeated training for each downstream task both communication-intensive and privacy-demanding. A promising alternative is to generate differentially private (DP) synthetic datasets that approximate the global distribution and can be reused across tasks. However, pretrained large language models (LLMs) often fail under domain shift, and federated finetuning is hindered by computational heterogeneity: only resource-rich clients can update the model, while weaker clients are excluded, amplifying data skew and the adverse effects of DP noise. We propose a flexible participation framework that adapts to client capacities. Strong clients perform DP federated finetuning, while weak clients contribute through a lightweight DP voting mechanism that refines synthetic text. To ensure the synthetic data mirrors the global dataset, we apply control codes (e.g., labels, topics, metadata) that represent each client’s data proportions and constrain voting to semantically coherent subsets. This two-phase approach requires only a single round of communication for weak clients and integrates contributions from all participants. Experiments show that our framework improves distribution alignment and downstream robustness under DP and heterogeneity.

Wang, Jiayi [ORNL]

Remark on Algorithm 1012: Computing Projections with Large Datasets

In ACM TOMS Algorithm 1012, the DELAUNAYSPARSE software is given for performing Delaunay interpolation in medium to high dimensions. When extrapolating outside the convex hull of the training set, DELAUNAYSPARSE calls the nonnegative least squares solver DWNNLS to compute projections onto the convex hull. However, DWNNLS and many other available sum-of-squares optimization solvers were not intended for usage with many variable problems, which result from the large training sets that are typical in machine learning applications. Thus, a new PROJECT subroutine is given, based on the highly customizable quadratic program solver BQPD. This solution is shown to be as robust as DELAUNAYSPARSE for projection onto both synthetic and real-world datasets, where other available solvers frequently fail. Although it is intended as an update for DELAUNAYSPARSE, due to the difficulty and prevalence of the problem, this solution is likely to be of external interest as well.

97 MATHEMATICS AND COMPUTING

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Generalized parton distributions from lattice QCD with asymmetric momentum transfer: Unpolarized quarks at nonzero skewness

We extend the formalism of asymmetric frames of reference for generalized parton distributions (GPDs) to the case of nonzero skewness, i.e., including longitudinal momentum transfer. The framework, based on Lorentz-invariant amplitudes and previously developed and numerically implemented for unpolarized, helicity and transversity GPDs at zero skewness, gives efficient access to a broad range of kinematics, making full mapping of GPDs from the lattice realistic. The general-skewness formalism is tested using lattice data with both transverse and longitudinal or only longitudinal momentum transfer, the latter being a special case with a reduced number of independent amplitudes. We extract the amplitudes in coordinate space and express the GPDs 𝐻 and 𝐸 in terms of these amplitudes. This is followed by reconstruction of quasidistributions and their matching to the light cone. We further identify and discuss the principal challenges for nonzero skewness GPDs.

Lattice QCD

Improving the estimate of higher-order moments from lidar observations near the top of the convective boundary layer

Abstract. Ground-based lidar data have proven extremely useful for profiling the convective boundary layer (CBL). Many groups have derived higher-order moments (e.g., variance, skewness, fluxes) from high-temporal-resolution lidar data using an autocovariance approach. However, these analyses are highly uncertain near the CBL top when the depth of the CBL (zi) is changing during the analysis period. This is because the autocovariance approach is usually applied to constant height levels and the character of the eddies is changing on either side of the changing CBL top. Here, a new approach is presented wherein the autocovariance analysis is performed on a normalized height grid, with a temporally smoothed zi. Output from a large eddy simulation model demonstrates that deriving higher-order moments from time series on a normalized height grid has better agreement with the slab-averaged quantities than the moments derived from the original height grid.

Rosenberger, Tessa E. (ORCID:0000000333205873)

Impact of the newly revised gravitational redshift of x-ray burster GS 1826-24 on the equation of state of supradense neutron-rich matter

Thanks to the recent advancement in producing rare isotopes and measuring their masses with unprecedented precision, the updated nuclear masses around the waiting-point nucleus 64 Ge in the rapid-proton capture process have led to a significant revision of the surface gravitational redshift of the neutron star (NS) in GS 1826-24 by refitting its x-ray burst light curve using Modules for Experiments in Stellar Astrophysics (MESA). The resulting NS compactness ξ is between 0.183 and 0.259 at 95% confidence level, and its upper boundary is significantly smaller than the maximum ξ previously known. Incorporating these new data within a comprehensive Bayesian statistical framework, we investigate its impact on the Equation of State (EOS) of supradense neutron-rich matter and the required spin frequency for GW190814’s minor m 2 with mass 2.59 ± 0.05⁢M ⊙ to be a rotationally stable pulsar. We found that the EOS of high-density symmetric nuclear matter (SNM) has to be softened significantly while the symmetry energy at supersaturation densities stiffened compared to our prior knowledge from earlier analyses using data from both astrophysical observations and terrestrial nuclear experiments. In particular, the skewness J 0 characterizing the stiffness of high-density symmetric nuclear matter (SNM) decreases significantly, while the slope L, curvature K sym , and skewness Jsym of nuclear symmetry energy all increase appreciably compared to their fiducial values. Here, we also found that the most probable spin rate for the m 2 to be a stable pulsar is very close to its mass-shedding limit once the revised redshift data from GS 1826-24 is considered, making the m 2 unlikely the most massive NS observed so far.

79 ASTRONOMY AND ASTROPHYSICS

Hybrid Star Models in the Light of New Multimessenger Data

Abstract Recent astrophysical mass inferences of compact stars HESS J1731-347 and PSR J0952-0607, with extremely small and large masses respectively, as well as the measurement of the neutron skin of Ca in the CREX experiment challenge and constrain the models of dense matter. We examine the concept of hybrid stars—objects containing quark cores surrounded by nucleonic envelopes—as models that account for these new data along with other inferences. We employ a family of 81 nucleonic equations of state (EOSs) with variable skewness and slope of symmetry energy at saturation density and a constant speed-of-sound EOS for quark matter. For each nucleonic EOS, a family of hybrid EOSs is generated by varying the transition density, the energy jump, and the speed of sound. These models are tested against the data from GW170817 and J1731-347, which favor low-density soft EOS and J0592-0607 and J0740+6620, which require high-density stiff EOS. The addition of J0592-0607's mass measurement to the constraints has no significant impact on the parameter space of the admissible EOS, but allows us to explore the potential effect of pulsars more massive than J0740+6620, if such exists. We then examine the occurrence of twin configurations and quantify the ranges of masses and radii that they can possess. It is shown that including J1731-347 data favors EOSs that predict low-mass twins with M ≲ 1.3 M ⊙ that can be realized if the deconfinement transition density is low. If combined with large speed of sound in quark matter such models allow for maximum masses of hybrid stars in 2.0–2.6 M ⊙ .

Astronomy & Astrophysics

Parametrization of Generalized Parton Distributions from 𝑡-Channel String Exchange in AdS Spaces

We introduce a string-based parametrization for nucleon quark and gluon generalized parton distributions (GPDs) that is valid for all skewness. Our approach leverages conformal moments, representing them as the sum of spin-𝑗 nucleon 𝐴-form factor and skewness-dependent spin-𝑗 nucleon 𝐷-form factor, derived from 𝑡-channel string exchange in AdS spaces consistent with Lorentz invariance and unitarity. This model-independent framework, satisfying the polynomiality condition due to Lorentz invariance, uses Mellin moments from empirical data to estimate these form factors. With just five Regge slope parameters, our method accurately produces various nucleon quark GPD types and symmetric nucleon gluon GPDs through pertinent Mellin-Barnes integrals. Our isovector nucleon quark GPD is in agreement with existing lattice data, promising to improve the empirical extraction and global analysis of nucleon GPDs in exclusive processes, by avoiding the deconvolution problem at any skewness, for the first time.

QCD phenomenology

Charged-particle multiplicity distributions over a wide pseudorapidity range in p–Pb collisions at $\sqrt{s_{NN}}$ = 5.02 TeV

This paper presents the primary charged-particle multiplicity distributions in proton–lead collisions at a centre-of-mass energy per nucleon–nucleon collision of $\sqrt{s_{NN}}$ = 5.02 TeV. The distributions are reported for non-single diffractive collisions in different pseudorapidity ranges. The measurements are performed using the combined information from the Silicon Pixel Detector and the Forward Multiplicity Detector of ALICE. The multiplicity distributions are parametrised with a double negative binomial distribution function which provides satisfactory descriptions of the distributions for all the studied pseudorapidity intervals. The data are compared to models and analyzed quantitatively, evaluating the first four moments (mean, standard deviation, skewness, and kurtosis). The shape evolution of the measured multiplicity distributions is studied in terms of KNO variables and it is found that none of the considered models reproduces the measurements. This paper also reports on the average charged-particle multiplicity, normalised by the average number of participating nucleon pairs, as a function of the collision energy. The multiplicity results are then compared to measurements made in proton–proton and nucleus–nucleus collisions across a wide range of collision energies.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Enhancing segmentation fairness through curriculum learning and progressive loss: a centralized and federated perspective on radiograph analysis

Bias in medical image segmentation can lead to unequal performance across demographic subgroups, raising concerns about fairness and reliability in clinical AI systems. While deep learning models have achieved high segmentation accuracy, ensuring equitable performance across race and gender remains a significant challenge, particularly in privacy-sensitive healthcare environments. This study investigates fairness-aware medical image segmentation for hip and knee radiographs using deep learning models evaluated in both centralized and Federated Learning (FL) settings. We introduce Curriculum Learning (CL) strategies and Progressive Loss (PL) functions to regulate sample difficulty during training. In addition, we propose two novel fairness-oriented federated learning algorithms, Federated Intersection over Union (FedIoU) and Federated Intersection over Union with Outlier Analysis (FedIoUoutlier). Experiments are conducted using multiple segmentation backbones and simulated multi-site data partitions derived from the Osteoarthritis Initiative dataset. Model performance is evaluated using Intersection over Union (IoU), IoU standard deviation, Skewed Error Ratio (SER), and Min-Max Disparity across race and gender subgroups. Statistical significance was verified using paired t-tests to compare per-sample IoU performance against baseline configurations. Across both hip and knee segmentation tasks, curriculum learning and progressive loss strategies consistently improved segmentation accuracy and reduced demographic performance disparities in centralized training. In federated settings, fairness-aware aggregation further enhanced performance. Notably, FedIoUoutlier combined with balanced curriculum learning and tiered progressive loss achieved the highest mean IoU while yielding the lowest SER and Min-Max Disparity, indicating improved fairness without sacrificing accuracy. In several configurations, federated models matched or exceeded the performance of optimized centralized models, with statistically significant improvements in per-sample IoU over baseline configurations. The results demonstrate that structured training strategies and fairness-aware federated aggregation can jointly improve accuracy, stability, and demographic fairness in medical image segmentation. By integrating curriculum learning, progressive loss, and novel FL algorithms, this work provides a practical pathway toward equitable and privacy-preserving AI systems for medical imaging.

97 MATHEMATICS AND COMPUTING

Rosenbluth-like separation of the $J/ψ$ near-threshold photoproduction: An access to the gluon gravitational form factors at high t

Here, we perform analysis of the near-threshold $J/\psi $ photoproduction data off the proton based on two theoretical approaches, GPD \cite{Guo3} and holographic \cite{Zahed2}, that represent the differential cross sections as powers of the skewness parameter with coefficients that depend only on the momentum transfer $t$. This allows to separate kinematically the corresponding coefficient functions, in much the same way as this is done for the electric and magnetic form factors using the Rosenbluth separation. We examine the independence of the extracted functions with the photon beam energy. These functions, under additional assumptions, are related to the proton's gluon Gravitational Form Factors (gGFFs). We compare the extracted functions with lattice calculations of the gGFFs in the region of $0.5<|t|<2$~GeV$^{2}$, where they overlap. Such analysis demonstrates the possibility of extracting some combinations of the gGFFs from the data at high $t$, complementary to the lattice calculations available in the low $t$ region. However, higher statistics are needed to more accurately check the predicted scaling behavior of the data and compare with the lattice results, thus testing and comparing the theoretical assumptions used in the GPD and holographic models.

Pentchev, Lubomir [Thomas Jefferson National Accel

Plan Position Indicator Hydrometeor Field Statistics (PPIHYD) Evaluation Data Product Version 1.0

The PPIHYD evaluation data product provides distinct hydrometeor field statistics calculated from U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility scanning radar plan position indicator (PPI) scans. These statistics include the equivalent reflectivity factor and Doppler spectral width percentiles, min/max values, and first four moments (mean, standard deviation, skewness, and kurtosis) of distinct hydrometeor features (clustered hydrometeor fields). Statistics also include morphological properties, water content and precipitation rate parameterization-based estimates, and thermodynamic properties interpolated using the Interpolated Sonde value-added product (INTERPSONDE VAP). The data set is organized in tabular form and is accompanied by mask arrays with corresponding indices. This straightforward file structure simplifies scanning radar data processing and renders this data set useful for process understanding and model evaluation studies. This report describes the data set and its processing algorithm and provides some examples.

54 ENVIRONMENTAL SCIENCES

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS

Dark Energy Survey Year 3 Results: Cosmological constraints from second- and third-order shear statistics

Here, we present a cosmological analysis of the third-order aperture mass statistic using Dark Energy Survey Year 3 (DES Y3) data. We perform a complete tomographic measurement of the three-point correlation function of the Y3 weak lensing shape catalog with the four fiducial source redshift bins. Building upon our companion methodology paper, we apply a pipeline that combines the two-point function ξ ± with the mass aperture skewness statistic ⟨ M ap 3 ⟩ , which is an efficient compression of the full shear three-point function. We use a suite of simulated shear maps to obtain a joint covariance matrix. By jointly analyzing ξ ± and ⟨ M ap 3 ⟩ measured from DES Y3 data with a Λ CDM model, we find S 8 = 0.780 ± 0.015 and Ω m = 0.26 6 - 0.040 + 0.039 , yielding 111% of figure-of-merit improvement in the Ω m - S 8 plane relative to ξ ± alone, consistent with expectations from simulated likelihood analyses. With a w CDM model, we find S 8 = 0.74 9 - 0.026 + 0.027 and w 0 = - 1.39 ± 0.31 , which gives an improvement of 22% on the joint S 8 - w 0 constraint. Our results are consistent with w 0 = - 1 . Our new constraints are compared to CMB data from the Planck satellite, and we find that with the inclusion of ⟨ M ap 3 ⟩ the existing tension between the datasets is at the level of 2.3 σ . We show that the third-order statistic enables us to self-calibrate the mean photometric redshift uncertainty parameter of the highest redshift bin with little degradation in the figure of merit. Our results demonstrate the constraining power of higher-order lensing statistics and establish ⟨ M ap 3 ⟩ as a practical observable for joint analyses in current and future surveys.

Gomes, R. C. H. [University of Pennsylvania] (ORCI

Comparing the predictive capabilities of subjective crackle metrics using leave-one-out cross validation

Leave-one-out cross validation was used to show the capability to predict crackling sound quality based on the use of the standard deviation of the time-varying sharpness and the skewness of the derivative of the pressure waveform. The predictive output of this method implies that the more sophisticated fits used to describe this data were not overfit, but instead exhibited substantial predictive value.

Swift, Stephen Hales [Sandia National Laboratories

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING

A unified neural-network framework for nucleon imaging from numerical simulations of QCD

Parton distributions encode the momentum-space structure and, in their generalizations, the spatial tomography of quarks and gluons inside hadrons, the building blocks of visible matter. We present a unified neural-network approach that learns these distributions directly from matrix elements calculated via numerical simulations of quantum chromodynamics (QCD) on the lattice by fitting two complementary inputs simultaneously: data matched to physical quantities via known momentum-space and coordinate-space formalisms. Utilizing data from both methods stabilizes the extraction and mitigates biases that can arise when either is used alone. We validate the method on controlled mock data and apply it to lattice-QCD matrix elements to extract parton distribution functions (PDFs). We show benefits of such an approach for determining the physical quantities. We further extend the framework to zero-skewness generalized parton distributions and demonstrate nucleon tomography within the same neural-network parameterization. Our results provide an adaptable and systematically improvable approach for extracting partonic distributions from Euclidean correlators. It can incorporate polarization, additional channels, and future experimental constraints from current and future facilities, such as the Electron-Ion Collider.

Hadronic Spectroscopy