Search NASA⌕ Search

SEARCH · Search NASA

Results for “variable selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

O Corona, where art thou? eROSITA’s view of UV-optical-IR variability-selected massive black holes in low-mass galaxies

Finding massive black holes (MBHs,M BH ≈ 10 4 –10 7 M ⊙ ) in the nuclei of low-mass galaxies $\left( {{M_*}\mathop {\mathop < \limits_ }\limits_ {{10}^{10}}{M_ \odot }} \right)$ is crucial to constrain seeding and growth of black holes over cosmic time, but it is particularly challenging due to their low accretion luminosities. Variability selection via long-term photometric ultraviolet, optical, or infrared (UVOIR) light curves has proved effective and identifies lower-Eddington ratios compared to broad and narrow optical spectral lines searches. In the inefficient accretion regime, X-ray and radio searches are effective, but they have been limited to small samples. Therefore, differences between selection techniques have remained uncertain. Here, we present the first large systematic investigation of the X-ray properties of a sample of known MBH candidates in dwarf galaxies. We extracted X-ray photometry and spectra of a sample of ~200 UVOIR variability-selected MBHs and significantly detected 17 of them in the deepest available SRG/eROSITA image, of which four are newly discovered X-ray sources and two are new secure MBHs. This implies that tens to hundreds of LSST MBHs will have SRG/eROSITA counterparts, depending on the seeding model adopted. Surprisingly, the stacked X-ray images of the many non-detected MBHs are incompatible with standard disk-corona relations, typical of active galactic nuclei, inferred from both the optical and radio fluxes. They are instead compatible with the X-ray emission predicted for normal galaxies. After careful consideration of potential biases, we identified that this X-ray weakness needs a physical origin. A possibility is that a canonical X-ray corona might be lacking in the majority of this population of UVOIR-variability selected low-mass galaxies or that unusual accretion modes and spectral energy distributions are in place for MBHs in dwarf galaxies. This result reveals the potential for severe biases in occupation fractions derived from data from only one waveband combined with SEDs and scaling relations of more massive black holes and galaxies.

Astronomy & Astrophysics↗

Forward variable selection enables fast and accurate dynamic system identification with Karhunen-Loève decomposed Gaussian processes

A promising approach for scalable Gaussian processes (GPs) is the Karhunen-Loève (KL) decomposition, in which the GP kernel is represented by a set of basis functions which are the eigenfunctions of the kernel operator. Such decomposed kernels have the potential to be very fast, and do not depend on the selection of a reduced set of inducing points. However KL decompositions lead to high dimensionality, and variable selection thus becomes paramount. This paper reports a new method of forward variable selection, enabled by the ordered nature of the basis functions in the KL expansion of the Bayesian Smoothing Spline ANOVA kernel (BSS-ANOVA), coupled with fast Gibbs sampling in a fully Bayesian approach. It quickly and effectively limits the number of terms, yielding a method with competitive accuracies, training and inference times for tabular datasets of low feature set dimensionality. Theoretical computational complexities are O ( N P 2 ) in training and O ( P ) per point in inference, where N is the number of instances and P the number of expansion terms. The inference speed and accuracy makes the method especially useful for dynamic systems identification, by modeling the dynamics in the tangent space as a static problem, then integrating the learned dynamics using a high-order scheme. The methods are demonstrated on two dynamic datasets: a ‘Susceptible, Infected, Recovered’ (SIR) toy problem, along with the experimental ‘Cascaded Tanks’ benchmark dataset. Comparisons on the static prediction of time derivatives are made with a random forest (RF), a residual neural network (ResNet), and the Orthogonal Additive Kernel (OAK) inducing points scalable GP, while for the timeseries prediction comparisons are made with LSTM and GRU recurrent neural networks (RNNs) along with the SINDy package.

Hayes, Kyle↗

Decision Tree for Variable Selection vs. Impact on Durability for Biomass and Biochar Burial Pathways [Slides]

Quantifying durability for lower-TRL BiCRS pathways has been challenging as limited data are available from real-world projects and long-term experiments, resulting in an overall lack of scientific consensus. We develop a decision tree that aims to summarize the current scientific understanding and state-of-the-art project experience. The decision tree can be used to (1) guide the selection of key variables and evaluate their relative impact on durability, (2) identify data and knowledge gaps for future research.

09 BIOMASS FUELS↗

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Robust Measurement of Stellar Streams around the Milky Way: Correcting Spatially Variable Observational Selection Effects in Optical Imaging Surveys

Observations of density variations in stellar streams are a promising probe of low-mass dark matter substructure in the Milky Way. However, survey systematics such as variations in seeing and sky brightness can also induce artificial fluctuations in the observed densities of known stellar streams. These variations arise because survey conditions affect both object detection and star–galaxy misclassification rates. To mitigate these effects, we use Balrog synthetic source injections in the Dark Energy Survey (DES) Y3 data to calculate detection rate variations and classification rates as functions of survey properties. We show that these rates are nearly separable with respect to survey properties and can be estimated with sufficient statistics from the synthetic catalogs. Applying these corrections reduces the standard deviation of relative detection rates across the DES footprint by a factor of 5, and our corrections significantly change the inferred linear density of the Phoenix stream when including faint objects. Additionally, for artificial streams with DES-like survey properties we are able to recover density power spectra with reduced bias. We also find that uncorrected power-spectrum results for Legacy Survey of Space and Time (LSST)-like data can be around 5 times more biased, highlighting the need for such corrections in future ground-based surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan↗

Efficiency of ML Anomaly Detection Triggers for Emerging Jets

Novel machine learning-based anomaly detection Level 1 (L1) triggers are currently under development at CMS, namely AXOL1TL and CICADA. The former employs a variational autoencoder, while the latter utilizes a convolutional autoencoder. These triggers aim to balance rate reduction with model independence, enabling the selection of potentially significant events that might be overlooked by traditional triggers relying on basic kinematic variable selections. Consequently, they have the potential to enhance signals indicative of physics beyond the Standard Model, such as those associated with emerging jets. Such signals are predicted by models featuring a composite dark sector where long-lived particles decay into Standard Model jets with displaced tracks and numerous vertices. This study evaluates the efficiency of these anomaly detection triggers in selecting events with emerging jets produced via the s-channel production of two dark quarks.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Efficiency of ML Anomaly Detection Triggers for Emerging Jets

Novel machine learning-based anomaly detection Level 1 (L1) triggers are currently under development at CMS, namely AXOL1TL and CICADA. The former employs a variational autoencoder, while the latter utilizes a convolutional autoencoder. These triggers aim to balance rate reduction with model independence, enabling the selection of potentially significant events that might be overlooked by traditional triggers relying on basic kinematic variable selections. Consequently, they have the potential to enhance signals indicative of physics beyond the Standard Model, such as those associated with emerging jets. Such signals are predicted by models featuring a composite dark sector where long-lived particles decay into Standard Model jets with displaced tracks and numerous vertices. This study evaluates the efficiency of these anomaly detection triggers in selecting events with emerging jets produced via the s-channel production of two dark quarks.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Efficiency of ML Anomaly Detection Triggers for Emerging Jets

Novel machine learning-based anomaly detection Level 1 (L1) triggers are currently under development at CMS, namely AXOL1TL and CICADA. The former employs a variational autoencoder, while the latter utilizes a convolutional autoencoder. These triggers aim to balance rate reduction with model independence, enabling the selection of potentially significant events that might be overlooked by traditional triggers relying on basic kinematic variable selections. Consequently, they have the potential to enhance signals indicative of physics beyond the Standard Model, such as those associated with emerging jets. Such signals are predicted by models featuring a composite dark sector where long-lived particles decay into Standard Model jets with displaced tracks and numerous vertices. This study evaluates the efficiency of these anomaly detection triggers in selecting events with emerging jets produced via the s-channel production of two dark quarks.

43 PARTICLE ACCELERATORS↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Optimization of direct air capture processes using reactive transport models of adsorption-desorption cycles

In this study, we develop and implement a reactive transport model in COMSOL Multiphysics® to address the challenges of direct air carbon capture. The model is validated against experimental data and used to simulate the cyclic steady state of the adsorption-desorption process. The optimization of this model is achieved through advanced trust-region methods integrated with Gaussian Processes. Key decision variables, including adsorption and desorption times, desorption temperature and pressure, input velocity, bed porosity, column length, and radius were optimized to minimize the capture cost. After optimization, a sensitivity analysis revealed the complex interplay between the decision variables and their effect on the specific energy and cost of removing the CO 2 . We optimized the capture cost while taking into account the trade-off between energy consumption and productivity. The resulting minimum capture cost was determined to be 265.2 $/t-CO 2 , which aligns with expected values reported in the literature. Numerical results suggest the effectiveness of the optimization strategies applied, and underscore the importance of simultaneous decision variable selection in improving the performance in direct air capture processes. We also extend the modeling approach to a 2D axisymmetric model to better visualize CO₂ uptake and temperature profiles, revealing significant radial gradients during the regeneration step. As a main drawback, this enhanced model comes with a computational cost approximately 40 times higher than that of the 1D model.

Adsorption-desorption process↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Revealing the complex chemistry of grain boundaries in K-doped BaFe 2 As 2 with atom probe tomography

Iron-based superconductors have attractive properties for high-field applications, but there is a lack of understanding of the effect of grain boundary chemistry on the in-field performance. The near atomic-scale resolution, ppm sensitivity and 3D analysis offered by atom probe tomography make it a powerful tool to investigate the nanoscale structure and chemistry of these defects in fine-grained K-doped BaFe 2 As 2 samples. A computational method to systematically extract and compare the Gibbsian interfacial excess of chemical species across grain boundaries has been explored in this work. The robustness of the method has been tested by evaluating the effects of selected variables on simulated APT datasets. The accuracy and precision of the calculated Gibbsian interfacial excess were found to be stable over a range of analysis conditions: varying grain boundary widths and detection efficiencies, spatial precisions below 1.5 nm, and bin widths between 1.2 and 1.6 nm. For the K-doped BaFe 2 As 2 samples studied, segregation of As, Ba, K and impurities of O, Na, and Sb were found at grain boundaries. The Gibbsian excess values were found to vary widely between different boundaries, showing the complexity of the grain boundary chemistry in this material. Possible links between the observed critical current density (Jc) of these samples and their nano- and micro-structure have also been investigated and discussed.

36 MATERIALS SCIENCE↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

sPHENIX heavy flavor jet tagging studies in p+p at $\sqrt{s_{NN}}=200~GeV$

Heavy-flavor jets, which are initiated from heavy quarks, are ideal probes for studying flavor dependent parton energy loss. We report on the performance of jet flavor tagging using two Neural Network Machine Learning (ML) models: the Long Short-Term Memory (LSTM) model and an Attention-based Neural Network, in simulations of 200 GeV p + p collisions. The tagging performance of bottom quark initiated jets with both ML models surpasses that of the traditional cut-based method. Technical details, including sample and kinematic variable selections, the machine learning training and testing setup with parameter tuning, and outcome comparisons, will be discussed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗