Search NASASearch

SEARCH · Search NASA

Results for “subset selections”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

“Which Projections Do I Use?” Strategies for Climate Model Ensemble Subset Selection Based on Regional Stakeholder Needs

Climate model (or earth system model) projections are increasingly used for climate adaptation planning and impact assessments. As part of this process, many end‐users evaluate a subset of downscaled climate projections without being aware of the implications of downscaling methodology for statistics or event outcomes. Approaches for determining a subset of global climate models to use often focus on values from the raw models, rather than from their downscaled counterparts, in other words assuming that the statistical distribution of the multi‐model ensemble does not change post downscaling. This study demonstrates that a downscaled ensemble will typically retain the change distribution as a raw ensemble, but individual models can differ dramatically post‐downscaling. We recommend that subset‐selection methods account for this possibility and that decision‐relevant downscaled climate projections provide proper descriptions of fitness‐for‐purpose and essential caveats, so that non‐specialists can interpret the results with an appropriate level of confidence.

54 ENVIRONMENTAL SCIENCES

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool for compressing large data matrices that arise in simulations of partial differential equations (PDEs), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. Here, to address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices in situ. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

Column subset selection

Rethinking the soil core microbiome

The concept of a core microbiome emerged from host-associated research to describe microbial members or functions conserved across clearly defined spatial, temporal, and biological boundaries. In soil- and plant-associated microbiome research, however, the term has increasingly shifted toward analytically defined subsets selected using study-specific thresholds or criteria. Synthesizing recent literature and cross-site analyses of bioenergy crop field soils, we show that the original biological meaning of the core microbiome has been blurred by dataset-specific analytical criteria. Taxa designated as ‘core’ were highly sensitive to methodological choices and often reflected explanatory value rather than conserved biological membership. Moreover, many studies that identify taxonomic ‘core’ members interpret their significance in functional terms, suggesting that functional conservation may be the biological interest. Taxonomic conservation may not be the most biologically meaningful target in highly heterogeneous soil and rhizosphere systems, where functional conservation may persist despite taxonomic turnover. Accordingly, ‘core microbiome’ should be reserved for microbial components explicitly demonstrated to be conserved across defined spatial, temporal, and environmental dimensions and linked to conserved ecological functions, while taxa selected for explanatory value are better described as ‘explanatory subsets of taxa’. Greater terminological precision will improve cross-study comparability and strengthen ecological inference in plant–soil microbiome research.

bioenergy crops

nuclear-score-maximization v1.0

This software library presents efficient and multithreaded implementations of matrix low rank approximation via column selection in C++17 code. The algorithms are described in Fornace, Mark, and Michael Lindsey. "Column and row subset selection using nuclear scores: algorithms and theory for Nystro m approximation, CUR decomposition, and graph Laplacian reduction." arXiv preprint arXiv:2407.01698 (2024). The presented methods are by-and-large ver novel, have provable approximation guarantees, multiple use-cases, and exhibit higher quality approximations on a variety of studied examples.

Fornace, Mark

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery

Influence of counterion substitution on the properties of imidazolium-based ionic liquid clusters

Due to their unique physiochemical properties that may be tailored for specific purposes, ionic liquids (ILs) have been investigated for various applications, including chemical separations, catalysis, energy storage, and space propulsion. The different cations and anions comprising ILs may be selected to optimize a range of desired properties, such as thermal stability, ionic conductivity, and volatility, leading to the designation of certain ILs as designer “green” solvents. The effect of counterions on the properties of ILs is of both fundamental scientific interest and technological importance. Herein, we report a systematic experimental and theoretical investigation of the size, charge, stability toward dissociation, and geometric/electronic structure of 1-ethyl-3-methyl imidazolium (EMIM)-based IL clusters containing two different atomic counterions (i.e., bromide [Br − ] and iodide [I − ]). This work extends our studies of EMIM + cations with atomic chloride (Cl − ) and molecular tetrafluoroborate (BF 4 − ) anions reported previously by Baxter et al. [Chem. Mater. 34, 2612 (2022)] and Zhang et al . [J. Phys. Chem. Lett. 11, 6844 (2020)], respectively. Distributions of anionic IL clusters were generated in the gas phase using electrospray ionization and characterized by high mass resolution mass spectrometry, energy-resolved collision-induced dissociation, and negative ion photoelectron spectroscopy experiments. The experimental results reveal anion-dependent trends in the size distribution, relative abundance, ionic charge state, stability toward dissociation, and electron binding energies of the IL clusters. Complementary global optimization theory provides molecular-level insights into the bonding and electronic structure of a selected subset of clusters, including their low energy structures and electrostatic potential maps, and how these fundamental characteristics are influenced by anion substitution. Collectively, our findings demonstrate how the fundamental properties of ILs, which determine their suitability for many applications, may be tuned by substituting counterions. These observations are critical in the sub-nanometer cluster size regime where phenomena do not scale predictably to the bulk phase, and each atom counts toward determining behavior.

cluster

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence

Uncertainty-informed selection of CMIP6 Earth System Model subsets for use in multisectoral and impact models

Earth system models (ESMs) and general circulation models (GCMs) are heavily used to provide inputs to sectoral impact and multisector dynamic models, which include representations of energy, water, land, economics, and their interactions. Therefore, representing the full range of model uncertainty, scenario uncertainty, and interannual variability that ensembles of these models capture is critical to the exploration of the future co-evolution of the integrated human–Earth system. The pre-eminent source of these ensembles has been the Coupled Model Intercomparison Project (CMIP). With more modeling centers participating in each new CMIP phase, the size of the model archive is rapidly increasing, which can be intractable for impact modelers to effectively utilize due to computational constraints and the challenges of analyzing large datasets. In this work, we present a method to select a subset of the latest phase, CMIP6, featuring models for use as inputs to a sectoral impact or multisector dynamics models, while prioritizing preservation of the range of model uncertainty, scenario uncertainty, and interannual variability in the full CMIP6 ensemble results. This method is intended to help impact modelers select climate information from the CMIP archive efficiently for use in downstream models that require global coverage of climate information. This is particularly critical for large-ensemble experiments of multisector dynamic models that may be varying additional features beyond climate inputs in a factorial design, thus putting constraints on the number of climate simulations that can be used. We focus on temperature and precipitation outputs of CMIP6 models, as these are two of the most used variables among impact models, and many other key input variables for impacts are at least correlated with one or both of temperature and precipitation (e.g., relative humidity). Besides preserving the multi-model ensemble variance characteristics, we prioritize selecting CMIP6 models in the subset that preserve the very likely distribution of equilibrium climate sensitivity values as assessed by the latest Intergovernmental Panel on Climate Change (IPCC) report. This approach could be applied to other output variables of climate models and, possibly when combined with emulators, offers a flexible framework for designing more efficient experiments on human-relevant climate impacts. It can also provide greater insight into the properties of existing CMIP6 models.

Snyder, Abigail C.

Tomographic Sparse View Selection Using the View Covariance Loss

Standard computed tomography (CT) reconstruction algorithms such as filtered back projection (FBP) and Feldkamp-Davis-Kress (FDK) require many views for producing high-quality reconstructions, which can slow image acquisition and increase cost in non-destructive evaluation (NDE) applications. Over the past 20 years, a variety of methods have been developed for computing high-quality CT reconstructions from sparse views. However, the problem of how to select the best views for CT reconstruction remains open. In this paper, we present a novel view covariance loss (VCL) function that measures the joint information of a set of views by approximating the normalized mean squared error (NMSE) of the reconstruction. We present fast algorithms for computing the VCL along with an algorithm for selecting a subset of views that approximately minimizes its value. Our experiments on simulated and measured data indicate that for a fixed number of views our proposed view covariance loss selection (VCLS) algorithm results in reconstructions with lower NRMSE, fewer artifacts, and greater accuracy than current alternative approaches.

Lin, Jingsong [Purdue University]

Optimizing resource allocation in Miscanthus breeding via sparse testing designs for genomic prediction

Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.

Miscanthus sacchariflorus (MSA)

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Plateau to River Model Predictive Simulations for All Ensemble Realizations to Support Modeling Work in Fiscal Year 2025

The purpose of this environmental calculation file (ECF) is to document predictions of flow and hydraulic head on the Central Plateau of the Hanford Site using the Plateau-to-River (P2R) Model (CP-57037, Model Package Report for the Plateau-to-River Model: Version 9.1). This calculation documents the simulation of the groundwater for the parent model domain of the P2R Model as a basis for use in other applications of the P2R Model. This application is unique from the standpoint that it will simulate all ensemble member models of the P2R Model whereas other applications may only utilize specific ensemble members. These simulations provide results that can be used in the process of selecting an appropriate subset of ensemble members for other applications.

54 ENVIRONMENTAL SCIENCES

GOES-16 Data for LASSO-CACTI Overview Paper

GOES-16 L1b satellite radiances have been obtained for the LASSO-CACTI case dates. Specifically, the period in the ARM subset is for select days in the period October 26, 2018 through March 15, 2019. These files were downloaded from Amazon Web Services using the GOES-2-Go library, https://blaylockbk.github.io/goes2go/_build/html/.

{"GOES-16 band 13",radiance,LASSO-CACTI}

Size-Resolved Chemical Composition of Particles Collected Using STAC at the Ground Site During the SAIL Campaign in Gunnison, Colorado

Aerosol particles were collected using a four-stage Size and Time-resolved Aerosol Collector (STAC) during the SAIL field campaign. Each stage of STAC separates particles into distinct aerodynamic size fractions with 50% cut-off diameters: Stage A: 2.27 µm Stage B: 0.615 µm Stage C: 0.421 µm Stage D: 0.119 µm Each stage provides both size- and time-resolved sampling, enabling investigation of particle composition across different atmospheric regimes. Only a subset of samples was selected for analysis based on prevailing meteorological conditions (e.g., temperature, humidity, and air-mass influence) to capture representative aerosol types under distinct weather patterns. Collected substrates were first examined under Scanning Electron Microscopy (SEM) to evaluate particle loading, morphology, and spatial distribution. Subsequently, Computer-Controlled Scanning Electron Microscopy with Energy-Dispersive X-ray Spectroscopy (CCSEM/EDX) was performed to obtain size-resolved elemental composition of individual particles. A rule-based classification scheme was applied to categorize particles into major compositional groups (e.g., biological, carbonaceous, dust, sulfate, Na-rich, and mixed types). This dataset provides high-resolution morphological and chemical information on atmospheric particles collected during the SAIL campaign, offering insights into the influence of meteorology on aerosol composition and mixing state.

Size and Time-resolved Aerosol Collector

“Beam à la carte”: Laser heater shaping for attosecond pulses in a multiplexed x-ray free-electron laser

Electron beam shaping allows the control of the temporal properties of x-ray free-electron laser pulses from femtosecond to attosecond timescales. Here, we demonstrate the use of a laser heater to shape electron bunches and enable the generation of attosecond x-ray pulses. We demonstrate that this method can be applied in a selective way, shaping a targeted subset of bunches while leaving the remaining bunches unchanged. This experiment enables the delivery of shaped x-ray pulses to multiple undulator beamlines, with pulse properties tailored to specialized scientific applications.

43 PARTICLE ACCELERATORS

A Novel Gene Stacking Method in Plant Transformation Utilizing Split Selectable Markers

Gene stacking, the process of introducing multiple genes into a single plant to enhance desired traits, is essential for plant genetic improvement through both conventional breeding and genetic transformation. In general, transformation-based gene stacking can be achieved through either co-transformation to simultaneously introduce multiple genes or sequential multi-round transformation. While co-transformation is generally faster and more efficient than sequential multi-round transformation, it often requires two selectable marker genes, which confer resistance to antibiotics, for selecting transgenic events. However, in most cases, there is only one best selectable marker gene for a specific plant species or genotype. Also, it is harder to optimize the concentrations of two antibiotics for co-transformation than using one antibiotic for selecting transgenic events. To overcome this challenge, we recently developed an innovative split selectable marker system for plant co-transformation, allowing the use of one selectable marker gene to select transgenic events. This method involves constructing two binary vectors, each carrying a subset of genes of interest and a partial fragment of the selectable marker gene, which is connected to a partial intein fragment. Following Agrobacterium -mediated co-transformation, plants harboring both binary vectors are selected using a single antibiotic, such as kanamycin. This split-marker system can be used to co-transform multiple genes into both herbaceous and woody plants, accelerating genetic improvement of polygenic traits or integrative improvement of multiple traits to simultaneously increase crop yield and quality.

59 BASIC BIOLOGICAL SCIENCES

Revealing Hidden Quinones Through Diagnostic MS² Fragmentation of Peptide–Quinone Adducts

Quinones are redox-active components of natural organic matter that mediate electron transfer and influence biogeochemical processes, but many quinones in pyrogenic organic matter (PyOM) remain unresolved because they ionize poorly by mass spectrometry. Here, we present a peptide-tagging approach to improve detection of cysteine-reactive electrophiles in PyOM, with quinones expected to be a dominant subset based on reaction chemistry and selectivity experiments. A cysteine-containing peptide was used to form Michael-addition adducts, enhancing electrospray ionization and enabling untargeted screening by high-performance liquid chromatography-high-resolution tandem mass spectrometry. The method was benchmarked with five quinone standards and applied to extracts from charred plant material as a discovery-level screen for cysteine-reactive targets. We identified 98 quinone-candidate adducts (mean neutral mass ~603 Da), of which more than 70% were not detectable in native MS1 data. Among formula-assigned features, hidden quinone candidates had median (O+N)/C of 0.391 and normalized oxidation state of carbon of -0.281, consistent with relatively low polarity and low oxidation state. These results reveal a previously inaccessible pool of hidden redox-active compounds in PyOM and provide a framework for prioritizing quinone-like electrophiles for confirmation and incorporation into models of fire-driven biogeochemical cycling.

LC-MS/MS

Demonstration of neutron identification in neutrino interactions in the MicroBooNE liquid argon time projection chamber

A significant challenge in measurements of neutrino oscillations is reconstructing the incoming neutrino energies. While modern fully-active tracking calorimeters such as liquid argon time projection chambers in principle allow the measurement of all final state particles above some detection threshold, undetected neutrons remain a considerable source of missing energy with little to no data constraining their production rates and kinematics. We present the first demonstration of tagging neutrino-induced neutrons in liquid argon time projection chambers using secondary protons emitted from neutron-argon interactions in the MicroBooNE detector. We describe the method developed to identify neutrino-induced neutrons and demonstrate its performance using neutrons produced in muon-neutrino charged current interactions. The method is validated using a small subset of MicroBooNE’s total dataset. The selection yields a sample with 60% of selected tracks corresponding to neutron-induced secondary protons. At this purity, the integrated efficiency is 8.4% for neutrons that produce a detectable proton.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS