Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sample selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Magnesium ions mitigate metastable states in the regulatory landscape of mRNA elements

Residing in the 5' untranslated region of the mRNA, the 2'-deoxyguanosine (2'-dG) riboswitch mRNA element adopts an alternative structure upon binding of the 2'-dG molecule, which terminates transcription. RNA conformations are generally strongly affected by positively charged metal ions (especially Mg 2+ ). We have quantitatively explored the combined effect of ligand (2'-dG) and Mg 2+ binding on the energy landscape of the aptamer domain of the 2'-dG riboswitch with both explicit solvent all-atom molecular dynamics simulations (99 μsec aggregate sampling for the study) and selective 2'-hydroxyl acylation analyzed by primer extension (SHAPE) experiments. We show that both ligand and Mg 2+ are required for the stabilization of the aptamer domain; however, the two factors act with different modalities. The addition of Mg 2+ remodels the energy landscape and reduces its frustration by the formation of additional contacts. In contrast, the binding of 2'-dG eliminates the metastable states by nucleating a compact core for the aptamer domain. Mg 2+ ions and ligand binding are required to stabilize the least stable helix, P1 (which needs to unfold to activate the transcription platform), and the riboswitch core formed by the backbone of the P2 and P3 helices. Mg 2+ and ligand also facilitate a more compact structure in the three-way junction region.

59 BASIC BIOLOGICAL SCIENCES↗

Formation of Organic Compounds Through Meteoritic Atmospheric Shock

This document is a Final Technical Report for DoE award DE-SC0023375 “Formation of Organic Compounds Through Meteoritic Atmospheric Shock”. The document includes a summary of topics studied, specific tasks completed, challenges, and results from the project. The main goal of this project was to investigate the production of organic molecules and/or complex inorganic precursor molecules in a plasma environment reminiscent of the environment surrounding meteoroids during atmospheric entries. The specific hypothesis tested in this project was that meteoroid ablation during the entry and the chemical reactions in the meteoroid plasma tail could have produced significant amounts of organics or precursor inorganics in the Early Earth’s atmosphere. Investigation of these processes is essential in understanding the origins of life on Earth and the search for life beyond our planet. This project was focused on a set of experiments conducted at the Utilizing the DIII-D tokamak in San Diego, CA. The experiments aimed to study the interaction of carbonaceous and silica materials (typically found in meteoroids) with mixtures of hot plasma gases (mimicking atmospheric entry conditions. The material samples and gas mixtures were selected to investigate the synthesis of the organic compound urea – a key ingredient in the origin of life – or one of its precursor, the complex inorganic compound ammonia.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

The mass profiles of dwarf galaxies from Dark Energy Survey lensing

We present a novel approach to extracting dwarf galaxies from photometric data to measure their average halo mass profile with weak lensing. We characterize their stellar mass and redshift distributions with a spectroscopic calibration sample. By combining the ${\sim} 5000\,\mathrm{deg}^2$ multiband photometry from the Dark Energy Survey and redshifts from the Satellites Around Galactic Analogs Survey with an unsupervised machine learning method, we select a low-mass galaxy sample spanning redshifts $z\lt 0.3$ and divide it into three mass bins. From low to high median mass, the bins contain [146 420, 330 146, 275 028] galaxies and have median stellar masses of $\log _{10}(M_*/\text{M}_\odot)=\left[8.52\substack{+0.57 -0.76},\, 9.02\substack{+0.50 -0.64},\, 9.49\substack{+0.50 -0.58}\right]$ . We measure the stacked excess surface mass density profiles, $\Delta \Sigma (R)$, of these galaxies using galaxy–galaxy lensing with a signal-to-noise ratio of [14, 23, 28]. Through a simulation-based forward-modelling approach, we fit the measurements to constrain the stellar-to-halo mass relation and find the median halo mass of these samples to be $\log _{10}(M_{\rm halo}/\text{M}_\odot)$ = [$10.67\substack{+0.2 -0.4}$, $11.01\substack{+0.14 -0.27}$, $11.40\substack{+0.08 -0.15}$]. The cold dark matter profiles are consistent with NFW (Navarro, Frenk, and White) profiles over scales ${\lesssim} 0.15 \, {h}^{-1}$ Mpc. We find that ${\sim} 20$ per cent of the dwarf galaxy sample are satellites. This is the first measurement of the halo profiles and masses of such a comprehensive, low-mass galaxy sample. The techniques presented here pave the way for extracting and analysing even lower mass dwarf galaxies and for more finely splitting galaxies by their properties with future photometric and spectroscopic survey data.

dark matter↗

Autonomous organic synthesis for redox flow batteries via flexible batch Bayesian optimization

Traditional trial-and-error methods for materials discovery are inefficient to meet the urgent demands posed by the rapid progression of climate change. This urgency has driven the increasing interest in integrating robotics and machine learning into materials research to accelerate experimental learning. However, idealized decision-making frameworks to achieve maximum sampling efficiency are not always compatible with high-throughput experimental workflows inside a laboratory. For multi-step chemical processes, differences in hardware capacities can complicate the digital framework by introducing constraints on the maximum number of samples in each step of the experiment, hence causing varying batch sizes in variable selection within the same batch. Therefore, designing flexible sampling algorithms is necessary to accommodate the multi-step synthesis with practical constraints unique to each high-throughput workflow. In this work, we designed and employed three strategies on a high-throughput robotic platform to optimize the sulfonation reaction of redox-active molecules used in flow batteries. Our strategies adapt to the multi-step experimental workflow, where their formulation and heating steps are separate, causing varying batch size requirements. By strategically sampling using clustering and mixed-variable batch Bayesian optimization, we were able to iteratively identify optimal conditions that maximize the yields. Our work presents a flexible approach that allows tailoring the machine learning decision-making to suit the practical constraints in individual high-throughput experimental platforms, followed by performing resource-efficient yield optimization using available open-source Python libraries.

Tamura, Clara [Univ. of Washington, Seattle, WA (U↗

Constraining cosmological parameters using the pairwise kinematic Sunyaev-Zel’dovich effect with CMB-S4 and future galaxy cluster surveys

We present a forecast of the pairwise kinematic Sunyaev-Zel’dovich (kSZ) measurement that will be achievable with the future CMB-S4 experiment. CMB-S4 is the next stage for ground-based cosmic microwave background experiments, with a planned wide-area survey that will observe approximately 50% of the sky. We construct a simulated sample of galaxy clusters that have been optically selected in a Legacy Survey of Space and Time–like survey and have spectroscopic redshifts. For this cluster sample, assuming the likelihood is Gaussian, we predict that CMB-S4 will reject the null hypothesis of zero pairwise kSZ signal at 36⁢𝜎. We estimate the effects of systematic uncertainties such as scatter in the mass-richness scaling relation and cluster miscentering. We find that these effects can reduce the signal-to-noise ratio of the CMB-S4 pairwise kSZ measurement by 20%. We explore the constraining power of the measured kSZ signal in combination with measurements of the galaxy clusters’ thermal SZ emission on two extensions to the standard cosmological model. The first extension allows the dark energy equation of state 𝑤 to vary. We find the CMB-S4 pairwise kSZ measurement yields a modest reduction in the uncertainty on 𝑤 by a factor of 1.36 over the Planck’s 2018 uncertainty. The second extension tests general relativity by varying the growth index 𝛾. In conclusion, we find that CMB-S4’s pairwise kSZ measurement will yield a 28⁢𝜎 constraint on 𝛾 and strongly constrain alternative theories of gravity.

79 ASTRONOMY AND ASTROPHYSICS↗

Identifying Missing Quasars from the DESI Bright Galaxy Survey

The Dark Energy Spectroscopic Instrument (DESI) cosmology survey includes a Bright Galaxy Survey (BGS), which will yield spectra for over 10 million bright galaxies (r < 20.2 AB mag). The resulting sample will be valuable for both cosmological and astrophysical studies. However, the star/galaxy separation criterion implemented in the nominal BGS target selection algorithm excludes quasar host galaxies in addition to bona fide stars. While this excluded population is comparatively rare (∼3–4 per square degrees), it may hold interesting clues regarding galaxy and quasar physics. Therefore, we present a target selection strategy that was implemented to recover these missing active galactic nuclei (AGN) from the BGS sample. The design of the selection criteria was both motivated and confirmed using spectroscopy. The resulting BGS-AGN sample is uniformly distributed over the entire DESI footprint. According to DESI survey validation data, the sample comprises 93% quasi-stellar objects (QSOs), 3% narrow-line AGN or blazars with a galaxy contamination rate of 2%, and a stellar contamination rate of 2%. Peaking around redshift z = 0.5, the BGS-AGN sample is intermediary between quasars from the rest of the BGS and those from the DESI QSO sample in terms of redshifts and AGN luminosities. The stacked spectrum is nearly identical to that of the DESI QSO targets, confirming that the sample is dominated by quasars. We highlight interesting small populations reaching z > 2, which are either faint quasars with nearby projected companions or very bright quasars with strong absorption features including the Lyα forest, metal absorbers, and/or broad absorption lines.

79 ASTRONOMY AND ASTROPHYSICS↗

Delamination-informed lifecycle decisions: A dielectric and machine learning framework for composite sorting and recycling

Composite materials are widely used in aerospace, marine, and automotive sectors due to their high strength-to-weight ratio and durability. However, their long-term reliability can be compromised by damage accumulation. Specifically, delamination initiation serves as a precursor to structural failure, which is often difficult to detect during damage inspection. Identifying and sorting delamination initiation in samples not only increases operational safety while providing critical information for end-of-life decisions, which influences both the service life extension value and the efficiency of fiber extraction during recycling. This research addresses two challenges: (1) developing a nondestructive, ex-situ framework to sort composite materials based on damage severity, particularly delamination, and (2) understanding how damage in composites influences resin removal during pyrolysis. Both experimental work and finite element analysis were performed to predict critical stress levels that are associated with delamination onset. Based on these results, three loading levels 50 %, 75 %, and 90 % of maximum stress, were selected for controlled experiments, generating composite samples with varying extents of damage for machine learning model training. Microscopic imaging of these samples confirmed the damage progression from matrix cracking to delamination, validating the computational predictions. We explored supervised machine learning using dielectric measurements to classify damage states. Preliminary results show an artificial neural network can identify early delamination which is a potential precursor to failure, with 94.44 % accuracy on our dataset. A parallel investigation into the effect of damage severity on pyrolysis recycling showed that heavily delaminated samples required significantly less energy for comparable matrix removal than undamaged samples.

dielectric variables↗

Dark Energy Survey Year 6 results: Clustering redshifts and importance sampling of self-organized-maps 𝑛⁡(𝑧) realizations for 3 × 2 ⁢pt samples

This work is part of a series establishing the redshift framework for the 3 × 2 ⁢pt analysis of the Dark Energy Survey Year 6 (DES Y6). For DES Y6, photometric redshift distributions are estimated using self-organizing maps (SOMs), calibrated with spectroscopic and many-band photometric data. To overcome limitations from color-redshift degeneracies and incomplete spectroscopic coverage, we enhance this approach by incorporating clustering-based redshift constraints (clustering-z, or WZ) from angular cross-correlations with BOSS and eBOSS galaxies and eBOSS quasar samples. We define a WZ likelihood and apply importance sampling to a large ensemble of SOM-derived 𝑛⁡(𝑧) realizations, selecting those consistent with the clustering measurements to produce a posterior sample for each lens and source bin. The analysis uses angular scales corresponding to 1.5–5 Mpc to optimize signal-to-noise ratio while mitigating modeling uncertainties and marginalizes over redshift-dependent galaxy bias and other systematics informed by the N-body simulation CARDINAL . While a sparser spectroscopic reference sample limits WZ constraining power at 𝑧 >1.1, particularly for source bins, we demonstrate that combining SOM with WZ improves redshift accuracy and enhances the overall cosmological constraining power of DES Y6. As a result, we estimate an improvement in 𝑆 8 of approximately 10% for cosmic shear and 3 ×2⁢pt analysis, primarily due to the WZ calibration of the source samples.

Cosmological parameters↗

Reconstruction and Selection of Neutrino Interactions in MicroBooNE using Deep Convolutional Neural Networks

In this document, we describe a new reconstruction workflow developed for the MicroBooNE experiment. It features the use of Deep Convolutional Neural Networks trained to recognize key structures within the data sufficient for the 3D reconstruction of neutrino interactions within the detector. As a test of the reconstruction utility, the products of the reconstruction workflow are used to select inclusive charged-current (CC) $\nu_e$ and $\nu_\mu$ interactions in both simulated and real MicroBooNE data. In simulation, our $\nu_e$ and $\nu_\mu$ selections achieve an efficiency of 57% and 68\%, respectively, with a purity of 91% and 96%, respectively. We find that these selections are competitive with the inclusive selections used for the most recent MicroBooNE LEE searches. In particular, the CC-$\nu_e$ inclusive selection efficiency improves by over 20% while also improving sample purity. As a first step in quantifying potential bias, the data and Monte Carlo expectati ons are compared for both selections using the MicroBooNE open data. Within statistical and systematic uncertainties, both the electron and muon CC-inclusive event samples agree. A comparison of the real data events chosen by our work and another reconstruction framework shows that the two analyses each identify a sizeable fraction of events the other does not. This suggests that future analyses integrating the strengths of each could lead to combined gains. This work demonstrates, for the first time on real LArTPC data, state-of-the-art neutrino interaction reconstruction centered around deep learning algorithms.

43 PARTICLE ACCELERATORS↗

Inhomogeneous Distribution and Coarsening of y″ Precipitates in a Ni-Based Superalloy and Their Effect on Creep

It has been reported that the addition of Nb or Ta in the high Ti/Al ratio alloys promotes the formation of y″ precipitates solely and greatly contributes to the elevated temperature strength and resistance to creep deformation. In this study, the y''-dominant INCONEL alloy 725 (IN725) variants (named M725-Nb/Ta) modified with a high level of Nb/Ta additions and high Ti/Al ratio was chosen as a model alloy for microstructural observation of y'' variants upon creep. Following the high temperature aging (HTA), y'' was found to be the main strengthening precipitate with no detectable y’ in both M725 alloys. After creep, the grip/gage sections of the crept M725-Nb/Ta exhibited preferential coarsening and inhomogeneous distribution of y''. That is, one or two of the variants were preferentially coarsened at the expense of another variant, and sandwich-like structures, comprising of cubic y′ precipitates with small y″ discs on each face, were formed. This variant selection behavior taking place in the grip section with no stress applied was different from the typical stress-induced variant selection in previous literature. However, the sample with prior high temperature exposure (700 °C for 500 hours) with varying y″ characteristics demonstrated a similar trend of creep when compared with those in the HTA sample, suggesting that the coarsening and the inhomogeneous distribution of y'' was not the determinant of creep life in bulk M725-Nb/Ta alloys.

Hung, Chang-Yu↗

Full-stack Quantification of Variability in Predicting Ion Transport Properties using Machine-learned Interatomic Potentials

Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is therefore crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods, and improving the MD sampling statistics.

36 MATERIALS SCIENCE↗

Harnessing large language models’ zero-shot and few-shot learning capabilities for regulatory research

Abstract Large language models (LLMs) are sophisticated AI-driven models trained on vast sources of natural language data. They are adept at generating responses that closely mimic human conversational patterns. One of the most notable examples is OpenAI's ChatGPT, which has been extensively used across diverse sectors. Despite their flexibility, a significant challenge arises as most users must transmit their data to the servers of companies operating these models. Utilizing ChatGPT or similar models online may inadvertently expose sensitive information to the risk of data breaches. Therefore, implementing LLMs that are open source and smaller in scale within a secure local network becomes a crucial step for organizations where ensuring data privacy and protection has the highest priority, such as regulatory agencies. As a feasibility evaluation, we implemented a series of open-source LLMs within a regulatory agency’s local network and assessed their performance on specific tasks involving extracting relevant clinical pharmacology information from regulatory drug labels. Our research shows that some models work well in the context of few- or zero-shot learning, achieving performance comparable, or even better than, neural network models that needed thousands of training samples. One of the models was selected to address a real-world issue of finding intrinsic factors that affect drugs' clinical exposure without any training or fine-tuning. In a dataset of over 700 000 sentences, the model showed a 78.5% accuracy rate. Our work pointed to the possibility of implementing open-source LLMs within a secure local network and using these models to perform various natural language processing tasks when large numbers of training examples are unavailable.

Biochemistry & Molecular Biology↗

Demonstration of new MeV-scale capabilities in large neutrino LArTPCs using ambient radiogenic and cosmogenic activity in MicroBooNE

Large neutrino liquid argon time projection chamber (LArTPC) experiments can broaden their physics reach by reconstructing and interpreting MeV-scale energy depositions, or blips, present in their data. We demonstrate new calorimetric and particle discrimination capabilities at the MeV energy scale using reconstructed blips in data from the MicroBooNE LArTPC at Fermilab. We observe a concentration of low-energy (<3 MeV) blips around fiberglass mechanical support struts along the time projection chamber edges with energy spectrum features consistent with the Compton edge of 2.614 MeV 208 Tl decay 𝛾 rays. These features are used to verify proper calibration of electron energy scales in MicroBooNE’s data to few percent precision and to measure the specific activity of 208 Tl in the fiberglass composing these struts, (11.7 ± 0.2⁢(stat) ± 3.1⁢(syst)) Bq/kg. Cosmogenically produced blips above 3 MeV in reconstructed energy are used to showcase the ability of large LArTPCs to distinguish between low-energy proton and electron energy depositions. An enriched sample of low-energy protons selected using this new particle discrimination technique is found to be smaller in data than in dedicated corsika cosmic-ray simulations, suggesting either incorrect corsika modeling of incident cosmic fluxes or particle transport modeling issues in geant4.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The DESI DR1 Peculiar Velocity Survey: Fundamental Plane Catalogue

Measurements of peculiar velocities in the local Universe are a powerful tool to study the nature of dark energy at low ($z < 0.1$) redshifts. Here we present the largest single set of $z<0.1$ peculiar velocity measurements to date, obtained using the Fundamental Plane (FP) of galaxies in the first data release (DR1) of the Dark Energy Spectroscopic Instrument (DESI). We describe the photometric and spectroscopic selection criteria used to define the sample, as well as extensive quality control checks on the photometry and velocity dispersion measurements. Additionally, we perform detailed systematics checks for the many analysis parameters in our pipeline. Our DESI DR1 catalogue contains FP-based distances and peculiar velocities for $98,292$ unique early-type galaxies, increasing the total number of $z < 0.1$ FP distances ever measured by a factor of $\sim2$. We achieve a precision of $26\%$ random error in our distance measurements which is comparable to previous surveys. A series of companion DESI papers use the distances and peculiar velocities presented in this paper to measure cosmological parameters.

Ross, C. E. [Queensland U.]↗

Analysis of Waste Material Feedstocks Using Laser-Induced Breakdown Spectroscopy and Machine Learning

Predicting properties such as heating value, ash fusion temperature, and mineral ash composition from Laser-Induced Breakdown Spectroscopy (LIBS) data can make gasifiers more flexible to different feedstocks. Understanding these feedstock properties in-situ improves feedstock conversion modelling methods that allow for consistent operation, higher carbon conversion, and reduced fouling and erosion rates. The purpose of this study is to demonstrate methods for model creation that take LIBS data as predictor features and estimate higher order material properties as a function of feedstock material properties. Six samples were chosen to represent a mixture of abundant and carbon rich waste materials. LIBS measurements were performed on these samples for elemental wavelengths and intensity values. Laboratory analytical results were obtained for each sample’s heating value, proximate and ultimate analysis, mineral ash composition, ash fusion temperatures, and viscosity temperatures. Thermal conductivity was measured using a HotDisk TPS 2500S. LIBS measurements were processed and used as predictor features for machine learning (ML) models to predict the sample’s material properties. Predictor feature selection algorithms, particularly minimum redundancy maximum relevance (mRMR), reduced the dimensionality of ML models. Many modelling methods such as Gaussian process regression (GPR), regression tree, neural networks (NN), and support vector machines (SVM) were demonstrated to be effective at predicting higher order properties; however, mRMR with GPR stood out as a clear winning combination.

01 COAL, LIGNITE, AND PEAT↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

In real-world applications like spectrum management and interference detection, dealing with unseen electromagnetic waveforms is critical. Although some methods attempt to simulate open set data using generator models, they face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. This results in difficulties capturing distinctive features across classes, especially in dynamic scenarios where new classes emerge. To detect unseen waveforms, we propose combining time and frequency domain features with cosine similarity loss to enhance feature distinctiveness and enabling more accurate predictions. This approach efficiently captures more comprehensive information than single-domain representations or approaches without cosine loss. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10\% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗