Search NASA⌕ Search

SEARCH · Search NASA

Results for “sample complexity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Petrology of Apollo 11 sample 10071 - A differentiated mini-igneous complex.

Sample 10071,33 is a thin section of Apollo 11 ferrobasalt showing an unusual dual texture. The major portion of the sample is very similar to other fine grained Apollo 11 basalts, but the thin section also includes materials with a distinct variolitic texture. The two areas are separated by a sharp boundary and the mineralogy and composition of the two textural types are quite distinct. The mineralogy and chemistry of the variolitic portion show it to be the product of rapid cooling of a liquid, intermediate between the typical Apollo 11 ferrobasalt and the associated Si and K-rich mesostasis. This liquid is the result of fractional crystallization of a magma of composition closely corresponding to the major portion of the 10071 system, followed by crystal-liquid separation. The sample provides strong and direct evidence for igneous differentiation on the lunar surface.

Drake, M. J.↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Microgravity Testing of a Surface Sampling System for Sample Return from Small Solar System Bodies

The return of samples from solar system bodies is becoming an essential element of solar system exploration. The recent National Research Council Solar System Exploration Decadal Survey identified six sample return missions as high priority missions: South-Aitken Basin Sample Return, Comet Surface Sample Return, Comet Surface Sample Return-sample from selected surface sites, Asteroid Lander/Rover/Sample Return, Comet Nucleus Sample Return-cold samples from depth, and Mars Sample Return [1] and the NASA Roadmap also includes sample return missions [2] . Sample collection methods that have been flown on robotic spacecraft to date return subgram quantities, but many scientific issues (like bulk composition, particle size distributions, petrology, chronology) require tens to hundreds of grams of sample. Many complex sample collection devices have been proposed, however, small robotic missions require simplicity. We present here the results of experiments done with a simple but innovative collection system for sample return from small solar system bodies.

Franzen, M. A.↗

Real classical shadows

Efficiently learning expectation values of a quantum state using classical shadow tomography has become a fundamental task in quantum information theory. In a classical shadows protocol, one measures a state in a chosen basis $\mathcal{W}$ after it has evolved under a unitary transformation randomly sampled from a chosen distribution $\mathcal{U}$. In this work we study the case where $\mathcal{U}$ corresponds to either local or global orthogonal Clifford gates, and $\mathcal{W}$ consists of real-valued vectors. Our results show that for various situations of interest, this ‘real’ classical shadow protocol improves the sample complexity over the standard scheme based on general Clifford unitaries. For example, when one is interested in estimating the expectation values of arbitrary real-valued observables, global orthogonal Cliffords typically decrease the required number of samples by a factor of two. More dramatically, for k-local observables composed only of real-valued Pauli operators, sampling local orthogonal Cliffords leads to a reduction by an exponential-in-k factor in the sample complexity over local unitary Cliffords. Finally, we show that by measuring in a basis containing complex-valued vectors, orthogonal shadows can, in the limit of large system size, exactly reproduce the original unitary shadows protocol.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Optimal Twirling Depth for Classical Shadows in the Presence of Noise

The classical shadows protocol is an efficient strategy for estimating properties of an unknown state p using a small number of state copies and measurements. In its original form, it involves twirling the state with unitaries from some ensemble and measuring the twirled state in a fixed basis. It was recently shown that for computing local properties, optimal sample complexity (copies of the state required) is remarkably achieved for unitaries drawn from shallow depth circuits composed of local entangling gates, as opposed to purely local (zero depth) or global twirling (infinite depth) ensembles. Here, we consider the sample complexity as a function of the depth of the circuit, in the presence of noise. We find that this noise has important implications for determining the optimal twirling ensemble. Under fairly general conditions, we (i) show that any single-site noise can be accounted for using a depolarizing noise channel with an appropriate damping parameter f, (ii) compute thresholds f th at which optimal twirling reduces to local twirling for Pauli operators, (iii) nth order Renyi entropies (n ≥2), and (iv) provide a meaningful upper bound t max on the optimal circuit depth for any finite noise strength f, which applies to observables and entanglement entropy measurements. In conclusion, these thresholds strongly constrain the search for optimal strategies to implement shadow tomography and are easily tailored to the experimental system at hand.

97 MATHEMATICS AND COMPUTING↗

The Average Spectrum Norm and Near-Optimal Tensor Completion

We propose the average spectrum norm to study the minimum number of measurements required to approximate a multidimensional array (i.e., sample complexity) via low-rank tensor recovery. Our focus is on the tensor completion problem, where the aim is to estimate a multiway array using a subset of tensor entries corrupted by noise. Our average spectrum norm-based analysis provides near-optimal sample complexities, exhibiting dependence on the ambient dimensions and rank that do not suffer from exponential scaling as the order increases.

97 MATHEMATICS AND COMPUTING↗

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum↗

3D Geologic Framework Modelling of the Los Alamos National Laboratory Site and Pajarito Plateau: Integrating a realistic 3D fault network and modelling subsurface relationships in a sparsely sampled and complex geologic region

The subsurface geology beneath the Pajarito Plateau is critical to understanding the seismic hazard of the Pajarito Fault System, yet our understanding of this geology is relatively poor. While previous 3D geologic framework models of the area have been created for the purposes of understanding hydrogeologic flow, they are inadequate for the purposes of understanding the Pajarito Fault System. The specific challenges of using oil and gas software for this purpose include: (1) the geologic complexities resulting from volcanism and tectonism; (2) a need for a high level of stratigraphic detail over a large area; (3) a near complete lack of seismic data; and (4) sparse wellbore data. Presented here is a workflow that handles these challenges of adapting commercially available software used by the oil and gas industries to this seismic hazard problem.

58 GEOSCIENCES↗

Evaluation of a Reference-Free Collision Cross Section Calibration Strategy for Proteomics Using SLIM-Based High-Resolution Ion Mobility Spectrometry–Mass Spectrometry

Ion mobility spectrometry (IMS) is a gas-phase analytical technique that separates ions with different sizes and shapes and is compatible with mass spectrometry (MS) to provide an additional separation dimension. The rapid nature of the IMS separation combined with the high sensitivity of MS-based detection and the ability to derive structural information on analytes in the form of the property collision cross section (CCS) makes IMS particularly well-suited for characterizing complex samples in -omics applications. In such applications, the quality of CCS from IMS measurements is critical to confident annotation of the detected components in the complex -omics samples. However, most IMS instrumentation in mainstream use requires calibration to calculate CCS from measured arrival times, with the most notable exception being drift tube IMS measurements using multifield methods. The strategy for calibrating CCS values, particularly selection of appropriate calibrants, has important implications for CCS accuracy, reproducibility, and transferability between laboratories. The conventional approach to CCS calibration involves explicitly defining calibrants ahead of data acquisition and crucially relies upon availability of reference CCS values. In this work, we present a novel reference-free approach to CCS calibration which leverages trends among putatively identified features and computational CCS prediction to conduct calibrations post-data acquisition and without relying on explicitly defined calibrants. We demonstrated the utility of this reference-free CCS calibration strategy for proteomics application using high-resolution structures for lossless ion manipulations (SLIM)-based IMS-MS. In conclusion, we first validated the accuracy of CCS values using a set of synthetic peptides and then demonstrated using a complex peptide sample from cell lysate.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying Sample Provenance From SEM/EDS Automated Particle Analysis via Few-Shot Learning Coupled With Similarity Graph Clustering

Automated particle analysis (APA) provides a vast amount of compositional data via energy-dispersive X-ray spectroscopy along with size and shape data via scanning electron microscopy for individual particles in a sample. In many instances, APA data are leveraged to support identification of the source of a sample based on the detection of particles of a specific composition. Often, the particles that provide context make up a minuscule portion of the sample. Additionally, the interpretation of complex samples can be difficult due to the diversity of compositions both in the mixture and within a particle. In this work, we demonstrate a method to compute and cluster similarity graphs that describe inter-particle relationships within a sample using a multi-modal few-shot learning neural network. Here, as a proof-of-concept, we show that samples known to have been exposed to gunshot residue can be distinguished from samples occasionally mistaken for gunshot residue. Our workflow builds upon standard APA techniques and data processing methods to unveil additional information in a readily interpretable and quantitatively comparable format.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Covariance operator estimation: Sparsity, lengthscale, and ensemble Kalman filters

This paper investigates covariance operator estimation via thresholding. For Gaussian random fields with approximately sparse covariance operators, we establish non-asymptotic bounds on the estimation error in terms of the sparsity level of the covariance and the expected supremum of the field. We prove that thresholded estimators enjoy an exponential improvement in sample complexity compared with the standard sample covariance estimator if the field has a small correlation lengthscale. As an application of the theory, we study thresholded estimation of covariance operators within ensemble Kalman filters.

Covariance operator estimation↗

High-resolution ion mobility based on traveling wave structures for lossless ion manipulation resolves hidden lipid features

Abstract High-resolution ion mobility (resolving power > 200) coupled with mass spectrometry (MS) is a powerful analytical tool for resolving isobars and isomers in complex samples. High-resolution ion mobility is capable of discerning additional structurally distinct features, which are not observed with conventional resolving power ion mobility (IM, resolving power ~ 50) techniques such as traveling wave IM and drift tube ion mobility (DTIM). DTIM in particular is considered to be the “gold standard” IM technique since collision cross section (CCS) values are directly obtained through a first-principles relationship, whereas traveling wave IM techniques require an additional calibration strategy to determine accurate CCS values. In this study, we aim to evaluate the separation capabilities of a traveling wave ion mobility structures for lossless ion manipulation platform integrated with mass spectrometry analysis (SLIM IM-MS) for both lipid isomer standards and complex lipid samples. A cross-platform investigation of seven subclass-specific lipid extracts examined by both DTIM-MS and SLIM IM-MS showed additional features were observed for all lipid extracts when examined under high resolving power IM conditions, with the number of CCS-aligned features that resolve into additional peaks from DTIM-MS to SLIM IM-MS analysis varying between 5 and 50%, depending on the specific lipid sub-class investigated. Lipid CCS values are obtained from SLIM IM ( TW(SLIM) CCS) through a two-step calibration procedure to align these measurements to within 2% average bias to reference values obtained via DTIM ( DT CCS). A total of 225 lipid features from seven lipid extracts are subsequently identified in the high resolving power IM analysis by a combination of accurate mass-to-charge, CCS, retention time, and linear mobility-mass correlations to curate a high-resolution IM lipid structural atlas. These results emphasize the high isomeric complexity present in lipidomic samples and underscore the need for multiple analytical stages of separation operated at high resolution. Graphical abstract

Reardon, Allison R. (ORCID:0000000165830134)↗

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning Functions Varying along a Central Subspace

Many functions of interest are in a high-dimensional space but exhibit low-dimensional structures. This paper studies regression of an s-Hölder function in $R^D$ which varies along a central subspace of dimension $d$ while $d \ll D$. A direct approximation of $f$ in $R^D$ with an accuracy $\varepsilon$ requires the number of samples in the order of $\varepsilon^{-(2s+D)/s}$. In this paper, we analyze the generalized contour regression (GCR) algorithm for the estimation of the central subspace and use piecewise polynomials for function approximation. GCR is among the best estimators for the central subspace, but its sample complexity is an open question. In this paper, we partially answer this questions by proving that if a variance quantity is exactly known, GCR leads to a mean squared estimation error of $O(n^{-1})$ for the central subspace. The estimation error of this variance quantity is also given in this paper. The mean squared regression error of $f$ is proved to be in the order of $(n/\log n)^{-\frac{2s}{2s+d}}$, where the exponent depends on the dimension of the central subspace instead of the ambient space . This result demonstrates that GCR is effective in learning the low-dimensional central subspace. We also propose a modified GCR with improved efficiency. Here, the convergence rate is validated through several numerical experiments.

97 MATHEMATICS AND COMPUTING↗

A Smart Vision-Aided RICH (Robotic Interface Control and Handling) System for VULCAN

High-flux neutron beams and high-efficiency detectors enable rapid neutron diffraction measurements at the Engineering Materials Diffractometer (VULCAN) at the Spallation Neutron Source (SNS), Oak Ridge National Laboratory (ORNL). To optimize beam time utilization, efficient sample exchange, alignment, and automated measurements are essential. Recent advances in artificial intelligence (AI) have expanded the capabilities of robotic systems. Here, we report the development of a Robotic Interactive Control and Handling (RICH) system for sample handling at VULCAN, designed to support high-throughput experiments and reduce overhead time. The RICH system employs a six-axis desktop robot integrated with AI-based computer vision models capable of recognizing and localizing samples in real time from instrument and depth-resolving cameras. Vision algorithms combine these detections to align samples with designated measurement positions or place them within complex sample environments such as furnaces. This integration of machine learning-assisted vision with robotic handling demonstrates the feasibility of autonomous sample detection and preparation, offering a pathway toward fully unmanned neutron scattering experiments.

automation↗

Testing Classical Properties from Quantum Data

Many properties of Boolean functions can be tested far more efficiently than the function itself can be learned. However, this dramatic advantage often disappears when testers are limited to random samples of ƒ instead of adaptively chosen queries to f. In this work we investigate the quantum version of this restriction: quantum algorithms that test properties of a Boolean function f solely from copies of either the function state |ƒ⟩ ∝ ∑ x |x, ƒ(x)⟩ or the phase state |(-1) ƒ ⟩ ∝ ∑ x (-1) ƒ(x) |x⟩. For monotonicity, symmetry, and triangle-freeness, we show passive quantum testers are unboundedly or super-polynomially better than their classical passive testing counterparts. They are competitive with classic query -based testers in each case. Our new testers use techniques beyond quantum Fourier sampling, and it turns out this is necessary: we show a certain class of bent functions can be tested from 𝒪(1) function states but has a sample complexity lower bound of 2 Ω(n) for any tester relying exclusively on Fourier and classical samples. Our passive quantum testers are competitive with classical query -based testers, but this isn't universal: we exhibit a testing problem that can be solved from 𝒪(1) classical queries but requires Ω(2 n/2 ) function state copies. The Forrelation problem provides a separation of the same magnitude in the opposite direction, so we conclude that quantum data and classical queries are "maximally incomparable" resources for testing. We also begin the study of lower bounds for testing from quantum data. For quantum monotonicity testing, we prove that the ensembles of [Goldreich et al., 2000; Black, 2024], which give exponential lower bounds for classical sample-based testing, do not yield any nontrivial lower bounds for testing from quantum data. New insights specific to quantum data will be required for proving copy complexity lower bounds for testing in this model.

Boolean Functions↗

In Situ Identification of Mineral Resources with an X-Ray-Optical "Hands-Lens" Instrument

The recognition of material resources on a planetary surface requires exploration strategies not dissimilar to those employed by early field geologists who searched for ore deposits primarily from surface clues. In order to determine the location of mineral ores or other materials, it will be necessary to characterize host terranes at regional or subregional scales. This requires geographically broad surveys in which statistically significant numbers of samples are rapidly scanned from a roving platform. To enable broad-scale, yet power-conservative planetary-surface exploration, we are developing an instrument that combines x-ray diffractometry (XRD), x-ray fluorescence spectrometry (XRF), and optical capabilities; the instrument can be deployed at the end of a rover's robotic arm, without the need for sample capture or preparation. The instrument provides XRD data for identification of mineral species and lithological types; diffractometry of minerals is conducted by ascertaining the characteristic lattice parameters or "d-spacings" of mineral compounds. D-spacings of 1.4 to 25 angstroms can be determined to include the large molecular structures of hydrated minerals such as clays. The XRF data will identify elements ranging from carbon (Atomic Number = 6) to elements as heavy as barium (Atomic Number = 56). While a sample is being x-rayed, the instrument simultaneously acquires an optical image of the sample surface at magnifications from lx to at least 50x (200x being feasible, depending on the sample surface). We believe that imaging the sample is extremely important as corroborative sample-identification data (the need for this capability having been illustrated by the experience of the Pathfinder rover). Very few geologists would rely on instrument data for sample identification without having seen the sample. Visual inspection provides critical recognition data such as texture, crystallinity, granularity, porosity, vesicularity, color, lustre, opacity, and so forth. These data can immediately distinguish sedimentary from igneous rocks, for example, and can thus eliminate geochemical or mineral ambiguities arising, say between arkose and granite. It would be important to know if the clay being analyzed was part of a uniform varve deposit laid down in a quiescent lake, or the matrix of a megabreccia diamictite deposited as a catastrophic impact ejecta blanket. The unique design of the instrument, which combines Debye-Scherrer geometry with elements of standard goniometry, negates the need for sample preparation of any kind, and thus negates the need for power-hungry and mechanically-complex sampling systems that would have to chip, crush, sieve, and mount the sample for x-ray analysis. Instead, the instrument is simply rested on the sample surface of interest (like a hand lens); the device can interrogate rough rock surfaces, coarse granular material, or fine rock flour. A breadboard version of the instrument has been deployed from the robotic arm of the Marsokhod rover in field trials at NASA Ames, where large vesicular boulders were x-rayed to demonstrate the functionality of the instrument design, and the ability of such a device to comply with constraints imposed by a roving platform. Currently under development is a flight prototype concept of this instrument that will weigh 0.3 kg, using about 4500 J of energy per sample analysis. It requires about 5 min. for XRD analysis, and about 30 min. for XRF interrogation. Its small mass and rugged design make it ideal for deployment on small rovers of the type currently envisaged for the exploration of Mars (e.g., Sojourner-scale platforms). The design utilizes a monolithic P-N junction photodiode pixel array for XRD, a Si PIN photodiode/avalanche photodiode system for XRF, and an endoscopic imaging camera system unobtrusively embedded between the detectors and the x-ray source (the endoscope with its board-mounted camera can be adapted for IR light in addition to visible wavelenths. A rugged, miniature (7 cu cm) x-ray source for the instrument has already been breadboarded.

Marshall, J.↗