Search NASA⌕ Search

SEARCH · Search NASA

Results for “RANDOM SAMPLE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Lessons from 18 Years of Hyperspectral Infrared Sounder Data

By the end of 2013 NASA and EUMETSAT will have accumulated more than 11 years of AIRS, 6 years of IASI and one year of CrIS data. All three instruments were nominally specified to support the NWC for short term weather forecasting with a five year lifetime, but continue to exceed the accuracy requirement needed for weather forecasting alone. This allows use of their data for a much broader range of applications, including the calibration of broad-band instruments in space and climate research. We illustrate calibration aspects with examples from AIRS, IASI and CrIS using spatially uniform clear conditions, simultaneous nadir overpasses and random nadir samples. The differences between AIRS, IASI and CrIS for the purpose of weather forecasting are small and we expect that the excellent forecast impact demonstrated by the combination of AIRS and IASI will be continued by the combination of CrIS and IASI. Clear data are useful for calibration, but contain no climate signal. The analysis of random nadir samples from AIRS and CrIS identifies larger biases for observation of extreme conditions, represented by 1% and 99%tile data than for non-extreme observations. This is relevant for climate analysis. Resolution of these differences require further work, since they can complicate the continuation of trends established by AIRS with CrIS data, at least for extrema. The unequaled stability of the AIRS data allows us to evaluate trends using random nadir sampled data. We see an increasing frequency in severe storms over land, a decreasing frequency over ocean. The 11 years of AIRS data are too short to tell if these trends are significant from a climate change viewpoint, or if they are parts of multi-decadal oscillations.

CRIS↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Randomized Adiabatic Quantum Linear Solver Algorithm with Optimal Complexity Scaling and Detailed Running Costs

Solving linear systems of equations is a fundamental problem with a wide variety of applications across many fields of science, and there is increasing effort to develop quantum linear solver algorithms. Subaşı et al. [Phys. Rev. Lett. 122, 060504 (2019)] proposed a randomized algorithm inspired by adiabatic quantum computing, based on a sequence of random Hamiltonian simulation steps, with suboptimal scaling in the condition number 𝜅 of the linear system and the target error 𝜖. Here we go beyond these results in several ways. Firstly, using filtering [Lin and Tong, Quantum 4, 361 (2020)] and Poissonization techniques [Cunningham and Roland, ArXiv:2406.03972 (2024)], the algorithm complexity is improved to the optimal scaling 𝑂⁡(𝜅⁢log (1/𝜖))—an exponential improvement in 𝜖, and a shaving of a log 𝜅 scaling factor in 𝜅. Secondly, the algorithm is further modified to achieve constant factor improvements, which are vital as we progress towards hardware implementations on fault-tolerant devices. We introduce a cheaper randomized walk operator method replacing Hamiltonian simulation—which also removes the need for potentially challenging classical precomputations; randomized routines are sampled over optimized random variables; circuit constructions are improved. We obtain a closed formula rigorously upper bounding the expected number of times one needs to apply a block-encoding of the linear system matrix to output a quantum state encoding the solution to the linear system. The upper bound is 837⁢𝜅 at 𝜖 = 10 −10 for Hermitian matrices.

97 MATHEMATICS AND COMPUTING↗

Model-based quantification of image quality

In 1982, Park and Schowengerdt published an end-to-end analysis of a digital imaging system quantifying three principal degradation components: (1) image blur - blurring caused by the acquisition system, (2) aliasing - caused by insufficient sampling, and (3) reconstruction blur - blurring caused by the imperfect interpolative reconstruction. This analysis, which measures degradation as the square of the radiometric error, includes the sample-scene phase as an explicit random parameter and characterizes the image degradation caused by imperfect acquisition and reconstruction together with the effects of undersampling and random sample-scene phases. In a recent paper Mitchell and Netravelli displayed the visual effects of the above mentioned degradations and presented subjective analysis about their relative importance in determining image quality. The primary aim of the research is to use the analysis of Park and Schowengerdt to correlate their mathematical criteria for measuring image degradations with subjective visual criteria. Insight gained from this research can be exploited in the end-to-end design of optical systems, so that system parameters (transfer functions of the acquisition and display systems) can be designed relative to each other, to obtain the best possible results using quantitative measurements.

Hazra, Rajeeb↗

Basin-Size Mapping: Prediction of Metastable Polymorph Synthesizability Across TaC–TaN Alloys

The sizes of the basins of attraction on the potential energy surface are helpful indicators in determining the experimental synthesizability of metastable phases. In principle, these basins can be controlled with changes in thermodynamic conditions such as composition, pressure, and surface energy. Herein, we use random structure sampling to computationally study how alloying smoothly perturbs basin of attraction sizes. The TaC 1-x N x pseudobinary is an ideal test system given the structural and polymorphic contrast of its parent compounds and their technological relevance as epitaxial substrates for Al 1-x Ga x N. While we find limited thermodynamic stability across all computationally observed phases, random structure sampling shows a significant composition region where the rocksalt basin dominates. As such, we predict the potential for the nonequilibrium synthesis of metastable rocksalt TaC 1-x N x alloys as substrates for Al 1-x Ga x N. At higher nitrogen concentrations, other low-energy metastable polymorphs emerge that continue to retain the hexagonal close packing suitable for III-N growth. Confidence in these trends was established through uncertainty quantification of the basin sizes and energy distributions; such analysis utilized the Beta and Dirichlet distributions. In conclusion, we also find (a) polymorph basin sizes can be rationalized in terms of energetic preferences for different coordination environments; and (b) basin sizes universally shrink with increasing nitrogen content, making the system more prone to amorphous growth.

36 MATERIALS SCIENCE↗

X-Ray Diffraction Reference Intensity Ratios of Amorphous and Poorly Crystalline Phases: Implications for CheMin on the Mars Science Laboratory

The CheMin instrument on the Mars Science Laboratory (MSL) rover Curiosity is an X-ray diffraction (XRD) and X-ray fluorescence (XRF) instrument capable of providing the mineralogical and chemical compositions of rocks and soils on the surface of Mars. CheMin uses a microfocus X-ray tube with a Co target, transmission geometry, and an energy-discriminating X-ray sensitive CCD to produce simultaneous 2-D XRD patterns and energy-dispersive X-ray histograms from powdered samples. Piezoelectric vibration of the cell is used to randomize the sample to reduce preferred orientation effects. Instrument details are provided in [1, 2, 3]. Analyses of rock and soil samples by the Mars Exploration Rovers (MER) show nanophase ferric oxide (npOx) is a significant component of the Martian global soil [4] and is thought to be one of the major contributing phases that the Curiosity rover will encounter if a soil sample is analyzed in Gale Crater. Because of the nature of this material, npOx will likely contribute to an X-ray amorphous or short-order component of a XRD pattern measured by the CheMin instrument.

Morris, R. V.↗

Certified randomness using a trapped-ion quantum processor

Although quantum computers can perform a wide range of practically important tasks beyond the abilities of classical computers, realizing this potential remains a challenge. An example is to use an untrusted remote device to generate random bits that can be certified to contain a certain amount of entropy. Certified randomness has many applications but is impossible to achieve solely by classical computation. Here we demonstrate the generation of certifiably random bits using the 56-qubit Quantinuum H2-1 trapped-ion quantum computer accessed over the Internet. Our protocol leverages the classical hardness of recent random circuit sampling demonstrations: a client generates quantum ‘challenge’ circuits using a small randomness seed, sends them to an untrusted quantum server to execute and verifies the results of the server. We analyse the security of our protocol against a restricted class of realistic near-term adversaries. Using classical verification with measured combined sustained performance of 1.1 × 10 18 floating-point operations per second across multiple supercomputers, we certify 71,313 bits of entropy under this restricted adversary and additional assumptions. Our results demonstrate a step towards the practical applicability of present-day quantum computers.

computer science↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

A systematic analytical framework for multi-source municipal solid waste characterization for energy recovery

Advancing municipal solid waste (MSW) management from disposal-oriented practices toward circular, value-driven systems requires standardized methodologies capable of identifying material composition and resource recoverable potential at the point of generation. Despite extensive research, MSW characterization remains fragmented due to inconsistences in sampling methodologies, waste sorting categories, and temporal coverage across previous studies which limit cross-site comparability, reproducibility, and constrain the reliable evaluation of potential resource recovery pathways. This lack of consistency has hindered the development of a unified framework for MSW characterization and resource assessment. This study introduces a standardized, field-validated protocol for MSW sampling and composition analysis that ensures consistent, traceable data across diverse waste sources. The protocol integrates randomized spatial sampling, systematic material sorting, and controlled subsampling for multi-site and multi-season field campaigns. Validation included MSW collection from residential, grocery, restaurant, and school MSW streams across five U.S. states, including Maryland, Idaho, Virginia, Ohio, and Mississippi, to demonstrate the protocol’s ability to identify source-based composition patterns relevant to resource recovery applications. Grocery and restaurant streams were dominated by food waste and high-moisture organics, while school waste contained higher paper content and residential waste showed greater heterogeneity. Aggregation into energy-relevant fractions highlighted practical recovery pathways via anaerobic digestion or gasification, supporting data-driven planning, policy, and circular economy strategies for sustainable waste management across waste sources.

09 BIOMASS FUELS↗

Masked Symbol Modeling for Demodulation of Oversampled Baseband Communication Signals in Impulsive Noise-Dominated Channels

Recent breakthroughs in natural language processing show that attention mech- anism in Transformer networks, trained via masked-token prediction, enables models to capture the semantic context of the tokens and internalize the grammar of language. While the application of Transformers to communication systems is a burgeoning field, the notion of context within physical waveforms remains under-explored. This paper addresses that gap by re-examining inter-symbol con- tribution (ISC) caused by pulse-shaping overlap. Rather than treating ISC as a nuisance, we view it as a deterministic source of contextual information embedded in oversampled complex baseband signals. We propose Masked Symbol Model- ing (MSM), a framework for the physical (PHY) layer inspired by Bidirectional Encoder Representations from Transformers methodology. In MSM, a subset of symbol-aligned samples is randomly masked, and a Transformer predicts the missing symbol identifiers using the surrounding “in-between” samples. Through this objective, the model learns the latent syntax of complex baseband waveforms. We illustrate MSM’s potential by applying it to the task of demodulating sig- nals corrupted by impulsive noise, where the model infers corrupted segments by leveraging the learned context. Our results suggest a path toward receivers that interpret, rather than merely detect communication signals, opening new avenues for context-aware PHY layer design.

Bedir, Oguz↗

Statistical methods for efficient design of community surveys of response to noise: Random coefficients regression models

Research studies of residents' responses to noise consist of interviews with samples of individuals who are drawn from a number of different compact study areas. The statistical techniques developed provide a basis for those sample design decisions. These techniques are suitable for a wide range of sample survey applications. A sample may consist of a random sample of residents selected from a sample of compact study areas, or in a more complex design, of a sample of residents selected from a sample of larger areas (e.g., cities). The techniques may be applied to estimates of the effects on annoyance of noise level, numbers of noise events, the time-of-day of the events, ambient noise levels, or other factors. Methods are provided for determining, in advance, how accurately these effects can be estimated for different sample sizes and study designs. Using a simple cost function, they also provide for optimum allocation of the sample across the stages of the design for estimating these effects. These techniques are developed via a regression model in which the regression coefficients are assumed to be random, with components of variance associated with the various stages of a multi-stage sample design.

Tomberlin, T. J.↗

Investigation of electrolyte measurement in diluted whole blood using spectroscopic and chemometric methods

The feasibility of using near-infrared (NIR) spectroscopy in combination with partial least-squares (PLS) regression was explored to measure electrolyte concentration in whole blood samples. Spectra were collected from diluted blood samples containing randomized, clinically relevant concentrations of Na+, K+, and Ca2+. Sodium was also studied in lysed blood. Reference measurements were made from the same samples using a standard clinical chemistry instrument. Partial least squares (PLS) was used to develop calibration models for each ion with acceptable results (Na+, R2 = 0.86, CVSEP = 9.5 mmol/L; K+, R2 = 0.54, CVSEP = 1.4 mmol/L; Ca2+, R2 = 0.56, CVSEP = 0.18 mmol/L). Slightly improved results were obtained using a narrower wavelength region (470-925 nm) where hemoglobin, but not water, absorbed indicating that ionic interaction with hemoglobin is as effective as water in causing measurable spectral variation. Good models were also achieved for sodium in lysed blood, illustrating that cell swelling, which is correlated with sodium concentration, is not required for calibration model development.

Non-NASA Center↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Remanent magnetization of lunar samples.

The remanent magnetization of samples returned from the moon by the Apollo 11 and 12 missions consists, in most cases, of two distinct components. An unstable component is readily removed upon alternating field (AF) demagnetization in fields less than 100 Oe and is considered to be an isothermal remanence acquired during or after return to earth. The second component is unaltered by demagnetization in fields up to 400 Oe. It is probably a thermoremanent magnetization due to cooling from above 800 C in the presence of a field of a few thousand gammas. Chips from individual rocks have the same direction of magnetization after demagnetization, while the directions of different samples are random. This again demonstrates the high stability. Our data imply that the moon experienced a magnetic field that lasted at least from about 3.0 to 3.8 b.y., which is the age of Apollo 11 and 12 samples. One explanation of the origin of this field is that the moon had a liquid core and a self-exciting dynamo early in its history.

Strangway, D. W.↗

Three-epoch VLBI observations of the nucleus in the lobe-dominated quasar 3C 334

VLBI observations at three epochs spanning 7 yr. reveals possible superluminal motion of the nucleus of the lobe-dominated quasar 3C 334. This is the third lobe-dominated quasar confirmed or suspected to be superluminal in a complete, flux density-limited sample of these objects, and the sixth including other surveys. The result continues the trend for lobe-dominated quasars to have lower superluminal speeds than core-dominated quasars, and coupled with other radio properties that are indicators of source orientation, it is consistent with the simple relativistic beaming model in a randomly oriented sample.

Hough, D. H.↗