Search NASA⌕ Search

SEARCH · Search NASA

Results for “generalization error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

418 records · Page 24

Spectroscopic Measurements and Models of Energy Deposition in the Substrate of Quantum Circuits by Natural Ionizing Radiation

Naturally occurring background radiation is a potential source of correlated decoherence events in superconducting qubits that will challenge error-correction schemes. In order to characterize the radiation environment in an unshielded laboratory representative of superconducting qubits’ environments, we performed broadband, spectroscopic measurements of background radiation events inside a millikelvin refrigerator. The spectrometer was designed to mimic the size and composition of a quantum circuit. Specifically, we measured the background radiation spectra in silicon substrates of two thicknesses, 500 and 1500 µm, and one area, 25 mm 2 . The observed spectra span energies from a few kilo-electron-volts up to nearly 10 MeV, are nearly featureless, and decrease in intensity by a factor of 40 000 between 100 keV and 3 MeV for the 500-µm substrate. We integrate the spectra to obtain the average event rates and deposited power levels. These quantities correspond to a rate of 0.023 events per second and a power of 4.9 keV s -1 , when counting events that deposit at least 40 keV for the 500-µm-thick substrate. We find that the cryogenic measurements are in good agreement with predictions based on simple measurements of the terrestrial gamma-ray flux outside the refrigerator, published models of cosmic-ray fluxes, a crude model of the cryostat, and radiation-transport simulations. This model requires no free parameters to predict the background radiation spectra in the silicon substrates. The agreement between measurements and predictions demonstrates that the model we present can be used to assess the relative contributions of terrestrial and cosmic-ray sources to background radiation interactions in silicon substrates of varying thickness. These spectroscopic measurements are performed with a novel combination of superconducting microresonators located on micromachined silicon islands that define the interaction volume with background radiation. The resonators transduce deposited energy to a readily detectable electrical signal. Microresonator readout closely resembles dispersive superconducting qubit readout, so similar devices—with or without micromachined islands—are suitable for integration with superconducting quantum circuits as detectors for background radiation events. For our specific laboratory conditions, we find that gamma-ray emissions from radioisotopes are responsible for the majority of events that deposit E < 1 ⁢Me⁢V. We present results demonstrating that the background radiation spectrum contains relevant contributions from cosmic-ray particles other than muons, particularly a tail of multi-mega-electron-volt events due to protons and neutrons. These observations suggest several paths to reducing the impact of background radiation on quantum circuits, supported by an empirically validated model for generating reliable predictions of radiation interactions with silicon substrates.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sampling Size Optimization for Bioburden Density Estimation in Planetary Protection

Planetary protection (PP) is a discipline that focuses on minimizing the biological contamination of spacecraft to ensure compliance with international policy. Precise estimation of bioburden - the total number of microbes in or on spacecraft hardware – and the bioburden density are of utmost importance for PP. Such estimation is the way concordance with requirements is demonstrated, and it is critical for quantifying the potential risk of inadvertently contaminating other planetary bodies. Although a suite of molecular techniques have been used to thoroughly characterize and profile the microbiome of various cleanroom environments and spacecraft, the gold standard remains the physical enumeration of microbes via culturing of samples directly taken from spacecraft and associated surfaces. However, due to technical, budgetary, and programmatic constraints, only a manageable portion (around 10%) of the entire spacecraft surface is directly sampled with cotton swabs or wipes. To generate the bioburden current best estimate (CBE) for components not directly verifiable, the accepted approach is to apply a NASA-defined bioburden estimate based on the components’ manufacturing or assembly environment. This approach utilizes a prespecified bioburden density estimation that applies a maximum value across the total surface area of the specified component. For hardware components that underwent similar assembly processes, an implied bioburden is adopted for all components, based on a direct verification of a representative component within the same lot. Once all components have a CBE, the bioburden estimates are generated. In previous publication [ 1], we have shown that statistical risks quantifying the accuracy of the estimates for sampled, prespecified, and implied components can be derived and ranked. For mean squared error (MSE) function, the risks are available analytically and hence a cost function can be obtained to optimize the risks with respect to the sampling area and sampling cost. Since the sampling area and sampling cost are two complimentary variables, their sum will have a well-defined minimum. This paper presents the multivariate optimization of the integrated risk of an empirical Bayes estimator to determine the optimal sampling schedule for a given number of components. It is assumed that given a number of components, N, the bioburden density for each component can either be sampled, implied, or prespecified. The multivariate optimization searches through different options to sample, imply or prespecify the bioburden density for a component, and account for the component’s surface area and cost of sampling. The idea of the optimization is based on the observation that the statistical risk of using an estimator is a monotonically decreasing function of the sampled area. The larger the sampled area, the lower the risk of using the estimator as the estimator becomes more and more accurate as the sampling area increases. On the other hand, the cost of sampling is monotonically increasing as the sampled surface grows. This makes the risk and total cost of sampling complimentary variables which can be counterbalanced to achieve an optimal overall value with respect to the sampled surface. In this paper, the integrated risk has been used to quantify the accuracy of the estimator. This risk has been selected because it depends on neither the true value of the parameter nor on the collected data. The cost of each sample was also available to obtain the total cost of sampling of N components. The paper will present the results based on computer-simulated data as well as the data collected during the InSight mission. The computer-simulated data have N components with randomly generated total areas and each component assigned to one of the three categories according to the method of estimating of bioburden density: sampled, implied, or prespecified. The cost of sampling is also available. The cost of sampling is estimated based on a cost model provided by the planetary protection group at JPL. For this paper, the overall cost was assumed to be a linear function of exposure. The optimization process finds the allocation of the components to the three categories that minimizes the tradeoff between integrated risk and total cost. For the InSight data, a set of components is selected representing all three categories, and optimization is performed to determine if the performed allocation was optimal or if a better allocation could have been obtained. To the best of our knowledge, this work is the first attempt not only perform an accurate estimation of bioburden density but also do it in an optimal way.

97 - MATHEMATICS AND COMPUTING↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Stacked reverberation mapping of high-redshift quasars in DESI. I. Feasibility analysis

The broad-line region of quasars has long been probed by reverberation mapping techniques that measure time lags between continuum and broad emission-line variations. Stacked reverberation mapping has been proposed as a less observationally expensive alternative to traditional methods. This ensemble approach also reduces biases from small-number statistics. The Dark Energy Spectroscopic Instrument (DESI) is conducting the most extensive spectroscopic survey of quasars to date. We create mock light curves emulating expected DESI quasar observations at redshifts $1.48\lt z\lt 5.2$ and luminosities $44.68 \le \log \lambda L_{1350 \mathring{\rm A}{}} / \mathrm{erg\, s^{-1}} \le 45.99$ to test stacked reverberation mapping feasibility using sparse spectroscopic data paired with well-sampled photometric data. The pipeline, using the lag estimation code JAVELIN (Just Another Vehicle for Estimating Lags In Nuclei), successfully recovers the simulated C IV lags within 1σ of the true values using spectroscopic light curves composed of only a few spectral epochs (2–10) with irregular cadences. We investigate how observational factors, including C IV flux error magnitude, number of stacked quasars, and spectral epoch count, affect performance. This work motivates a pathway for future stacked reverberation mapping projects with large-scale spectroscopic surveys of quasars having $\ge 2$ spectroscopic observations. Our results suggest an economical alternative for constraining and extending the radius–luminosity relation to higher redshifts and luminosities. Subsequently, this relation can be employed more reliably in single-epoch black hole mass measurements and quasar cosmology in these distant regimes.

quasars: general, quasars: supermassive black hole↗