Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Computational toolkit for predicting thickness of 2D materials using machine learning and autogenerated dataset by large language model

The thickness of 2D materials not only plays a crucial role in determining the performance of nanoelectronic and optoelectronic devices but also introduces complexities in predicting volume-dependent properties, such as energy storage capacity, due to the intrinsic vacuum within these materials. Although a plethora of experimental techniques, including but not limited to optical contrast, Raman spectroscopy, nonlinear optical spectroscopy, near-field optical imaging, and hyperspectral imaging, facilitate the measurement of 2D material thickness, comprehensive data for many materials remain elusive. Over the past decade, the exponential proliferation of 2D materials and their heterostructures has outstripped the capabilities of conventional experimental and computational approaches. In this evolving landscape, machine learning (ML) has emerged as an indispensable tool, offering a scalable approach to augment these traditional methodologies. Addressing the critical gap, we introduce THICK2D—Thickness Hierarchy Inference and Calculation Kit for 2D Materials. This Python-based computational framework harnesses an autogenerated thickness database, developed using large language models, and advanced ML algorithms to facilitate the rapid and scalable estimation of material thickness, relying solely on crystallographic data. To demonstrate the utility and robustness of THICK2D, we successfully used the toolkit to predict the thickness of more than 8000 2D-based materials, sourced from two extensive 2D materials databases. THICK2D is disseminated as an open-source utility, accessible on GitHub at https://github.com/gmp007/THICK2D, and archived on Zenodo at https://10.5281/zenodo.11216648.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Overcoming sparse datasets with multi-task learning as applied to high entropy alloys

Abstract The design of novel High Entropy Alloys for use in high-temperature applications is an area of active interest due to their potential to provide exceptional properties compared to conventional alloys. Since the increased popularity of machine learning, an important cog in the design process has been training surrogate models on alloy properties. However, these Single-Task models are trained on individual mechanical properties and do not take advantage of the relatedness between properties. Multi-Task models can capture the interdependencies between tasks, leading to potentially more accurate predictions for all tasks. In this paper, we investigate if Multi-Task models can show improvement over Single-Task models when used for predicting the mechanical properties of these alloys. To ensure fair evaluation between the models, we apply L 0 regularization and skip connections to the models, which allows them to adjust the number of model parameters and depth for optimal performance. We find that the Multi-Task models can leverage task relationships to perform better than Single-Task models, especially for high amounts of missing data in the tasks. Furthermore, adding simple auxiliary targets can boost Multi-Task performance even further despite not being effective as input descriptors to single-task models themselves. We anticipate that the proposed strategies can achieve more accurate predictions and consequently enable better design capabilities for such data-constrained domains without incurring much additional computational cost.

Debnath, Arindam (ORCID:0000000194274499)↗

Dark Energy Survey: Galaxy sample for the baryonic acoustic oscillation measurement from the final dataset

In this paper, we present and validate the galaxy sample used for the analysis of the baryon acoustic oscillation (BAO) signal in the Dark Energy Survey (DES) Y6 data. The definition is based on a color and redshift-dependent magnitude cut optimized to select galaxies at redshifts higher than 0.6, while ensuring a high-quality photo- z determination. The optimization is performed using a Fisher forecast algorithm, finding the optimal i -magnitude cut to be given by i < 19.64 + 2.894 z ph . For the optimal sample, we forecast an increase in precision in the BAO measurement of ∼ 25 % with respect to the Y3 analysis. Our BAO sample has a total of 15,937,556 galaxies in the redshift range 0.6 < z ph < 1.2 , and its angular mask covers 4 , 273.42 deg 2 to a depth of i = 22.5 . We validate its redshift distributions with three different methods: directional neighborhood fitting algorithm (DNF), which is our primary photo- z estimation; direct calibration with spectroscopic redshifts from VIPERS, which is a spectroscopic galaxy sample that overlaps with our BAO sample and is complete within our selection cuts; and clustering redshift using SDSS galaxies. The fiducial redshift distribution is a combination of these three techniques performed by modifying the mean and width of the DNF distributions to match those of VIPERS and clustering redshift. In this paper, we also describe the methodology used to mitigate the effect of observational systematics, which is analogous to the one used in the Y3 analysis. This paper is one of the two dedicated to the analysis of the BAO signal in DES Y6. In its companion paper, we present the angular diameter distance constraints obtained through the fitting to the BAO scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey: A 2.1% measurement of the angular baryonic acoustic oscillation scale at redshift z eff = 0.85 from the final dataset

Here, we present the angular diameter distance measurement obtained with the baryonic acoustic oscillation (BAO) feature from galaxy clustering in the completed Dark Energy Survey, consisting of six years (Y6) of observations. We use the Y6 BAO galaxy sample, optimized for BAO science in the redshift range 0.6 < z <1.2, with an effective redshift at z eff = 0.85 and split into six tomographic bins. The sample has nearly 16 million galaxies over 4,273 square degrees. Our consensus measurement constrains the ratio of the angular distance to sound horizon scale to D M ⁡(z eff )/r d = 19.51 ± 0.41 (at 68.3% confidence interval), resulting from comparing the BAO position in our data to that predicted by planck Λ⁢CDM via the BAO shift parameter α =(D M /r d )/(D M /r d ) PLANCK . To achieve this, the BAO shift is measured with three different methods, angular correlation function (ACF), angular power spectrum (APS), and projected correlation function (PCF), obtaining α = 0.952 ± 0.023, 0.962 ± 0.022, and 0.955 ± 0.020, respectively, which we combine to α = 0.957 ± 0.020, including systematic errors. When compared with the Λ⁢CDM model that best fits planck data, this measurement is found to be 4.3% and 2.1⁢σ below the angular BAO scale predicted. To date, it represents the most precise angular BAO measurement at z > 0.75 from any survey and the most precise measurement at any redshift from photometric surveys. The analysis was performed blinded to the BAO position, and it is shown to be robust against analysis choices, data removal, redshift calibrations, and observational systematics.

79 ASTRONOMY AND ASTROPHYSICS↗

Measurement of the B 8 solar neutrino flux using the full SNO + water phase dataset

The SNO+ detector operated initially as a water Cherenkov detector. The implementation of a sealed cover gas system midway through water data taking resulted in a significant reduction in the activity of 222 Rn daughters in the detector and allowed the lowest background to the solar electron scattering signal above 5 MeV achieved to date. This paper reports an updated SNO+ water phase 8 B solar neutrino analysis with a total livetime of 282.4 days and an analysis threshold of 3.5 MeV. The 8 B solar neutrino flux is found to be (2.3⁢2$^{+0.18}_{-0.17}⁢$(stat)$^{+0.07}_{-0.05}$⁢(syst))×10 6 cm -2 s -1 assuming no neutrino oscillations, or (5.3⁢6$^{+0.41}_{-0.39}⁢$(stat)+$^{0.17}_{-0.16}$⁢(syst))×10 6 cm -2 s -1 assuming standard neutrino oscillation parameters, in good agreement with both previous measurements and standard solar model calculations. The electron recoil spectrum is presented above 3.5 MeV.

79 ASTRONOMY AND ASTROPHYSICS↗

Search for a Sub-eV Sterile Neutrino Using Daya Bay’s Full Dataset

This Letter presents results of a search for the mixing of a sub-eV sterile neutrino with three active neutrinos based on the full data sample of the Daya Bay Reactor Neutrino Experiment, collected during 3158 days of detector operation, which contains 5.55 × 106 reactor $\overline{v}$ e candidates identified as inverse beta-decay interactions followed by neutron capture on gadolinium. The analysis benefits from a doubling of the statistics of our previous result and from improvements of several important systematic uncertainties. No significant oscillation due to mixing of a sub-eV sterile neutrino with active neutrinos was found. Exclusion limits are set by both Feldman-Cousins and CLs methods. Light sterile neutrino mixing with sin 2⁡ 2⁢θ 14 ≳ 0.01 can be excluded at 95% confidence level in the region of 0.01 eV 2 ≲ |Δ⁢$m$$^{2}_{41}$| ≲ 0.1 eV 2 . This result represents the world-leading constraints in the region of 2 × 10 –4 eV 2 ≲ |Δ⁢$m$$^{2}_{41}$| ≲ 0.2 eV 2 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Constraints on lepton number violation with the 2 tonne · year CUORE dataset

Matter-antimatter asymmetry underlines the incompleteness of the current understanding of particle physics. Neutrinoless double-beta decay (0νββ) may help explain this asymmetry while unveiling the Majorana nature of the neutrino. The CUORE (Cryogenic Underground Observatory for Rare Events) experiment searches for 0νββ of 130 Te using a tonne-scale cryogenic calorimeter operated at milli-kelvin temperatures. We report no evidence of 0νββ and place a lower limit on the half-life of T 1/2 > 3.5 × 10 25 years (90% credibility interval) with over 2 tonne·years of TeO 2 exposure. Finally, the tools and techniques developed for this result and the 5-year stable operation of nearly 1000 detectors demonstrate crucial infrastructure for future-generation experiments capable of searching for 0νββ across multiple isotopes.

Adams, D. Q. [Department of Physics and Astronomy,↗

Avian Activity Classification Using Recurrent Networks to Fuse Videos with Metadata on Imbalanced Datasets

Activity classification plays a crucial role in various real-life scenarios involving both humans and animals. There is an increasing need for precise activity classification focused on avian-solar interactions, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no work to develop an automated system to monitor and classify these avian-solar interactions. All current methods rely on human observers, which is time and human resources costly and subject to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification problems, this paper develops a recurrent neural network-based model to automatically classify six avian activities around solar energy facilities. Our proposed model integrates critical feature engineering metadata with video frame data, enabling improved learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate the efficacy of our model in detecting and classifying different activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.

Avian activity classification; bidirectional LSTM;↗

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗