Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Adversarial Binaries: AI-guided Instrumentation Methods for Malware Detection Evasion

Adversarial binaries are executable files that have been altered without loss of function by an AI agent in order to deceive malware detection systems. Progress in this emergent vein of research has been constrained by the complex and rigid structure of executable files. Although prior work has demonstrated that these binaries deceive a variety of malware classification models which rely on disparate feature sets, a consensus as to the best approach has not been reached, either in terms of the optimization algorithms or the instrumentation methods. Furthermore, although inconsistencies in the data sets, target classifiers, and functionality verification methods make head-to-head comparisons difficult, here we extract lessons learned and make recommendations for future research.

malware obfuscation↗

Conversational Grid Storage: Bridging Rucio and LLMs with Model Context Protocol

Experiments at Fermilab use Rucio to handle datasets that can be up to exabyte scale. However, navigating through Rucio’s syntax-heavy Command Line Interface (CLI) is a major workflow obstruction for researchers who just want to check quotas, track data identifiers (DIDs), or locate data sets. This project introduces a natural language interface. By building a containerized Model Context Protocol (MCP) server, an AI agent is created that translates plain English queries into data operations.

Akella, Kashyap [Fermilab; Illinois U., Urbana (ma↗

Image Deconvolution and Point-spread Function Reconstruction with STARRED: A Wavelet-based Two-channel Method Optimized for Light-curve Extraction

We present starred, a point-spread function (PSF) reconstruction, two-channel deconvolution, and light-curve extraction method designed for high-precision photometric measurements in imaging time series. An improved resolution of the data is targeted rather than an infinite one, thereby minimizing deconvolution artifacts. In addition, starred performs a joint deconvolution of all available data, accounting for epoch-to-epoch variations of the PSF and decomposing the resulting deconvolved image into a point source and an extended source channel. The output is a high-signal-to-noise-ratio, high-resolution frame combining all data and the photometry of all point sources in the field of view as a function of time. Of note, starred also provides exquisite PSF models for each data frame. We showcase three applications of starred in the context of the imminent LSST survey and of JWST imaging: (i) the extraction of supernovae light curves and the scene representation of their host galaxy; (ii) the extraction of lensed quasar light curves for time-delay cosmography; and (iii) the measurement of the spectral energy distribution of globular clusters in the "Sparkler," a galaxy at redshift z = 1.378 strongly lensed by the galaxy cluster SMACS J0723.3-7327. starred is implemented in jax, leveraging automatic differentiation and graphics processing unit acceleration. This enables the rapid processing of large time-domain data sets, positioning the method as a powerful tool for extracting light curves from the multitude of lensed or unlensed variable and transient objects in the Rubin-LSST data, even when blended with intervening objects.

79 ASTRONOMY AND ASTROPHYSICS↗

Data Analytics for Catalysis Predictions: Are We Ready Yet?

Catalysis informatics has received tremendous attention in recent years as a tool to design catalysts and discover unique descriptors that capture the relationships between chemical properties and catalytic performance. One of the stop-gaps in understanding catalytic effects, which is often ignored and limits the deployment of data science tools, relates to the lack of uniform data. The catalytic cleavage of C–X (X= H, C, N, and O) bonds is relevant to many fundamental catalytic processes. In this Perspective, we performed data analytics on four groups of C–X cleavage reactions that are common in production, upcycling, or reactive separation: the C–C cleavage in cyclopropyl alcohol, the C–H cleavage in hydroacylation reactions, the C–O cleavage in β-O-4 linkages, and the C–N cleavage in amides, using experimental data collected from the literature to understand their underlying correlations. Experimental variables of high impact are identified for each reaction by dimensionality reduction methods. We highlight the urgent need for experimental data sets that include full details on the reaction conditions, such as reagent concentration, reaction temperature, or time in machine-readable forms. We discuss the potential improvement of the data of these reactions and promising approaches such as autonomous experiments to fill the gaps in unbiased experimental data. Finally, we also address the early stage consideration of separation aspects in the experimental design of efficient catalytic systems for these fundamental examples of chemical reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dark Energy Survey Supernova 5YR Data Release

This data set was generated by the Dark Energy Survey (DES) team, and was used for the cosmological analysis in the DES key supernova cosmology paper (arXiv:2401.02929). When using this data please cite both the data release paper (Sanchez et al. 2024, ApJ, arXiv:2406.05046) and the DES key paper (Dark Energy Survey Collaboration 2024, ApJL, arXiv:2401.02929). At DES-SN5YR Github Repo (https://github.com/des-science/DES-SN5YR), we provide a comprehensive list of references for the different data products (SALT3 light-curve model, calibration, DCR corrections, SN classification, SN distances and uncertainties covariance matrix). If you have questions, post an issue on the DES-SN5YR Github repository and/or get in touch with the contact people indicated on the Github repository.

Contractor, Fermilab [Fermilab]↗

Unique & challenging aspects of plutonium metal standards exchange program for actinide measurements

The Los Alamos National Laboratory exchange program is the only program of its kind for the distribution of plutonium (Pu) standards materials with a range of impurity contents to multiple laboratories for destructive measurements of elemental concentration. This paper discusses statistical methods used to address challenges in Pu metal exchange data by way of two case studies. Challenges include how to evaluate a data set when a large fraction of the values are minimum detection limits (MDLs), and how to determine potential outliers with limited in-formation on the true spread of the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dark energy survey year 3 results: likelihood-free, simulation-based w CDM inference with neural compression of weak-lensing map statistics

We present simulation-based cosmological wcold dark matter (wCDM) inference using dark energy survey year 3 weak-lensing maps, via neural data compression of weak-lensing map summary statistics: power spectra, peak counts, and direct map-level compression/inference with convolutional neural networks (CNN). Using simulation-based inference, also known as likelihood-free or implicit inference, we use forward-modelled mock data to estimate posterior probability distributions of unknown parameters. This approach allows all statistical assumptions and uncertainties to be propagated through the forward-modelled mock data; these include sky masks, non-Gaussian shape noise, shape measurement bias, source galaxy clustering, photometric redshift uncertainty, intrinsic galaxy alignments, non-Gaussian density fields, neutrinos, and non-linear summary statistics. We include a series of tests to validate our inference results. This paper also describes the Gower Street simulation suite: 791 full-sky pkdgrav3 dark matter simulations, with cosmological model parameters sampled with a mixed active-learning strategy, from which we construct over 3000 mock dark energy survey lensing data sets. For wCDM inference, for which we allow –1 < w < –$\frac{1}{3}$⁠, our most constraining result uses power spectra combined with map-level (CNN) inference. Using gravitational lensing data only, this map-level combination gives Ω m = 0.283$^{+0.020}_{–0.027}$⁠, S 8 = 0.804$^{+0.025}_{–0.017⁠}$, and w < –0.80 (with a 68 per cent credible interval); compared to the power spectrum inference, this is more than a factor of two improvement in dark energy parameter (Ω⁠ DE , w⁠) precision.

79 ASTRONOMY AND ASTROPHYSICS↗

WHOLESCALE - Water & Hole Observations Leverage Effective Stress Calculations And Lessen Expenses (Final Technical Report 2020 - 2024)

The WHOLESCALE acronym stands for Water & Hole Observations Leverage Effective Stress Calculations and Lessen Expenses. The goal of the WHOLESCALE project is to simulate the spatial distribution and temporal evolution of stress in the geothermal system at San Emidio in Nevada, United States. To reach this goal, the WHOLESCALE team has developed a methodology to incorporate and interpret data from four methods of measurement into a multi-physics model that couples thermal, hydrological, and mechanical (T H-M) processes. The WHOLESCALE team has applied this methodology at the San Emidio geothermal field, located ~100 km north of Reno, Nevada in the northwestern Basin and Range province. The WHOLESCALE team includes 30 individuals working at two universities, two national laboratories, and one industry partner. Two master-degree students and five post-doctoral researchers have gained professional experience and earned partial financial support via the WHOLESCALE project. The WHOLESCALE team has taken advantage of the perturbations created by changes in pumping operations during planned shutdowns in 2016, 2021, and 2022 to infer temporal changes in the state of stress in the geothermal system at San Emidio, Nevada, U.S. The WHOLESCALE results support the working hypothesis that increasing pore-fluid pressure reduces the effective normal stress acting across fault zones. During normal operations, pumping in deep production wells decreases fluid pressures and thus increases the effective normal stresses on faults, reducing microseismicity. During planned shutdowns, the cessation of production increases pore-fluid pressure and reduces effective normal stress. The WHOLESCALE products generated during the 4-year period between 2020 and 2024 include: three articles published in the open-access, peer-reviewed scientific literature, two master’s theses, 20 presentations or papers at scientific conferences, and 17 data sets available on public repositories. The WHOLESCALE project has been completed in two phases that included three performance periods separated by two Go/No-go Stage Gate Reviews. Tasks were classified by data type (i.e., Geologic Structure, Borehole, Geodesy, Hydrology, Seismology, and Modeling). The first phase of the project started July 31, 2020 and included ongoing project coordination (Task 1), a project kickoff (Task 2), analysis of existing data (Task 3), development of the initial stress model & deployment design (Task 4), and Go/No-go Decision Point #1 (Task 5). Phase II began with implementing the 2022 deployment (Task 6), followed by Go/No-go Decision Point #2 (Task 7) The remainder of Phase II consisted of analyzing data collected during deployment (Task 8), calibration of the stress model on all observations (Task 9), and the Final Review (August 23, 2024) & Reporting (Task 10).

15 GEOTHERMAL ENERGY↗

L-PBF High-Throughput Data Pipeline Approach for Multi-modal Integration

Abstract Metal-based additive manufacturing requires active monitoring solutions for assessing part quality. Multiple sensors and data streams, however, generate large heterogeneous data sets that are impractical for manual assessment and characterization. In this work, an automated pipeline is developed that enables feature extraction from high-speed camera video and multi-modal data analysis. The framework removes the need for manual assessment through the utilization of deep learning techniques and training models in a weakly supervised paradigm. We demonstrate this pipeline’s capability over 700,000 high-speed camera frames. The pipeline successfully extracts melt pool and spatter geometries and links them to corresponding pyrometry, radiography, and processparameter information. 715 individual prints are examined to reveal melt pool areas that exceeds 0.07 mm 2 and pyrometry signal over a threshold (375 pyrometry units) were more likely to have defects. These automated processes enable massive throughput of characterization techniques.

36 MATERIALS SCIENCE↗

Non-intrusive reduced-order modeling for dynamical systems with spatially localized features

This work presents a non-intrusive reduced-order modeling framework for dynamical systems with spatially localized features characterized by slow singular value decay. The proposed approach builds upon two existing methodologies for reduced and full-order non-intrusive modeling, namely Operator Inference (OpInf) and sparse Full-Order Model (sFOM) inference. We decompose the domain into two complementary subdomains that exhibit fast and slow singular value decay. The dynamics of the subdomain exhibiting slow singular value decay are learned with sFOM while the dynamics with intrinsically low dimensionality on the complementary subdomain are learned with OpInf. The resulting, coupled OpInf-sFOM formulation leverages the computational efficiency of OpInf and the high resolution of sFOM, and thus enables fast non-intrusive predictions for conditions beyond those sampled in the training data set. A novel regularization technique with a closed-form solution based on the Gershgorin disk theorem is introduced to promote stable sFOM and OpInf models. We also provide a data-driven indicator for subdomain selection and ensure solution smoothness over the interface via a post-processing interpolation step. We evaluate the efficiency of the approach in terms of offline and online speedup through a quantitative, parametric computational cost analysis. We demonstrate the coupled OpInf-sFOM formulation for two test cases: a one-dimensional Burgers’ model for which accurate predictions beyond the span of the training snapshots are presented, and a two-dimensional parametric model for the Pine Island Glacier ice thickness dynamics, for which the OpInf-sFOM model achieves an average prediction error on the order of 1% with an online speedup factor of approximately 8$\times$ compared to the numerical simulation.

42 ENGINEERING↗

Quantitative Imaging of Cobalt Phthalocyanine Distribution on Carbon Nanotubes: A Deep Learning Approach to Catalyst Characterization

Electrochemical reduction of carbon dioxide (CO 2 ) offers a pathway to valuable products, with catalysts playing a crucial role. This study investigates the distribution of cobalt tetraaminophthalocyanine (CoPc-NH 2 ) immobilized on carbon nanotubes (CNTs), utilizing high-angle annular dark-field scanning transmission electron microscopy (HAADF-STEM) to characterize CoPc-NH 2 distribution. A challenge in the quantitative HAADF-STEM analysis is the introduction of bias from manual Co atom identification. To address this, we developed and trained a convolutional neural network (CNN) using a data set generated from images of CoPc-NH 2 /CNT samples with varying Co loadings. The CNN, implemented in TensorFlow and Keras, facilitated Co atom detections. Analysis of the CNN-generated data confirmed a correlation between Co loading and surface density, consistent with findings from UV–vis spectroscopy. Furthermore, the application of Ripley’s L(d) function highlighted the presence of slight Co atom clustering. Furthermore, this work demonstrates the utility of the combined HAADF-STEM and CNN approach for providing spatially resolved information about catalyst distribution on nonplanar supports, revealing structural details that are typically lost through other characterization methods.

HAADF-STEM↗

BLADE: An Automated Framework for Classifying Light Curves from the Center for Near-Earth Object Studies Fireball Database

Fireballs (bolides) are high-energy luminous phenomena produced when meteoroids and small asteroids enter Earth’s atmosphere at hypersonic speeds, often resulting in fragmentation or complete disintegration accompanied by significant energy release. The resulting bolide light curves capture temporal brightness variations as these objects traverse increasingly dense atmospheric layers, providing essential information on meteoroid entry dynamics, fragmentation behavior, and atmospheric energy deposition processes. The Center for Near-Earth Object Studies’ (CNEOS) continuously expanding fireball database offers a globally comprehensive archive of bolide events, including light curves and associated metadata. Events associated with infrasound detections allow direct correlations between acoustic signatures and light curve features, therefore enabling detailed analyses of fragmentation dynamics and energy deposition. Here, we introduce Bolide Light-curve Analysis and Discrimination Explorer (BLADE), a robust and high-fidelity framework specifically designed to analyze bolide light curves for objects detected from space. BLADE incorporates a processing pipeline integrating Savitzky–Golay filtering, prominence-based peak detection, and gradient analysis, enabling systematic identification and classification of fragmentation events and their associated energy release characteristics. Preliminary results demonstrate that BLADE reliably distinguishes distinct bolide behaviors, providing an objective, scalable methodology for characterization and analysis of large bolide light curve data sets. This foundational work establishes a novel pathway for advanced bolide research, with promising applications in planetary defense and global atmospheric monitoring. Future research should adopt an integrative approach combining CNEOS optical data with complementary infrasound measurements, further clarifying relationships between bolide energy deposition and acoustic signatures, thus refining our understanding of meteoroid and asteroid atmospheric entry processes.

Asteroids↗

Comprehensive review of 2 β decay half-lives

Here, the double-beta (2 β )-decay is the rarest nuclear physics process, and its experimental half-lives (T 1/2 ) exceed the age of the Universe from nine to fourteen orders of magnitude. Double-beta decay was observed, and its half-life was measured in 14 parent nuclei using direct, radiochemical, and geochemical methods. The decay observables are analyzed using the Evaluated Nuclear Structure Data File (ENSDF) procedures, and the recommended T 1/2 were deduced. Using the calculated values of phase factors, the effective nuclear matrix elements were extracted and compared with available data. Thousands of theoretical and experimental works have been dedicated to these topics in the last 85 years, and we present two data sets of recommended values to encapsulate the results.

2β-decay↗

Impact of composition and symmetry energy on the temperature of quasiprojectiles simulated with antisymmetrized molecular dynamics

The equation of state describes the emergent physical properties of matter. Experimental data is needed to help constrain the equation of state for nuclear matter. These constraints can help distinguish between an “asy-stiff” and an “asy-soft” equation of state, which has astrophysical implications. One path to help constrain the models is to analyze the nuclear caloric curve; some experiments have shown dependence on neutron excess, and may thus be sensitive to the asymmetry. A difference in the caloric curve based on the asymmetry of the reconstructed quasiprojectile (QP) had been observed using 70 Zn on 70 Zn at 35 MeV/nucleon taken with the NIMROD array. Antisymmetrized molecular dynamics calculations were performed for the same system and deexcited with gemini++. Both Gogny (asy-soft) and Gogny-as (asy-stiff) data sets were generated. The particles were then filtered based on detector geometric acceptance and thresholds. From the accepted particles, the excitation energy and temperature were calculated in the same way as for experimental data. Additionally, filter effects on the observed nuclear caloric curves were investigated. A tendency for the asy-stiff nuclear caloric curves to have higher temperatures than their asy-soft counterparts was observed for a number of probes. In addition, some probes may show sensitivity to the reconstructed composition of the QP, but this is inconclusive due to high statistical fluctuations and a large dependence on the exact method of event selection.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Comprehensive Review of 2$β$ Decay Half-Lives

The double-beta (2β)-decay is the rarest nuclear physics process, and its experimental half-lives (T 1/2 ) exceed the age of the Universe from nine to fourteen orders of magnitude. Double-beta decay was observed, and its half-life was measured in 14 parent nuclei using direct, radiochemical, and geochemical methods. The decay observables are analyzed using the Evaluated Nuclear Structure Data File (ENSDF) procedures, and the recommended T 1/2 were deduced. Using the calculated values of phase factors, the effective nuclear matrix elements were extracted and compared with available data. Thousands of theoretical and experimental works have been dedicated to these topics in the last 85 years, and we present two data sets of recommended values to encapsulate the results.

2β-decay↗

Exploration of signal processing methods for superconducting magnet and quench data

Quenching is the phenomenon of a superconducting magnetic material carrying current transitioning into a regular conducting material. This may cause severe and irreparable damage to the superconductor due to Joule heating. The Magnet Department at Fermi National Accelerator Laboratory (FNAL) has acquired experimental data through quench antenna arrays that are recorded when the quench is detected. These data are in terms of voltage signals that are sampled at 100kHz for several minutes. There are multiple channels and each channel provides a data set of more than 20 million observations, while there is one channel, called the trigger channel which shows the time when quench is detected. Despite some advancements that were made including machine learning, data complexity still shadows the progress. In this work, we studied a multi-resolution analysis of the quench antenna data through the Haar wavelet transform. In particular, we applied the maximally overlapped discrete w avelet transform (MODWT) of a suitable level L to the given data and then projected it onto the wavelet basis. This decomposes a given signal (Original data) $x ϵ \mathbb{R}^N$ into $L + 1$ subspaces of $\mathbb{R}^N$. One of the subspaces called the approximation, captures the trend of the signal, and the others, called the details, capture the fluctuations at different frequency bands. This decomposition provides a clear trend of the data at a suitable level and also various activities (spikes) are seen in the details of the decomposition at every level. These spikes might reveal some information about the quench under investigation but in any case, give information about magnet behavior. Also, this decomposition is seen to be very useful in removing noise present in the data due to the source or mechanism of the experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data for "A 13-year record indicates differences in the duration and depth of soil carbon accrual among potential bioenergy crops"

Data sets for material included in "A 13-year record indicates differences in the duration and depth of soil carbon accrual among potential bioenergy crops" by Kantola et al., 2025, in Global Change Biology Bioenergy. Data include soil organic carbon (SOC), carbon stable isotope ratios, annual belowground biomass, and annual post-harvest litter for four crops, maize/soybean, miscanthus, switchgrass, and prairie, between 2008 and 2021.

bioenergy crops↗

Denoising Seismic Waveforms Using a Wavelet-Transform-Based Machine-Learning Method

Seismic waveform data recorded at stations can be thought of as a superposition of the signal from a source of interest and noise from other sources. Frequency‐based filtering methods for waveform denoising do not result in desired outcomes when the targeted signal and noise occupy similar frequency bands. Recently, denoising techniques based on deep‐learning convolutional neural networks (CNNs), in which a recorded waveform is decomposed into signal and noise components, have led to improved results. These CNN methods, which use short‐time Fourier transform representations of the time series, provide signal and noise masks for the input waveform. These masks are used to create denoised signal and designaled noise waveforms, respectively. However, advancements in the field of image denoising have shown the benefits of incorporating discrete wavelet transforms (DWTs) into CNN architectures to create multilevel wavelet CNN (MWCNN) models. The MWCNN model preserves the details of the input due to the good time–frequency localization of the DWT. In this report we use a data set of over 382,000 constructed seismograms recorded by the University of Utah Seismograph Stations network to compare the performance of CNN and MWCNN‐based denoising models. Evaluation of both models on constructed test data shows that the MWCNN model outperforms the CNN model in the ability to recover the ground‐truth signal component in terms of both waveform similarity and preservation of amplitude information. Model evaluation of real‐world data shows that both the CNN and MWCNN models outperform standard band‐pass filtering (BPF; average improvement in signal‐to‐noise ratio of 9.6 and 19.7 dB, respectively, with respect to BPF). Evaluation of continuous data suggests the MWCNN denoiser can improve both signal detection capabilities and phase arrival time estimates.

58 GEOSCIENCES↗