Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Long‐Term Large‐Scale Atmospheric Forcing Data From Three‐Dimensional Constrained Variational Analysis for the ARM SGP Site

Here, this study presents a long‐term three‐dimensional large‐scale forcing data set (VARANAL3D) derived from the three‐dimensional constrained variational analysis (3DCVA) method at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) site from 2004 to 2018. Building on the same input data sets as the conventional continuous forcing data set (VARANAL), VARANAL3D maintains overall consistency in domain‐averaged fields while introducing spatial variability, offering critical insights into the influence of mesoscale synoptic systems on cloud‐related processes. Evaluations are conducted across four cloud and precipitation regimes: Clear‐sky, Shallow‐clouds, Afternoon‐precipitation, and Nocturnal‐precipitation, presenting high consistency of the domain‐mean forcing data sets while emphasizing the role of subdomain forcing variability particularly in precipitating regimes. Single column model (SCM) simulations demonstrate that subdomain VARANAL3D forcing improves cloud and precipitation representation, with the ensemble outperforming domain‐mean forcing in three cloudy and precipitating regimes. Overall, these results highlight VARANAL3D's value for investigating the impacts of spatial variability of large‐scale forcing on atmospheric processes. The VARANAL3D data set provides new opportunities for evaluating model physics, advancing the development of scale‐aware parameterizations and deepening our understanding of cloud and precipitation dynamics.

Environmental sciences↗

Low Activity Tritium Detection in CCDs Using Deep Learning Techniques

Here, this study explores the use of charge-coupled devices (CCDs) for detecting low-energy beta particles from tritium decay - a critical signal for nuclear safety, nuclear nonproliferation, and environmental monitoring. We employ a dual approach utilizing both measured CCD data and detailed Geant4 simulations. Our analysis compares classical techniques with advanced deep learning methods, including convolutional neural networks (CNNs), autoencoders trained exclusively on tritium data, and preliminary studies on boosted decision trees (BDTs). The CNN, trained on mixed signal/background datasets, demonstrates superior classification performance, while the autoencoder shows the potential of unsupervised, background-agnostic strategies when background characteristics are poorly defined. These results highlight the excellent sensitivity achievable thanks to the background rejection made possible by information-rich CCD data, paving the way for improved portable tritium monitoring.

Autoencoder↗

An R Shiny graphical user interface for analyzing, visualizing, and interpreting high precision mass spectrometric data

There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2 , IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

Labone, Elizabeth↗

An R shiny graphical user interface for highprecision mass spectrometric data analysis

• There is currently a lack of software that meets the needs for the analysis of raw data produced by modern isotope ratio mass spectrometers for both R&D and routine use at SRNL and other US national labs • Needs to accommodate multiple isotope systems, instruments, and manufacturers • Include modern statistical methods and handling/visualization of uncertainty • Flexible software with transparent (no “black box”) and reproducible methods • This project is inspired by existing discipline-specific data analysis software (e.g., Tripoli1 , ET_Redux2, IsoplotR3) used in the geochemical community • Our goal is to build an open source data analysis software package that focuses on flexibility, transparency, and reproducibility

LABONE, ELIZABETH↗

Ramp-release experiments for strength measurements: Strain-rate dependence

This paper presents an enhanced analysis method for investigating material properties at high strain rates, extending the capability of established experimental techniques to gain more information. The ramp-release method has been applied to many experiments reported at high (≈10 5 − 10 6 s −1 ) strain-rates. More recent data gathered at the National Ignition Facility (NIF) has enabled higher (≈10 8 s −1 ) strain-rates to be studied. Here, we present an initial application of ramp-release analysis to NIF ramp-compression data, illustrating both the opportunities and the practical challenges of extending these methods to laser-driven platforms. The higher strain-rates accessed at the NIF mean that there is more strain-rate enhancement to strength, and the experimental configuration means that this enhancement is more readily seen in the data. This is enabled by the capability of avoiding peak-compression attenuation through the sample thickness with a designed hold period made possible by the pulse-shaping capability of NIF. We propose that this combination of experimental conditions and an enhanced analysis method enables the strain-rate enhancement to strength to be studied, and potentially for this to inform physics models at smaller scales than the continuum.

36 MATERIALS SCIENCE↗

LABQ3: Bayesian method for quantification of mineral compositions and nano-scale elemental mapping of 3D synchrotron XCT data

Quantitative analysis of mineral compositions is essential in understanding geochemical, mineralogical and environmental processes. Fine-resolution 3D imaging is widely done using synchrotron X-ray computed tomography (XCT), but existing analyses are limited to visualization and segmentation. This paper presents a new method, Linear Attenuation Bayesian Quantitative 3D-mapper (LABQ3), based on the linearity of X-ray attenuation with respect to elemental concentrations. To address the random variability in attenuation measurements, LABQ3 employs Bayesian decision theory to minimize classification error, using reference attenuation distributions from scans of pure mineral standards. To demonstrate LABQ3 and test its performance, we studied precipitated carbonate samples. XCT scans were done at multiple energies using the transmission X-ray microscope (TXM) at beamline 32-ID-C of the Advanced Photon Source at Argonne National Laboratory. The reconstructed 3D images have a voxel size of 20 nm. Analyses revealed rich nano-scale compositional heterogeneity within individual particles. A mixture of calcium and cadmium produced an overall stoichiometric composition of (Ca 0.78 ,Cd 0.22 )CO 3 , with some voxels containing nearly pure CdCO 3 . The addition of zinc led to an overall stoichiometric composition of 33% Ca, 28% Cd, 39% Zn, with a nearly pure CaCO 3 core and compositional zonation through the rim. These compositional gradients are related to temporal sequences of carbonate mineral formation where Cd precipitated at the beginning in (Ca,Cd)CO 3 , while Cd and Zn precipitated at the end in (Ca, Cd,Zn)CO 3 . Results differ from bulk analyses using Inductively Coupled Plasma-Mass Spectrometry (ICP-MS), showing that LABQ3 provides particle-specific insights. LABQ3 distinguishes itself by quantifying chemical compositions along a continuum, making it different from XCT analyses based on segmentation. LABQ3 allows simultaneous acquisition of morphology and chemical composition in 3D, facilitating the interpretation of chemical gradients of trace elements, quantification of solid solution compositions, inferences about temporal sequences of mineral precipitation, and addressing other concerns about solid-phase chemistry.

58 GEOSCIENCES↗

Measurement of 𝑑 2⁢ 𝜎/𝑑⁢|$\vec{q}$|⁢𝑑⁢𝐸 avail in charged current 𝜈 𝜇 -nucleus interactions at ⟨𝐸 𝜈 ⟩=1.86 GeV using the NOvA Near Detector

Double- and single-differential cross sections for inclusive charged-current 𝜈 𝜇 -nucleus scattering are reported for the kinematic domain 0 to 2 GeV/𝑐 in three-momentum transfer and 0 to 2 GeV in available energy, at a mean 𝜈 𝜇 energy of 1.86 GeV. The measurements are based on an estimated 995,760 𝜈 𝜇 charged-current (CC) interactions in the scintillator medium of the NOvA Near Detector. The subdomain populated by 2-particle-2-hole (2p2h) reactions is identified by the cross section excess relative to predictions for 𝜈 𝜇 -nucleus scattering that are constrained by a data control sample. Models for 2-particle-2-hole processes are rated by 𝜒 2 comparisons of the predicted-versus-measured 𝜈 𝜇 CC inclusive cross section over the full phase space and in the restricted subdomain. Shortfalls are observed in neutrino generator predictions obtained using the theory-based València and SuSAv2 2p2h models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measurements of Pion and Muon Nuclear Capture at Rest on Argon in the LArIAT Experiment

We report the measurement of the final-state products of negative pion and muon nuclear capture at rest on argon by the LArIAT experiment at the Fermilab Test Beam Facility. We measure a population of isolated MeV-scale energy depositions, or blips, in 296 LArIAT events containing tracks from stopping low-momentum pions and muons. The average numbers of visible blips are measured to be 0.74 ± 0.19 and 1.86 ± 0.17 near muon and pion track endpoints, respectively. The 3.6⁢𝜎 statistically significant difference in blip content between muons and pions provides the first demonstration of a new method of pion-muon discrimination in neutrino liquid argon time projection chamber experiments. LArIAT Monte Carlo simulations predict substantially higher average blip counts for negative muon (1.22 ± 0.08) and pion (2.34 ± 0.09) nuclear captures. We attribute this difference to geant4’s inaccurate simulation of the nuclear capture process.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cadmus v1.0

This software package is for distributed computation of persistent (co)homology. It is primarily intended for Topological Data Analysis (TDA) audience and scientists who use TDA methods and apply them to datasets that are too big to handle on a single machine. There are very few codes available for this; Cadmus outperforms DIPHA, if the user is interested in cohomology. The accompanying paper explaining the algorithm was accepted to ALENEX 26.

Nigmetov, Arnur [Lawrence Berkeley National Labora↗

Pivotal trial characteristics and types of endpoints used to support Food and Drug Administration rare disease drug approvals between 2013 and 2022

Background/aims Rare disease drug development faces unique challenges, such as genotypic and phenotypic heterogeneity within small patient populations and a lack of established outcome measures for conditions without previously successful drug development programs. These challenges complicate the process of selecting the appropriate trial endpoints and conducting clinical trials in rare diseases. In this descriptive study, we examined novel drug approvals for non-oncologic rare diseases by the U.S. Food and Drug Administration’s Center for Drug Evaluation and Research over the past decade and characterized key regulatory and trial design elements with a focus on the primary efficacy endpoint utilized as the basis of approval. Methods Using the Food and Drug Administration’s Data Analysis Search Host database, we identified novel new drug applications and biologics license applications with orphan drug designation that were approved between 2013 and 2022 for non-oncologic indications. From Food and Drug Administration review documents and other external databases, we examined characteristics of pivotal trials for the included drugs, such as therapeutic area, trial design, and type of primary efficacy endpoints. Differences in trial design elements associated with primary efficacy endpoint type were assessed such as randomization and blinding. Then, we summarized the primary efficacy endpoint types utilized in pivotal trials by therapeutic area, approval pathway, and whether the disease etiology is well defined. Results One hundred and seven drugs that met our inclusion criteria were approved between 2013 and 2022. Assessment of the 107 drug development programs identified 150 pivotal trials that were subsequently analyzed. The pivotal trials were mostly randomized (80%) and blinded (69.3%). Biomarkers (41.1%) and clinical outcomes (42.1%) were commonly utilized as primary efficacy endpoints. Analysis of the use of clinical trial design elements across trials that utilized biomarkers, clinical outcomes, or composite endpoints did not reveal statistically significant differences. The choice of primary efficacy endpoint varied by the drug’s therapeutic area, approval pathway, and whether the indicated disease etiology was well defined. For example, biomarkers were commonly selected as primary efficacy endpoints in hematology drug approvals (70.6%), whereas clinical outcomes were commonly selected in neurology drug approvals (69.6%). Further, if the disease etiology was well defined, biomarkers were more commonly used as primary efficacy endpoints in pivotal trials (44.7%) than if the disease etiology was not well defined (27.3%). Discussion In the past 10 years, numerous novel drugs have been approved to treat non-oncologic rare diseases in various therapeutic areas. To demonstrate their efficacy for regulatory approval, biomarkers and clinical outcomes were commonly utilized as primary efficacy endpoints. Biomarkers were not only frequently used as surrogate efficacy endpoints in accelerated approvals, but also in traditionally approved rare disease drugs. The choice of primary efficacy endpoints varied by therapeutic area, approval pathway, and understanding of disease etiology.

Hong, Kyungwan [Rare Diseases Team, Office of New ↗

ML-based Micro-CT SOFC Microstructure Models (from Kent 2026 Microstructural Augmentation paper)

Overview -------------------------- This repository contains datasets from the manuscript **"Enhanced Generalizability to Deep-Learning Quantification of 3D Microstructural Characteristics through Microstructurally Aware Augmentation of Scarce Data"** (*William F. Kent, Rochan Bajpai, Rachel C. Kurchin, William K. Epting, Harry W. Abernathy, Paul A. Salvador. Submitted 2026*). The methods are also described in the dissertation **Data Intensive Analysis of Solid Oxide Cell Microstructures** (*Doctoral dissertation, Carnegie Mellon University, 2025*). The datasets here are trained convolutional neural network (CNN) models for predicting key microstructural properties of solid oxide cell (SOC) electrodes from low-res, 2-channel 3D images, as well as some helpful code. The parameters for input images are provided in the paper. Sample data is provided in the file `Combined_anode_aug_dual_1k_examples` - that particular data was used to train `anode_all_aug.pth` and will work most accurately with that model. Please familiarize yourself with all caveats on accuracy and applicability, as detailed in the associated paper. Usage -------------------------- The basic usage is as follows, assuming `model_fn` is the path to the .pth file, and `X` is 2-channel input image(s) of the proper dimensions (either one image of shape `[2,12,24,24]`, or a batch of N input images of shape `[N,2,12,24,24]`): from CNN_inferencer import load_model_for_inference model = load_model_for_inference(model_fn) y_predicted = model(X) The model object automatically handles input scaling and output de-scaling based on the way the models were trained - in other words, pass in a 2-channel micro-CT image, and it will output microstructural property values in real units. ## Other model object attributes Note that model has useful attributes other than its forward pass model(X). * `model.output_descaler` - returns the output descaler object. Model does the de-scaling when generating inferences, but you may want to re-use this de-scaler on other values to e.g. compare predictions to ground truth from already-scaled training data. * `model.prop_names` - Gives the property names of the predicted y values, in order. Only exists if there's an output scaler as part of the model object, which there will be in the models provided here. ## Usage with sample data Here is a short script to use with the included sample data. from CNN_inferencer import display_predictions, load_model_for_inference, calculate_mape, parity_plot import h5py import numpy as np model_fn = 'anode_all_aug.pth' data_fn = 'Combined_anode_aug_dual_1k_examples.h5' N_samples = 200 figure_outdir = '.' model = load_model_for_inference(model_fn) with h5py.File(data_fn,'r') as f: XX = f['X'] #These are the 2-channel 3D images yy = f['y'] #These are the ground-truth microstructural properties, but they have been scaled for training - need to de-scale below N = XX.shape[0] #How many images total in the input data file #Run inferences on N_samples random samples from XX. #Run in a batch, much more efficient than one at a time. ii = np.random.choice(N,N_samples,replace=False) ii.sort() y_pred = model(XX[ii]) #Get the original/true (but normalized/scaled) values from the training dataset... #Because they were normalized, they are not in real units yet. So let's also de-scale them using model.output_scaler. y_true = model.output_scaler.transform(yy[ii]) #Let's display actual values for just 5 random ones for i in np.random.choice(N_samples,5,replace=False): display_predictions(y_true[i], y_pred[i], model.prop_names) #Make parity plots for each property (ground truth vs predicted values) #Also label each plot with the mean abs. percent error (MAPE) of the predicted values for i,key in enumerate(model.prop_names): mape = calculate_mape(y_true[:,i], y_pred[:,i]) parity_plot(y_true[:,i], y_pred[:,i], figure_outdir, key, extra_title=f' ({mape:.2f}% MAPE)')

3D microstructure↗

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)↗

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

Increasing the Reproducibility and Replicability of Supervised AI/ML in the Earth Systems Science by Leveraging Social Science Methods

Artificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision-making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well-documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision-making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.

58 GEOSCIENCES↗

3D Continuous Forcing Dataset from 3D Constrained Variational Analysis at SGP

The continuous 3D large-scale forcing (VARANAL3D) data set derived from 3D constrained variational analysis (3DCVA) extends the conventional constrained variational analysis method by incorporating multiple sub-columns within the analysis domain. This advancement introduces spatial variability into the large-scale forcing fields, thereby enriching the data set’s applicability. The VARANAL3D data set spans from 2004 to 2018 and covers a region of 5˚×4.5˚ domain around the ARM SGP site. The analysis domain is divided into 10×9 sub-columns with 0.5˚ resolution. The 3D large-scale forcing data provides necessary variables to drive and evaluate single-column models (SCM), cloud-resolving models (CRM) ,and large-eddy simulations (LES), as well as information for testing model sensitivity to spatial variability of the large-scale forcing data, facilitating more rigorous testing and refinement of physical processes in SCM/CRM/LES.

54 ENVIRONMENTAL SCIENCES↗

Updated ASME design correlations and qualification plan for powder bed fusion 316H stainless steel

This report provides an update on the Advanced Materials and Manufacturing Technologies (AMMT) program effort to qualify Laser-Powder Bed Fusion (L-PBF) 316H stainless steel for use with the ASME Boiler & Pressure Vessel Code Section III, Division 5 rules. The report summarizes progress in testing and characterizing L-PBF material at elevated temperatures by providing preliminary design data for L-PBF 316H and by comparing the elevated temperature performance of the L-PBF material to wrought and conventional fusion welded 316H. The report then updates the initial AMMT qualification plan for L-PBF 316H, originally developed in 2023, to update the accelerated qualification strategy adopted in that plan to account for the new high temperature test data. The report also explores a few methods for further accelerating the qualification process using machine learning techniques to supplement the more conventional, empirical analysis methods typically used by ASME to correlate and extrapolate time-dependent material test data.

36 MATERIALS SCIENCE↗

Investigation of Americium-Containing Phosphates, Silicates, Borates, Molybdates, and Fluorides Synthesized via High-Temperature Flux Crystal Growth

The crystal chemistry of americium-containing extended structures was investigated, and several classes of americium-containing solid-state oxide materials were obtained in single-crystal form via high-temperature flux crystal growth. This enabled the structural characterization of rare examples of ternary, quaternary, and penternary americium-containing silicates K 3 Am (Si 2 O 7 ) and Cs 6 Am 2 Si 21 O 48 , phosphates Na 3 Am (PO 4 ) 2 and K 3 Am (PO 4 ) 2 , borates Ba 3 Am 2 (BO 3 ) 4 and AmBO 3 , borate halides Ca 5 Am(BO 3 ) 4 Cl, molybdates Li 0.5 Am 0.5 MoO 4 , and fluorides CsAm 2 F 7 . Using these crystallographic data, the ionic radii of Am 3+ with coordination numbers of six (0.975 Å), seven (1.052 Å), and nine (1.162 Å) were established. A maximum entropy method (MEM) analysis was performed on the single-crystal X-ray diffraction data that were collected for K 3 Nd(PO 4 ) 2 /K 3 Am(PO 4 ) 2 , K 3 NdSi 2 O 7 /K 3 AmSi 2 O 7 , and NdBO 3 /AmBO 3 , to qualitatively compare the ionicities of the Nd–O and Am–O bonds. In conclusion, Raman spectroscopy data were collected on single crystals of K 3 Am(PO 4 ) 2 and compared to the calculated Raman spectrum of K 3 Am(PO 4 ) 2 obtained from DFT calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗