Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data-analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Jas4pp — A data-analysis framework for physics and detector studies

This paper describes the Jas4pp framework for exploring physics cases and for detector-performance studies of future particle collision experiments. Jas4pp is a multi-platform Java program for numeric calculations, scientific visualization in 2D and 3D, storing data in various file formats and displaying collision events and detector geometries. It also includes complex data-analysis algorithms for function minimization, regression analysis, event reconstruction (such as jet reconstruction), limit settings and other libraries widely used in particle physics. The framework can be used with several scripting languages, such as Python/Jython, Groovy and JShell. Several benchmark tests discussed in the paper illustrate significant improvements in the performance of the Groovy and JShell scripting languages compared to the standard Python implementation in C. Furthermore, the improvements for numeric computations in Java are attributed to recent enhancements in the Java Virtual Machine.

97 MATHEMATICS AND COMPUTING↗

Constraining Hamiltonians from chiral effective field theory with neutron-star data

Multi-messenger observations of neutron stars (NSs) and their mergers have placed strong constraints on the dense-matter equation of state (EOS). The EOS, in turn, depends on microscopic nuclear interactions that are described by nuclear Hamiltonians. These Hamiltonians are commonly derived within chiral effective field theory (EFT). Ideally, multi-messenger observations of NSs could be used to directly inform our understanding of EFT interactions, but such a direct inference necessitates millions of model evaluations. This is computationally prohibitive because each evaluation requires us to calculate the EOS from a Hamiltonian by solving the quantum many-body problem with methods such as auxiliary-field diffusion Monte Carlo (AFDMC), which provides very accurate and precise solutions but at a significant computational cost. Additionally, we need to solve the stellar structure equations for each EOS which further slows down each model evaluation by a few seconds. In this work, we combine emulators for AFDMC calculations of neutron matter, built using parametric matrix models, and for the stellar structure equations, built using multilayer perceptron neural networks, with the PyCBC data-analysis framework to enable a direct inference of coupling constants in an EFT Hamiltonian using multi-messenger observations of NSs. We find that astrophysical data can provide informative constraints on two-nucleon couplings despite the high densities probed in NS interiors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Planck 2018 results. V. CMB power spectra and likelihoods

We describe the legacy Planck cosmic microwave background (CMB) likelihoods derived from the 2018 data release. The overall approach is similar in spirit to the one retained for the 2013 and 2015 data release, with a hybrid method using different approximations at low ( ℓ < 30) and high ( ℓ ≥ 30) multipoles, implementing several methodological and data-analysis refinements compared to previous releases. With more realistic simulations, and better correction and modelling of systematic effects, we can now make full use of the CMB polarization observed in the High Frequency Instrument (HFI) channels. The low-multipole EE cross-spectra from the 100 GHz and 143 GHz data give a constraint on the ΛCDM reionization optical-depth parameter τ to better than 15% (in combination with the TT low- ℓ data and the high- ℓ temperature and polarization data), tightening constraints on all parameters with posterior distributions correlated with τ . We also update the weaker constraint on τ from the joint TEB likelihood using the Low Frequency Instrument (LFI) channels, which was used in 2015 as part of our baseline analysis. At higher multipoles, the CMB temperature spectrum and likelihood are very similar to previous releases. A better model of the temperature-to-polarization leakage and corrections for the effective calibrations of the polarization channels (i.e., the polarization efficiencies) allow us to make full use of polarization spectra, improving the ΛCDM constraints on the parameters θ MC , ω c , ω b , and H 0 by more than 30%, and n s by more than 20% compared to TT-only constraints. Extensive tests on the robustness of the modelling of the polarization data demonstrate good consistency, with some residual modelling uncertainties. At high multipoles, we are now limited mainly by the accuracy of the polarization efficiency modelling. Using our various tests, simulations, and comparison between different high-multipole likelihood implementations, we estimate the consistency of the results to be better than the 0.5 σ level on the ΛCDM parameters, as well as classical single-parameter extensions for the joint likelihood (to be compared to the 0.3 σ levels we achieved in 2015 for the temperature data alone on ΛCDM only). Minor curiosities already present in the previous releases remain, such as the differences between the best-fit ΛCDM parameters for the ℓ < 800 and ℓ > 800 ranges of the power spectrum, or the preference for more smoothing of the power-spectrum peaks than predicted in ΛCDM fits. These are shown to be driven by the temperature power spectrum and are not significantly modified by the inclusion of the polarization data. Overall, the legacy Planck CMB likelihoods provide a robust tool for constraining the cosmological model and represent a reference for future CMB observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating End-to-End Deep Learning for Particle Reconstruction using CMS open data

Machine learning algorithms are gaining ground in high energy physics for applications in particle and event identification, physics analysis, detector reconstruction, simulation and trigger. Currently, most data-analysis tasks at LHC experiments benefit from the use of machine learning. Incorporating these computational tools in the experimental framework presents new challenges. This paper reports on the implementation of the end-to-end deep learning with the CMS software framework and the scaling of the end-to-end deep learning with multiple GPUs. The end-to-end deep learning technique combines deep learning algorithms and low-level detector representation for particle and event identification. We demonstrate the end-to-end implementation on a top quark benchmark and perform studies with various hardware architectures including single and multiple GPUs and Google TPU.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Protocols and methodologies for acquiring and analyzing critical-current versus longitudinal-strain data in Bi 2 Sr 2 CaCu 2 O 8+x wires

Abstract In the literature on Bi 2 Sr 2 CaCu 2 O 8+ x (Bi-2212) superconducting wires, it is evident that measurement protocols for transport critical-current I c versus longitudinal strain ϵ and definitions of the so-called ‘strain limit’ are generally dissimilar. Yet, values obtained for the ‘strain limit’ are frequently assimilated to being those of the irreversible strain limit ϵ irr , regardless of the I c degradation-criterion used to define it. In effect, ϵ irr should correspond specifically to the I c ( ϵ ) irreversibility onset , where crack formation in Bi-2212 filaments presumably starts. Because I c ( ϵ ) degradation remains progressive over a fairly wide strain range beyond ϵ irr , the different I c degradation-criteria in use do not yield to the same result and, thus, are not equivalent from metrology perspective. Indeed, in studying densified samples of a modern Bi-2212 round wire, we found ϵ irr ≈ 0.4% and ϵ 5% ≈ 0.6% ( ϵ 5% being the strain where I c degrades by 5%). In this paper, we outline and suggest I c ( ϵ )-measurement protocols and data-analysis methodologies in the hope to converge the various approaches taken for studying Bi-2212 strain properties and, thus, remove related result discrepancies. A unified approach would enable more objective data comparisons among laboratories and among different Bi-2212 conductors. It would pave the way for more rigorous studies of effects potentially associated with wire design, powder, heat treatments, and other such parameters on the conductor’s strain properties.

protocols↗

Sub-m s−1 upper limits from a deep HARPS-N radial-velocity search for planets orbiting HD 166620 and HD 144579

ABSTRACT Minimizing the impact of stellar variability in radial velocity (RV) measurements is a critical challenge in achieving the 10 cm s−1 precision needed to hunt for Earth twins. Since 2012, a dedicated programme has been underway with HARPS-N, to conduct a blind RV rocky planets search (RPS) around bright stars in the Northern hemisphere. Here we describe the results of a comprehensive search for planetary systems in two RPS targets, HD 166620 and HD 144579. Using wavelength-domain line-profile decorrelation vectors to mitigate the stellar activity and performing a deep search for planetary reflex motions using a trans-dimensional nested sampler, we found no significant planetary signals in the data sets of either of the stars. We validated the results via data-splitting and injection recovery tests. Additionally, we obtained the 95th percentile detection limits on the HARPS-N RVs. We found that the likelihood of finding a low-mass planet increases noticeably across a wide period range when the inherent stellar variability is corrected for using scalpelsU-vectors. We are able to detect planet signals with Msin i ≤ 1 M⊕ for orbital periods shorter than 10 d. We demonstrate that with our decorrelation technique, we are able to detect signals as low as 54 cm s−1, which brings us closer to the calibration limit of 50 cm s−1 demonstrated by HARPS-N. Therefore, we show that we can push down towards the RV precision required to find Earth analogues using high-precision radial velocity data with novel data-analysis techniques.

Anna John, A. (ORCID:0000000217156939)↗

Efficient lattice QCD computation of radiative-leptonic-decay form factors at multiple positive and negative photon virtualities

In previous work [D. Giusti, Methods for high-precision determinations of radiative-leptonic decay form factors using lattice QCD, Phys. Rev. D 107, 074507 (2023)], we showed that form factors for radiative leptonic decays of pseudoscalar mesons can be determined efficiently and with high precision from lattice QCD using the “three-dimensional (3D) method,” in which three-point functions are computed for all values of the current insertion time and the time integral is performed at the data-analysis stage. Here, we demonstrate another benefit of the 3D method: the form factors can be extracted for any number of nonzero photon virtualities from the same three-point functions at no extra cost. We present results for the $D_s → ℓνγ*$ vector form factor as a function of photon energy and photon virtuality, for both positive and negative virtuality, for a single ensemble with 340 MeV pion mass and 0.11 fm lattice spacing. In our analysis, we separately consider the two different time orderings and the different quark flavors in the electromagnetic current. We discuss in detail the behavior of the unwanted exponentials contributing to the three-point functions, as well as the choice of fit models and fit ranges used to remove them for various values of the virtuality. While positive photon virtuality is relevant for decays to multiple charged leptons, negative photon virtuality suppresses soft contributions and is of interest in QCD-factorization studies of the form factors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Pixel Anomaly Detection Tool : a user-friendly GUI for classifying detector frames using machine-learning approaches

Data collection at X-ray free electron lasers has particular experimental challenges, such as continuous sample delivery or the use of novel ultrafast high-dynamic-range gain-switching X-ray detectors. This can result in a multitude of data artefacts, which can be detrimental to accurately determining structure-factor amplitudes for serial crystallography or single-particle imaging experiments. Here, a new data-classification tool is reported that offers a variety of machine-learning algorithms to sort data trained either on manual data sorting by the user or by profile fitting the intensity distribution on the detector based on the experiment. This is integrated into an easy-to-use graphical user interface, specifically designed to support the detectors, file formats and software available at most X-ray free electron laser facilities. The highly modular design makes the tool easily expandable to comply with other X-ray sources and detectors, and the supervised learning approach enables even the novice user to sort data containing unwanted artefacts or perform routine data-analysis tasks such as hit finding during an experiment, without needing to write code.

47 OTHER INSTRUMENTATION↗

X-ray scattering based scanning tomography for imaging and structural characterization of cellulose in plants

X-ray and neutron scattering have long been used for structural characterization of cellulose in plants. Due to averaging over the illuminated sample volume, these measurements traditionally overlooked the compositional and morphological heterogeneity within the sample. Here, a scanning tomographic imaging method is described, using contrast derived from the X-ray scattering intensity, for virtually sectioning the sample to reveal its internal structure at a resolution of a few micrometres. This method provides a means for retrieving the local scattering signal that corresponds to any voxel within the virtual section, enabling characterization of the local structure using traditional data-analysis methods. This is accomplished through tomographic reconstruction of the spatial distribution of a handful of mathematical components identified by non-negative matrix factorization from the large dataset of X-ray scattering intensity. Joint analysis of multiple datasets, to find similarity between voxels by clustering of the decomposed data, could help elucidate systematic differences between samples, such as those expected from genetic modifications, chemical treatments or fungal decay. The spatial distribution of the microfibril angle can also be analyzed, based on the tomographically reconstructed scattering intensity as a function of the azimuthal angle.

36 MATERIALS SCIENCE↗

Data from: "Towards CONUS-Wide ML-Augmented Conceptually-Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics"

This data package was generated to support the manuscript “Towards CONUS-Wide Machine Learning-Augmented Conceptually Interpretable Modeling of Catchment-Scale Precipitation-Storage-Runoff Dynamics.” It provides input files, model outputs, plotting data, scripts, notebooks, and documentation used to develop, evaluate, and reproduce Mass-Conserving Perceptron (MCP)-based hydrologic modeling experiments across 513 selected Catchment Attributes and Meteorology for Large-sample Studies in the United States (CAMELS-US) basins. The files are organized by modeling component and analysis purpose, including rainfall–runoff experiments, snow module experiments, coupled hydrologic-snow experiments, Long Short-Term Memory (LSTM) benchmark results, model skill metrics, initialization and epoch records, cell-state normalization files, Akaike Information Criterion (AIC)-based model comparison files, and data used to generate manuscript figures. Tabular files can be opened using standard spreadsheet software or Python/R data-analysis tools. Python scripts, Jupyter notebooks, and selected MATLAB scripts are included for model execution, postprocessing, plotting, and statistical analysis. Quality assurance and quality control were conducted through the source-data selection and modeling workflow. Meteorological forcing, streamflow, and static catchment attributes were derived from the CAMELS-US dataset, and snow water equivalent data were derived from the University of Arizona (UA) Snow Water Equivalent dataset. Selected basins and time periods were screened during the associated research workflow to avoid missing observations or poor-quality cases. Static geospatial features were processed primarily using Quantum Geographic Information System (QGIS) and Geospatial Data Abstraction Library (GDAL) workflows. Additional details are provided in the associated manuscript and documentation.

ESS-DIVE CSV File Formatting Guidelines Reporting ↗

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING↗

Robust Dark Energy Constraints with the Dark Energy Spectroscopic Survey (Final Technical Report)

This project developed and applied advanced theoretical, computational, and data-analysis methodologies to extract robust and precise cosmological constraints from the Dark Energy Spectroscopic Instrument (DESI). The work focused on maximizing the scientific return of DESI through optimized survey strategy, novel higher-order clustering statistics, improved modeling of small-scale structure, and rigorous mitigation of observational systematics. Over the award period, the project made substantial contributions to DESI science planning, produced new methods for bispectrum and three-point correlation function analyses, advanced constraints on primordial non-Gaussianity, and delivered widely used software tools. The project also played a major role in training graduate students and a postdoctoral researcher who contributed directly to DESI key projects. The results have significantly enhanced the cosmological reach of DESI and provide a strong foundation for future surveys such as DESI-II and Stage-V experiments.

79 ASTRONOMY AND ASTROPHYSICS↗

The Planetary Ephemeris Program: Capability, Comparison, and Open Source Availability

We describe for the first time in scientific literature the Planetary Ephemeris Program (PEP), an open-source general-purpose astrometric data-analysis program. We discuss, in particular, the implementation of pulsar timing analysis, which was recently upgraded in PEP to handle more options. This implementation was done independently of other pulsar programs, with minor exceptions that we discuss. We illustrate the implementation of this capability by comparing the postfit residuals from the analyses of time-of-arrival observations by both PEP and Tempo2. The comparison shows substantial agreement: 22 ns rms differences for 1065 pulse time-of-arrival measurements for the millisecond pulsar in a binary system, PSR J1909-3744 (pulse period 2.947108 ms; full width half maximum of pulse 43 μs), for epochs in the interval from 2002 December to 2011 February.

42 ENGINEERING↗

Kinematic Variables and Feature Engineering for Particle Phenomenology

Kinematic variables have been playing an important role in collider phenomenology, as they expedite discoveries of new particles by separating signal events from unwanted background events and allow for measurements of particle properties such as masses, couplings, spins, etc. For the past 10 years, an enormous number of kinematic variables have been designed and proposed, primarily for the experiments at the Large Hadron Collider, allowing for a drastic reduction of high-dimensional experimental data to lower-dimensional observables, from which one can readily extract underlying features of phase space and develop better-optimized data-analysis strategies. We review these recent developments in the area of phase space kinematics, summarizing the new kinematic variables with important phenomenological implications and physics applications. We also review recently proposed analysis methods and techniques specifically designed to leverage the new kinematic variables. As machine learning is nowadays percolating through many fields of particle physics including collider phenomenology, we discuss the interconnection and mutual complementarity of kinematic variables and machine learning techniques. We finally discuss how the utilization of kinematic variables originally developed for colliders can be extended to other high-energy physics experiments including neutrino experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A discrete integral transform for rapid spectral synthesis

Accurate synthetic spectra that rely on large Line-By-Line (LBL)-databases are used in a wide range of applications such as high temperature combustion, atmospheric re-entry, planetary surveillance and laboratory plasmas. Conventionally synthetic spectra are calculated by computing a lineshape for every spectral line in the database and adding those together, which may take multiple hours for large databases. In this paper we propose a new approach for spectral synthesis based on an integral transform: the synthetic spectrum is calculated as the integral over the product of a Voigt profile and a newly proposed three-dimensional “lineshape distribution function”, which is a function of spectral position and Gaussian- & Lorentzian width coordinates. A fast discrete version of this transform based on the Fast Fourier Transform (FFT) is proposed, which improves performance compared to the conventional approach by several orders of magnitude while maintaining accuracy. Strategies that minimize the discretization error are discussed. A Python implementation of the method is compared against state-of-the-art spectral code RADIS, and is since adopted as RADIS's default synthesis method. The synthesis of a benchmark CO2 spectrum consisting of 1.8 M spectral lines and 200k spectral points took only 3.1 s using the proposed method (1011 lines × spectral points/s), a factor ~300 improvement over the state-of-the-art, with the relative improvement generally increasing for higher number of lines and/or number of spectral points. Finally, an experimental GPU-implementation of the method was also benchmarked, which demonstrated another 2~3 orders performance increase, achieving up to 5 ∙ 10 14 lines × spectral points/s.

42 ENGINEERING↗