Search NASA⌕ Search

SEARCH · Search NASA

Results for “massive datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

48 records · Page 3

Sensitivity to low-mass WIMPs with an improved liquid argon ionization response model within the DarkSide program

Dark matter detection experiments using liquid argon rely on a precise characterization of the ionization response to nuclear recoils, especially in the keV energy range relevant for light dark matter interactions. In this work, we present a comprehensive analysis that combines new measurements from the ReD setup, part of the DarkSide experimental program, with calibration data from DarkSide-50, as well as results from the ARIS and SCENE experiments. These combined datasets enable improved constraints on atomic screening effects in the modeling of the ionization response of liquid argon to nuclear recoils. The analysis is performed within the Thomas-Imel recombination framework adopted in previous DarkSide studies, and is here further constrained by the inclusion of ReD data, which allow the screening function to be determined from calibration measurements. By including the updated ionization model into the DarkSide-50 analysis framework, we obtain stronger exclusion limits on low-mass weakly interacting massive particle (WIMP) interactions, setting new world-leading constraints in the 1 – 3 GeV / c 2 WIMP mass range. Finally, we recast the sensitivity projections for the upcoming DarkSide-20k detector, demonstrating a significantly enhanced discovery potential for low-mass dark matter candidates.

Acerbi, F. [Fond. Bruno Kessler, Trento]↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Dark Matter Search Results from 4.2 Tonne−Years of Exposure of the LUX-ZEPLIN (LZ) Experiment

We report results of a search for nuclear recoils induced by weakly interacting massive particle (WIMP) dark matter using the LUX-ZEPLIN (LZ) two-phase xenon time projection chamber. This analysis uses a total exposure of 4.2 ±0.1 tonne-years from 280 live days of LZ operation, of which 3.3 ± 0.1 tonne-years and 220 live days are new. A technique to actively tag background electronic recoils from 214 Pb 𝛽 decays is featured for the first time. Enhanced electron-ion recombination is observed in two-neutrino double electron capture decays of 124 Xe, representing a noteworthy new background. After removal of artificial signal-like events injected into the dataset to mitigate analyzer bias, we find no evidence for an excess over expected backgrounds. World-leading constraints are placed on spin-independent (SI) and spin-dependent WIMP-nucleon cross sections for masses ≥9 GeV/𝑐 2 . The strongest SI exclusion set is 2.2×10 −48 cm 2 at the 90% confidence level and the best SI median sensitivity achieved is 5.1 ×10 −48 cm 2 , both for a mass of 40 GeV/𝑐 2 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Investigating performance and variability of NIF ICF experiments with deep learning

The parameter space involved in designing an inertial confinement fusion shot at the National Ignition Facility (NIF) is massively multi-dimensional and the cost of a single shot makes a comprehensive set of sensitivity studies in the laboratory impractical. The use of machine learning to overcome these challenges has gained popularity and has had several successful applications by the scientific community. We extend on these efforts by training a neural network (NN) on information about the experimental design, engineering elements, and drive asymmetry to predict with uncertainty the neutron yield of an experiment. We find the measured and model predicted values are in good agreement, with an R 2 value of 0.91 for a randomly selected test dataset. Almost all the predicted 95% credible intervals contain the corresponding measured value for both training and test datasets. We identify correlations picked up by the NN between the shot design, yield, and variability and use them to motivate shot sensitivity studies. The first shot to exceed the Lawson-like ignition criteria (N210808) was conducted at the NIF and subsequent shots studied the design’s robustness. In a follow-up shot to N210808, our model predicts capsule quality to be the main performance degradation mechanism that prevented the shot from repeating previous performance levels. Shot N221204 was the first shot to exceed a target energy gain of 1. Our model predicts increased yield with reduced coast time for a N221204 study and greater variability for designs with lower peak powers at constant yield. The model’s fast prediction speed and uncertainty prediction are useful for identifying interesting design paths that could warrant further investigation with conventional simulations to search for robust high yield designs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES↗

Investigation of low-energy particle remnants in high-energy collisions at the LHC with a skipper-CCD detector

We deployed the Mobile Skipper Testing Apparatus ∼33 m away from the Compact Muon Solenoid collision point, the first skipper-CCD detector probing low-energy particles produced in high-energy collisions at the Large Hadron Collider. In this work, we search for beam-related events using data collected in 2024 during beam-on and beam-off periods. The dataset corresponds to integrated luminosities of 113.3 fb −1 and 1.54 nb −1 for the proton-proton and Pb-Pb collision periods, respectively. We report observed event rates in a model-independent framework across two ionization regions: ≤ 20⁢𝑒 − and > 20⁢𝑒 − . For the low-energy region, we perform a likelihood analysis to test the null hypothesis of no beam-correlated signal. We found no significant correlation during proton-proton and Pb-Pb collisions. For the high-energy region, we present the energy spectra for both collision periods and compare event rates for images with and without luminosity. We observe a slight increase in the event rate following the Pb-Pb collisions, coinciding with a rise in the single-electron rate, which will be investigated in future work. Using the low-energy proton-proton results, we place 95% confidence level constraints on the mass-millicharge parameter space of millicharged particles. Overall, the results in this work demonstrate the viability of skipper-CCD technology to explore new physics at high-energy colliders and motivate future searches with more massive detectors.

Cervantes-Vergara, Brenda A. [Fermi National Accel↗

Multiprobe cosmology from the abundance of SPT clusters and DES galaxy clustering and weak lensing

Cosmic shear, galaxy clustering, and the abundance of massive halos each probe the large-scale structure of the Universe in complementary ways. We present cosmological constraints from the joint analysis of the three probes, building on the latest analyses of the lensing-informed abundance of clusters identified by the South Pole Telescope (SPT) and of the auto- and cross-correlation of galaxy position and weak lensing measurements (3 × 2 ⁢pt) in the Dark Energy Survey (DES). We consider the cosmological correlation between the different tracers and we account for the systematic uncertainties that are shared between the large-scale lensing correlation functions and the small-scale lensing-based cluster mass calibration. Marginalized over the remaining Λ cold dark matter (Λ ⁢CDM) parameters (including the sum of neutrino masses) and 52 astrophysical modeling parameters, we measure Ω m = 0.300 ± 0.017 and 𝜎 8 = 0.797 ± 0.026. Compared to constraints from Planck primary cosmic microwave background (CMB) anisotropies, our constraints are only 15% wider with a probability to exceed of 0.22 (1.2⁢𝜎) for the two-parameter difference. We further obtain 𝑆 8 ≡𝜎 8 ⁢(Ω m /0.3) 0.5 = 0.796 ± 0.013 which is lower than the Planck measurement at the 1.6⁢𝜎 level. The combined SPT cluster, DES 3 ×2 ⁢pt, and Planck datasets mildly prefer a nonzero positive neutrino mass, with a 95% upper limit ∑ 𝑚 𝜈 < 0.25 eV on the sum of neutrino masses. Assuming a 𝑤⁢CDM model, we constrain the dark energy equation of state parameter 𝑤 = −1.1⁢5$^{+0.23}_{−0.17}$ and when combining with Planck primary CMB anisotropies, we recover 𝑤 = −1.2⁢0$^{+0.15}_{−0.09}$, a 1.7⁢𝜎 difference with a cosmological constant. The precision of our results highlights the benefits of multiwavelength multiprobe cosmology and our analysis paves the way for upcoming joint analyses of next-generation datasets.

79 ASTRONOMY AND ASTROPHYSICS↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Search for cascade decays of charged sleptons and sneutrinos in final states with three leptons and missing transverse momentum in 𝑝⁢𝑝 collisions at $\sqrt{𝑠} = 13$ TeV with the ATLAS detector

A search for cascade decays of charged sleptons and sneutrinos using final states characterized by three leptons (electrons or muons) and missing transverse momentum is presented. The analysis is based on a dataset with 140 fb −1 of proton-proton (pp) collisions at a center-of-mass energy of $\sqrt{𝑠} = 13$ TeV recorded by the ATLAS detector at the Large Hadron Collider. This paper focuses on a supersymmetric scenario that is motivated by the muon anomalous magnetic moment observation, dark-mattter relic density abundance, and electroweak naturalness. A mass spectrum involving light Higgsinos and heavier sleptons with a bino at intermediate mass is targeted. No significant deviation from the Standard Model expectation is observed. This search enables us to place stringent constraints on this model, excluding at the 95% confidence level charged slepton and sneutrino masses up to 450 GeV when assuming a lightest neutralino mass of 100 GeV and mass-degenerate selectrons, smuons and sneutrinos.

extensions of Higgs sector↗