Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Modeling the Behavior of Complex Aqueous Electrolytes Using Machine Learning Interatomic Potentials: The Case of Sodium Sulfate

Understanding the structure and thermodynamics of solvated ions is essential for advancing applications in electrochemistry, water treatment, and energy storage. While ab initio molecular dynamics methods are highly accurate, they are limited by short accessible time and length scales whereas classical force fields struggle with accuracy. Herein, we explore the structure and thermodynamics of complex monovalent-divalent ion pairs using Na 2 SO 4 (aq) as a case study by applying a machine learning interatomic potential (MLIP) trained on density functional theory (DFT) data. Our MLIP-based approach reproduces key bulk properties such as density and radial distribution functions of water. We provide the hydration structure of the sodium and sulfate ions in the 0.1–2 M concentration range and the one-dimensional and two-dimensional potentials of mean force for the sodium–sulfate ion pairing at the low concentration limit (0.1 M), which are inaccessible to DFT. At low concentrations, the sulfate ion is strongly solvated, leading to the stabilization of solvent-separated ion pairs over contact ion pairs. Minimum energy pathway analysis revealed that coordinating two sodium ions with a sulfate ion is a multistep process whereby the sodium ions coordinate to the sulfate ion sequentially. Finally, we demonstrate that MLIPs allow the study of solvated ions beyond simple monovalent pairs with DFT-level accuracy in their low concentration limit (0.1 M) via statistically converged properties from ns-long simulations.

anions↗

Tetracene Functionalized Si(111) Achieves Enhanced Solar-to-Chemical Energy Conversion via Molecular Acceptor States

The properties of semiconductor|liquid interfaces play a critical role in determining the efficiency of solar-to-hydrogen (STH) conversion. Here, we investigate how molecular functionalization of Si(111) and Si(111)|TiO 2 surfaces impacts photoelectrochemical (PEC) hydrogen production efficiency. We find that functionalization of ∼3% of the atop sites of Si(111) with either 9-anthracene (Anth) or 5-tetracene (Tet), with the remaining sites passivated by methyl groups, provides substrates with high electronic quality and low surface oxide densities, as determined by X-ray photoelectron spectroscopy (XPS) measurements. Surface photovoltage (SPV) spectroscopy shows that surfaces modified with Anth or Tet exhibit an increased photovoltage, with Tet-functionalized surfaces yielding an additional 192 meV relative to methyl-terminated Si(111), indicating improved charge separation for Si-Tet. Further improvement in onset potential was achieved by replacing a nitrogen-containing TiO 2 atomic layer deposition (ALD) precursor (TDMAT) with a precursor lacking nitrogen (TTIP), which eliminates the parasitic defect band in the TiO 2 overlayer (p-Si(111)-Tet|TTIP-TiO 2 |Pt: V OC = +0.283 ± 0.041 V vs RHE). Density functional theory (DFT) analysis demonstrates that compared with Anth-modified Si(111), the Tet-modified surface exhibits more hybridized Si(111)-Tet states closer to the silicon band edges. Mercury contact current–voltage (I–V, dark) measurements quantified the relative interfacial density of states of Si-Tet, Si-Anth and Si-Me surfaces─revealing that the interfacial state density was highest for Si-Tet. This suggests that such hybridized interfaces serve to capture better photoexcited charge, which enables facile electron transfer to molecular acceptors in solution. Altogether, the data indicate that beneficial hybrid molecular LUMO surface states interacting with the Si conduction band edge results in improved hydrogen evolution (HER) performance for p-Si devices.

Group theory↗

Scattering-based structural inversion of soft materials via Kolmogorov–Arnold networks

Small-angle scattering techniques are indispensable tools for probing the structure of soft materials. However, traditional analytical models often face limitations in structural inversion for complex systems, primarily due to the absence of closed-form expressions of scattering functions. To address these challenges, we present a machine learning framework based on the Kolmogorov–Arnold Network (KAN) for directly extracting real-space structural information from scattering spectra in reciprocal space. This model-independent, data-driven approach provides a versatile solution for analyzing intricate configurations in soft matter. By applying the KAN to lyotropic lamellar phases and colloidal suspensions—two representative soft matter systems—we demonstrate its ability to accurately and efficiently resolve structural collectivity and complexity. Here, our findings highlight the transformative potential of machine learning in enhancing the quantitative analysis of soft materials, paving the way for robust structural inversion across diverse systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Effect of Time Window and Spectral Measurement Options on Empirical Green’s Function Analysis Using DAS Array and Seismic Stations

The recorded seismic waveform is a convolution of event source term, path term, and station term. Removing high-frequency attenuation due to path effect is a challenging problem. Empirical Green’s function (EGF) method uses nearly collocated small earthquakes to correct the path and station terms for larger events recorded at the same station. However, this method is subject to variability due to many factors. Here, we focus on three events that were well recorded by the seismic network and a rapid response distributed acoustic sensing (DAS) array. Using a suite of high-quality EGF events, we assess the influence of time window, spectral measurement options, and types of data on the spectral ratio and relative source time function (RSTF) results. Increased number of tapers (from 2 to 16) tends to increase the measured corner frequency and reduce the source complexity. Extended long time window (e.g., 30 s) tends to produce larger variability of corner frequency. The multitaper algorithm that simultaneously optimizes both target and EGF spectra produces the most stable corner-frequency measurements. The stacked spectral ratio and RSTF from the DAS array are more stable than two nearby seismic stations, and are comparable to stacked results from the seismic network, suggesting that DAS array has strong potential in source characterization.

58 GEOSCIENCES↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

DESI DR1 Lyα 1D power spectrum: the optimal estimator measurement

The one-dimensional power spectrum P 1D of Lyα forest offers rich insights into cosmological and astrophysical parameters, including constraints on the sum of neutrino masses, warm dark matter models, and the thermal state of the intergalactic medium. We present the measurement of P 1D using the optimal quadratic maximum likelihood estimator applied to over 300,000 Lyα quasars from Data Release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI) survey. This sample represents the largest to date for P 1D measurements and is larger than the Extended Baryon Oscillation Spectroscopic Survey (eBOSS) by a factor of 1.7. We conduct a meticulous investigation of instrumental and analysis systematics and quantify their impact on P 1D . This includes the development of a cross-exposure estimator that eliminates the need to model the pipeline noise and has strong potential for future P 1D measurements. We also present new insights into metal contamination through the 1D correlation function. Using a fitting function we measure the evolution of the Lyα forest bias with high precision: b F (z) = (-0.218 ± 0.002) × ((1 + z)/4) 2.96±0.06 . In a companion validation paper, we substantially extend our previous suite of CCD image simulations to quantify the pipeline's exquisite performance accurately. In another companion paper, we present DR1 P 1D measurements using the Fast Fourier Transform (FFT) approach to power spectrum estimation. These two measurements produce a forest bias parameter that differs by 2.2 sigma. However, our model is simplistic, so this disagreement will be investigated in future work.

Lyman alpha forest↗

Search for photons above 10 18 eV by simultaneously measuring the atmospheric depth and the muon content of air showers at the Pierre Auger Observatory

The Pierre Auger Observatory is the most sensitive instrument to detect photons with energies above 1 0 17 eV . It measures extensive air showers generated by ultrahigh energy cosmic rays using a hybrid technique that exploits the combination of a fluorescence detector with a ground array of particle detectors. The signatures of a photon-induced air shower are a larger atmospheric depth of the shower maximum ( X max ) and a steeper lateral distribution function, along with a lower number of muons with respect to the bulk of hadron-induced cascades. In this work, a new analysis technique in the energy interval between 1 and 30 EeV ( 1 EeV = 1 0 18 eV ) has been developed by combining the fluorescence detector-based measurement of X max with the specific features of the surface detector signal through a parameter related to the air shower muon content, derived from the universality of the air shower development. No evidence of a statistically significant signal due to photon primaries was found using data collected in about 12 years of operation. Thus, upper bounds to the integral photon flux have been set using a detailed calculation of the detector exposure, in combination with a data-driven background estimation. The derived 95% confidence level upper limits are 0.0403, 0.01113, 0.0035, 0.0023, and 0.0021 km − 2 sr − 1 yr − 1 above 1, 2, 3, 5, and 10 EeV, respectively, leading to the most stringent upper limits on the photon flux in the EeV range. Compared with past results, the upper limits were improved by about 40% for the lowest energy threshold and by a factor 3 above 3 EeV, where no candidates were found and the expected background is negligible. The presented limits can be used to probe the assumptions on chemical composition of ultrahigh energy cosmic rays and allow for the constraint of the mass and lifetime phase space of super-heavy dark matter particles. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

ICE Calculator 2.0: Final Report for Phase 1 of the National Initiative to Update the Interruption Cost Estimate (ICE) Calculator

In 2021, Berkeley Lab and Resource Innovations, Inc. launched the “ICE 2.0 Initiative” – a national study to refresh the underlying data and enhance the functionality of the ICE Calculator. The Initiative involves Berkeley Lab contracting with sponsoring utilities to administer identical, updated and comprehensive interruption cost surveys to statistically representative samples of each utility’s customers. Berkeley Lab and Resource Innovations then pool the survey results across the utilities and use them to update the analytical engines that drive the ICE Calculator. The ICE 2.0 Initiative is being conducted in phases. Each phase involves the administration of interruption cost surveys to the customers of sponsoring utilities, followed by an update to the ICE Calculator based on analysis of the pooled survey results. This report describes the activities and findings from Phase 1 of the ICE 2.0 Initiative. Phase 1 was sponsored by eight utilities: American Electric Power, Commonwealth Edison, Dominion Energy, Duke Energy, DTE Electric, Exelon, National Grid, and Puget Sound Energy. Phase 1 involved 11 customer interruption cost survey activities representing a total of 24 electricity distribution service territories, 23 of them located in the Eastern and Midwestern regions of the U.S. and one located in the Pacific Northwest. ICE 2.0 vs. 1.0 Comparison This memorandum compares customer power interruption costs estimated using the recently updated Interruption Cost Estimate (ICE) Calculator (“ICE 2.0”) to the original ICE Calculator (“ICE 1.0”). ICE 1.0 was developed in 2009 based on 15 independent power interruption cost surveys conducted by 10 electric utilities between 1989 and 2012. ICE 2.0 was developed in 2025 through a national initiative based on a consistent set of power interruption cost surveys and 11 surveying efforts conducted across 24 electric utility service territories between 2022 and 2024.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Improving the User Interface of the DeepLynx Data Warehouse

DeepLynx is an open-source ontology-based data warehouse created by INL to support the creation and life cycle of digital engineering projects, with a particular emphasis on digital twins [1]. Digital twins are systems that represent physical assets and process in a real-time digital environment [1]. Most well-known commercial data warehouses use Graphical User Interfaces (GUIs) for users to interact with their systems [3]. Limited publications have addressed the design of these interfaces and understanding of their target users. The current users and development team acknowledge the need to improve the current UI, not just for aesthetics but to improve functionality and workflow of DeepLynx. Traditional data warehouse users are developers, data scientists and business analysts [2]. DeepLynx users have a vast range of experience using data warehouses, and diverse roles, including engineers, scientists and management positions. Because there is a broader audience of target users for DeepLynx than a typical data warehouse, it is essential that DeepLynx has a useable and intuitive user interface. To achieve this the team performed human-computer interaction methods, including a Heuristic Evaluation of current UI using Neilsen’s Usability Heuristic, create personas based on current users by designing a user survey, data analysis and develop of personas. Followed by a redesign of the UI following using Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design in industry standard software Figma. Lastly a Heuristic Evaluation of new UI design, using Neilsen’s Usability Heuristic and User testing of redesign UI and have a group of users complete a Thinking Aloud Test of the new UI. Preliminary results of the Heuristic Evaluation of current UI arise issue with Consistency and Standards, Visibility of System Status, Match System and Real World and Recognition Rather than Recall. These issues were addressed in the proposed redesign by applying Neilsen’s Usability Heuristic and Norman’s Principles of Interactive Design. Next steps include formalized list of lessons learned and design implications for future publications.

97 MATHEMATICS AND COMPUTING↗

Accurate and reliable thermochemistry by data analysis of complex thermochemical networks using Active Thermochemical Tables: the case of glycine thermochemistry

Active Thermochemical Tables (ATcT) were successfully used to resolve the existing inconsistencies related to the thermochemistry of glycine, based on statistically analyzing and solving a thermochemical network that includes >3350 chemical species interconnected by nearly 35 000 thermochemically-relevant determinations from experiment and high-level theory. Here, the current ATcT results for the 298.15 K enthalpies of formation are −394.70 ± 0.55 kJ mol −1 for gas phase glycine, −528.37 ± 0.20 kJ mol −1 for solid α-glycine, −528.05 ± 0.22 kJ mol −1 for β-glycine, −528.64 ± 0.23 kJ mol −1 for γ-glycine, −514.22 ± 0.20 kJ mol −1 for aqueous undissociated glycine, and −470.09 ± 0.20 kJ mol −1 for fully dissociated aqueous glycine at infinite dilution. In addition, a new set of thermophysical properties of gas phase glycine was obtained from a fully corrected nonrigid rotor anharmonic oscillator (NRRAO) partition function, which includes all conformers. Corresponding sets of thermophysical properties of α-, β-, and γ-glycine are also presented.

Active Thermochemical Tables↗

Data for Sugar Accumulation Enhancement in Sorghum Stem is Associated with Reduced Reproductive Sink Strength and Increased Phloem Unloading Activity

Sweet sorghum has emerged as a promising source of bioenergy mainly due to its high biomass and high soluble sugar yield in stems. Studies have shown that loss-of-function Dry locus alleles have been selected during sweet sorghum domestication, and decapitation can further boost sugar accumulation in sweet sorghum, indicating that the potential for improving sugar yields is yet to be fully realized. To maximize sugar accumulation, it is essential to gain a better understanding of the mechanism underlying the massive accumulation of soluble sugars in sweet sorghum stems in addition to the Dry locus. We performed a transcriptomic analysis upon decapitation of near-isogenic lines for mutant (d, juicy stems, and green leaf midrib) and functional (D, dry stems and white leaf midrib) alleles at the Dry locus. Our analysis revealed that decapitation suppressed photosynthesis in leaves, but accelerated starch metabolic processes in stems. SbbHLH093 negatively correlates with sugar levels supported by genotypes (DD vs. dd), treatments (control vs. decapitation), and developmental stages post anthesis (3d vs.10d). D locus gene SbNAC074A and other programmed cell death-related genes were down regulated by decapitation, while sugar transporter-encoding gene SbSWEET1A was induced. Both SbSWEET1A and Invertase 5 were detected in phloem companion cells by RNA in situ assay. Loss of the SbbHLH093 homolog, AtbHLH093, in Arabidopsis led to a sugar accumulation increase. This study provides new insights into sugar accumulation enhancement in bioenergy crops, which can be potentially achieved by reducing reproductive sink strength and enhancing phloem unloading.

Transcriptomics↗

ICE Calculator 2: Final Report for Phase 1 and 2 of the National Initiative to Update the Interruption Cost Estimate (ICE) Calculator

ICE 2.0 Phase 2 Final Report In 2021, Berkeley Lab and Resource Innovations, Inc. launched the “ICE Calculator 2 Initiative” – a national study to refresh the underlying data and enhance the functionality of the ICE Calculator. The Initiative involves Berkeley Lab contracting with sponsoring utilities to administer identical, updated and comprehensive interruption cost surveys to statistically representative samples of each utility’s customers. Berkeley Lab and Resource Innovations then pool the survey results across the utilities and use them to update the analytical engines that drive the ICE Calculator. The ICE Calculator 2 Initiative is being conducted in phases. Each phase involves the administration of interruption cost surveys to the customers of sponsoring utilities, followed by an update to the ICE Calculator based on analysis of the pooled survey results. This report describes the activities and findings from Phase 1 and 2 of the ICE Calculator 2 Initiative. Phase 1 was sponsored by eight utilities: American Electric Power, Commonwealth Edison, Dominion Energy, Duke Energy, DTE Electric, Exelon, National Grid, and Puget Sound Energy. Phase 2 was sponsored by six utilities: Empire District Electric Company, Evergy Missouri, Pacific Gas & Electric, San Diego Gas & Electric, Southern California Edison, and Union Electric. Phase 1 and 2 involved 15 customer interruption cost survey activities representing a total of 30 electricity distribution service territories. ICE Calculator Version 2.0 and 2.2 Comparison This memorandum describes–at a high-level–the improvements in interruption cost estimates for the version 2.2 of the ICE Calculator (released February 2026) compared to version 2.0 (released in April 2025). Version 2.2 of the ICE Calculator corresponds to Phase 2 of the initiative, while version 2.0 corresponds to Phase 1. The improvements in version 2.2 result from both a significant increase in the number of customer responses that have been collected and the identification of seven additional factors that help estimate customer interruption costs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Sensitivity of an integrated experiment to uncertainty in the high explosive equations of state

Traditionally, hydrodynamics simulations are performed with a single equation-of-state (EOS) to describe each material. These EOSs typically have a physics-informed functional form with adjustable parameters that are calibrated in order to replicate small-scale data. However, because the calibration data have uncertainty and there are typically inherent degeneracies in fitting the EOS, there are actually multiple EOSs that might be consistent with calibration data. In this work, we perform uncertainty quantification (UQ) for the reactant and product equations of state for the high explosive PBX 9501 to yield an ensemble of EOSs that match the uncertain small-scale calibration data. We then simulate an experiment of an explosively formed penetrator repeatedly with different EOSs to both validate the UQ analysis and determine the effects of EOS uncertainty on the prediction of quantities of interest in the experiment. In general, we find good agreement between the simulation predictions and the experimental measurements, and we identify an EOS variable that contributes most directly to the spread in the predictions as the EOSs are varied.

36 MATERIALS SCIENCE↗

Ultra-Long Distance BOTDA Sensor System Employing Hybrid Amplification and Advanced Noise Reduction Techniques

Brillouin Optical Time Domain Analysis (BOTDA) sensor system play a pivotal role in distributed sensing, which enables precise measurements of strain and temperature across extensive fiber lengths. Nonetheless, challenges emerge as distances grow due to signal attenuation and noise interference resulting in measurement errors. This research offers a comprehensive strategy to extend the sensing range of BOTDA systems beyond 10’s of kilometers while maintaining high spatial resolutions. Such enhanced sensing is realized through the integration of distributed Raman amplification, inline amplification using erbium-doped fiber amplifiers (EDFA), and advanced noise reduction techniques. Leveraging inherent redundancy in measured data as a function of frequency and fiber distance, the non-local means (NLM) filter removes noise while preserving essential physical information. This approach proves particularly advantageous in BOTDA systems, where accurate measurement of Brillouin scattering signals is paramount for long-range sensing, while concurrently safeguarding high spatial resolutions. In summary, this research has shown a holistic exploration of extending BOTDA's distance sensing capabilities up to 150 km with spatial resolutions of 8 meters.

Bhatta, Hari↗

NOvA joint $ν^{e}$ + $ν^{µ}$ oscillation results in neutrino and antineutrino modes

NOvA is an experiment devoted to studying neutrino oscillations in the NuMI neutrino beam from FNAL (USA). It is a long-baseline experiment consisting of two functionally identical, finely granulated detectors which are separated by 810 km of Earth crust and sited at 14 mrad off the beam axis. By measuring the transition probabilities P(νμ→νe \nu_\mu \rightarrow \nu_e ) and P(νμ→νμ \nu_\mu \rightarrow \nu_\mu ) NOvA is able to extract oscillation parameters: Δm232 \Delta m^2_{32} , mixing angle θ23 \theta_{23} , CP violating phase δCP \delta_{CP} and neutrino mass hierarchy. This analysis will be the first to include both neutrino (9⋅1020 9 \cdot 10^{20} POT) and antineutrino (7⋅1020 7 \cdot 10^{20} POT) data, which helps to resolve degeneracies in the oscillation probability. In this poster, the lastest NOvA oscillation results will be discussed, as well as important intermediate steps in νe \nu_e analysis like signal and background predictions, projected sensitivities in upcoming analyses will be presented.

Back, Ashley [Iowa State U.]↗

guppy i : a code for reducing the storage requirements of cosmological simulations

ABSTRACT As cosmological simulations have grown in size, the permanent storage requirements of their particle data have also grown. Even modest simulations present a major logistical challenge for the groups which run these boxes and researchers without access to high performance computing facilities often need to restrict their analysis to lower quality data. In this paper, we present guppy, a compression algorithm and code base tailored to reduce the sizes of dark matter-only cosmological simulations by approximately an order of magnitude. guppy is a ‘lossy’ algorithm, meaning that it injects a small amount of controlled and uncorrelated noise into particle properties. We perform extensive tests on the impact that this noise has on the internal structure of dark matter haloes, and identify conservative accuracy limits which ensure that compression has no practical impact on single-snapshot halo properties, profiles, and abundances. We also release functional prototype libraries in C, Python, and Go for reading and creating guppy data.

79 ASTRONOMY AND ASTROPHYSICS↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗