Search NASA⌕ Search

SEARCH · Search NASA

Results for “data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

COMPASS-FME Synoptic Site Characterization

This dataset contains soil biogeochemical and physicochemical characterization data for the COMPASS-FME synoptic sites.This dataset also contains data for the paper Patel et al. 2025 "Transition zones at the changing coastal terrestrial-aquatic interface", https://doi.org/10.1029/2025JG008978.Coastal soils are a significant but highly uncertain component of global biogeochemical cycles. These systems experience unique spatial and temporal variability in biogeochemical processes, driven by wetland-to-upland gradients and hydrological fluctuations. We studied drivers of coastal soil variability (a) at regional scales and (b) across transects from upland forest to wetland, in two contrasting regions — Lake Erie, a freshwater lacustrine system, and Chesapeake Bay, a saltwater estuarine system. Salinity-related analytes were a key driver of soil variability, not just in the saltwater system, but surprisingly, also in the freshwater system. We had hypothesized linear trends in biogeochemical parameters along the TAI – however, contrary to expectations, transition soils were not consistently intermediate between upland and wetland endmembers; the non-monotonic trends of carbon, phosphorus, iron along our transects suggest that these are key analytes to study in our regions. Rapidly changing soil factors across coastal gradients provide insights into which soil processes may act as precursors to ecosystem shifts. Our comprehensive soil characterization across the coastal transects provides essential data for mechanistic modeling of ecosystem dynamics.The data are provided as processed, csv files. Raw data and processing scripts can be accessed on GitHub (https://github.com/COMPASS-DOE/cmps-soil_characterization).A note on the nomenclature: the experimental design represents three points along the coastal gradient -- upland, transition, and wetland. "wetland" is referred to as "marsh" in the corresponding paper. The two terms can be used interchangeably for the sites in this study.

54 ENVIRONMENTAL SCIENCES↗

Chemistry and Metallurgy Research Building Historical Perchlorate Data and Discussion

The Chemistry and Metallurgy Research Building (CMR), located at TA-03-0029, was built in 1952 and was occupied by Los Alamos National Laboratory (LANL) employees in 1953. The facility was built to perform actinide analytical chemistry and material characterization in support of LANL and Department of Energy (DOE) mission of stockpile stewardship and research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Effect of alloying on intrinsic ductility in WTaCrV high entropy alloys

Tungsten (W) exhibits desirable properties for extreme applications, such as the divertor in magnetic fusion reactors, but its practicality remains limited due to poor formability and insufficient irradiation resistance. In this work, we study the intrinsic ductility of body-centered cubic WTaCrV based high entropy alloys (HEAs), which are known to exhibit excellent irradiation resistance. The ductility evaluations are carried out using a criterion based on the competition between the critical stress intensity factors for emission (K Ie ) and cleavage (K Ic ) in the {110} slip planes and {110} crack planes, which are evaluated within the linear elastic fracture mechanics framework and computed using density functional theory. The results suggest that increasing the alloying concentrations of V and reducing the concentrations of W can significantly improve the ductility in these HEAs. The elastic anisotropy for these HEAs is analyzed using the Zener anisotropy ratio and its correlation with the concentration of W in the alloys is studied. Results indicate that these alloys tend to be fairly isotropic independently from the concentration of W in them. The computed data for the elastic constants of these HEAs is also compared against available experimental data. The results are in good agreement, validating the robustness and accuracy of the computational methods. Multiple phenomenological ductility metrics were also computed and analyzed against the analytical model. Some metrics, mainly the surrogate D parameter, show a good correlation with the Rice model. The potential of these empirical metrics to serve as surrogate screening models for optimizing the compositional space is also discussed.

Anisotropy↗

Ecosystem Water‐Saving Timescale Varies Spatially With Typical Drydown Length

Abstract Stomatal optimization theory is a commonly used framework for modeling how plants regulate transpiration in response to the environment. Most stomatal optimization models assume that plants instantaneously optimize a reward function such as carbon gain. However, plants are expected to optimize over longer timescales given the rapid environmental variability they encounter. There are currently no observational constraints on these timescales. Here, a new stomatal model is developed and is used to analyze the timescales over which stomatal closure is optimized. The proposed model assumes plants maximize carbon gain subject to the constraint that they cannot draw down soil moisture below a critical value. The reward is integrated over time, after being weighted by a discount factor that represents the timescale ( τ ) that a plant considers when optimizing stomatal conductance to save water. The model is simple enough to be analytically solvable, which allows the value of τ to be inferred from observations of stomatal behavior under known environmental conditions. The model is fitted to eddy covariance data in a range of ecosystems, finding the value of τ that best predicts the dynamics of evapotranspiration at each site. Across 82 sites, the climate metrics with the strongest correlation to τ are measures of the average number of dry days between rainfall events. Values of τ are similar in magnitude to the longest such dry period encountered in an average year. The results here shed light on which climate characteristics shape spatial variations in ecosystem‐level water use strategy.

54 ENVIRONMENTAL SCIENCES↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Progress of Gas Injection EOR Surveillance in the Bakken Unconventional Play—Technical Review and Machine Learning Study

Although considerable laboratory and modeling activities were performed to investigate the enhanced oil recovery (EOR) mechanisms and potential in unconventional reservoirs, only limited research has been reported to investigate actual EOR implementations and their surveillance in fields. Eleven EOR pilot tests that used CO2, rich gas, surfactant, water, etc., have been conducted in the Bakken unconventional play since 2008. Gas injection was involved in eight of these pilots with huff ‘n’ puff, flooding, and injectivity operations. Surveillance data, including daily production/injection rates, bottomhole injection pressure, gas composition, well logs, and tracer testing, were collected from these tests to generate time-series plots or analytics that can inform operators of downhole conditions. A technical review showed that pressure buildup, conformance issues, and timely gas breakthrough detection were some of the main challenges because of the interconnected fractures between injection and offset wells. The latest operation of co-injecting gas, water, and surfactant through the same injection well showed that these challenges could be mitigated by careful EOR design and continuous reservoir monitoring. Reservoir simulation and machine learning were then conducted for operators to rapidly predict EOR performance and take control actions to improve EOR outcomes in unconventional reservoirs.

Energy & Fuels↗

Precise measurements of W - and Z -boson transverse momentum spectra with the ATLAS detector using pp collisions at $\sqrt{s} = 5.02$ TeV and 13 TeV

This paper describes measurements of the transverse momentum spectra of W and Z bosons produced in proton–proton collisions at centre-of-mass energies of $\sqrt{s}$ = 5.02 TeV and $\sqrt{s}$ = 13 TeV with the ATLAS experiment at the Large Hadron Collider. Measurements are performed in the electron and muon channels, W → $\ell$$v$ and Z → $\ell$$\ell$ ($\ell$ = e or μ), and for W events further separated by charge. The data were collected in 2017 and 2018, in dedicated runs with reduced instantaneous luminosity, and correspond to 255 and 338 pb -1 at $\sqrt{s}$ = 5.02 TeV and 13 TeV, respectively. These conditions optimise the reconstruction of the W-boson transverse momentum. The distributions observed in the electron and muon channels are unfolded, combined, and compared to QCD calculations based on parton shower Monte Carlo event generators and analytical resummation. The description of the transverse momentum distributions by Monte Carlo event generators is imperfect and shows significant differences largely common to W - , W + and Z production. The agreement is better at $\sqrt{s}$ = 5.02 TeV, especially for predictions that were tuned to Z production data at $\sqrt{s}$ = 7 TeV. Higher-order, resummed predictions based on DYTurbo generally match the data best across the spectra. Distribution ratios are also presented and test the understanding of differences between the production processes.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics↗

Stress intensity factor models using mechanics-guided decomposition and symbolic regression

The finite element method can be used to compute accurate stress intensity factors (SIFs) for cracks with complex geometries and boundary conditions. In contrast, handbook solutions act as surrogate SIF models that provide significantly faster evaluation times. However, the development of conventional surrogate SIF models relies on manual development based on low-order parameterizations. This limits surrogate model accuracy and generalizability. Here, in this paper, we develop a framework for the automated development of mechanics-guided handbook SIF solutions by using interpretable machine learning via genetic programming for symbolic regression (GPSR). Formalizing the mechanics-based approach of Raju and Newman, SIF training data is decomposed into multiple subsets. This decomposition enables parallel GPSR model development of subfunctions, each of which accounts for specific geometrical corrections with respect to a known analytical model. Using this mechanics-based approach with GPSR allows for equations to be learned with improved accuracy and reduced complexity relative to the Raju Newman equations while maintaining the inherent interpretability of mathematical expressions. In this paper, we present equations that match the complexity of the Raju Newman equations while having reduced error, as well as equations with similar errors and reduced complexity.

42 ENGINEERING↗

STREAMS guidelines: standards for technical reporting in environmental and host-associated microbiome studies

The interdisciplinary nature of microbiome research, coupled with the generation of complex multi-omics data, makes knowledge sharing challenging. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines provide a checklist for the reporting of study information, experimental design and analytical methods within a scientific manuscript on human microbiome research. Here, in this Consensus Statement, we present the standards for technical reporting in environmental and host-associated microbiome studies (STREAMS) guidelines. The guidelines expand on STORMS and include 67 items to support the reporting and review of environmental (for example, terrestrial, aquatic, atmospheric and engineered), synthetic and non-human host-associated microbiome studies in a standardized and machine-actionable manner. Based on input from 248 researchers spanning 28 countries, we provide detailed guidance, including comparisons with STORMS, and case studies that demonstrate the usage of the STREAMS guidelines. In conclusion, STREAMS, like STORMS, will be a living community resource updated by the Consortium with consensus-building input of the broader community.

59 BASIC BIOLOGICAL SCIENCES↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry↗

Conservative projection-based data-driven model order reduction of a fluid-kinetic spectral solver

Kinetic simulations are computationally intensive due to six-dimensional phase space discretization. Many kinetic spectral solvers use the asymmetrically weighted Hermite expansion due to its conservation and fluid-kinetic coupling properties, i.e., the lower-order Hermite moments capture and describe the macroscopic fluid dynamics, and higher-order Hermite moments describe the microscopic kinetic dynamics. We leverage this structure by developing a parametric data-driven reduced-order model based on the proper orthogonal decomposition, which projects the higher-order kinetic moments while retaining the fluid moments intact. We demonstrate analytically and numerically that the method ensures local and global mass, momentum, and energy conservation. The numerical results show that the proposed method effectively replicates the high-dimensional spectral simulations at a fraction of the computational cost and memory, as validated on the weak Landau damping and two-stream instability benchmark problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Visual Analytics of Performance of Quantum Computing Systems and Circuit Optimization

Driven by potential exponential speedups in business, security, and scientific scenarios, interest in quantum computing is surging. This interest feeds the development of quantum computing hardware, but several challenges arise in optimizing application performance for hardware metrics (e.g., qubit coherence and gate fidelity). In this work, we describe a visual analytics approach for analyzing the performance properties of quantum devices and quantum circuit optimization. Our approach allows users to explore spatial and temporal patterns in quantum device performance data and it computes similarities and variances in key performance metrics. Detailed analysis of the error properties characterizing individual qubits is also supported. We also describe a method for visualizing the optimization of quantum circuits. The resulting visualization tool allows researchers to design more efficient quantum algorithms and applications by increasing the interpretability of quantum computations.

Chae, Junghoon↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗

SigTime: Learning and Visually Explaining Time Series Signatures

Understanding and distinguishing temporal patterns in time series data is essential for scientific discovery and decision-making. For example, in biomedical research, uncovering meaningful patterns in physiological signals can improve diagnosis, risk assessment, and patient outcomes. However, existing methods for time series pattern discovery face major challenges, including high computational complexity, limited interpretability, and difficulty in capturing meaningful temporal structures. Here, to address these gaps, we introduce a novel learning framework that jointly trains two Transformer models using complementary time series representations: shapelet-based representations to capture localized temporal structures and traditional feature engineering to encode statistical properties. The learned shapelets serve as interpretable signatures that differentiate time series across classification labels. Additionally, we develop a visual analytics system—SigTime—with coordinated views to facilitate exploration of time series signatures from multiple perspectives, aiding in useful insights generation. We quantitatively evaluate our learning framework on eight publicly available datasets and one proprietary clinical dataset. Additionally, we demonstrate the effectiveness of our system through two usage scenarios along with the domain experts: one involving public ECG data and the other focused on preterm labor analysis.

97 MATHEMATICS AND COMPUTING↗

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗