Search NASA⌕ Search

SEARCH · Search NASA

Results for “data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

FREDA: A Web Application for the Processing, Analysis, and Visualization of Fourier‐Transform Mass Spectrometry Data

The high-resolution measurement capability of Fourier-transform mass spectrometry (FT-MS) has made it a necessity for exploring the molecular composition of complex organic mixtures, like soil, plant, aquatic, and petroleum samples. This demand has driven a need for informatics tools to explore and analyze FT-MS data in a robust and reproducible manner. FREDA is an interactive web application developed to enable spectrometrists to format, process, and explore their FT-MS data without the need for statistical programming expertise. FREDA was built to explore outputs from a molecular identification tool, like CoreMS, and provide a suite of methods to filter data, compute chemical properties of peaks, statistically compare samples and groups of samples, conduct exploratory data analysis, and download the results with a report detailing all steps conducted. To demonstrate the utility of FREDA, an example analysis was conducted using FT-MS data from a soil microbiology study of samples collected in two different soil depths at the Sphagnum bog forest north of Grand Rapids, Minnesota. Differences between the two depths are observed using Kendrick, Gibbs free energy, and van Krevelen plots. G-tests are used to quantify a significant difference between the groups. All analyses and plotting are conducted using only the FREDA application. FREDA is an open-source and readily available web application that allows users to explore and make statistically valid conclusions about their FT-MS data. The application is available online (https://map.emsl.pnnl.gov/app/freda) with a tutorial web series (https://youtu.be/k5HLE2kNSBY?si=yB6sGoyvzxrFf5MP) and freely accessible code on Github (https://github.com/EMSL-Computing/FREDA).

47 OTHER INSTRUMENTATION↗

Time-resolved electrical potential pump – X-ray photoelectron spectroscopy probe developments for investigating dynamic processes occurring at electrochemical interfaces

Electrode–electrolyte interfaces are of critical importance in several fields, including renewable energy, corrosion, and environmental chemistry. However, investigating these interfaces under operational conditions poses considerable challenges due to the limitations of the instrumentation employed. While recent advancements in in situ and operando techniques have enhanced our comprehension of the steady-state properties of solid-liquid interfaces, the dynamic behaviors of these systems remain inadequately explored. This study introduces a time-resolved X-ray photoelectron spectroscopy (XPS) technique designed to capture transient reaction intermediates and charging dynamics at electrified interfaces. The presented proof-of-principle study demonstrates that electrochemical processes, represented by an equivalent electrical circuit (EEC) model, can be probed and understood using square wave voltage pulses of a potentiostat synchronized to the modified data acquisition of an XPS setup. This method offers a valuable alternative to traditional pump–probe techniques, facilitating the investigation of a broader range of electrochemical systems. A dedicated software package for analyzing time- and energy-resolved XPS with a focus on extracting parameters of the EEC is geared towards benchmarking different EECs in future real-world electrochemical experiments.

Electrochemistry↗

Searches for New Long-Lived Particles and Upgrade to the ATLAS Inner Detector (Final Technical Report)

The search for new fundamental particles is one of the defining goals of the Large Hadron Collider (LHC). The discovery of the Higgs Boson by the ATLAS and CMS collaborations provided the capstone of the Standard Model of particle physics, but outstanding questions remain. Why does the Higgs boson have a mass of 125 GeV when its natural mass would be many orders of magnitude larger? Is there a universal symmetry which unites all three forces described by the Standard Model? Can that symmetry be extended to include gravity? Is dark matter, evidenced by astronomical observations, made of a particle that interacts via Standard Model forces with the rest of matter? Together, these motivations provide compelling arguments that new physical processes await discovery. This project addressed some outstanding questions about the fundamental particles and their interactions with the ATLAS experiment at the Large Hadron Collider. In particular, the project improved the discovery potential for new, long- lived particles produced via electroweak processes in proton-proton collisions and set world-leading limits on their existence for certain values of their potential mass and lifetime. To achieve this, the project developed new data analysis methods, developed new triggers to select events with new long-lived particles during data-taking of the ATLAS experiment, and analyzed the largest proton–proton collision dataset ever produced. The project also supported significant development of the data acquisition software for the upgrade to the ATLAS inner detector, the Inner TracKer (ITk). The upgrade of the ATLAS inner detector is essential to the success of the entire Phase II physics program on ATLAS. Personnel supported by the project provided support for integration, assembly, and testing of the inner two layers of the ITk pixel system during its prototype and pre-production phase. Four PhD students and two post-doctoral scholars were supported by the grant and received invaluable scientific training as part of the research endeavor. The students and postdocs gained essential professional skills in the areas of advanced data analysis techniques, statistical analysis of data and simulation, programming in C++ and Python, hardware and instrumentation development, and presentation and collaboration skills. Additionally, approximately ten undergraduate students supported through other funding sources participated in research activities synergistic with the goals of this project, receiving essential mentorship from the personnel supported by this project.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Extended dark energy analysis using DESI DR2 BAO measurements

We conduct an extended analysis of dark energy constraints, in support of the findings of the Dark Energy Spectroscopic Instrument (DESI) second data release cosmology key paper, including DESI data, Planck cosmic microwave background observations, and three different supernova compilations. Using a broad range of parametric and nonparametric methods, we explore the dark energy phenomenology and find consistent trends across all approaches, in good agreement with the 𝑤 0⁢ 𝑤 𝑎⁢ CDM (cold dark matter) key paper results. Even with the additional flexibility introduced by nonparametric approaches, such as binning and Gaussian processes, we find that extending Λ⁢ CDM to include a two-parameter 𝑤⁡(𝑧) is sufficient to capture the trends present in the data. Finally, we examine three dark energy classes with distinct dynamics, including quintessence scenarios satisfying 𝑤 ≥ −1, to explore what underlying physics can explain such deviations. The current data indicate a clear preference for models that feature a phantom crossing; although alternatives lacking this feature are disfavored, they cannot yet be ruled out. Our analysis confirms that the evidence for dynamical dark energy, particularly at low redshift (𝑧 ≲ 0.3), is robust and stable under different modeling choices.

79 ASTRONOMY AND ASTROPHYSICS↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Tree-level carbon stock estimations across diverse species using multi-source remote sensing integration

Forests are critical carbon sinks, and remote sensing has been increasingly widely used for forest monitoring and biomass estimations. However, species-specific tree-level studies remain limited. In this study, we demonstrated the feasibility of integrating UAV-based LiDAR with high-resolution optical satellite imagery (0.5 m) to estimate biomass for individual trees across different species. The proposed method accurately estimated biomass for 53 trees (R² = 0.82, rRMSE = 0.44), with species-specific datasets, showing an average 25.2% increase in R² and a 14.8% reduction in rRMSE. A novel vegetation index combining forest structure parameters with vegetation indices (VIs) was developed using high-resolution multispectral satellite data (3 m) to explore its relationship with individual tree biomass. Combining forest structural parameters with VIs further improved estimation accuracy, achieving an R²of 0.89 and an rRMSE of 0.34. Species-specific datasets show an 11.6% increase in R²compared to methods without VIs, and a 22.2% improvement over methods using only VIs. SHapley Additive exPlanations (SHAP) analysis shows that the volume feature played a key role in model performance and remained stable throughout the training process. Altogether, the proposed approach enhances individual tree biomass and carbon sink estimations, showing great potential for large-scale precise forest carbon monitoring using multi-source remote sensing data.

59 BASIC BIOLOGICAL SCIENCES↗

Sulfur isotopes from the Paleoproterozoic Francevillian Basin record multigenerational pyrite formation, not depositional conditions

Bulk-rock sulfur isotope data from pyrite in the ~2.1 billion-year sedimentary rocks of the Francevillian Basin, Gabon, have underpinned ideas about initial oxygenation of Earth’s surface environments and eukaryote evolution. In this paper, we show, using micro-scale analytical methods, that the bulk sulfur isotope record represents progressive diagenetic modification. Our findings indicate no significant change in microbial sulfur cycling processes and seawater sulfate composition throughout that initial phase of atmosphere-ocean oxygenation of Paleoproterozoic time. This offers an alternative view of Earth system evolution during the transition from an anoxic to an oxic state and highlights the need for a judicious reappraisal of conceptual models using sulfur isotope data as primary depositional signals linked to global-scale biogeochemical processes.

58 GEOSCIENCES↗

A Comparison of Machine Learning Methods of Association Tested on Dense Nodal Arrays

The association of phase picks to form events is one of the fundamental components of seismology. Large and dense sensor networks, such as >1000 geophone arrays (and distributed acoustic sensing), offer unique challenges in association due to the vast numbers of observations and high likelihood of errant picks. In addition, the large number of stations can greatly increase the time it takes to perform the association. For this reason, machine learning (ML) methods might provide a more optimal method of association for such networks. In this work, we examine how well ML methods (e.g., Gaussian mixture model association, PhaseLink, and Graph Earthquake Neural Interpretation Engine) can incorporate dense seismic arrays into regional networks and how well they handle the increasing numbers of stations. Here, we test their capabilities on two dense seismic deployments, one within Rock Valley Nevada (52 nodes and a 9-station sparse local network), and the LArge-n Seismic Survey in Oklahoma dense nodal array (>1800 vertical-component geophones). Processing data from these two different styles of dense seismic deployments allows testing of how the ML algorithms can merge array data with a broader regional network, how they deal with poorly picked phases, and how they handle anthropogenic noise. We compare the ML-associated bulletins to those obtained using the Rapid Earthquake Association and Location algorithm, a more traditional method of association. We find that there are very small differences in results between the methods for small networks (<100 stations) with low pick rates. For large networks (>1000), there are enough errant picks that some of the ML methods start to create false events out of noise. We also find that the ML methods vary in computation time significantly but are all faster than the traditional method tested here.

58 GEOSCIENCES↗

Evaluating downscaled products with expected hydroclimatic co-variances

Abstract. There has been widespread adoption of downscaled products amongst practitioners and stakeholders to ascertain risk from climate hazards at the local scale (e.g., ∼ 5 km resolution). Such products must nevertheless be consistent with physical laws to be credible and of value to users. Here we evaluate statistically and dynamically downscaled products by examining local co-evolution of downscaled temperature and precipitation during convective and frontal precipitation events (two mechanisms testable with just temperature and precipitation). We find that two widely used statistical downscaling techniques (Localized Constructed Analogs version 2, LOCA2, and Seasonal Trends and Analysis of Residuals Empirical Statistical Downscaling Model, STAR-ESDM) generally preserve expected co-variances during convective precipitation events over the historical and future projected intervals as compared to European Centre for Medium-Range Weather Forecasts Reanalysis v5 (ERA5) and two observation-based data products (Livneh and nClimGrid-Daily). However, both techniques dampen future intensification of frontal precipitation that is otherwise robustly captured in global climate models (i.e., prior to downscaling) and with process-based dynamical downscaling across five different regional climate models. In the case of LOCA2, this leads to appreciable underestimation of future frontal precipitation event intensity. This study is one of the first to quantify a likely ramification of the stationarity assumption underlying statistical downscaling methods and identify a phenomenon where projections of future change diverge depending on data production method employed. Finally, our work proposes expected co-variances during convective and frontal precipitation as useful evaluation diagnostics that can be universally applied to a wide range of statistically downscaled products.

54 ENVIRONMENTAL SCIENCES↗

Open‐Source Anaerobic Digestion Modeling Platform, Anaerobic Digestion Model No. 1 Fast (ADM1F)

An open‐source modeling platform, called Anaerobic Digestion Model No. 1 Fast (ADM1F), is introduced to achieve fast and numerically stable simulations of anaerobic digestion processes. ADM1F is compatible with an iPython interface to facilitate model configuration, simulation, data analysis, and visualization. Faster simulations and more stable results are accomplished by implementing an advanced open‐source library of numerical methods called Portable Extensive Toolkit for Scientific Computation (PETSc) to solve the ADM1 system of equations. Leveraging PETSc, ADM1F can consistently complete a steady‐state simulation under 0.2 s, over 99% faster than a benchmark ADM1 model implemented with MATLAB while achieving agreement of model outputs within 1% of those obtained with the benchmark model. For dynamic simulations, however, ADM1F has a computational speed advantage only when the influent characteristics update more frequently than every 4 h. The ability of ADM1F to be useful as a tool to study anaerobic digestion systems is demonstrated through two example implementations of ADM1F: (1) a two‐phase co‐digestion scenario evaluating the impact of the organic loading rate and the substrate composition on reactor performance and stability, and (2) a conventional digester scenario assessing the effectiveness of recovery strategies after disruptions that led to instability. These examples demonstrate how the high simulation speed and the convenience of the iPython interface allow ADM1F to complete complex analyses within minutes, much faster than computational strategies currently reported in the literature.

anaerobic co-digestion↗

Desmearing Bonse–Hart USANS data using Bayesian Gaussian process regression

Ultra-small-angle neutron scattering (USANS) enables access to micrometer-scale structures but is intrinsically affected by strong, anisotropic resolution smearing arising from slit-geometry optics. As a result, recovery of the intrinsic scattering intensity constitutes an ill-posed inverse problem, and commonly used iterative desmearing methods lack rigorous uncertainty quantification. We present a Bayesian desmearing framework for slit-geometry USANS based on Gaussian process regression. In this approach, the scattering intensity is modeled as a smooth random function, and the instrumental point spread function is incorporated explicitly as a forward operator. The resulting formulation yields a closed-form maximum a posteriori solution with well-defined credibility intervals. Computational benchmarks and experimental validation using combined USANS and small-angle neutron scattering (SANS) measurements demonstrate that the framework enables stable desmearing, suppresses experimental noise, and preserves physically meaningful structural features under realistic conditions.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION↗

An AI-driven framework for evaluating local and state authorities’ permitting processes

The demand for new energy infrastructure is increasing across the United States, but heterogenous permitting processes and embedded requirements across different local jurisdictions can cause project delays, increase “soft costs,” and hinder developer expansion. This study analyzes the variability in local permitting requirements across the U.S. and develops a quantitative approach to describe their clarity and effectiveness in enabling infrastructure project development. By using an Energy Language Model (ELM), a large language model (LLM) for energy technologies, we systematically gathered permitting information from nearly 300 state-, county-, and city-level documents, creating a structured dataset of requirements and procedures on an unprecedented scale and speed. Our analysis revealed that local (city and county) permitting requirement documents are underrepresented compared to state-level guidance documents, which can impede timely and cost-effective installation of new electric infrastructure. Our validation process showed that the final database has an accuracy of approximately 95%. We, further, created a new quantitative method to score permitting requirements for clarity and efficiency, with electric vehicle supply equipment as an initial use case. The average local permitting document scored a 1.8 out of 5, which we interpret as meaning that half of the requirements developers face when installing electric infrastructure are ambiguous, increasing both cost and time. We also created a “Generalized Permit Process”, highlighting common procedural steps and identifying specific opportunities for municipalities to improve their documentation. This research establishes a systematic and scalable framework for evaluating the complexities of local infrastructure permitting processes by combining LLM-powered data collection and quantitative scoring. The framework enables policymakers and developers to identify and mitigate procedural bottlenecks, with the expectation that these improvements can accelerate application review and approval, reduce project costs, and expedite connection to utility distribution grids. As a foundational approach for streamlining local project development processes, this study’s methods are intended to be extended to a wide range of energy applications.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Low latency optical-based mode tracking with machine learning deployed on FPGAs on a tokamak

Active feedback control in magnetic confinement fusion devices is desirable to mitigate plasma instabilities and enable robust operation. Optical high-speed cameras provide a powerful, non-invasive diagnostic and can be suitable for these applications. Here, in this study, we process high-speed camera data, at rates exceeding 100 kfps, on in situ field-programmable gate array (FPGA) hardware to track magnetohydrodynamic (MHD) mode evolution and generate control signals in real time. Our system utilizes a convolutional neural network (CNN) model, which predicts the n = 1 MHD mode amplitude and phase using camera images with better accuracy than other tested non-deep-learning-based methods. By implementing this model directly within the standard FPGA readout hardware of the high-speed camera diagnostic, our mode tracking system achieves a total trigger-to-output latency of 17.6 μs and a throughput of up to 120 kfps. This study at the High Beta Tokamak-Extended Pulse (HBT-EP) experiment demonstrates an FPGA-based high-speed camera data acquisition and processing system, enabling application in real-time machine-learning-based tokamak diagnostic and control as well as potential applications in other scientific domains.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗

CNN-Based Phase Fault Classification in Real and Simulated Power Systems Data

This study proposes a convolutional neural network (CNN)–based two-step phase fault detection and identification method to classify anomalies in the power grid signal. Specifically, the first step checks the fault’s existence and determines the need for the second step. Subsequently, in the case of anomalies in the power grid signal, the second step identifies the type of fault, including line-to-line, single-line-to-ground, double-line-to-ground, and triple-line. Accordingly, the CNN architecture is both designed for the classification layers and trained with simulated data. To provide maximum prediction accuracy with minimum processing time, this study investigates the combinations of various feature extraction (FE) techniques, such as fast Fourier transform (FFT), amplitude and phase (AP), auto-correlation function, power spectral density, and wavelet transform (WT). Consequently, simulated and real-world results demonstrate that the proposed two-step method outperforms conventional one-step techniques, with the best performance obtained by using the combination of AP-AP, AP-WT, FFT-AP, and FFT-WT–based FE methods.

Alaca, Ozgur↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗