Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Elastic Changepoint Detection for Globally-indexed Functional Time Series Data with Climate Applications

Changepoint detection is a vital tool in the application of climate data analysis. Numerous types of climate observation data are most properly represented by functional time series, implying a need for accurate changepoint detection methods applicable to functional time series data. Such data taken at a global scale often contain both spatial heterogeneity and dependence as well as phase (time) misalignment. In this report, we present methods which can detect spatially-dependent changepoints while allowing different estimates of change time and change strength depending on location. Additionally, we provide extensions to this spatially-predicted model which controls for phase variability among observations. Our methods provide the ability to detect a single change, or control for epidemic changes (where a “return-to-normal” change is more likely to be detected than the initial change). We showcase results analyzing the June 1991 eruption of Mt. Pinatubo, where our methods demonstrate the ability to accurately detect both single and epidemic changepoints even in the presence of strong seasonal variability. We find that our spatially-predicted model improves the detection of relevant changepoints versus methods which do not take spatial information into account, and we find that controlling for phase variability helps to control the false discovery rate during the detection process.

54 ENVIRONMENTAL SCIENCES↗

Robustness of topological persistence in knowledge distillation for wearable sensor data

Topological data analysis (TDA) has shown great success in various applications involving wearable sensor data. However, there are difficulties in leveraging topological features in machine learning and wearable sensors because of the large time consumption and computational resources required to extract the features. To address this problem, knowledge distillation (KD) is utilized to generate a small model and accommodate topological features with persistence image (PI) representations from the raw time series data. Deploying topological knowledge in KD enables the student to achieve better performance compared to the one trained solely on raw time series data. However, it is not yet known if there are coherent characteristics for topological features in PI, which can aid in improving the performance during KD. In this paper, we investigate the suitability and challenges of utilizing topological features in KD for wearable sensor data, thereby contributing to the advancement of the field. Our study explores the impact of transferred topological features by comparing the Teacher-to-Student framework with Multiple Teachers-to-Student where teachers utilize both time series data and persistence images obtained by TDA as inputs. Additionally, we conduct a rigorous examination of topological knowledge effects by testing under various corruptions, knowledge types, and learning strategies in the context of human activity recognition tasks. Our analysis of topological features in KD presents the optimal strategy for incorporating these features. This study includes datasets of varying scales, window lengths, and activity classes, providing a comprehensive evaluation. Our results demonstrate that leveraging topological features in KD to enhance performance across databases.

97 MATHEMATICS AND COMPUTING↗

The composition of gases from a diffusion flame above longleaf pine needle fuel beds

The gas and tar composition of a diffusion flame from longleaf pine needles is currently poorly understood and more data are needed to fill in the gap between pyrolysis data and smoke plume data, thus improving physical and chemical modeling of wildland smoke formation. A pilot experiment to measure light gas and tar composition of such a flame is described for three flame regions: persistent flame (flame base), intermittent flame, and smoke plume. Flame gases from 24 experimental fires were collected in canisters and analyzed using EPA method TO-14A for CO 2 , CO, H 2 , CH 4 , and C 2 to C 7 hydrocarbon gases. Condensed gas (tar) samples were collected and analyzed using GC/MS. Other light gases were measured using FTIR spectroscopy. Results from compositional data analysis suggest significant differences in (relative) concentration of compounds detected in the three regions of the flame. Statistical tests for differences in flame zones were performed using the canister data: Concentration of hydrocarbons relative to CO and CO 2 decreased from the persistent flame zone above the pyrolyzing needles through the intermittent flame region into the flame-free plume. This was likely due to both chemical reactions (oxidation) occurring in the flame as well as the introduction of air into the flame/plume by entrainment.

Biomass↗

Protocols and methodologies for acquiring and analyzing critical-current versus longitudinal-strain data in Bi 2 Sr 2 CaCu 2 O 8+x wires

Abstract In the literature on Bi 2 Sr 2 CaCu 2 O 8+ x (Bi-2212) superconducting wires, it is evident that measurement protocols for transport critical-current I c versus longitudinal strain ϵ and definitions of the so-called ‘strain limit’ are generally dissimilar. Yet, values obtained for the ‘strain limit’ are frequently assimilated to being those of the irreversible strain limit ϵ irr , regardless of the I c degradation-criterion used to define it. In effect, ϵ irr should correspond specifically to the I c ( ϵ ) irreversibility onset , where crack formation in Bi-2212 filaments presumably starts. Because I c ( ϵ ) degradation remains progressive over a fairly wide strain range beyond ϵ irr , the different I c degradation-criteria in use do not yield to the same result and, thus, are not equivalent from metrology perspective. Indeed, in studying densified samples of a modern Bi-2212 round wire, we found ϵ irr ≈ 0.4% and ϵ 5% ≈ 0.6% ( ϵ 5% being the strain where I c degrades by 5%). In this paper, we outline and suggest I c ( ϵ )-measurement protocols and data-analysis methodologies in the hope to converge the various approaches taken for studying Bi-2212 strain properties and, thus, remove related result discrepancies. A unified approach would enable more objective data comparisons among laboratories and among different Bi-2212 conductors. It would pave the way for more rigorous studies of effects potentially associated with wire design, powder, heat treatments, and other such parameters on the conductor’s strain properties.

protocols↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Quantitative insights for diagnosing performance bottlenecks in lithium–sulfur batteries

Lithium–sulfur (Li–S) batteries hold significant promise for electric vehicles and aviation due to their high energy density and cost-effectiveness. However, understanding the root causes of performance degradation remains a formidable challenge, as the interplay of multiple factors obscures key failure mechanisms. A major limitation has been the inability to quantify soluble sulfur species within practical detection limits accurately and to correlate electrochemical processes with associated physical inventory changes. Here, we introduce the high-performance liquid chromatography-ultraviolet spectroscopy and gas chromatography sequential characterization (HUGS) toolkit, capable of precisely quantifying seven distinct sulfur and polysulfide species at concentrations as low as 40 ppb. HUGS has been successfully applied to practical coin and pouch cells without requiring cell modification. Furthermore, our self-developed software, Dr HUGS, enhanced the data analysis speed by over 30 times, enabling multi-source data integration and delivering comprehensive analysis results within minutes. Using HUGS, we identify significant capacity losses from inactive lithium and sulfur during initial cycles and sulfide-rich solid–electrolyte interphase (SEI) formation on the anode during later cycles. Notably, our findings reveal that soluble polysulfides have minimal contributions to capacity loss, challenging long-standing assumptions. Moreover, HUGS demonstrates that constant-pressure setups in Li–S pouch cells improve compositional uniformity compared to constant-gap configurations. For sulfurized polyacrylonitrile (SPAN) cathodes, unique issues such as non-sulfide SEI formation and lithium pulverization are observed, which can be mitigated through localized high-concentration electrolytes to enhance lithium inventory retention. By enabling precise quantification of critical inventory components, HUGS provides transformative insights into failure mechanisms across various electrolytes and cathode chemistries, guiding rational design strategies for next-generation energy storage systems.

25 ENERGY STORAGE↗

21 cm Power Spectrum Analysis of North Celestial Pole Observations with the Tianlai Dish Pathfinder Array

The Tianlai Dish Pathfinder Array (TDPA) is a radio interferometer designed to test techniques for 21 cm intensity mapping in the post-reionization Universe as a means of measuring large-scale cosmic structure. Using nine nights of observations targeting the North Celestial Pole field, totaling approximately 107 hr of integration time, we analyze data in the frequency range 700–800 MHz (corresponding to redshift z ∼ 0.9). We do the data format conversion, radio frequency interference flagging, calibration, imaging and point source subtraction, and foreground removal via Singular Value Decomposition. The spherically averaged power spectrum Δ 2 (k) is obtained. Furthermore, this work successfully establishes and validates a comprehensive data analysis framework for the TDPA. We identify key improvements including sky model refinement, increased integration time, and pipeline optimization that will enable future detection of the 21 cm signal through auto-correlation and cross-correlation with optical galaxy surveys.

cosmology: large-scale structure of universe↗

An Approach to Dynamic Human Reliability Analysis and Its Data Collection Framework

Human reliability analysis (HRA) is a method for evaluating human errors in a variety of complex systems such as nuclear power plants, military systems, aircraft, and chemical plants. Most HRA methods currently used by regulatory institutes or utilities are called static HRA and are carried out by simple worksheets or simple calculators. To date, there are many unsolved or intrinsic challenges in static HRA. For example, existing static HRA does not realistically model and evaluate human actions as they would be performed at actual systems. There is no method with HRA to objectively estimate the time required for human actions despite being essential to HRA processes. In addition, many HRA methods still rely on a dataset generated prior to the 1980s, from unrelated industry experience or simply from expert judgment. Accordingly, this study attempted to research how to overcome the challenges of existing HRA via dynamic risk assessment (a.k.a., simulation-based or computation-based risk assessment) techniques. First, this study developed a dynamic HRA method, named as PRocedure-based Investigation Method of EMRALD Risk Assessment – HRA (PRIMERA-HRA). The PRIMERA-HRA mainly concentrates on providing HRA analysts with specific guidelines on how to reasonably model human actions, assign human reliability data and evaluate output of simulation within a dynamic probabilistic risk assessment tool, called as Event Modeling Risk Assessment using Linked Diagrams (EMRALD). Second, this study also developed a module for performance shaping factors (i.e., the key concept in HRA quantification) applicable to dynamic HRA, then implemented it based on PRIMERA-HRA within the EMRALD tool. Third, this study developed an HRA data collection framework to support dynamic HRA, called as Simplified Human Error Experimental Program (SHEEP). Originally, the SHEEP study aimed to support static HRA and its data collection, but recently extended the scope to the new technologies such as dynamic HRA or HRA for advanced reactors. SHEEP focuses on the use of data collected from simplified simulators to complement—but not replace—data collection studies using full-scope simulators and actual operators. To date, many experiments were conducted under the SHEEP framework. Multiple analyses, such as human performance analysis, human error analysis, task complexity analysis, learning effect analysis and time distribution analysis, were also carried out using the collected data. Then, based on the major insights, an approach to inferring full-scope data based on simplified simulator data was proposed. The PRIMERA-HRA and SHEEP research are expected to evaluate human actions more realistically than existing static HRA, provide an opportunity to collect more HRA data with reasonable cost and labor, then contribute to enhance the quality of HRA.

99 - GENERAL AND MISCELLANEOUS↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

In this work, we examine the use of high-order tensor decompositions to analyze degradation pathways emerging from accelerated stress testing of silicon photovoltaic (PV) modules. Matrix-based decompositions are powerful tools for studying two-dimensional data arrays and form the foundation of a host of classical data analysis techniques. Tensors are high-order extrapolations of matrices that are able to account for more parameter dimensions, and a variety of tensor decomposition methods have been developed that similarly seek to extend insights from matrix decompositions to higher dimensions. Applying and interpreting tensor decomposition methods to sequences of PV module image data, we seek to uncover and isolate different degradation modes occurring from accelerated stress testing procedures. Further, we consider the contributions of different modes to PV module performance degradations.

data analysis↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Analytical methods for online data quality assessment

This chapter provides a comprehensive overview of the main steps for algorithmic sensor signal quality assessment, which can enhance the decision-making process for water resource recovery facility (WRRF) operation and optimization. It introduces the concept of redundancy as the basis for data quality assessment. It also explains the typical data processing pipeline, which consists of preliminary analysis, data pre-processing, and specific algorithmic approaches. Each of these processes is presented and discussed in three separate sections. Importantly, this chapter introduces the main approaches for data quality assessment, provides guidelines for selecting the most suitable one and the key performance indicators to evaluate them and explains how to collect metadata through such an algorithmic approach.

Aguado, Daniel↗

Preparation for DM searches with high Q SRF cavities

This work focuses on the preparatory for two dark matter searches based on high Q SRF cavities. In the context of the SERAPH experiment I participated to experimental work of characterizing the cavity at mK temperature and subsequently analyzed the data collected. For the SHADE experiment, I worked on the preparation of the data analysis starting from the collection of simulated data and built a flexible framework to analyze them. The goal of this report is to present my contribution to these two experiments that are being investigated at Fermilab by the Physics Sensing group at SQMS.

47 OTHER INSTRUMENTATION↗

Electrification Analysis: Manhattan Beer

This one-page highlight details the key takeaways from a project that utilized NREL's Fleet Research, Energy Data, and Insights (FleetREDI) data analysis pipeline, the Manhattan Beer Electrification Project. This project determined that Class-8 beverage distribution trucks operating in Manhattan show substantial electrification potential due to daily driving distances below 50 miles and low average speeds of 22mph or less. Their duty cycle needs can often be met by even modestly sized batteries and charging infrastructure. Vulnerable communities near their routes would benefit from fleet electrification.

ADVANCED PROPULSION SYSTEMS↗

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.

97 MATHEMATICS AND COMPUTING↗

EFIT‐AI: Machine Learning and Artificial Intelligence Assisted Equilibrium Reconstruction for Tokamak Experiments and Burning Plasmas (Final Report)

The EFIT-AI project is creating a modern advanced equilibrium reconstruction code suitable for tokamak experiments of burning plasmas. EFIT [1,2] was the first and is the most extensively used equilibrium reconstruction code in the world. This project builds on the production-level experience and adds key elements as follows. 1. A Model Order Reduction (MOR) version of the two-dimensional (2D) Grad-Shafranov equation solver (EFIT-MORNN) using physics-informed neural networks. 2. Improved optimization and data analysis capabilities using a Bayesian framework enhanced with machine learning. 3. A MOR version of the three-dimensional (3D) perturbed equilibrium reconstruction tool.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

A Comparative Analysis of Infrastructure-Based Perception Sensors for Intelligent Transportation Systems

The rise of privatized and public investment in smart city infrastructure and intelligent transportation systems has generated a heightened demand for perception sensors that effectively track and detect objects while being reliable in diverse weather and lighting conditions. This growing demand for perception sensors has accelerated their development and enhanced their capabilities. With these new capabilities, it is challenging to determine the most suitable sensing unit to use in each situation. Therefore, it is essential to have a comprehensive understanding of the benefits and limitations of each sensing unit to effectively leverage their capabilities. The purpose of this paper is to provide a detailed evaluation of various perception sensors. Additionally, this paper will demonstrate the benefits of combining multiple perception sensors, which complement each other by addressing data gaps inherent to single-sensor systems, to facilitate the creation of a digital twin that models the real world. The Infrastructure, Perception, and Control (IPC) team will conduct data analysis using data collected through field testing at traffic intersections in Colorado Springs, Colorado, to make comparisons between sensors. This research aims to provide clear and concise information about modern perception systems, which will support the development of intelligent transportation systems.

33 ADVANCED PROPULSION SYSTEMS↗

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering↗