Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Frontal Slice Approaches for Tensor Linear Systems

Inspired by the row and column action methods for solving large-scale linear systems, in this work, we explore the use of frontal slices for solving tensor linear systems. In particular, this paper presents a novel approach for using frontal slices of a tensor $\mathcal{A}$ to solve tensor linear systems $\mathcal{A} ∗\mathcal{X} = \mathcal{B}$ where ∗ denotes the $t$-product. In addition, we consider variations of this method, including cyclic, block, and randomized approaches, each designed to optimize performance in different operational contexts. Our primary contribution lies in the development and convergence analysis of these methods. Experimental results on synthetically generated and real-world data, including applications such as image and video deblurring, demonstrate the efficacy of our proposed approaches and validate our theoretical findings.

Luo, Hengrui↗

SpectraCodec: A Hilbert curve-based method for encoding metadata in mass spectra for machine learning applications (SpectraCodec) v1

Machine learning approaches to mass spectrometry (MS) data analysis require structured metadata for optimal performance. However, current MS file formats necessitate external metadata sources, creating integration challenges that impede analytical workflows. Here, we present a novel approach for encoding metadata directly within mzML files using one-hot encoding of ASCII characters mapped via Hilbert space-filling curves. This strategy embeds metadata in the first spectrum's m/z-intensity space, ensuring persistence with the primary data, eliminating the need for external metadata files, and maintaining compatibility with existing MS software. We demonstrate that the Hilbert curve mapping efficiently utilizes the two-dimensional spectral space while maintaining robust data recovery. This method offers a practical solution for machine learning applications in mass spectrometry by ensuring metadata and spectral data remain unified through all stages of analysis.

Bowen, Benjamin [Lawrence Berkeley National Labora↗

The relationship between stress, anxiety and eating behavior among Chinese students: a cross-sectional study

Background The expansion of higher education and the growing number of college students have led to increased awareness of mental health issues such as stress, anxiety, and eating disorders. In China, the educational system and cultural expectations contribute to the stress experienced by college students. This study aims to clarify the role of anxiety as a mediator in the relationship between stress and eating behaviors among Chinese college students. Methods This study utilized data from the 2021 Psychology and Behavior Investigation of Chinese Residents, which included 1,672 college students under the age of 25. The analysis methods comprised descriptive statistics, t -tests, Pearson correlation analyses, and mediation effect analysis. Results The findings indicate that Chinese college students experience high levels of stress, with long-term stress slightly exceeding short-term stress. Both types of stress were positively correlated with increased anxiety and the adoption of unhealthy eating behaviors. Anxiety was identified as a significant mediator, accounting for 28.3% of the relationship between long-term stress and eating behavior (95% CI = 0.058–0.183). The mediation effect of short-term stress on eating behavior through anxiety was also significant, explaining 61.4% of the total effect (95% CI = 0.185–0.327). Conclusion The study underscores the importance of stress management and mental health services for college students. It recommends a comprehensive approach to reducing external pressures, managing anxiety, and promoting healthy eating behaviors among college students. Suggestions include expanding employment opportunities, providing career guidance, enhancing campus and societal support for holistic development, strengthening mental health services, leveraging artificial intelligence technologies, educating on healthy lifestyles, and implementing targeted health promotion programs.

Chai, Yulin↗

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dataset of mechanically induced thermal runaway measurement and severity level on Li-ion batteries

The deployment of Li-ion batteries covers a wide range of energy storage applications, from mobile phones, e-bikes, electric vehicles (EV) and stationary energy storage systems. However, safety issue such as thermal runaway is always one of the most important concerns to prevent Li-ion batteries from further market penetration. A standardized single-side indentation test protocol was developed to mechanically induce an internal short-circuit. The cell voltage, compressive load, indenter stroke, and temperature at the indentation point are measured in time series. The test data of each cell, along with cell parameters such as dimensions, mass, chemistry, state of charge (SOC), capacity, are integrated together to calculate a thermal runaway severity score from 0 to100. Complete data collection process including the original measured record, test method, severity score calculation scheme is presented in this article. The thermal runaway severity analysis and the more than 100 tested Li-ion battery records provide a good data source for further comparison and ranking of thermal runaway risks.

25 ENERGY STORAGE↗

Radioisotope Identification with List-Mode Gamma Ray Data: A rigorous assessment on the value of temporal information applied to radioisotope identification.

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with “confuser” sources, or spectra with similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research rigorously examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and information theory. We further propose a basic classification model that can utilize spectral or temporal data (or both) to determine if the incorporation of temporal information can improve radioisotope identification. The findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Enhancing weak lensing redshift distribution characterization by optimizing the Dark Energy Survey Self-Organizing Map Photo-z method

Characterization of the redshift distribution of ensembles of galaxies is pivotal for large scale structure cosmological studies. In this work, we focus on improving the Self-Organizing Map (SOM) methodology for photometric redshift estimation (SOMPZ), specifically in anticipation of the Dark Energy Survey Year 6 (DES Y6) data. This data set, featuring deeper and fainter galaxies than DES Year 3 (DES Y3), demands adapted techniques to ensure accurate recovery of the underlying redshift distribution. We investigate three strategies for enhancing the existing SOM-based approach used in DES Y3: 1) Replacing the Y3 SOM algorithm with one tailored for redshift estimation challenges; 2) Incorporating $\textit{g}$-band flux information to refine redshift estimates (i.e. using $\textit{griz}$ fluxes as opposed to only $\textit{riz}$); 3) Augmenting redshift data for galaxies where available. These methods are applied to DES Y3 data, and results are compared to the Y3 fiducial ones. Our analysis indicates significant improvements with the first two strategies, notably reducing the overlap between redshift bins. By combining strategies 1 and 2, we have successfully managed to reduce redshift bin overlap in DES Y3 by up to 66$\%$. Conversely, the third strategy, involving the addition of redshift data for selected galaxies as an additional feature in the method, yields inferior results and is abandoned. Our findings contribute to the advancement of weak lensing redshift characterization and lay the groundwork for better redshift characterization in DES Year 6 and future stage IV surveys, like the Rubin Observatory.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Classification of Vehicle Movements at Signalized Intersections Using Vehicle Trajectories

Accurate vehicle movement classification through signalized intersections is of paramount importance to the analysis of intersection performance and the optimization of traffic control strategies. Conventional techniques for tracking vehicle turning movements depend on infrastructure-based strategies like human counts, loop detectors, and video analytics, all of which are costly, prone to errors, and spatially constrained. High-frequency trajectory data can be utilized to determine vehicle movement patterns in a scalable and infrastructure-independent method due to the adoption of connected vehicles (CVs). In recent years, several studies have utilized connected vehicle data to generate performance measures. Most of the trajectory-based performance measures approaches, however, require map matching-i.e., extracting geospatial references from maps to identify the movements that individual vehicles make at a signalized intersection. These approaches are often time-consuming and hinder scalability since geographic features need to be provided for an analysis to be conducted. Map matching methods are prone to errors as different map versions change these geographic features. This research presents a novel automatic classification pipeline that uses CV trajectory data to classify vehicle movements at signalized crossings, specifically pass-through left-turn and right-turn maneuvers. The process starts by filtering trips that cross a spatial bounding box that has been defined at the target intersection. Approach and departure headings for each trajectory crossing the boundary are computed and are clustered together to identify dominant movements. The proposed algorithm is used to classify the movement of vehicles at 10 intersections in the state of California, and the results indicate that the algorithm can classify movements at these intersections with varying traffic volumes and road network configurations, all in a map-less framework with no need for conflation of vehicle trajectories to a digital base map.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multi-facility analysis using metered power data to quantify MRI energy use and utility bill costs across scanner operating modes

This study quantifies the energy consumption of magnetic resonance imaging (MRI) scanners across discrete operating modes during routine clinical workflows, based solely on electrical power measurements. Although previous studies have investigated MRI energy consumption within single hospitals or specific clinical settings, this research provides a broader and more systematic analysis. Researchers analyzed electrical power data and applied a previously developed semi-automatic method for identifying MRI operating modes using load duration curves for 20 MRI scanners across four different U.S. healthcare facilities, encompassing outpatient, inpatient, and mixed-use clinical settings. A key innovation is the inclusion of localized hourly utility rates to estimate costs, a parameter absent in prior literature. Key findings indicate significant variability in energy and cost profiles between weekdays and weekends. Scanner characteristics, including magnet strength, manufacturer, vintage, location, and clinical setting, influenced average daily energy consumption and power thresholds for operating modes. Notably, the clinical setting of a scanner predominantly determines its energy use. For example, the scanners in outpatient facilities consumed more energy. The breakdown of energy usage and costs by operating modes showed scanners spend between 61% and 93% of their time in nonproductive modes, with one outlier spending 34%. Average daily energy use for the scanners in the study ranged from 160 to 1069 kWh, with energy costs ranging from $\$$9 to $\$$149. This study uses an existing framework to quantify MRI energy behavior, leading to insights that can enable improved performance and cost savings across different healthcare environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

TOFHunter—unlocking rapid untargeted screening of inductively coupled plasma–time-of-flight–mass spectrometry data

This study provides an overview of a newly developed open source program written in Python, TOFHunter, which permits the rapid and untargeted screening of inductively coupled plasma (ICP)-time-of-flight (TOF)-mass spectrometry (MS) datasets. ICP-TOF-MS is an analytical tool capable of providing quasi simultaneous detection of all nuclides from Li to Pu. This capability has triggered an increase in studies investigating single-particle analysis in which the TOF-MS provides correlated elemental/isotopic signatures on a particle basis in time. Similarly, laser ablation mapping has seen rapid growth owing to ICP-TOF-MS's capacity to handle fast washout times (<10 ms) while providing a broad nuclide coverage. The caveat to this broad mass coverage and high time resolution comes in the form of large, overwhelming datasets. With datasets typically on the scale of gigabytes, it is easy for a user to only focus on very targeted analytes; however, this focus diminishes the opportunity offered by the TOF-MS detector. TOFHunter applies chemometric methods, principal component analysis (PCA), and interesting features finder (IFF) on ICP-TOF-MS data, allowing for investigation of correlations, major and minor variance sources, and sample screening. The unique spectra identified by the (IFF) are used to generate a list of mass peaks, which are then matched with both nuclides and potential interferences before being exported for the user to investigate. Several case studies are discussed herein, demonstrating TOFHunter's ability to screen aqueous injections, single-particle/single-cell analysis, and probe laser ablation mapping files for unique regions of interest.

47 OTHER INSTRUMENTATION↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗

Density estimation via measure transport: Outlook for applications in the biological sciences

Abstract One among several advantages of measure transport methods is that they allow or a unified framework for processing and analysis of data distributed according to a wide class of probability measures. Within this context, we present results from computational studies aimed at assessing the potential of measure transport techniques, specifically, the use of triangular transport maps, as part of a workflow intended to support research in the biological sciences. Scenarios characterized by the availability of limited amount of sample data, which are common in domains such as radiation biology, are of particular interest. We find that when estimating a distribution density function given limited amount of sample data, adaptive transport maps are advantageous. In particular, statistics gathered from computing series of adaptive transport maps, trained on a series of randomly chosen subsets of the set of available data samples, leads to uncovering information hidden in the data. As a result, in the radiation biology application considered here, this approach provides a tool for generating hypotheses about gene relationships and their dynamics under radiation exposure.

gene expression data↗

CMPLE: Correlation Modeling to Decode Photosynthesis Using the Minorize–Maximize Algorithm

In plant genomic experiments, correlations among various biological traits (phenotypes) give new insights into how genetic diversity may have tuned biological processes to enhance fitness under diverse conditions. Consequently, knowing how the correlations are affected by genetic (G) and environmental (E) factors helps develop climate-resilient plants. However, the current literature lacks any method for assessing the effect of predictors on pairwise correlations among multiple phenotypes together with easily interpretable model parameters. To address this need, we propose to model pairwise correlations directly in terms of G and E and develop a computationally efficient inference procedure. Two major novelties in our methodology are (1) the use of a composite pairwise likelihood method to avoid the positive definiteness restriction on the correlation matrix and (2) the use of a novel Minorize–Maximize (MM) algorithm for the efficient estimation of a large number of parameters. The proposed method shows excellent numerical performance on synthetic datasets. Here, the analysis of the motivating data on cowpea reveals that the rates of solar energy storage by photosynthesis (the aggregate trait) are differentially affected by different genetic loci through two distinct processes: “photoinhibition” which results from photodamage caused by excess light, and “photoprotection” which protects plants from photodamage but also results in energy loss.

Correlation modeling↗

Non-propagating structures and propagating waves in solar wind turbulence revealed by simulations and observations

Structures and waves are common features of solar wind turbulence at various scales. The interplay between structures and waves is important for processes such as the turbulent energy cascade, plasma heating, and particle scattering. Our understanding of turbulence has been advanced by not only new space missions and numerical simulations, but also techniques that have been developed to interpret the rapidly growing turbulence data. We review basic models of turbulence with a specific focus on the analysis methods for understanding magnetic structures and waves. MHD and kinetic waves in single-spacecraft time series measurements can be identified through mode decomposition or their characteristic polarization signatures. The structures in this paper are considered as zero-frequency, non-propagating or convected modes embedded in the solar wind. The synergy between observations and simulations is most evident in the application of spatial-temporal analysis to multi-spacecraft observation and turbulence simulations. The spatial-temporal analysis has greatly improved our understanding of structures and waves in turbulence. We conclude by discussing prospects for future research.

79 ASTRONOMY AND ASTROPHYSICS↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

What more can be done with XPS? Highly informative but underused approaches to XPS data collection and analysis

Because of the importance of surfaces and interfaces in many scientific and technological areas, the use of x-ray photoelectron spectroscopy (XPS) has been growing exponentially. Although XPS is being used to obtain useful information about the surface composition of samples, much more information about materials and their properties can be extracted from XPS data than commonly obtained. This paper describes some of the areas where alternative analysis methods or experimental design can obtain information about the near-surface region of a sample, often information not available in other ways. Experienced XPS analysts are familiar with many of these methods, but they may not be known to new or casual XPS users, and sometimes, they have not been used because of an inappropriately assumed complexity. The information available includes optical, electronic, and electrical properties; nanostructure; expanded chemical information; and enhanced analysis of biological materials and solid/liquid interfaces. Many of these analyses can be conducted on standard laboratory XPS systems, with either no or relatively minor system alterations. Topics discussed include (1) considerations beyond the “traditional” uniform surface layer composition calculation, (2) using the Auger parameter to determine a sample property, (3) use of the D parameter to identify sp 2 and sp 3 carbon information, (4) information from the XPS valence band, (5) using cryocooling to expand range of samples that can be analyzed and minimize damage, and (6) using electrical potential effects on XPS signals to extract chemically resolved electrical measurements including band alignment and electrical property information.

Baer, Donald R. [Pacific Northwest National Labora↗

2017 Puget Sound Regional Travel Study

# 2017 Puget Sound Regional Travel Study The 2017 Puget Sound Regional Travel Study collected household- and person-level activity and travel pattern information from residents throughout the Puget Sound Regional Council's four-county region in Washington State. It followed the [2014-2015 Puget Sound Regional Travel Study](https://www.nrel.gov/transportation/secure-transportation-data/tsdc-puget-sound-travel-study), starting a planned six-year data collection that includes 2019 and 2021. The multiyear program's goal is to maintain an updated source of household travel behavior data that: - Supports modeling and planning needs - Facilitates trend analysis over time - Allows for regular study design updates to integrate evolving data collection methods and emerging travel behaviors and transportation issues. ## Data Collection Agency The Puget Sound Regional Council conducted the study. ## Methodology The 2017 study featured both the design and administration of a one-day household travel diary (approximately 80% of households before data cleaning) and a seven-day smartphone global positioning system (GPS) diary (approximately 20% of households before data cleaning). It combined data collection methods, including smartphone, online, and telephone. The survey design included several stages to recruit and collect data about households, their members, and their travel behaviors during the assigned travel period. ## Survey Records Survey records include a total of 6,254 participants. ## More Information For more information, see the [survey documentation](https://www.nrel.gov/media/docs/libraries/tsdc/zip/tsdc-2017-puget-sound-travel-study-documentation.zip?sfvrsn=45991f2f_1). ## Transportation Data The data set contains a demographic and socioeconomic composition of 6,254 people from 3,285 households in the Puget Sound regional area, as well as detailed information on the travel behavior of each household for a designated 24-hour period. The survey logged over 508 thousand vehicle miles of travel by participants during 52,492 trips. Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗