Search NASA⌕ Search

SEARCH · Search NASA

Results for “high dimensional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Visualizing Temporal Topic Embeddings with a Compass

—Dynamic topic modeling is useful at discovering the development and change in latent topics over time. However, present methodology relies on algorithms that separate document and word representations. This prevents the creation of a meaningful embedding space where changes in word usage and documents can be directly analyzed in a temporal context. This paper proposes an expansion of the compass-aligned temporal Word2Vec methodology into dynamic topic modeling. Such a method allows for the direct comparison of word and document embeddings across time in dynamic topics. This enables the creation of visualizations that incorporate temporal word embeddings within the context of documents into topic visualizations. In experiments against the current state-of-the-art, our proposed method demonstrates overall competitive performance in topic relevancy and diversity across temporal datasets of varying size. Simultaneously, it provides insightful visualizations focused on temporal word embeddings while maintaining the insights provided by global topic evolution, advancing our understanding of how topics evolve over time.

Cluster analysis↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗

Exploiting correlations in multi-coincidence Coulomb explosion patterns for differentiating molecular structures using machine learning

Coulomb explosion imaging (CEI) is a powerful technique for capturing the real-time motion of individual atoms during ultrafast photochemical reactions. CEI generates high-dimensional data with naturally embedded correlations that allow mapping the coordinated motion of nuclei in molecules. This enables reliable separation of competing reaction pathways and makes this approach uniquely suited for characterizing weak reaction channels. However, rich information contained in experimental CEI patterns remains largely underexploited due to challenges in visualizing correlations between multiple observables in multi-dimensional parameter space. Here we present a new approach to CEI of intermediate-sized polyatomic molecules, detecting up to eight ionic fragments in coincidence and leveraging machine-learning-based analysis to identify patterns and correlations in the resulting high-dimensional momentum-space data, enabling robust molecular structure identification and differentiation. Our approach provides high-dimensional background-free data encoding exceptionally rich structural information and establishes an automated, scalable framework for extracting insightful information from the data. As a demonstration, we apply this method to image and distinguish dichloroethylene isomers, showcasing its potential for broader applications in molecular imaging. Our results pave the way for channel-specific analysis of ultrafast structural dynamics in chemically relevant systems, particularly for disentangling mixed reaction pathways and detecting contributions from weak channels and minority species.

Chemical Physics (physics.chem-ph)↗

Image processing software for imaging spectrometry

The paper presents a software system, Spectral Analysis Manager (SPAM), which has been specifically designed and implemented to provide the exploratory analysis tools necessary for imaging spectrometer data, using only modest computational resources. The basic design objectives are described as well as the major algorithms designed or adapted for high-dimensional images. Included in a discussion of system implementation are interactive data display, statistical analysis, image segmentation and spectral matching, and mixture analysis.

Mazer, Alan S.↗

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics↗

Cometary ion flow variations at Comet P/Halley as observed by the Giotto IMS experiment

The moments of the cometary-ion distributions are determined through a three-dimensional analysis of the Giotto IMS high-intensity spectrometer (HIS) data. The spectrometer is described, with emphasis on its angle analyzer and mass analyzer. The method of data analysis is outlined, with ion-flow vectors and temperatures being addressed. The results of the water group ion-flow profile are presented, and it is noted that, after crossing the cometopause region, the ions become gradually colder. At cometocentric distances larger than 130,000 km, the cometary-ion temperature is found to be in the area of 100 eV or higher, and derivations of the flow parameters are uncertain. The ion temperature and the flow speed become lower by about 50 eV after crossing the magnetic pile-up boundary. It is concluded that the observed velocity and temperature profiles can be explained on the basis of charge exchange processes.

Kettmann, G.↗

High dimensional reflectance analysis of soil organic matter

Recent breakthroughs in remote-sensing technology have led to the development of high spectral resolution imaging sensors for observation of earth surface features. This research was conducted to evaluate the effects of organic matter content and composition on narrowband soil reflectance across the visible and reflective infrared spectral ranges. Organic matter from four Indiana agricultural soils, ranging in organic C content from 0.99 to 1.72 percent, was extracted, fractionated, and purified. Six components of each soil were isolated and prepared for spectral analysis. Reflectance was measured in 210 narrow bands in the 400- to 2500-nm wavelength range. Statistical analysis of reflectance values indicated the potential of high dimensional reflectance data in specific visible, near-infrared, and middle-infrared bands to provide information about soil organic C content, but not organic matter composition. These bands also responded significantly to Fe- and Mn-oxide content.

Henderson, T. L.↗

An Aerodynamic Analysis of a Mixed Flow Turbine

The aerodynamic performance of a high-work Mixed Flow Turbine (MFT) is computed and compared with experimental data. A three dimensional (3-D) viscous analysis is applied to the single stage MFT geometry with a relatively long upstream transition duct. Predicted vane surface static pressures and circumferentially averaged spanwise quantities at stator and rotor exits agree favorably with data. Compared to the results of axisymmetric flow analysis from design intent, the 3-D computation agrees much better especially in the endwall regions where throughflow prediction fails to assess the loss mechanism properly. Potential sources of performance loss such as tip leakage and secondary flows are also properly captured by the analysis.

Kim, Chan M.↗

Discovering System Health Anomalies Using Data Mining Techniques

We present a data mining framework for the analysis and discovery of anomalies in high-dimensional time series of sensor measurements that would be found in an Integrated System Health Monitoring system. We specifically treat the problem of discovering anomalous features in the time series that may be indicative of a system anomaly, or in the case of a manned system, an anomaly due to the human. Identification of these anomalies is crucial to building stable, reusable, and cost-efficient systems. The framework consists of an analysis platform and new algorithms that can scale to thousands of sensor streams to discovers temporal anomalies. We discuss the mathematical framework that underlies the system and also describe in detail how this framework is general enough to encompass both discrete and continuous sensor measurements. We also describe a new set of data mining algorithms based on kernel methods and hidden Markov models that allow for the rapid assimilation, analysis, and discovery of system anomalies. We then describe the performance of the system on a real-world problem in the aircraft domain where we analyze the cockpit data from aircraft as well as data from the aircraft propulsion, control, and guidance systems. These data are discrete and continuous sensor measurements and are dealt with seamlessly in order to discover anomalous flights. We conclude with recommendations that describe the tradeoffs in building an integrated scalable platform for robust anomaly detection in ISHM applications.

Sriastava, Ashok, N.↗

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

In this work, we examine the use of high-order tensor decompositions to analyze degradation pathways emerging from accelerated stress testing of silicon photovoltaic (PV) modules. Matrix-based decompositions are powerful tools for studying two-dimensional data arrays and form the foundation of a host of classical data analysis techniques. Tensors are high-order extrapolations of matrices that are able to account for more parameter dimensions, and a variety of tensor decomposition methods have been developed that similarly seek to extend insights from matrix decompositions to higher dimensions. Applying and interpreting tensor decomposition methods to sequences of PV module image data, we seek to uncover and isolate different degradation modes occurring from accelerated stress testing procedures. Further, we consider the contributions of different modes to PV module performance degradations.

data analysis↗

A comparison of spectral mixture analysis an NDVI for ascertaining ecological variables

In this study, we compare the performance of spectral mixture analysis to the Normalized Difference Vegetation Index (NDVI) in detecting change in a grassland across topographically-induced nutrient gradients and different management schemes. The Konza Prairie Research Natural Area, Kansas, is a relatively homogeneous tallgrass prairie in which change in vegetation productivity occurs with respect to topographic positions in each watershed. The area is the site of long-term studies of the influence of fire and grazing on tallgrass production and was the site of the First ISLSCP (International Satellite Land Surface Climatology Project) Field Experiment (FIFE) from 1987 to 1989. Vegetation indices such as NDVI are commonly used with imagery collected in few (less than 10) spectral bands. However, the use of only two bands (e.g. NDVI) does not adequately account for the complex of signals making up most surface reflectance. Influences from background spectral variation and spatial heterogeneity may confound the direct relationship with biological or biophysical variables. High dimensional multispectral data allows for the application position of techniques such as derivative analysis and spectral curve fitting, thereby increasing the probability of successfully modeling the reflectance from mixed surfaces. The higher number of bands permits unmixing of a greater number of surface components, separating the vegetation signal for further analyses relevant to biological variables.

Wessman, Carol A.↗

A new method for mapping multidimensional data to lower dimensions

A multispectral mapping method is proposed which is based on the new concept of BEND (Bidimensional Effective Normalised Difference). The method, which involves taking one sample point at a time and finding the interrelationships between its features, is found very economical from the point of view of storage and processing time. It has good dimensionality reduction and clustering properties, and is highly suitable for computer analysis of large amounts of data. The transformed values obtained by this procedure are suitable for either a planar 2-space mapping of geological sample points or for making grayscale and color images of geo-terrains. A few examples are given to justify the efficacy of the proposed procedure.

Gowda, K. C.↗

Ares I and Ares I-X Stage Separation Aerodynamic Testing

The aerodynamics of the Ares I crew launch vehicle (CLV) and Ares I-X flight test vehicle (FTV) during stage separation was characterized by testing 1%-scale models at the Arnold Engineering Development Center s (AEDC) von Karman Gas Dynamics Facility (VKF) Tunnel A at Mach numbers of 4.5 and 5.5. To fill a large matrix of data points in an efficient manner, an injection system supported the upper stage and a captive trajectory system (CTS) was utilized as a support system for the first stage located downstream of the upper stage. In an overall extremely successful test, this complex experimental setup associated with advanced postprocessing of the wind tunnel data has enabled the construction of a multi-dimensional aerodynamic database for the analysis and simulation of the critical phase of stage separation at high supersonic Mach numbers. Additionally, an extensive set of data from repeated wind tunnel runs was gathered purposefully to ensure that the experimental uncertainty would be accurately quantified in this type of flow where few historical data is available for comparison on this type of vehicle and where Reynolds-averaged Navier-Stokes (RANS) computational simulations remain far from being a reliable source of static aerodynamic data.

Pinier, Jeremy T.↗

Affine Transformations to Enable Machine Learning for Semi-Quantitative EDS Analysis

Energy Dispersive X-ray Spectroscopy (EDS) is an essential technique for determining elemental concentrations and distributions within microstructures, critical for materials discovery, optimization, and qualification. However, most published EDS data is qualitative because current quantitative EDS analysis methods require extensive calibration and post-processing, limiting their practicality and widespread adoption. This work seeks to establish a framework for accelerated EDS characterization and spectrum analysis that can leverage ML to analyze correlations between various elemental compositions and resulting EDS spectra. The complex physics and data result in a high-dimensional problem that grows exponentially with the number of elements in the system and the complexity of the spectrum analysis. ML provides a way to compute and optimize the results of this highly dimensional problem in a flexible way to tailor it to the user’s specific needs and material system. However, the framework emphasizes transparency through a strictly mathematical affine transformation, so the analysis remains understandable and reviewable to facilitate adoption by the scientific community. While currently implemented methods are simplistic and unvalidated, further development and demonstration of this framework could enable high-throughput, accurate, and accessible EDS characterization.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗