Search NASA⌕ Search

SEARCH · Search NASA

Results for “t-SNE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Galactic ArchaeoLogIcaL ExcavatiOns (GALILEO) II. t-SNE portrait of local fossil relics and structures

Based on high-quality Apache Point Observatory Galactic Evolution Experiment (APOGEE) DR17 and Gaia DR3 data for 1742 red giants stars within 5 kpc of the Sun and not rotating with the Galactic disk (V φ < 100 km s -1 ), we used the nonlinear technique of unsupervised analysis t-Distributed Stochastic Neighbor Embedding (t-SNE) to detect coherent structures in the space of ten chemical-abundance ratios: [Fe/H], [O/Fe], [Mg/Fe], [Si/Fe], [Ca/Fe], [C/Fe], [N/Fe], [Al/Fe], [Mn/Fe], and [Ni/Fe]. Additionally, we obtained orbital parameters for each star using the nonaxisymmetric gravitational potential GravPot16. Seven structures are detected, including Splash, Gaia-Sausage-Enceladus (GSE), the high-α heated-disk population, N-C-O peculiar stars, and inner disk-like stars, plus two other groups that did not match anything previously reported in the literature, here named Galileo 5 and Galileo 6 (G5 and G6). These two groups overlap with Splash in [Fe/H], with G5 having a lower metallicity than G6, and they are both between GSE and Splash in the [Mg/Mn] versus [Al/Fe] plane, with G5 being in the α-rich in situ locus and G6 on the border of the α-poor in situ one. Nonetheless, their low [Ni/Fe] hints at a possible ex situ origin. Their orbital energy distributions are between Splash and GSE, with G5 being slightly more energetic than G6. We verified the robustness of all the obtained groups by exploring a large range of t-SNE parameters, applying it to various subsets of data, and also measuring the effect of abundance errors through Monte Carlo tests.

79 ASTRONOMY AND ASTROPHYSICS↗

Structuring Nutrient Yields throughout Mississippi/Atchafalaya River Basin Using Machine Learning Approaches

To minimize the eutrophication pressure along the Gulf of Mexico or reduce the size of the hypoxic zone in the Gulf of Mexico, it is important to understand the underlying temporal and spatial variations and correlations in excess nutrient loads, which are strongly associated with the formation of hypoxia. This study’s objective was to reveal and visualize structures in high-dimensional datasets of nutrient yield distributions throughout the Mississippi/Atchafalaya River Basin (MARB). For this purpose, the annual mean nutrient concentrations were collected from thirty-three US Geological Survey (USGS) water stations scattered in the upper and lower MARB from 1996 to 2020. Eight surface water quality indicators were selected to make comparisons among water stations along the MARB over the past two decades. Principal component analysis (PCA) was used to comprehensively evaluate the nutrient yields across thirty-three USGS monitoring stations and identify the major contributing nutrient loads. The results showed that all samples could be analyzed using two main components, which accounted for 81.6% of the total variance. The PCA results showed that yields of orthophosphate (OP), silica (SI), nitrate–nitrites (NO 3 -NO 2 ), and total suspended sediment (TSS) are major contributors to nutrient yields. It also showed that land-planted crops, density of population, domestic and industrial discharges, and precipitation are fundamental causes of excess nutrient loads in MARB. These factors are of great significance for the excess nutrient load management and pollution control of the Mississippi River. It was found that the average nutrient yields were stable within the sub-MARB area, but the large nitrogen yields in the upper MARB and the large phosphorus yields in the lower MARB were of great concern. t-distributed stochastic neighbor embedding (t-SNE) revealed interesting nonlinear and local structures in nutrient yield distributions. Clustering analysis (CA) showed the detailed development of similarities in the nutrient yield distribution. Moreover, PCA, t-SNE, and CA showed consistent clustering results. This study demonstrated that the integration of dimension reduction techniques, PCA, and t-SNE with CA techniques in machine learning are effective tools for the visualization of the structures of the correlations in high-dimensional datasets of nutrient yields and provide a comprehensive understanding of the correlations in the distributions of nutrient loads across the MARB.

54 ENVIRONMENTAL SCIENCES↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗

Unsupervised machine learning for unbiased chemical classification in X-ray absorption spectroscopy and X-ray emission spectroscopy

Here we report a comprehensive computational study of unsupervised machine learning for extraction of chemically relevant information in X-ray absorption near edge structure (XANES) and in valence-to-core X-ray emission spectra (VtC-XES) for classification of a broad ensemble of sulphorganic molecules. By progressively decreasing the constraining assumptions of the unsupervised machine learning algorithm, moving from principal component analysis (PCA) to a variational autoencoder (VAE) to t-distributed stochastic neighbour embedding (t-SNE), we find improved sensitivity to steadily more refined chemical information. Surprisingly, when embedding the ensemble of spectra in merely two dimensions, t-SNE distinguishes not just oxidation state and general sulphur bonding environment but also the aromaticity of the bonding radical group with 87% accuracy as well as identifying even finer details in electronic structure within aromatic or aliphatic sub-classes. We find that the chemical information in XANES and VtC-XES is very similar in character and content, although they unexpectedly have different sensitivity within a given molecular class. We also discuss likely benefits from further effort with unsupervised machine learning and from the interplay between supervised and unsupervised machine learning for X-ray spectroscopies. Our overall results, i.e., the ability to reliably classify without user bias and to discover unexpected chemical signatures for XANES and VtC-XES, likely generalize to other systems as well as to other one-dimensional chemical spectroscopies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Rheological Properties of Small-Molecular Liquids at High Shear Strain Rates

Molecular-scale understanding of rheological properties of small-molecular liquids and polymers is critical to optimizing their performance in practical applications such as lubrication and hydraulic fracking. We combine nonequilibrium molecular dynamics simulations with two unsupervised machine learning methods: principal component analysis (PCA) and t-distributed stochastic neighbor embedding (t-SNE), to extract the correlation between the rheological properties and molecular structure of squalane sheared at high strain rates (10 6 –10 10 s -1 ) for which substantial shear thinning is observed under pressures P ϵ 0.1–955 MPa at 293 K. Intramolecular atom pair orientation tensors of 435 × 6 dimensions and the intermolecular atom pair orientation tensors of 61 × 6 dimensions are reduced and visualized using PCA and t-SNE to assess the changes in the orientation order during the shear thinning of squalane. Dimension reduction of intramolecular orientation tensors at low pressures P = 0.1,100 MPa reveals a strong correlation between changes in strain rate and the orientation of the side-backbone atom pairs, end-backbone atom pairs, short backbone-backbone atom pairs, and long backbone-backbone atom pairs associated with a squalane molecule. At high pressures P ≥ 400 MPa, the orientation tensors are better classified by these different pair types rather than strain rate, signaling an overall limited evolution of intramolecular orientation with changes in strain rate. Dimension reduction also finds no clear evidence of the link between shear thinning at high pressures and changes in the intermolecular orientation. The alignment of squalane molecules is found to be saturated over the entire range of rates during which squalane exhibits substantial shear thinning at high pressures.

36 MATERIALS SCIENCE↗

Dimensionality reduction using elastic measures

With the recent surge in big data analytics for hyperdimensional data, there is a renewed interest in dimensionality reduction techniques. In order for these methods to improve performance gains and understanding of the underlying data, a proper metric needs to be identified. This step is often overlooked, and metrics are typically chosen without consideration of the underlying geometry of the data. Here, in this paper, we present a method for incorporating elastic metrics into the t-distributed stochastic neighbour embedding (t-SNE) and Uniform Manifold Approximation and Projection (UMAP). We apply our method to functional data, which is uniquely characterized by rotations, parameterization and scale. If these properties are ignored, they can lead to incorrect analysis and poor classification performance. Through our method, we demonstrate improved performance on shape identification tasks for three benchmark data sets (MPEG-7, Car data set and Plane data set of Thankoor), where we achieve 0.77, 0.95 and 1.00 F1 score, respectively.

97 MATHEMATICS AND COMPUTING↗

Non-intrusive reduced order modeling of natural convection in porous media using convolutional autoencoders: Comparison with linear subspace techniques

Natural convection in porous media is a highly nonlinear multiphysical problem relevant to many engineering applications (e.g., the process of CO 2 sequestration). Here, we extend and present a non-intrusive reduced order model of natural convection in porous media employing deep convolutional autoencoders for the compression and reconstruction and either radial basis function (RBF) interpolation or artificial neural networks (ANNs) for mapping parameters of partial differential equations (PDEs) on the corresponding nonlinear manifolds. To benchmark our approach, we also describe linear compression and reconstruction processes relying on proper orthogonal decomposition (POD) and ANNs. Further, we present comprehensive comparisons among different models through three benchmark problems. The reduced order models, linear and nonlinear approaches, are much faster than the finite element model, obtaining a maximum speed-up of 7 × 10 6 because our framework is not bound by the Courant–Friedrichs–Lewy condition; hence, it could deliver quantities of interest at any given time contrary to the finite element model. Our model’s accuracy still lies within a relative error of 7% in the worst-case scenario. We illustrate that, in specific settings, the nonlinear approach outperforms its linear counterpart and vice versa. We hypothesize that a visual comparison between principal component analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE) could indicate which method will perform better prior to employing any specific compression strategy.

97 MATHEMATICS AND COMPUTING↗

A computer vision algorithm for interpreting lacustrine carbonate textures at Searles Valley, USA

Investigations of the paleohydrologies of pluvial lake systems have often employed lake carbonate deposits called “tufa” that grow subaqueously and can be preserved long after the drying of the lake. For this reason, tufa have been used as a proxy for minimum lake level. However, they exhibit a variety of textures that hold the potential to reveal richer paleoclimatological information. With the goal of determining if tufa texture can be used as a proxy for lake environment, this study investigates the textures of tufa at Mono Lake, California in comparison to the fossil tufa in Searles Valley, California. While observations in the last century suggest that the tufa in the Mono basin grew in waters similar to the modern, the tufa at Searles formed during the last glacial period, when the Great Basin contained a system of pluvial lakes on the scale of the modern Great Lakes. The tufa at both basins have been observed to have a range of classifiable textures, and new methods of inspecting visual data could be informative about what factors control these textures. To this end, a t-Distributed Stochastic Neighbor Embedding (t-SNE) algorithm is used to project images of the tufa at Searles and Mono into a coordinate space, allowing for simple, quantitative comparisons of the visual similarity of textures. In this work, the textures of tufa at Searles are compared to each other, as well as to the tufa at Mono. This study performs a robust assessment of the feasibility of Mono Lake as a modern analogue for Searles Valley. It finds that there is a justifiable basis for the comparison of certain fossil facies at Searles to the tufa at Mono, significant progress towards the goal of using texture as a metric for the environment in which tufa formed.

58 GEOSCIENCES↗

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure↗

Uncertainty quantification for molecular property predictions with graph neural architecture search

Graph Neural Networks (GNNs) have emerged as a prominent class of data-driven methods for molecular property prediction. However, a key limitation of typical GNN models is their inability to quantify uncertainties in the predictions. This capability is crucial for ensuring the trustworthy use and deployment of models in downstream tasks. To that end, we introduce AutoGNNUQ, an automated uncertainty quantification (UQ) approach for molecular property prediction. AutoGNNUQ leverages architecture search to generate an ensemble of high-performing GNNs, enabling the estimation of predictive uncertainties. Our approach employs variance decomposition to separate data (aleatoric) and model (epistemic) uncertainties, providing valuable insights for reducing them. In our computational experiments, we demonstrate that AutoGNNUQ outperforms existing UQ methods in terms of both prediction accuracy and UQ performance on multiple benchmark datasets, and generalizes well to out-of-distribution datasets. Additionally, we utilize t-SNE visualization to explore correlations between molecular features and uncertainty, offering insight for dataset improvement. AutoGNNUQ has broad applicability in domains such as drug discovery and materials science, where accurate uncertainty quantification is crucial for decision-making.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

StarHorse results for spectroscopic surveys and Gaia DR3: Chrono-chemical populations in the solar vicinity, the genuine thick disk, and young alpha-rich stars

The Gaia mission has provided an invaluable wealth of astrometric data for more than a billion stars in our Galaxy. The synergy between Gaia astrometry, photometry, and spectroscopic surveys gives us comprehensive information about the Milky Way. Using the Bayesian isochrone-fitting code StarHorse, we derive distances and extinctions for more than 10 million unique stars listed in both Gaia Data Release 3 and public spectroscopic surveys: 557 559 in GALAH+ DR3, 4 531 028 in LAMOST DR7 LRS, 347 535 in LAMOST DR7 MRS, 562 424 in APOGEE DR17, 471 490 in RAVE DR6, 249 991 in SDSS DR12 (optical spectra from BOSS and SEGUE), 67 562 in the Gaia-ESO DR5 survey, and 4 211 087 in the Gaia RVS part of the Gaia DR3 release. StarHorse can increase the precision of distance and extinction measurements where Gaia parallaxes alone would be uncertain. We used StarHorse for the first time to derive stellar ages for main-sequence turnoff and subgiant branch stars, around 2.5 million stars, with age uncertainties typically around 30%; the uncertainties drop to 15% for subgiant-branch-only stars, depending on the resolution of the survey. With the derived ages in hand, we investigated the chemical-age relations. In particular, the α and neutron-capture element ratios versus age in the solar neighbourhood show trends similar to previous works, validating our ages. We used the chemical abundances from local subgiant samples of GALAH DR3, APOGEE DR17, and LAMOST MRS DR7 to map groups with similar chemical compositions and StarHorse ages, using the dimensionality reduction technique t-SNE and the clustering algorithm HDBSCAN. We identify three distinct groups in all three samples, confirmed by their kinematic properties: the genuine chemical thick disk, the thin disk, and a considerable number of young alpha-rich stars (427) that are also a part of the delivered catalogues. We confirm that the genuine thick disk’s kinematics and age properties are radically different from those of the thin disk and compatible with high-redshift (z ≈ 2) star-forming disks with high dispersion velocities. We also find a few extra chemical populations in GALAH DR3 thanks to the availability of neutron-capture element information.

79 ASTRONOMY AND ASTROPHYSICS↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗

Few-shot Learning for Post-disaster Structure Damage Assessment

Automating post-disaster damage assessment with remote sensing data is critical for faster surveys of structures impacted by natural disasters. One significant obstacle to training state-of-the-art deep neural networks to support this automation is that large quantities of labelled data are often required. However, obtaining those labels is particularly unrealistic to support post-disaster damage assessment in a timely manner. Few-shot learning methods could help to mitigate this by reducing the amount of labelled data required to successfully train a model while achieving satisfactory results. To this end, we explore a feature reweighting method to the YOLOv3 object detection architecture to achieve few-shot learning of damage assessment models on the xBD dataset. Our results show that the feature reweighting approach yield improved mAP over the baseline with significantly fewer labelled samples. In addition, we use t-SNE to analyze the class-specific reweighting vectors generated by the reweighting module in order to evaluate their inter-class and intra-class similarity. We find that the vectors form clusters based on class, and that these clusters overlap with visually similar classes. Those results show the potential to employ this few-shot learning strategy for rapid damage assessment with post-event remote sensing images.

Bowman, Jordan↗

Machine learning-based analysis of COVID-19 pandemic impact on US research networks

Here in this study we explore how fallout from the changing public health policy around COVID-19 has changed how researchers access and process their science experiments. Using a combination of techniques from statistical analysis and machine learning, we conduct a retrospective analysis of historical network data for a period around the stay-at-home orders that took place in March 2020. Our analysis takes data from the entire ESnet infrastructure to explore DOE high-performance computing (HPC) resources at OLCF, ALCF, and NERSC, as well as User sites such as PNNL and JLAB. We look at detecting and quantifying changes in site activity using a combination of t-Distributed Stochastic Neighbor Embedding (t-SNE) and decision tree analysis. Our findings bring insights into the working patterns and impact on data volume movements, particularly during late-night hours and weekends.

97 MATHEMATICS AND COMPUTING↗

AI-powered topic modeling: comparing LDA and BERTopic in analyzing opioid-related cardiovascular risks in women

Topic modeling is a crucial technique in natural language processing (NLP), enabling the extraction of latent themes from large text corpora. Traditional topic modeling, such as Latent Dirichlet Allocation (LDA), faces limitations in capturing the semantic relationships in the text document although it has been widely applied in text mining. BERTopic, created in 2022, leveraged advances in deep learning and can capture the contextual relationships between words. In this work, we integrated Artificial Intelligence (AI) modules to LDA and BERTopic and provided a comprehensive comparison on the analysis of prescription opioid-related cardiovascular risks in women. Opioid use can increase the risk of cardiovascular problems in women such as arrhythmia, hypotension etc. 1,837 abstracts were retrieved and downloaded from PubMed as of April 2024 using three Medical Subject Headings (MeSH) words: “opioid,” “cardiovascular,” and “women.” Machine Learning of Language Toolkit (MALLET) was employed for the implementation of LDA. BioBERT was used for document embedding in BERTopic. Eighteen was selected as the optimal topic number for MALLET and 23 for BERTopic. ChatGPT-4-Turbo was integrated to interpret and compare the results. The short descriptions created by ChatGPT for each topic from LDA and BERTopic were highly correlated, and the performance accuracies of LDA and BERTopic were similar as determined by expert manual reviews of the abstracts grouped by their predominant topics. The results of the t-SNE (t-distributed Stochastic Neighbor Embedding) plots showed that the clusters created from BERTopic were more compact and well-separated, representing improved coherence and distinctiveness between the topics. Our findings indicated that AI algorithms could augment both traditional and contemporary topic modeling techniques. In addition, BERTopic has the connection port for ChatGPT-4-Turbo or other large language models in its algorithm for automatic interpretation, while with LDA interpretation must be manually, and needs special procedures for data pre-processing and stop words exclusion. Therefore, while LDA remains valuable for large-scale text analysis with resource constraints, AI-assisted BERTopic offers significant advantages in providing the enhanced interpretability and the improved semantic coherence for extracting valuable insights from textual data.

Research & Experimental Medicine↗

Uncovering Structure–Conductivity Relationships in Anion Exchange Membranes (AEMs) Using Interpretable Machine Learning

Anion exchange membranes (AEMs) play a vital role in the performance of water electrolyzers and fuel cells, yet their discovery and optimization remain challenging due to the complexity of structure–property relationships. In this study, we introduce a machine learning framework that leverages conditional graph neural networks (cGNNs) and descriptor-based models and a hybrid graph neural network (HGARE) to predict and interpret ionic conductivity. The descriptor-based pipeline employs principal component analysis (PCA), ablation, and SHAP analysis to identify factors governing anion conductivity, revealing electronic, topological, and compositional descriptors as key contributors. Beyond prediction, dimensionality reduction and clustering are performed by employing t-SNE and KMeans as well as SOM, which reveal distinct membranes clusters, some of which were enriched with high anion conductivity. Among graph-based approaches, the graph convolutional (GCN) achieved strong predictive performance, while the Hybrid Graph Autoencoder-Regressor Ensemble (HGARE) achieved the highest accuracy. Additionally, atom-level saliency maps from GCN provide spatial explanations for conductive behavior, revealing the importance of polarizable and flexible regions. This work contributes to the accelerated and data-driven design of high-performance AEMs.

Naghshnejad, Pegah [Department of Chemical Enginee↗

Development of a large-mass, low-threshold detector system with simultaneous measurements of athermal phonons and scintillation light

For this work, we have combined two low-threshold detector technologies to develop a large-mass, low-threshold detector system that simultaneously measures the athermal phonons in a sapphire detector while an adjacent silicon high-voltage detector detects the scintillation light from the sapphire detector. This detector system could provide event-by-event discrimination between electron and nuclear events due to the difference in their scintillation light yield. While such systems with simultaneous phonon and light detection have been demonstrated earlier with smaller detectors, our system is designed to provide a large detector mass with high amplification for the limited scintillation light. Future work will focus on at least an order of magnitude improvement in the light collection efficiency by having a highly reflective detector housing and custom phonon mask design to maximize light collection by the silicon high-voltage detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Nitrogen Metabolism in Pseudomonas putida : Functional Analysis Using Random Barcode Transposon Sequencing

Pseudomonas putida KT2440 has long been studied for its diverse and robust metabolisms, yet many genes and proteins imparting these growth capacities remain uncharacterized. Using pooled mutant fitness assays, we identified genes and proteins involved in the assimilation of 52 different nitrogen containing compounds. To assay amino acid biosynthesis, 19 amino acid drop-out conditions were also tested. From these 71 conditions, significant fitness phenotypes were elicited in 672 different genes including 100 transcriptional regulators and 112 transport-related proteins. We divide these conditions into 6 classes, and propose assimilatory pathways for the compounds based on this wealth of genetic data. To complement these data, we characterize the substrate range of three promiscuous aminotransferases relevant to metabolic engineering efforts in vitro. So we examine the specificity of five transcriptional regulators, explaining some fitness data results and exploring their potential to be developed into useful synthetic biology tools. In addition, we use manifold learning to create an interactive visualization tool for interpreting our BarSeq data, which will improve the accessibility and utility of this work to other researchers.

59 BASIC BIOLOGICAL SCIENCES↗