Search NASA⌕ Search

SEARCH · Search NASA

Results for “cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

Cosmological constraints from a joint DESI DR1 Full-Shape and DR2 BAO

We present a cosmological analysis combining full-shape (FS) clustering measurements from the Dark Energy Spectroscopic Instrument (DESI) DR1 with baryon acoustic oscillation (BAO) measurements from DESI DR2. To achieve a robust combination that accounts for the correlation between the two data releases, we employ the ShapeFit compression method and estimate the joint covariance using EZmocks. This compressed approach inherently mitigates the prior volume effects that have previously dominated Bayesian constraints from DESI data with minimal external priors. Consequently, we obtain — for the first time within a Bayesian framework — reliable DESI-only constraints on extensions to ΛCDM using only a Big Bang Nucleosynthesis prior on the baryon density and a wide prior on the spectral index. In flat ΛCDM, we find Ω m = 0.3035 ± 0.0085, h = 0.6876 ± 0.0059, and σ 8 = 0.822 ± 0.034. For the w 0 w a CDM dynamical dark energy model, we measure w 0 = -0.49 ± 0.25 and w a = -1.52 ± 0.77, improving constraints by ∼ 30% relative to the analogous DR1 measurement and reducing the discrepancy with ΛCDM to 1.4σ when compared to BAO only analyses. We also report competitive limits on the sum of neutrino masses and spatial curvature. This work demonstrates that the ShapeFit compression provides a prior-robust and computationally efficient pathway to constrain beyond-ΛCDM physics with large-scale structure.

baryon acoustic oscillations↗

Assembly bias and local Primordial non-Gaussianity from DESI DR1 quasars

The analysis of the large-scale clustering of quasars (QSO) observed by the Dark Energy Spectroscopic Instrument (DESI) represents a promising avenue for constraining local Primordial non-Gaussianity (PNG), parameterized by f NL . The signal to be constrained is the scale-dependent bias induced in the 2-point clustering of the considered tracer sample. The resulting constraints on f NL , however, are fully degenerate with the local PNG bias parameter b ϕ , dependent on the assembly bias parameter p. Using IllustrisTNG hydrodynamical simulations, we select a QSO sample reflecting the selection criteria and properties of DESI QSOs, and provide a robust prior for p, and thus for b ϕ , building on the findings of Fondi et al. 2024. We find a distribution with mean p̅ ≃ 1.4 with weak redshift dependence, stable to selection noise and consistent with the expected recent merger history typical of quasar-hosting halos. By comparing with the CAMELS simulations we demonstrate that this prior is robust to astrophysical assumptions and cosmic variance. Finally, applying this prior to the DESI DR1 dataset, we derive updated constraints on local PNG, obtaining f NL = -3.3±9.2.

cosmological parameters from LSS↗

Genetic and Functional Diversity Help Explain Pathogenic, Weakly Pathogenic, and Commensal Lifestyles in the Genus Xanthomonas

The genus Xanthomonas has been primarily studied for pathogenic interactions with plants. However, besides host and tissue-specific pathogenic strains, this genus also comprises nonpathogenic strains isolated from a broad range of hosts, sometimes in association with pathogenic strains, and other environments, including rainwater. Based on their incapacity or limited capacity to cause symptoms on the host of isolation, nonpathogenic xanthomonads can be further characterized as commensal and weakly pathogenic. This study aimed to understand the diversity and evolution of nonpathogenic xanthomonads compared to their pathogenic counterparts based on their cooccurrence and phylogenetic relationship and to identify genomic traits that form the basis of a life history framework that groups xanthomonads by ecological strategies. We sequenced genomes of 83 strains spanning the genus phylogeny and identified eight novel species, indicating unexplored diversity. While some nonpathogenic species have experienced a recent loss of a type III secretion system, specifically the hrp2 cluster, we observed an apparent lack of association of the hrp2 cluster with lifestyles of diverse species. We performed association analysis on a large data set of 337 Xanthomonas strains to explain how xanthomonads may have established association with the plants across the continuum of lifestyles from commensals to weak pathogens to pathogens. Presence of distinct transcriptional regulators, distinct nutrient utilization and assimilation genes, transcriptional regulators, and chemotaxis genes may explain lifestyle-specific adaptations of xanthomonads.

59 BASIC BIOLOGICAL SCIENCES↗

Polyyne production is regulated by the transcriptional regulators PgnC and GacA in Pseudomonas protegens Pf-5

ABSTRACT Polyynes produced by bacteria have promising applications in agriculture and medicine due to their potent antimicrobial activities. Polyyne biosynthetic genes have been identified inPseudomonasandBurkholderia. However, the molecular mechanisms underlying the regulation of polyyne biosynthesis remain largely unknown. In this study, we used a soil bacteriumPseudomonas protegensPf-5, which was recently reported to produce polyyne called protegenin, as a model to investigate the regulation of bacterial polyyne production. Our results show that Pf-5 controls polyyne production at both the pathway-specific level and a higher global level. Mutation ofpgnC, a transcriptional regulatory gene located in the polyyne biosynthetic gene cluster, abolished polyyne production. Gene expression analysis revealed that PgnC directly activates the promoter of polyyne biosynthetic genes. The production of polyyne also requires a global regulator GacA. Mutation ofgacAdecreased the translation of PgnC, which is consistent with the result thatpgnCleader mRNA bound directly to RsmE, an RNA-binding protein negatively regulated by GacA. These results suggest that GacA induces the expression of the PgnC regulator, which in turn activates polyyne biosynthesis. Additionally, the polyyne-producing strain of Pf-5, but not the polyyne-nonproducing strain, could inhibit a broad spectrum of bacteria including both Gram-negative and Gram-positive bacteria. IMPORTANCE Antimicrobial metabolites produced by bacteria are widely used in agriculture and medicine to control plant, animal, and human pathogens. Although bacteria-derived polyynes have been identified as potent antimicrobials for decades, the molecular mechanisms by which bacteria regulate polyyne biosynthesis remain understudied. In this study, we found that polyyne biosynthesis is directly activated by a pathway-specific regulator PgnC, which is induced by a global regulator GacA through the RNA-binding protein RsmE inPseudomonas protegens. To our knowledge, this work is the first comprehensive study of the regulatory mechanisms of bacterial polyyne biosynthesis at both pathway-specific level and global level. The discovered molecular mechanisms can help us optimize polyyne production for agricultural or medical applications.

Biotechnology & Applied Microbiology↗

Machine Learning for Anomaly Detection in Neural Network Security and SRF Cavities

This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.

Ferguson, Hal [Old Dominion University]↗

Origin and modal petrography of Luna 24 soils

Petrographic modal analyses of polished grain mounts of fractions in the 20 to 250 micron size range from Luna 24 soil samples are presented and used to infer the nature and relative contributions of source rocks. It is found that more than 90% of the identifiable rock fragments are mare basalts, with about 11% of the soil consisting of the crystalline form. Soil breccias, which make up nearly 10% of the soil, are found to be immature. Electron probe analysis of glass particles reveals principle clusters conforming to anorthosite, anorthositic gabbro and mare basalts. More than half of the soil is composed of monomineralic particles, with pyroxene as the most abundant mineral. It is concluded that 85% of the regolith is derived from local mare basalts and gabbros and about 10% is derived from early cumulates of local mare basalt magma. Highland sources are considered to contribute not more than 3% of the regolith.

Basu, A.↗

Carbon and nitrogen abundances in red giant stars in the globular cluster 47 Tucanae

The effects of changes in temperature, gravity, overall metal abundance, and carbon and nitrogen abundances have been investigated for model stellar spectra and colors representing globular-cluster giants of moderate metal deficiency. The results are presented in the form of spectral atlases and theoretical color-color diagrams. Using these results, approximate abundances of carbon and nitrogen have been derived for some red giant stars in 47 Tuc, from intermediate- and low-dispersion spectra and from intermediate- and narrow-band photometry. In all the normal giants studied, nitrogen is overabundant by up to about a factor of 5 (the precise value depends on the adopted carbon abundance), with different enhancements for different giants. The observational material is not sufficient to distinguish between a normal carbon abundance and a slight carbon depletion for the giant-branch stars, but carbon appears to be somewhat depleted in stars on the asymptotic giant branch. A most probable value of M/H = -0.8 for the overall cluster metal abundance is suggested from analysis of Stromgren photometry of red horizontal-branch stars.

Dickens, R. J.↗

The discovery of an O subdwarf in the globular cluster NGC 6712

Spectral data for C49 in the globular cluster NGC 6712 are examined with reference to evolutionary models for low-mass stars that are contracting into white dwarfs. The spectrum of this star is found to be almost identical to that of the first sdO star found in a globular cluster, star A in NGC 6397. Analysis of the spectral data using nonlocal thermodynamic equilibrium models gives surface parameters T not less than 50,000 K, log g = 5.5 plus or minus 0.5, and N(He)/N(H) approximately equal to 0.1. The results are consistent with the models considered.

Remillard, R. A.↗

Content-addressable read/write memories for image analysis

The commonly encountered image analysis problems of region labeling and clustering are found to be cases of search-and-rename problem which can be solved in parallel by a system architecture that is inherently suitable for VLSI implementation. This architecture is a novel form of content-addressable memory (CAM) which provides parallel search and update functions, allowing speed reductions down to constant time per operation. It has been proposed in related investigations by Hall (1981) that, with VLSI, CAM-based structures with enhanced instruction sets for general purpose processing will be feasible.

Snyder, W. E.↗

Complete positive ion, electron, and ram negative ion measurements near Comet Halley (COPERNIC) plasma experiment for the European Giotto Mission

Participation of U.S. scientists on the COPERNIC (COmplete Positive ions, Electrons and Ram Negative Ion measurements near Comet Halley) plasma experiment on the Giotto mission is described. The experiment consisted of two detectors: the EESA (electron electrostatic analyzer) which provided three-dimensional measurements of the distribution of electrons from 10 eV to 30 keV, and the PICCA (positive ion cluster composition analyzer) which provided mass analysis of positively charged cold cometary ions from mass 10 to 210 amu. In addition, a small 3 deg wide sector of the EESA looking in the ram direction was devoted to the detection of negatively charged cold cometary ions. Both detectors operated perfectly up to near closest approach (approx. 600 km) to Halley, but impacts of dust particles and neutral gas on the spacecraft contaminated parts of the data during the last few minutes. Although no flight hardware was fabricated in the U.S., The U.S. made very significant contributions to the hardware design, ground support equipment (GSE) design and fabrication, and flight and data reduction software required for the experiment, and also participated fully in the data reduction and analysis, and theoretical modeling and interpretation. Cometary data analysis is presented.

Lin, Robert P.↗

The all-sky extragalactic X-ray foreground

It is argued that the local gravitational dipole implied by the 50 brightest X-ray clusters of galaxies at z greater than 0.013 considered by Lahav et al. (1989) is relatively small compared with that inferred from the only three clusters at lower redshifts. Recent dipole analysis of the X-ray flux from bright AGN observed with HEAO-1A2 indicates that they are indeed strong tracers of this matter. The implications of this for the very pronounced large-scale foreground anisotropies to be measured via low-redshift AGN resolved in more sensitive all-sky surveys are explored. For the total extragalactic X-ray sky, i.e., including the relatively large Cosmic X-ray Background as well as the contribution of the resolved sources, the dipole moment is small compared to the monopole. Because of the large peculiar velocity of the Local Group, it is found that the currently estimated value for this total dipole moment can be used to set severe constraints on the volume emissivity arising from all X-ray sources within the present epoch.

Boldt, Elihu↗

Galaxy Cluster Masses at Moderate Redshifts

The masses of galaxy clusters are dominated by dark matter, and a robust determination of their masses has the potential of indicating how much dark matter exists on large scales in the universe, and the cosmological parameter Omega. X-ray observations of galaxy clusters provide a direct measure of both the gas mass in the intra-cluster medium, and also the total gravitating mass of the cluster. We used new and archival ROSAT observations to measure these quantities for a sample of intermediate redshift clusters which have also been subject to intensive dynamical studies, in order to compare the mass estimates from different methods. We used data from 14 of the CNOC cluster sample at 0.18 less than z less than 0.55 for this study. A direct comparison of dynamical mass estimates from Carlberg, Yee & Ellingson (1997) yielded surprisingly good results. The X-ray/dynamical mass ratios have a mean of 0.96+/- 0.10, indicating that for this sample, both methods are probably yielding very robust mass estimates. Comparison with mass estimates from gravitational lensing studies from the literature showed a small systematic with weak lensing estimates, and large discrepancies with strong lensing estimates. This latter is not surprising, given that these measurement are made close to the central core, where optical and X-ray estimates are less certain, and where substructure and the effects of individual galaxies will be more pronounced. These results are presented in Lewis, Ellingson, Morris/Carlberg, 1998, submitted to the Astrophysical Journal. (Note that Lewis is Ellingson's Ph.D. thesis, who received direct support from this grant and is using this investigation as part of his thesis.) Three additional papers are in preparation. The first provides a comparison of the mass profiles as measured in X- rays and in galaxy dynamics. These profiles are difficult to determine for individual clusters, and are subject to asphericity and other individual quirks of each cluster. However, a composite profile for each method will allow us to test our assumptions of hydrostatic/dynamical equilibrium in the sample as a whole. A second paper provides a more detailed look at the cluster MS0906+11, which is a merging system. A third paper authord with J. Mohr at U. Chicago will invetigate the size-temoerature relationship for intermediate redshift clusters and its impications on cluster formation and cosmology. Future work on these data will include comparisons of the cluster galaxy populations and the extent of the intra-cluster medium, and a more homogeneous analysis of gravitational lensing.

Ellingson, E.↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗

Interpretable Categorization of Heterogeneous Time Series Data

We analyze data from simulated aircraft encounters to validate and inform the development of a prototype aircraft collision avoidance system. The high-dimensional and heterogeneous time series dataset is analyzed to discover properties of near mid-air collisions (NMACs) and categorize the NMAC encounters. Domain experts use these properties to better organize and understand NMAC occurrences. Existing solutions either are not capable of handling high-dimensional and heterogeneous time series datasets or do not provide explanations that are interpretable by a domain expert. The latter is critical to the acceptance and deployment of safety-critical systems. To address this gap, we propose grammar-based decision trees along with a learning algorithm. Our approach extends decision trees with a grammar framework for classifying heterogeneous time series data. A context-free grammar is used to derive decision expressions that are interpretable, application-specific, and support heterogeneous data types. In addition to classification, we show how grammar-based decision trees can also be used for categorization, which is a combination of clustering and generating interpretable explanations for each cluster. We apply grammar-based decision trees to a simulated aircraft encounter dataset and evaluate the performance of four variants of our learning algorithm. The best algorithm is used to analyze and categorize near mid-air collisions in the aircraft encounter dataset. We describe each discovered category in detail and discuss its relevance to aircraft collision avoidance.

Drones↗