Search NASA⌕ Search

SEARCH · Search NASA

Results for “Algorithms and data structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Bayesian desmearing algorithm for Bonse–Hart USANS with anisotropic scattering

Ultra-small-angle neutron scattering (USANS) using Bonse–Hart optics provides micrometer-scale structural insights but suffers from severe slit-geometry smearing. While well-established for isotropic systems, quantitative desmearing of anisotropic data remains a challenge because conventional corrections break down for non-radial scattering. In this work, we address this by developing a resolution-aware Bayesian framework that explicitly incorporates anisotropy via an affine deformation to the scattering pattern, guided by the principle of parsimony. This results in orientation-resolved point-spread functions that enable a self-consistent determination of both the resolution and deformation parameters. Using Gaussian process regression with uncertainty quantification and a probabilistic correction for multiple scattering, we demonstrate the framework’s effectiveness through numerical benchmarks and experimental studies of a stretched polymer melt. Our approach enables the seamless integration of SANS and USANS data, facilitating quantitative structural analysis of deformed materials at nanometer to micrometer scales.

36 MATERIALS SCIENCE↗

The DESI DR1 peculiar velocity survey: growth rate measurements from the maximum likelihood fields method

We present the constraint on the growth rate of structure from the combination of DESI DR1 BGS sample, Fundamental Plane, and Tully-Fisher peculiar velocity catalogues using the maximum likelihood fields method. The combined catalogue contains 415,523 galaxy redshifts and 76,616 peculiar velocity measurements. To handle the large amount of data in the DESI DR1 peculiar velocity catalogue, we significantly improve the computational efficiency by rewriting the algorithm with JAX. After removing outliers and Tully-Fisher galaxies that are affected by systematics, we find fσ 8 = 0.483 -0.043 +0.080 (stat) ± 0.018(sys), consistent within 1σ with the power spectrum and correlation function analysis using the same dataset. Combining all three measurements with appropriate correlations, the consensus measurement is fσ 8 (z eff = 0.07) = 0.450±0.055, consistent with Planck +ΛCDM cosmology (fσ 8 = 0.449±0.008). Combining with the high redshift growth rate of structure measurements from DESI ShapeFit, the constraint on the growth index is γ = 0.58±0.11, consistent with GR.

cosmic flows↗

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

FAIR Data and Interpretable AI Framework for Architectured Metamaterials

Our interdisciplinary effort successfully generated FAIR (Findable, Accessible, Interoperable, and Reusable) benchmark datasets for mechanical metamaterials while introducing a novel Artificial Intelligence (AI) framework known as Learning Refined Compositional Rules (LRCR). This framework was specifically designed to bridge the gap across varying computational length scales and extract the underlying physical mechanisms that connect a material's structural geometry to its bulk acoustic properties. Historically, the discovery of such structured materials relied heavily on human intuition or opaque, black-box optimization algorithms that were difficult to generalize. By combining interpretable machine learning techniques with rigorous experimental validation, this project established clear, generalizable design guidelines for tuning wave dispersion and controlling vibrations. Ultimately, the public availability of these structured datasets and algorithms will significantly reduce computational costs and accelerate the design of advanced multi-functional acoustic devices, offering broad societal impacts across fields like aerospace engineering, telecommunications, and biomedical implant design.

36 MATERIALS SCIENCE↗

Harmonic analysis of discrete tracers of large-scale structure

It is commonplace in cosmology to analyze fields projected onto the celestial sphere, and in particular density fields that are defined by a set of points e.g. galaxies. When performing an harmonic-space analysis of such data (e.g. an angular power spectrum) using a pixelized map one has to deal with aliasing of small-scale power and pixel window functions. We compare and contrast the approaches to this problem taken in the cosmic microwave background and large-scale structure communities, and advocate for a direct approach that avoids pixelization. We describe a method for performing a pseudo-spectrum analysis of a galaxy data set and show that it can be implemented efficiently using well-known algorithms for special functions that are suited to acceleration by graphics processing units (GPUs). The method returns the same spectra as the more traditional map-based approach if in the latter the number of pixels is taken to be sufficiently large and the mask is well sampled. The method is readily generalizable to cross-spectra and higher-order functions. It also provides a convenient route for distributing the information in a galaxy catalog directly in harmonic space, as a complement to releasing the configuration-space positions and weights, and a route to spectral apodization. Finally, we make public a code enabling the application of our method to existing and upcoming datasets.

79 ASTRONOMY AND ASTROPHYSICS↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Efficient analysis of small-angle scattering curves for large biomolecular assemblies using Monte Carlo methods

Structure elucidation from small-angle scattering curves of large biomolecular assemblies is notoriously challenging. This is because the simulation of high-resolution features in the structure of large macromolecular assemblies, such as de novo protein assemblies, is computationally demanding when it needs to cover a broad range of length scales. Conventional methods, such as the numerical approximation to the Debye equation or the use of spherical harmonics, do not scale well as the size of the assembly increases, which limits their application to small structures (e.g. individual proteins). This work explores the effectiveness of a Monte Carlo method to simulate and fit scattering curves for large biomolecular assemblies spanning over ranges covering atomic and molecular detail (e.g. spacing and orientation of proteins in an assembly) as well as large-scale (hundreds of nanometres) features. Owing to its speed and scalability, it can be combined with a fitting algorithm to extract structural features from experimental small-angle scattering curves in biomolecular assemblies that are otherwise intractable for interpretation. This work first demonstrates the effectiveness of the tool using experimental small-angle X-ray scattering (SAXS) data from tile-like proteins that assemble into 1D tube-like macromolecular structures. Here, the diameter distribution of tubes is extracted from SAXS fits, and this is quantitatively compared with distributions from electron microscopy. SAXS data are also obtained from 2D sheet-like protein assemblies, and the proposed method is used to quantify structural features such as the separation distance between protein building blocks and the flexing of the sheet. An open-source implementation of the methodology is provided for use in a broad range of biological systems involving multi-scale scattering analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Elastic and resonance structures of the nucleon from the hadronic tensor in lattice QCD: Implications for neutrino-nucleon scattering and hadron physics

We compute the Euclidean hadronic tensor from charge density operators and extract elastic and resonance structures by employing exponential fits to the four-point function correlator, as well as a Bayesian reconstruction inverse algorithm to obtain the corresponding spectral density for qualitative comparison. We present the determination of the nucleon’s Sachs electric form factor using the hadronic tensor formalism and verify that it is consistent with that from the conventional three-point function calculation. Beyond the elastic peak, we observe a structure located approximately 0.5–0.7 GeV above the nucleon mass in the Bayesian reconstruction. This structure is interpreted as a mixture of the Roper resonance [𝑁⁡(1440)], and states with both positive and negative parities in this mass region, as well as multihadron states. Assuming the observed structure is dominated by 𝐽 𝑃 = 1/2 ± states, we extract the transition electric form factor 𝐺$^*_𝐸$⁡(𝑄 2 ) and the corresponding longitudinal helicity amplitude 𝑆 1/2 ⁡(𝑄 2 ), and compare them with those determined from the CLAS experimental data of nucleon-to-Roper transition. Although fitting to the four-point correlation function or using the inverse algorithm does not resolve individual resonances, it nevertheless enables the determination of total inclusive lepton–nucleon scattering cross sections in appropriate energy bins. This lattice QCD calculation presents the first major step toward studying the inclusive 𝑁 → 𝑋 contributions within the hadronic tensor formalism.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The gravitational lensing imprints of DES Y3 superstructures on the CMB: a matched filtering approach

Low-density cosmic voids gravitationally lens the cosmic microwave background (CMB), leaving a negative imprint on the CMB convergence |$\kappa$|⁠. This effect provides insight into the distribution of matter within voids, and can also be used to study the growth of structure. We measure this lensing imprint by cross-correlating the Planck CMB lensing convergence map with voids identified in the Dark Energy Survey Year 3 (DES Y3) data set, covering approximately 4200 deg|$^2$| of the sky. We use two distinct void-finding algorithms: a 2D void-finder that operates on the projected galaxy density field in thin redshift shells, and a new code, Voxel, which operates on the full 3D map of galaxy positions. We employ an optimal matched filtering method for cross-correlation, using the Marenostrum Institut de Ciències de l’Espai N-body simulation both to establish the template for the matched filter and to calibrate detection significances. Using the DES Y3 photometric luminous red galaxy sample, we measure |$A_\kappa$|⁠, the amplitude of the observed lensing signal relative to the simulation template, obtaining |$A_\kappa = 1.03 \pm 0.22$| (⁠|$4.6\sigma$| significance) for Voxel and |$A_\kappa = 1.02 \pm 0.17$| (⁠|$5.9\sigma$| significance) for 2D voids, both consistent with Lambda cold dark matter expectations. We additionally invert the 2D void-finding process to identify superclusters in the projected density field, for which we measure |$A_\kappa = 0.87 \pm 0.15$| (⁠|$5.9\sigma$| significance). The leading source of noise in our measurements is Planck noise, implying that data from the Atacama Cosmology Telescope, South Pole Telescope and CMB-S4 will increase sensitivity and allow for more precise measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

ED-cPSD: Fast Phase-Size Distribution via Sequential Erosion-Dilation

The Erosion-Dilation continuous Phase-Size Distribution, ED-cPSD, is an application for calculating continuous pore and particle-size distribution from digital reconstructions and/or image-based structural data. It is based on the erosion-dilation continuous phase-size distribution method. A continuous size distribution is a measure of the probability density of finding a particle or pore of a certain size. These distributions are of interest in any field of study involving porous media, including but not limited to electrochemistry, petroleum engineering, geology, and food science. The algorithm behind the software provides a computationally efficient way to calculate phase-size distributions for large domains. For a 3D battery electrode reconstruction with 1.3 x 10 8 voxels, the particle size distribution is derived in under 2 min on a desktop, while also retaining flexibility and computational efficiency for HPC-scale multi-threading. The software can handle structures with over 10 9 voxels. The algorithm is roughly 280 times faster than a previous version on the same task.

Characterization↗

Multitiered computational methodology for extracting three-dimensional rotational diffusion coefficients from x-ray photon correlation spectroscopy data without structural information

X-ray photon correlation spectroscopy (XPCS) is a powerful technique for analyzing particle systems by investigating their dynamics in suspensions across a broad range of temporal and spatial scales. This is done by illuminating samples with coherent x-ray beams and calculating the correlation function of the obtained x-ray scattering images. XPCS is uniquely suited for studying Brownian dynamics, consisting of translational and rotational diffusion. While traditional XPCS image analysis techniques can extract translational diffusion components, they are unable to estimate rotational diffusion coefficients. Here, we introduce a methodology that combines the angular-temporal cross-correlation analysis and a algorithmic framework called Multi-Tiered Estimation for Correlation Spectroscopy in 3D for estimating three-dimensional rotational diffusion coefficients from XPCS images of three-dimensional particle systems. We demonstrate our methodology for extracting rotational diffusion coefficients from XPCS data by applying it to simulated noisy x-ray images of systems of crossing nanotubes and proteins that evolve under translational and rotational Brownian motion for different diffusion rates. Furthermore, our results show that our approach determines rotational diffusion coefficients within a few percent error.

97 MATHEMATICS AND COMPUTING↗

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Two datasets are better than one: method of double moments for 3D reconstruction in cryo-EM

Cryo-electron microscopy is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic projection images of randomly oriented particles. We introduce a new data fusion framework, termed the method of double moments, which reconstructs molecular structures from two instances of the second-order moment of projection images obtained under distinct orientation distributions: one uniform, the other non-uniform and unknown. We prove that these moments generically uniquely determine the underlying structure, up to a global rotation and reflection, and we develop a convex-relaxation-based algorithm that achieves accurate recovery using only second-order statistics. Our results demonstrate the advantage of collecting and modeling multiple datasets under different experimental conditions, illustrating that leveraging dataset diversity can substantially enhance reconstruction quality in computational imaging tasks.

Kam’s method↗

Space‐Time Causal Discovery in Earth System Science: A Local Stencil Learning Approach

Causal discovery tools enable scientists to infer meaningful relationships from observational data, spurring advances in fields as diverse as biology, economics, and climate science. Despite these successes, the application of causal discovery to space-time systems remains immensely challenging due to the high-dimensional nature of the data. For example, in climate sciences, modern observational temperature records over the past few decades regularly measure thousands of locations around the globe. To address these challenges, we introduce Causal Space-Time Stencil Learning (CaStLe), a novel meta-algorithm for discovering causal structures in complex space-time systems. CaStLe leverages regularities in local space-time dependencies to learn governing global dynamics. This local perspective eliminates spurious confounding and drastically reduces sample complexity, making space-time causal discovery practical and effective. For causal discovery, CaStLe flexibly accepts any appropriately adapted time series causal discovery algorithm to recover local causal structures. These advances enable causal discovery of geophysical phenomena that were previously unapproachable, including non-periodic, transient phenomena such as volcanic eruption plumes. Regularities in local space-time dependencies are transformed into informative spatial replicates, which actually improve CaStLe's performance when applied to ever-larger spatial grids. We successfully apply CaStLe to discover the atmospheric dynamics governing the climate response to the 1991 Mount Pinatubo volcanic eruption. We provide validation experiments to demonstrate the effectiveness of CaStLe over existing causal-discovery frameworks on a range of geophysics-inspired benchmarks while identifying the method's limitations and domains where its assumptions may not hold.

Nichol, J. Jake [Univ. of New Mexico, Albuquerque,↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Advancing the Frontiers of Deep Learning for Low-Dose 3D Cone-Beam CT Reconstruction

X-ray computed tomography (CT) is an important noninvasive medical imaging modality for studying the structural details of internal organs. Image reconstruction in CT is an inverse problem of recovering an object's internal structure from the absorption profile of X-ray beams (sinogram) measured using a detector. The classical variational approach for CT reconstruction minimizes an energy functional using an appropriate iterative algorithm. Motivated by the success of deep learning (DL), researchers have begun to leverage training data and enhanced computing capabilities in recent years to produce high-fidelity reconstructed images. Nonetheless, much of the academic research in DL algorithms for CT has focused primarily on the two-dimensional setting (with simplified forward operators and noise model) for proofs-of-concept, and a comprehensive benchmarking of various classical and data-driven CT reconstruction approaches has not beenundertaken. The key objective of our CT reconstruction grand challenge was to promote methodological advancements for both classical and DL-based approaches for clinical CT with a reasonably accurately simulated 3D CT forward operator and noise model. We have utilized the publicly available LIDC-IDRI dataset and simulated sinograms and FDK images corresponding to two dose levels (clinical- and low-dose, constituting two tracks of the challenge) starting from the normal-dose images as the ground truth. In this paper, we summarize the motivation, context, and results of our challenge, and highlight the future research directions in DL for clinical CT.

X-ray tomography↗