Search NASA⌕ Search

SEARCH · Search NASA

Results for “scene modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Dark Energy Survey Supernova Program: Light Curves and 5 Yr Data Release

We present griz photometric light curves for the full 5 yr of the Dark Energy Survey Supernova (DES-SN) program, obtained with both forced point-spread function photometry on difference images (DiffImg) performed during survey operations, and scene modelling photometry (SMP) on search images processed after the survey. This release contains 31,636 DiffImg and 19,706 high-quality SMP light curves, the latter of which contain 1635 photometrically classified SNe that pass cosmology quality cuts. This sample spans the largest redshift (z) range ever covered by a single SN survey (0.1 < z < 1.13) and is the largest single sample from a single instrument of SNe ever used for cosmological constraints. We describe in detail the improvements made to obtain the final DES-SN photometry and provide a comparison to what was used in the 3 yr DES-SN spectroscopically confirmed Type Ia SN sample. We also include a comparative analysis of the performance of the SMP photometry with respect to the real-time DiffImg forced photometry and find that SMP photometry is more precise, more accurate, and less sensitive to the host-galaxy surface brightness anomaly. The public release of the light curves and ancillary data can be found at github.com/des-science/DES-SN5YR and doi:10.5281/zenodo.12720777.

79 ASTRONOMY AND ASTROPHYSICS↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Application of automated iterative target detection for standoff hyperspectral imaging

The utility of hyperspectral imaging (HSI) has been well established for a wide array of applications but has generated a need for automated screening of high volumes of large HSI cubes. We report two important automated algorithms for more efficient standoff processing: atmospheric correction and target detection. The atmospheric correction method is based on a fast asymmetric least squares approach that is applied on a pixel-by-pixel basis. Here, the correction can be applied to entire images without manually identifying regions of interest and utilizes only in-scene information, no ancillary modeling of the atmosphere is required. An iterative target detection approach is also introduced which demonstrates faster speeds relative to moving window approaches. The target detection algorithm classifies each pixel as true target detections, near target detections, clutter, and no-calls. The algorithms were tested on forty images of twenty-two solid mineral targets placed at a 14-meter standoff distance allowing general observations on expected detection performance for a variety of minerals. In addition to identifying anomalous pixels, the inclusion of “no-calls” reduced the number of false detections significantly.

47 OTHER INSTRUMENTATION↗

Combining geometric-optical and spectral invariants theories for modeling canopy fluorescence anisotropy

The spectral invariants theory ( p -theory) has received much attention in the field of quantitative remote sensing over the past few decades and has been adopted for modeling of canopy solar-induced chlorophyll fluorescence (SIF). However, the spectral invariant properties (SIP) in simple analytical formulae have not been applied for modeling canopy fluorescence anisotropy primarily because they are parameterized in terms of leaf total scattering, which precludes the differentiation between forward and backward leaf SIF emissions. In this study, we have developed the canopy-SIP SIF model by combining geometric-optical (GO) theory to account for asymmetric leaf SIF forward and backward emissions at the first-order scattering and by modeling multiple scattering based on the p-theory, thus avoiding the dependence on radiative transfer models. The applicability of the model simulations especially over 3D heterogeneous canopies was improved by incorporating canopy structure through multi-angular clumping index, and by modeling single scattering from the four components of the scene in view according to the GO approach. The results show good consistency with both the state-of-the-art SIF models and multi-angular field SIF observations over grass and chickpea canopies. Further, the coefficient of determination (R²) between the simulated SIF and field measurements was 0.75 (red) and 0.74 (far-red) for chickpea, and 0.65 (both red and far-red) for grass. The average relative error was approximately 3% for 1D homogeneous scenes when comparing the canopy-SIP SIF model simulations to the SCOPE model simulations, and around 4% for the 3D heterogeneous scene when comparing to the LESS model simulations. The results indicate that the proposed approach for separating asymmetric leaf SIF emissions is a robust way to keep a balance between satisfactory simulation accuracy and efficiency. Model simulations suggest that neglecting the leaf SIF asymmetry can lead to an underestimation of canopy red SIF by 6.3% to 42.6% for various leaf biochemical and canopy structural parameters. This study presents a simple but efficient analytical approach for canopy fluorescence modeling, with potential for large-scale canopy fluorescence simulations.

3D heterogeneous structure↗

A Real2Sim Digital Twin Pipeline for Photorealistic Robot Simulation: Evaluating VLA Policy Deployment on a Bimanual Mobile Robot

Digital twins that are automatically constructed from robot sensor data offer a promising pathway for scalable Real2Sim and Sim2Real transfer. However, it remains an open question whether photorealistic reconstruction alone is sufficient to support reliable deployment of vision-language-action (VLA) policies. We present a generative-AI-assisted Real2Sim pipeline that generates simulation-ready digital twins from real-world RGB observations with minimal manual intervention. The pipeline uses prompted segmentation to isolate scene components and a generative 3D model to directly produce simulation assets, eliminating the need for traditional multi-view reconstruction or manual 3D modeling.\r\nTo evaluate simulation fidelity, we deploy and compare policies from two VLA models in both the real robot and the reconstructed\r\nsimulation under identical tasks and initial conditions. We compare joint-level action trajectories and analyze how divergence evolves over time in closed-loop execution. Although the reconstructed environments are visually accurate, we observe increasing trajectory divergence during closedloop operation. These results indicate that photorealistic reconstruction alone is insufficient to preserve closed-loop control behavior\r\nin VLA policies, particularly in contact-rich manipulation settings where small perceptual errors compound over time.

97 MATHEMATICS AND COMPUTING↗

Factorized visual representations in the primate visual system and deep neural networks

Object classification has been proposed as a principal objective of the primate ventral visual stream and has been used as an optimization target for deep neural network models (DNNs) of the visual system. However, visual brain areas represent many different types of information, and optimizing for classification of object identity alone does not constrain how other information may be encoded in visual representations. Information about different scene parameters may be discarded altogether (‘invariance’), represented in non-interfering subspaces of population activity (‘factorization’) or encoded in an entangled fashion. In this work, we provide evidence that factorization is a normative principle of biological visual representations. In the monkey ventral visual hierarchy, we found that factorization of object pose and background information from object identity increased in higher-level regions and strongly contributed to improving object identity decoding performance. We then conducted a large-scale analysis of factorization of individual scene parameters – lighting, background, camera viewpoint, and object pose – in a diverse library of DNN models of the visual system. Models which best matched neural, fMRI, and behavioral data from both monkeys and humans across 12 datasets tended to be those which factorized scene parameters most strongly. Notably, invariance to these parameters was not as consistently associated with matches to neural and behavioral data, suggesting that maintaining non-class information in factorized activity subspaces is often preferred to dropping it altogether. Thus, we propose that factorization of visual scene information is a widely used strategy in brains and DNN models thereof.

59 BASIC BIOLOGICAL SCIENCES↗

Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images

Machine learning models are known to be vulnerable to adversarial attacks, but prior works have mostly focused on single-modalities. With the rise of large multi-modal models (LMMs) like CLIP, which combine vision and language capabilities, new vulnerabilities have emerged. However, these multimodal targeted attacks aim to completely change the model's output to what the adversary wants. In many realistic scenarios, an adversary might seek to make only subtle modifications to the output, so that the changes go unnoticed by downstream models or even by humans. We introduce Hiding-in-Plain-Sight (HiPS) attacks, a novel class of adversarial attacks that subtly modifies model predictions by selectively concealing target object(s), as if the target object was absent from the scene. We propose two HiPS attack variants, HiPS-cls and HiPS-cap, and demonstrate their effectiveness in transferring to downstream image captioning models, such as CLIP-Cap, for targeted object removal from image captions.

Daw, Arka [ORNL] (ORCID:0009000633191271)↗

Sensitivity to low-mass WIMPs with an improved liquid argon ionization response model within the DarkSide program

Dark matter detection experiments using liquid argon rely on a precise characterization of the ionization response to nuclear recoils, especially in the keV energy range relevant for light dark matter interactions. In this work, we present a comprehensive analysis that combines new measurements from the ReD setup, part of the DarkSide experimental program, with calibration data from DarkSide-50, as well as results from the ARIS and SCENE experiments. These combined datasets enable improved constraints on atomic screening effects in the modeling of the ionization response of liquid argon to nuclear recoils. The analysis is performed within the Thomas-Imel recombination framework adopted in previous DarkSide studies, and is here further constrained by the inclusion of ReD data, which allow the screening function to be determined from calibration measurements. By including the updated ionization model into the DarkSide-50 analysis framework, we obtain stronger exclusion limits on low-mass weakly interacting massive particle (WIMP) interactions, setting new world-leading constraints in the 1 – 3 GeV / c 2 WIMP mass range. Finally, we recast the sensitivity projections for the upcoming DarkSide-20k detector, demonstrating a significantly enhanced discovery potential for low-mass dark matter candidates.

Acerbi, F. [Fond. Bruno Kessler, Trento]↗

Interferometric focal planes

We propose arrays of integrated interferometers to characterize the mutual intensity on focal planes. While focal coherence measurement does not increase the aperture-limited spatial bandpass, it can increase Shannon information capacity relative to irradiance measurement by increasing the number of degrees of freedom per Nyquist sample. We describe a sampling model for interferometric measurement using arrays of 2-port Mach-Zehnder interferometers and show within this model that interferometric focal planes enable more accurate estimation of prototypical scene parameters.

Brady, David J. (ORCID:0000000156552478)↗

Deep Learning Scene Classification Experiments in Automatic Detection of Slums on Planetscope Imagery

Population growth is increasingly happening in slum settlements of the large urban centers in the Global South. The term "slum" encompasses a wide range of communities, located mostly in underserved areas, and often exhibiting distinct structural and functional informalities with a relatively high concentration of marginalized populations. To address the issues confronting slums for effective planning and development, including the realistic estimation of the resident population, identifying them accurately is fundamental. Given the disagreements over a universal definition, diverse characteristic features, and socio-political limitations, global detection of slums is a veritable challenge. In this paper, we present experiments in slum detection using a scene classification algorithm and 3-meter spatial resolution satellite imagery. We train and evaluate the model for slum detection in Mumbai, India for the year 2023 and test the temporal generalization of the trained model on Mumbai in 2020 and 2018. In addition, we explore the pathways toward geographic generalization to Kolkata and Delhi (India). We discuss several limitations in the workflow and model, situate our findings in the existing literature, and suggest improvements and alternatives. With this, we establish baseline methods and experiments as a first step towards developing an image-based global slum detection framework and algorithm. This work adds to the community discussion on methods, data challenges, and open questions related to the detection of slums globally. With this research, we hope to improve our understanding of human settlements, especially in critical areas, improve population estimates, and help measure progress towards the sustainable development goals.

Arndt, Jacob↗

Selective Amnesia using Contrastive Subnet Erasure for Class Level Unlearning in Vision Models

We study concept-level forgetting in pretrained vision models: removing an entire semantic category so the system no longer recognizes that object in unseen images and contexts, rather than merely forgetting specific training examples. Prior work either applies blunt global projections or fine-tunes parameters, which can introduce collateral damage to unrelated features, add compute, and become unstable as forgetting strength increases. We introduce Contrastive Subnet Erasure (CSE), a training-free, encoder-centric edit that targets a compact set of channels most responsible for the class and attenuates them in a calibrated manner. The modification is algebraically folded into the subsequent layer, yielding no inference-time overhead and leaving task heads unchanged. To evaluate whether forgetting generalizes beyond the data used to specify the class, we introduce a cross dataset protocol in which the class is defined on a source dataset and performance is measured on a disjoint target dataset drawn from a different distribution with no shared images. This setup tests whether the model still fails to recognize the object when it looks different or appears in new scenes, and it helps avoid overfitting to patterns in the source dataset. Across CIFAR 10, CIFAR 100, and ImageNet under this protocol, CSE achieves stronger forgetting of the target class while better preserving non target utility than existing baselines in both single class and multi class settings. Overall, CSE provides a simple, stable, and deployment-ready mechanism for class-level unlearning in vision.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Data-Driven Invertible Neural Surrogates of Atmospheric Transmission

We present Data-Driven Invertible Neural Surrogates of Atmospheric transmission, or DINSAT. DINSAT is a novel framework for inferring an atmospheric transmission profile from a spectral scene. This framework leverages a lightweight, physics-based simulator that is automatically tuned -- by virtue of autodifferentiation and differentiable programming -- to construct a surrogate atmospheric profile to model the observed data. The framework has utility in (i) performing atmospheric correction, (ii) recasting spectral data between various modalities (e.g. radiance and reflectance at the surface and at the sensor), and (iii) inferring atmospheric transmission profiles, such as absorbing bands and their relative magnitudes. We demonstrate the utility of these methods by performing a canonical atmospheric correction task for the purposes of further analysis - in this case, target detection within a scene.

Koch, James V.↗

Orbital-Radar v1.0.0: a tool to transform suborbital radar observations to synthetic EarthCARE cloud radar data

The Earth Cloud, Aerosol and Radiation Explorer (EarthCARE) satellite developed by the European Space Agency (ESA) and the Japan Aerospace Exploration Agency (JAXA) launched in May 2024 carries a novel 94 GHz cloud profiling radar (CPR) with Doppler capability. This work describes the open-source instrument simulator Orbital-Radar, which transforms high-resolution radar data from field observations or forward simulations of numerical models to CPR primary measurements and uncertainties. The transformation accounts for sampling geometry and surface effects. We demonstrate Orbital-Radar's ability to provide realistic CPR views of typical cloud and precipitation scenes. The presented case studies show small-scale convection, marine stratus clouds, and Arctic mixed-phase cloud cases. These results provide valuable insights into the capabilities and challenges of the EarthCARE CPR mission and its advantages over the CloudSat CPR. Finally, Orbital-Radar allows for evaluating kilometre-scale numerical weather prediction models with EarthCARE CPR observations. So, Orbital-Radar can generate calibration and validation (Cal/Val) data sets already pre-launch. Nevertheless, an evaluation of synthetic CPR output data to accurate EarthCARE CPR data is missing.

54 ENVIRONMENTAL SCIENCES↗

Image Deconvolution and Point-spread Function Reconstruction with STARRED: A Wavelet-based Two-channel Method Optimized for Light-curve Extraction

We present starred, a point-spread function (PSF) reconstruction, two-channel deconvolution, and light-curve extraction method designed for high-precision photometric measurements in imaging time series. An improved resolution of the data is targeted rather than an infinite one, thereby minimizing deconvolution artifacts. In addition, starred performs a joint deconvolution of all available data, accounting for epoch-to-epoch variations of the PSF and decomposing the resulting deconvolved image into a point source and an extended source channel. The output is a high-signal-to-noise-ratio, high-resolution frame combining all data and the photometry of all point sources in the field of view as a function of time. Of note, starred also provides exquisite PSF models for each data frame. We showcase three applications of starred in the context of the imminent LSST survey and of JWST imaging: (i) the extraction of supernovae light curves and the scene representation of their host galaxy; (ii) the extraction of lensed quasar light curves for time-delay cosmography; and (iii) the measurement of the spectral energy distribution of globular clusters in the "Sparkler," a galaxy at redshift z = 1.378 strongly lensed by the galaxy cluster SMACS J0723.3-7327. starred is implemented in jax, leveraging automatic differentiation and graphics processing unit acceleration. This enables the rapid processing of large time-domain data sets, positioning the method as a powerful tool for extracting light curves from the multitude of lensed or unlensed variable and transient objects in the Rubin-LSST data, even when blended with intervening objects.

79 ASTRONOMY AND ASTROPHYSICS↗

Monitoring installation of partially occluded subassemblies in modular construction factories using BIM, ray tracing, and computer vision

Modular and offsite construction methods are being increasingly adopted due to the advantages they offer in terms of project completion time, quality, and energy-efficiency. Despite these advantages, the current state of monitoring systems in modular construction factories highly relies on labor-intensive, subjective, and error-prone observational methods. A large body of research has aimed to automate the monitoring process using an array of sensors, such as IMUs and RFIDs, during the past two decades. Recently, computer vision-based methods have gained increasing interest as a non-intrusive technology to monitor the process inside modular construction factories. However, partial occlusion challenges have impeded their practical application on a large scale. This challenge is specifically important for monitoring the installation of subassemblies since they can obstruct the view of the monitoring camera, especially those that enable long-term monitoring like closed-circuit television (CCTV) fixed-view surveillance cameras. Here, this paper aims to address this challenge by proposing a novel computer vision-based method to monitor the installation of new subassemblies inside modular factories in highly occluded scenes. The proposed methodology identifies the subassemblies in the CCTV video footage using computer vision, analyzes the occlusions using BIM and ray casting techniques, and estimates the progress of assembly by comparing the BIM model with the detected subassemblies in the video. The proposed methodology was successfully validated on surveillance videos captured from a volumetric modular construction factory in the U.S., achieving 93% accuracy in identifying the installation of subassemblies. The results from this research show that the integration of BIM and computer vision is a promising method for monitoring the installation processes inside modular factories under severe occlusion.

97 MATHEMATICS AND COMPUTING↗

Surrogate Distributed Radiological Sources—Part III: Quantitative Distributed Source Reconstructions

In this third part of a multi-paper series, we present quantitative image reconstruction results from aerial measurements of eight different surrogate distributed gamma-ray sources on flat terrain. Here, we show that our quantitative imaging methods can accurately reconstruct the expected shapes, and, after appropriate calibration, the absolute activity of the distributed sources. We conduct several studies of imaging performance versus various measurement and reconstruction parameters, including detector altitude and raster pass spacing, data and modeling fidelity, and regularization type and strength. The imaging quality performance is quantified using various quantitative image quality metrics. Our results confirm the utility of point source arrays as surrogates for truly distributed radiological sources, and advance the quantitative capabilities of Scene Data Fusion gamma-ray imaging methods.

Airborne survey↗

Role of depth in optical diffractive neural networks

Free-space all-optical diffractive neural networks have emerged as promising systems for neuromorphic scene classification. Understanding the fundamental properties of these systems is important to establish their ultimate performance. Here we consider the case of diffraction by subwavelength apertures and study the behavior of the system as a function of the number of diffractive layers by employing a co-design modeling approach. We show that adding depth allows the system to achieve high classification accuracies with a reduced number of diffractive features compared to a single layer, but that it does not allow the system to surpass the performance of an optimized single layer. The improvement from depth is found to be limited to the first few layers. These properties originate from the constraints imposed by the physics of light, in particular the weakening electric field with distance from the aperture.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Competition response of cloud supersaturation explains diminished Twomey effect for smoky aerosol in the tropical Atlantic

The Twomey effect brightens clouds by increasing aerosol concentrations, which activates more droplets and decreases cloud supersaturation in response to more competition for water vapor. To quantify this competition response, we used marine low cloud observations in clean and smoky conditions at Ascension Island in the tropical South Atlantic during the Layered Aerosol Smoke Interactions with Cloud (LASIC) campaign. These observations show similar increases in droplet number for increased accumulation-mode particles from surface-based and satellite cloud retrievals, demonstrating the importance of below-cloud aerosol measurements for retrieving aerosol–cloud interactions (ACI) in clean and smoky aerosol conditions. Four methods for estimating cloud supersaturation from aerosol–cloud measurements were compared, with cloud scene-based and parcel-based methods showing sufficient variability for a strong dependence on both aerosol accumulation number concentration and cloud-base updraft velocities. Decomposing aerosol-related changes in cloud albedo and optical depth shows the calculated competition response accounts for dampening the activation response by 12 to 35%, explaining the diminished Twomey effect at high aerosol concentrations observed for smoky conditions at LASIC and previously around the world. This result was consistent for independent supersaturation retrievals by cloud scene-based droplet number and cloud condensation nuclei and parcel-based multimode size-resolving Lagrangian methods. Translating aerosol effects to local radiative forcing with clean conditions as a proxy for preindustrial and smoky conditions for present-day showed that the competition response reduces cooling from the Twomey radiative forcing by 12 to 35%, providing an essential process-specific constraint for improving the representation of aerosol competition in climate model simulation of indirect aerosol forcing.

54 ENVIRONMENTAL SCIENCES↗