Search NASA⌕ Search

Engineering topics

Ambrozio Dias, Philipe

Publications and source records attributed to Ambrozio Dias, Philipe.

OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery

While the pretraining of Foundation Models (FMs) for remote sensing (RS) imagery is on the rise, models remain restricted to a few hundred million parameters. Scaling models to billions of parameters has been shown to yield unprecedented benefits including emergent abilities, but requires data scaling and computing resources typically not available outside industry R&D labs. In this work, we pair high-performance computing resources including Frontier supercomputer, America's first exascale system, and high-resolution optical RS data to pretrain billion-scale FMs. Our study assesses performance of different pretrained variants of vision Transformers across image classification, semantic segmentation and object detection benchmarks, which highlight the importance of data scaling for effective model scaling. Moreover, we discuss construction of a novel TIU pretraining dataset, model initialization, with data and pretrained models intended for public release. By discussing technical challenges and details often lacking in the related literature, this work is intended to offer best practices to the geospatial community toward efficient training and benchmarking of larger FMs.

Ambrozio Dias, Philipe↗

Conditional Experts for Improved Building Damage Assessment Across Satellite Imagery View Angles

Rapid building damage assessment (BDA) is vital in guiding disaster response missions and estimating population distribution across impacted areas. While commercial satellite imagery providers have enabled near-daily monitoring of the Earth, near-realtime assessment of disaster scenarios frequently requires analysis of off-nadir imagery, as satellites are often far from impacted areas for at-nadir post-event imaging to occur Such scenarios are, however, underrepresented in existing BDA datasets and methodologies. With this motivation, we investigate generalization capabilities of current BDA practices across overhead view-angles and strategies for their improvement. Using a labeled dataset of images capturing conflict-related damages, we first train a baseline BDA architecture using imbalanced and balanced datasets with respect to view-angle. Then, we explore conditional convolutions parameterized on image features, image nadir, and their combination as a mechanism for conditioning on view-angles. Experiments demonstrate the limitations of current practice and the potential of conditional mechanisms to increase model robustness to view-angle variations.

Ambrozio Dias, Philipe↗

Introducing SpaceNet 9 - Cross-Modal Satellite Imagery Registration for Natural Disaster Responses

Computer vision algorithms are increasingly leveraged to accelerate geospatial analysis for disaster response and recovery. As the diversity of remote sensing imagery grows with optical, SAR, and other modalities, a perquisite for analytics is cross-modal image registration. There is a high potential to harness computer vision for this pre-processing requirement toward enabling downstream analytics such as heterogeneous change detection, automated feature extraction, and data fusion. Advancement in these areas has the potential to simplify data wrangling tasks and further accelerate disaster response timelines. The SpaceNet 9 challenge (launching in mid-2024) focuses on addressing the cross-modal image registration problem and demonstrating the utility of such modules on earthquake impacted scenarios. This paper describes the motivation for the SpaceNet 9 and provides a first overview of the dataset, the baseline algorithm, and implications for seeking cross-modal image registration in Earth observation. Code is available at https://github.com/SpaceNetChallenge/SpaceNet9.

Hansch, Ronny↗

Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation Models

The design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications.

Arndt, Jacob↗

Pretraining Billion-Scale Geospatial Foundational Models on Frontier

As AI workloads increase in scope, generalization capability becomes challenging for small task-specific models and their demand for large amounts of labeled training samples increases. On the contrary, Foundation Models (FMs) are trained with internet-scale unlabeled data via self-supervised learning and have been shown to adapt to various tasks with minimal fine-tuning. Although large FMs have demonstrated significant impact in natural language processing and computer vision, efforts toward FMs for geospatial applications have been restricted to smaller size models, as pretraining larger models requires very large computing resources equipped with state-of-the-art hardware accelerators. Current satellite constellations collect 100+TBs of data a day, resulting in images that are billions of pixels and multimodal in nature. Such geospatial data poses unique challenges opening up new opportunities to develop FMs. We investigate billion scale FMs and HPC training profiles for geospatial applications by pretraining on publicly available data. We studied from end-to-end the performance and impact in the solution by scaling the model size. Our larger 3B parameter size model achieves up to 30% improvement in top1 scene classification accuracy when comparing a 100M parameter model. Moreover, we detail performance experiments on the Frontier supercomputer, America's first exascale system, where we study different model and data parallel approaches using PyTorch's Fully Sharded Data Parallel library. Specifically, we study variants of the Vision Transformer architecture (ViT), conducting performance analysis for ViT models with size up to 15B parameters. By discussing throughput and performance bottlenecks under different parallelism configurations, we offer insights on how to leverage such leadership-class HPC resources when developing large models for geospatial imagery applications.

Tsaris, Aristeidis (aris)↗

Towards Geospatial Knowledge Graph Infused Neuro-Symbolic AI for Remote Sensing Scene Understanding

Deep learning has proven its effectiveness in numerous tasks for remote sensing scene understanding. However there is an increasing interest to explore fusion of domain-specific background information to the deep neural network to further improve its performance. Remote sensing researchers are also working towards developing models that generalize and adapt to multiple applications. Generalization challenges coupled with the scarcity of large corpora of high-quality noise-free labelled data, have together fueled an interest for leveraging background information. Knowledge graphs serve as excellent choice to represent domain-specific information in a structured, standardized and extensible manner. Integrating symbolic knowledge representations in the form of Knowledge Graph Embedding (KGE) to perform neuro-symbolic reasoning is an emerging research direction promising significant impacts. This vision paper seeks to position ideas and provoke early thoughts toward advancing neuro-symbolic artificial intelligence in the context of geospatial challenges. Specifically, it conceptualizes and elaborates on an architecture for infusing geospatial knowledge from knowledge graph in a deep neural network pipeline. As guiding case studies - land-use land-cover classification, object detection and instance segmentation can benefit from infusing spatio-contextual information with remote sensing imagery. The discussion further reflects on and articulates the challenges and explainable AI opportunities anticipated when scaling and maintaining large-scale geospatial knowledge graphs.

Potnis, Abhishek↗

Scaling Automatic Vector Data Alignment to Satellite Imagery

Given the tremendous volume of accessible Earth Observation (EO) data, there is a need to develop scalable Geospatial Artificial Intelligence (GeoAI) solutions for time-sensitive applications. Scalability in this context refers to rapidly processing large-scale EO data using high performance computing resources. Accurate mapping of the built environment from remote sensing (RS) imagery has been one of the crucial components in GeoAI workflows for a wide spectrum of humanitarian applications. Derived vector data of built environment is often leveraged for disaster preparedness and response activities. However, factors such as differences in ortho-rectification, atmospheric conditions and human error, results in spatial misalignment between vector data and the timely available RS imagery. Model training for downstream tasks such as object detection, change analysis, etc., is negatively impacted due to such spatial misalignment. Although there has been progress towards automatic alignment of vector data, the lack of scalability remains an open research challenge. This paper proposes to leverage parallel computing to optimize an automatic vector data alignment workflow. It further employs CPU-level multi-core parallelism for improving the performance of the workflow for scalable built environment mapping. We report observations and discuss findings from the preliminary experiments performed on the Summit Supercomputer.

Potnis, Abhishek↗

An Agenda for Multimodal Foundation Models for Earth Observation

Archives of remote sensing (RS) data are increasing swiftly as new sensing modalities with enhanced spatiotemporal resolution become operational. While promising new breakthroughs, the sheer volume of RS archives stretches the limits of human analysts and existing AI tools, as most models are: i) limited to single data modalities; ii) task-specific; iii) heavily reliant on labeled data. The emerging Foundation Models (FMs) have the potential to address these limitations. Trained on vast unlabeled datasets through self-supervised learning, FMs enable generic feature extraction that facilitate specialization to a wide variety of downstream tasks. This paper describes a vision towards an FM for multimodal Earth Observation data (FM4EO), discussing key building blocks and open challenges. We put particular emphasis on multimodal reasoning, a topic underexplored in EO. Our ultimate goal is a practical path toward FM4EO with capacity to unlock breakthroughs in few-shot learning scenarios, multimodal geographic knowledge integration, synthesis, and hypothesis generation.

Ambrozio Dias, Philipe↗

TOWARDS RAPID RESPONSE UPDATES OF POPULATIONS AT RISK

Understanding population at risks has been a focus of the LandScan program through its development of population estimates. With advancements in computer vision, deep learning technologies and access to High Performance Computing (HPC) and high resolution imagery, population estimates are now modeled at the building level. However, when those patterns are disrupted, rapid updates to population distribution estimates are needed to support humanitarian aid and response. Oak Ridge National Laboratory (ORNL) recently adapted an existing deep learning building footprint extraction model in development of a scalable approach to Building Damage Assessments (BDA). This new opportunity opens the possibility of automating BDA to support rapid population distribution estimate updates for geographic areas involved in geopolitical conflicts or natural events for humanitarian aid and response or where to focus recovery efforts. In addition, incorporate social surveys to further model human behavior under conflict or other scenarios that disrupt normal patterns of life.

Urban, Marie↗

Semi-automated Design of Artificial Intelligence Earth Science Models

Prediction and observation of water cycles involve not only patterns isolated in space and time, but rather modeling complex spatio-temporal relationships across multiple sources of data and domains. For instance, Evapotranspiration (ET) and Leaf Area Indexes (LAI) are two critical components in DOE’s Energy Exascale Earth System Model (E3SM). Accurate assessments of ET and LAI are critical for understanding hydrological processes, deforestation, crop yield, and irrigation impacts. However, current ET estimates for global simulations are available at very coarse spatial resolution. They are usually derived from satellite data based on broad plant functional types (PFT), which fail to capture the fine-scale variations due to change in vegetation type across the globe. Within this context and in light of the data-model integration challenges highlighted in the EESSD Strategic Plan, the new era of AI model development for geosciences calls for data-driven methods that provide domain scientists with estimations of parameters such as PFT and LAI in an efficient, interpretable, and easy-to-operate manner.

Ambrozio Dias, Philipe↗

Embedding Ethics and Trustworthiness for Sustainable AI in Earth Sciences: Where Do We Begin?

As in many other research domains, Artificial Intelligence (AI) techniques have been increasing their footprint in Earth Sciences to extract meaningful information from the large amount of high-detailed data available from multiple sensor modalities. While on the one hand the existing success cases endorse the great potential of AI to help address open challenges in ES, on the other hand on-going discussions and established lessons from studies on the sustainability, ethics and trustworthiness of AI must be taken into consideration if the community is to ensure that its research efforts move into directions that effectively benefit the society and the environment. In this paper, we discuss insights gathered from a brief literature review on the subtopics of AI Ethics, Sustainable AI, AI Trustworthiness and AI for Earth Sciences in an attempt to identify some of the promising directions and key needs to successfully bring these concepts together.

Ambrozio Dias, Philipe↗

Distributed Training for High Resolution Images: A Domain and Spatial Decomposition Approach

In this work we developed two Pytorch libraries using the PyTorch RPC interface for distributed deep learning approaches on high resolution images. The spatial decomposition library allows for distributed training on very large images, which otherwise wouldn’t be possible on a single GPU. The domain parallelism library allows for distributed training across multiple domain unlabeled data, by leveraging the domain separation architecture. Both of those libraries where tested on the Summit supercomputer at Oak Ridge National Laboratory at a moderate scale.

Tsaris, Aristeidis (aris)↗