Search NASASearch

SEARCH · Search NASA

Results for “Deep generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Checkerboard Patterns in E3SMv2 and E3SM-MMFv2

An unphysical checkerboard pattern is identified in E3SMv2 and E3SM-MMF that is detectable across a wide range of timescales, from instantaneous snapshots to multi-year averages. A detection method is developed to quantify characteristics of the checkerboard signal by cataloguing all possible configurations of the eight adjacent neighbors for each cell on the model's cubed sphere grid using daily mean data. The checkerboard pattern is only found in cloud-related quantities, such as precipitation and liquid water path. Instances of pure and partial checkerboard are found to occur more often in E3SMv2 and E3SM-MMF when compared to satellite data regridded to the model grid. Continuous periods of partial checkerboard state are found to be more persistent in both models compared to satellite data, with E3SM-MMF exhibiting more persistence than E3SMv2. The checkerboard signal in E3SMv2 is found to be a direct consequence of the recently added deep convective trigger condition based on dynamically generated CAPE (DCAPE). In E3SM-MMF the checkerboard signal is found to be associated with the “trapping” of cloud-scale fluctuations within the embedded cloud-resolving model. Solutions to remedy this issue are discussed.

E3SMv2

AI-Based Analytics and Energy Modeling Framework for Characterizing Urban Energy Systems

Developing location-specific district energy models is essential for understanding energy patterns and supporting efficient management and planning decisions. However, accurately characterizing these models remains challenging due to gaps in building characteristics and labor-intensive traditional modeling workflows. To address these challenges, we develop an AI-based framework that integrates top-down and bottom-up building energy data to automate urban energy model characterization. The framework trains multimodal deep learning models using heterogeneous ResStockTM datasets to infer missing building characteristics from varying levels of known information and generate simulation-ready inputs for district-scale energy modeling. It also employs a conditioning-based injection approach to generate ”what-if” scenarios, enabling users to explore retrofit, efficiency, and technology-upgrade pathways. Integrated within URBANoptTM, a bottom-up district energy modeling platform for simulating co-located buildings, the framework infers detailed building-level inputs required for bottom-up simulations. Both localized and generalized AI models are developed to learn relationships across categorical, numerical, and time-series data, enabling reconstruction of missing attributes and generation of targeted upgrade scenarios. We demonstrate this methodology on a residential neighborhood in Baltimore, MD, assessing internal consistency against ResStock reference data and URBANopt simulation, and comparing selected attributes against real-world building characteristics. Results show strong overall predictive accuracy in data completion and scenario generation, with localized and generalized models offering complementary trade-offs between precision and scalability. Overall, our automated framework streamlines energy modeling and provides a reliable framework for urban building energy characterization.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS

Further Constraints and Uncertainties on the Deep Seismic Structure of the Moon

The Apollo Passive Seismic Experiment (APSE) consisted of four 3-component seismometers deployed between 1969 and 1972, that continuously recorded lunar ground motion until late 1977. The APSE data provide a unique opportunity for investigating the interior of a planet other than Earth, generating the most direct constraints on the elastic structure, and hence the thermal and compositional evolution of the Moon. Owing to the lack of far side moonquakes, past seismic models of the lunar interior were unable to constrain the lowermost 500 km of the interior. Recently, array methodologies aimed at detecting deep lunar seismic reflections found evidence for a lunar core, providing an elastic model of the deepest lunar interior consistent with geodetic parameters. Here we study the uncertainties in these models associated with the double array stacking of deep moonquakes for imaging deep reflectors in the Moon. We investigate the dependency of the array stacking results on a suite of parameters, including amplitude normalization assumptions, polarization filters, assumed velocity structure, and seismic phases that interfere with our desired target phases. These efforts are facilitated by the generation of synthetic seismograms at high frequencies (approx. 1Hz), allowing us to directly study the trade-offs between different parameters. We also investigate expected amplitudes of deep reflections relative to direct P and S arrivals, including predictions from arbitrarily oriented focal mechanisms in our synthetics. Results from separate versus combined station stacking help to establish the robustness of stacks. Synthetics for every path geometry of data were processed identically to that done with data. Different experiments were aimed at examining various processing assumptions, such as adding random noise to synthetics and mixing 3 components to some degree. The principal stacked energy peaks put forth in recent work persist, but their amplitude (which maps into reflector impedance contrast) and timing (which maps into reflector depth) depend on factors that are not well constrained -- most notably, the velocity structure of the overlying lunar interior. Thus, while evidence for the lunar core remains strong, the depths of imaged reflectors have associated uncertainties that will require new seismic data and observations to constrain. These results strongly advocate further investigations on the Moon to better resolve the interior (e.g., Selene missions), for the Moon apparently has a rich history of construction and evolution that is inextricably tied to that of Earth.

Lin, Pei-Ying Patty

From clutter to clarity: Emergent neural operators via questionnaire metrics

Real-world datasets in chemical engineering and bioengineering processes—such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials—can often be unlabeled or disorganized, rendering the training of existing supervised learning models ineffective at learning the underlying dynamics. To salvage these datasets for decision-making, we first seek to obtain clarity from the cluttered data. Here, we present a framework for developing “structural” generative models, discovering emergent equations, and constructing efficient emulators from scrambled datasets by integrating unsupervised organizational learning techniques (Questionnaires) with advanced deep learning architectures (Deep Hidden Physics Models and Deep Operator Networks). Our approach is demonstrated on two illustrative model systems: (a) a 1D advection–diffusion partial differential equation representing a winding underground pipe and (b) an ensemble of Stuart–Landau oscillators, an agent-based system of coupled ordinary differential equations. In both cases, we successfully reconstruct meaningful spatial, temporal, and parameter embeddings from scrambled data, enabling good predictions of system dynamics. As a result, we highlight the framework’s potential for broader applications, enabling data-driven system identification in fields with inherently disorganized or hidden parameter spaces.

42 ENGINEERING

STFM: Accurate Spatio-Temporal Fusion Model for Weather Forecasting

Meteorological prediction is crucial for various sectors, including agriculture, navigation, daily life, disaster prevention, and scientific research. However, traditional numerical weather prediction (NWP) models are constrained by their high computational resource requirements, while the accuracy of deep learning models remains suboptimal. In response to these challenges, we propose a novel deep learning-based model, the Spatiotemporal Fusion Model (STFM), designed to enhance the accuracy of meteorological predictions. Our model leverages Fifth-Generation ECMWF Reanalysis (ERA5) data and introduces two key components: a spatiotemporal encoder module and a spatiotemporal fusion module. The spatiotemporal encoder integrates the strengths of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), effectively capturing both spatial and temporal dependencies. Meanwhile, the spatiotemporal fusion module employs a dual attention mechanism, decomposing spatial attention into global static attention and channel dynamic attention. This approach ensures comprehensive extraction of spatial features from meteorological data. The combination of these modules significantly improves prediction performance. Experimental results demonstrate that STFM excels in extracting spatiotemporal features from reanalysis data, yielding predictions that closely align with observed values. In comparative studies, STFM outperformed other models, achieving a 7% improvement in ground and high-altitude temperature predictions, a 5% enhancement in the prediction of the u/v components of 10 m wind speed, and an increase in the accuracy of potential height and relative humidity predictions by 3% and 1%, respectively. This enhanced performance highlights STFM’s potential to advance the accuracy and reliability of meteorological forecasting.

54 ENVIRONMENTAL SCIENCES

FORECASTOR – II. Simulating galaxy surveys with the Cosmological Advanced Survey Telescope for Optical and UV Research

The Cosmological Advanced Survey Telescope for Optical and UV Research (CASTOR) is a planned flagship space telescope, covering the blue-optical and UV part of the spectrum. Here, we introduce the CASTOR image simulator, a python GalSim package-based script capable of generating mock CASTOR images from an input catalogue. We generate example images from the CASTOR Wide, Deep, and Ultra-Deep surveys using simulated lightcones from the Santa Cruz semi-analytic model. We make predictions for the performance of these surveys by comparing galaxies that are extracted from each image using Source Extractor to the input catalogue. We find that the Wide, Deep, and Ultra-Deep surveys will be 75 per cent complete for point sources down to $\sim 27$, 29, and 30 mag, respectively, in the UV, u, and g filters, with the UV-split and u-split filters reaching a shallower depth. With a large area of $\sim 2200$ deg$^2$, the Wide survey will detect hundreds of millions of galaxies out to $z\sim 4$, mostly with $M_\ast \gtrsim 10^{9}\,{\rm M}_{\odot }$. The Ultra-Deep survey will probe to $z\sim 5$, detecting galaxies with $M_\ast \gtrsim 10^{7}{\rm M}_{\odot }$. These galaxy samples will enable precision measurements of the distribution of star formation in the cosmic web, connecting the growth of stellar mass to the assembly of dark matter haloes over two thirds of the history of the Universe, and other core goals of CASTOR’s legacy surveys. These image simulations and the tools developed to generate them will be a vital planning tool to estimate CASTOR’s performance and iterate the telescope and survey designs prior to launch.

79 ASTRONOMY AND ASTROPHYSICS

Development of a Phasor Diagram Creator to Visualize the Piston and Displacer Forces in an Advanced Stirling Convertor

The steady state, nearly sinusoidal behavior of the components in a Free Piston Stirling Engine allows for visualization of the forces in the system using phasor diagrams. Based on Newton's second law, F=ma, any phasor diagrams modeling a given component in a system should close if all of the acting forces have been considered. Since the Advanced Stirling Radioisotope Generator (ASRG), currently being developed for future NASA deep space missions, is made up of such nearly sinusoidally oscillating components, its phasor diagrams would also be expected to close. A graphical user interface (GUI) has been written in MATLAB by taking user input data, passing it to Sage, a 1-D thermodynamic modeling program used to model the Stirling convertor, running Sage and then automatically plotting the phasor diagrams. Using this software tool, the effect of varying different Sage inputs on the phasor diagrams was determined. The parameters varied were piston amplitude, hot end temperature, cold end temperature, operating frequency, and displacer spring constant. By using these phasor diagrams, better insight can be gained as to why the convertor operates the way that it does.

Saha, Dipanjan

Unlocking the Mysteries of the Moon’s Shadowed Regions

The Moon poles host large quantities of water-ice deposits in the permanently shadowed regions (PSRs), which are vital for enabling sustainable human space exploration, making these regions high-priority targets for upcoming Artemis missions [1]. Unfortunately, today, the best available orbital lunar imagery [2, 3] lacks the meter-scale resolution and signal needed to understand the geomorphology and trafficability of PSRs, complicating the planning and execution of future missions seeking to explore PSRs. We have developed an image enhancement tool called HORUS (Hyper-effective nOise Removal Unet Software) [4, 5], designed to enhance LRO Narrow-Angle Camera (NAC) optical low-light imagery of permanently shadowed regions by effectively removing the CCD-related, photon, and other residual noises that corrupt the images. The tool is composed of two deep learning neural networks trained on environmental metadata and real and synthetic imagery, the latter generated by a physical noise model (LROC). We demonstrated that HORUS effectively produces low-noise, high-resolution images (~1.5m/px), achieving a 5 to 10x improvement over existing long-exposure images of PSRs. HORUS allows scientists and engineers to identify geomorphic features (e.g., craters and boulders) in shadowed regions as small as 3 meters across as well as to peek inside of small shadowed regions, for the first time. The tool was deployed and thoroughly validated for NASA's VIPER mission [6], where it was applied to 20 candidate target regions across the lunar South Pole. Additionally, we conducted different approaches to validate the resulting HORUS-processed images. With HORUS denoised images, VIPER scientists can increase their confidence on what surface features (previously unseen) exist in the shadowed regions, helping them plan rover traverses more safely and efficiently (e.g., Fig. 1) In this manuscript, we will describe how VIPER scientists are utilizing HORUS denoised images to extract new information from the terrain and increase their confidence in what surface features exist in the shadowed regions. In combination with other high-resolution images and digital elevation maps, HORUS images are helping the team analyze potential lading and science sites, as well as planning traverses more safely and efficiently (e.g., Fig. 1). Additionally, we will describe how HORUS tool unlocks a broad range of scientific and exploration applications to other Artemis and CPLS missions to the lunar poles, including (but not limited to) geomorphic analysis, change detection, surface hazard detection, and terrain relative navigation.

artificial intelligence

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya

Model-Based Engineering Design for Trade Space Exploration throughout the Design Cycle

This paper presents ongoing work to standardize model-based system engineering as a complement to point design development in the conceptual design phase of deep space missions. It summarizes two first steps towards practical application of this capability within the framework of concurrent engineering design teams and their customers. The first step is standard generation of system sensitivities models as the output of concurrent engineering design sessions, representing the local trade space around a point design. A review of the chosen model development process, and the results of three case study examples, demonstrate that a simple update to the concurrent engineering design process can easily capture sensitivities to key requirements. It can serve as a valuable tool to analyze design drivers and uncover breakpoints in the design. The second step is development of rough-order- of-magnitude, broad-range-of-validity design models for rapid exploration of the trade space, before selection of a point design. At least one case study demonstrated the feasibility to generate such models in a concurrent engineering session. The experiment indicated that such a capability could yield valid system-level conclusions for a trade space composed of understood elements. Ongoing efforts are assessing the practicality of developing end-to-end system-level design models for use before even convening the first concurrent engineering session, starting with modeling an end-to-end Mars architecture.

models

Development of a Phasor Diagram Creator to Visualize the Piston and Displacer Forces in an Advanced Stirling Convertor

The steady-state, nearly sinusoidal behavior of the components in a free-piston Stirling engine allows for visualization of the forces in the system using phasor diagrams. Based on Newton's second law, F = ma, any phasor diagrams modeling a given component in a system should close if all of the acting forces have been considered. Since the Advanced Stirling Radioisotope Generator (ASRG), currently being developed for future NASA deep space missions, is made up of such nearly sinusoidally oscillating components, its phasor diagrams would also be expected to close. A graphical user interface (GUI) has been written in MATLAB (MathWorks), which takes user input data, passes it to Sage (Gedeon Associates), a one-dimensional thermodynamic modeling program used to model the Stirling convertor, runs Sage, and then automatically plots the phasor diagrams. Using this software tool, the effect of varying different Sage inputs on the phasor diagrams was determined. The parameters varied were piston amplitude, hot-end temperature, cold-end temperature, operating frequency, and displacer spring constant. These phasor diagrams offer useful insight into convertor operation and performance.

Saha, Dipanjan

Deep-learning-enhanced assessment of wellbore barrier effectiveness in geologic storage systems with intermediate aquifers

For geologic systems where carbon dioxide (CO 2 ) is injected underground, existing wells represent potential pathways for fluid migration. Here, this study introduces a novel deep learning model to quantify the likelihood and potential magnitude of fluid migration through wellbores at sites with intermediate aquifers or thief zones between the injection units and underground drinking water sources. Synthetic datasets, generated using reservoir simulations, captured a wide range of subsurface conditions, well attributes, operational parameters, and fluid migration scenarios. Among the regression models developed to predict brine and CO 2 leakage rates and CO 2 saturations along leaky wellbores, convolutional neural network (CNN) outperformed both Light Gradient Boosting Machine and deep neural network. Additionally, a CNN-based classification model was created to predict whether brine and CO 2 would leak along a wellbore, further improving performance over regression alone. The best models were integrated into the National Risk Assessment Partnership Open-source Integrated Assessment Model for rapid, stochastic assessment of storage system containment and leakage risks. A case study demonstrated the model’s ability to simulate fluid migration through existing wells with multiple intermediate aquifers. This computationally efficient wellbore model offers value in support of site performance evaluation and risk-informed decision making by stakeholders.

CO2 leakage

Development and Evaluation of a General Drag Model for Gas-Solid Flows via Deep Learning

This project presents the development and evaluation of a general drag model for gas–solid multiphase flows using deep learning techniques. A comprehensive database of more than 4,000 experimental and numerical data points for spherical and non spherical particles was compiled, incorporating geometric features such as sphericity, aspect ratio, and orientation. Several predictive approaches—including traditional em pirical correlations, machine learning, and deep neural networks—were benchmarked, with the proposed Drag Coefficient Correlation-aided Deep Neural Network (DCC DNN) demonstrating superior accuracy. To account for particle–particle interactions, additional drag data were generated using CFD-based simulations of packed and flu idized beds, leading to the development of a retrained model capable of incorporat ing volume fraction effects. Integration of the trained model with the MFiX CFD solver was achieved using FTorch, enabling drag predictions during discrete element method (DEM) simulations. Validation against experimental data for single particles and fluidized beds confirmed the model’s improved predictive ability, particularly for non-spherical geometries. While the model performed strongly under fluidized con ditions, limitations remained in unfluidized regimes, suggesting a need for expanded datasets. Overall, this study demonstrates the feasibility of combining deep learning with physics-informed CFD to improve drag modeling for gas–solid flows, with promis ing implications for scaling multiphase simulations in industrial applications.

42 ENGINEERING

A unified large language model–based framework for heterogeneous PV image diagnosis

With advances in imaging technologies, modern photovoltaic (PV) systems generate large volumes of heterogeneous image data, including visible, electroluminescence (EL), and infrared (IR) images. Existing PV image analysis models, particularly deep learning approaches, are typically task-specific and lack cross-modality generalization. To address this limitation, this paper proposes an open-source large language model (LLM)–based unified framework for heterogeneous PV image diagnostics. Through task-aware diagnostic prompting, the framework enables analysis of visible, EL, and IR images within a single pipeline, supporting both zero-shot and few-shot inference and binary and multiclass classification. It is compatible with state-of-the-art multimodal LLMs, including ChatGPT, Gemini, Claude, Qwen, and CLIP. The framework is evaluated on PV module condition classification (clean, soiling, snow, hail, and bird droppings) using visible images, cell crack detection using EL images, and hotspot detection using IR images. GPT-5.1 in few-shot mode achieves the best performance, with classification accuracy exceeding 97.3%. Open-source models such as Qwen and CLIP also deliver competitive results on visible images (around 90% accuracy), though their performance is more limited on EL and IR modalities. On the full ELPV dataset, the framework achieves 83.5% zero-shot accuracy, within 2.8% of the supervised CNN baseline, confirming scalability to larger benchmarks. Practical aspects such as reproducibility, response latency, and confidence estimation are systematically analyzed. The framework operates across PV image modalities without modality- or task-specific training, making it well suited as a rapid pre-screening tool to support downstream detailed diagnostics. A benchmark dataset of diverse labeled PV images is also released.

Li, Baojie

pop-cosmos : redshifts and physical properties of KiDS-1000 galaxies

ABSTRACT Principled Bayesian inference of galaxy properties has not previously been performed for wide-area weak-lensing surveys with millions of sources. We address this gap by applying the pop-cosmos generative model to perform spectral energy distribution (SED) fitting for 4 million KiDS (Kilo-Degree Survey)-1000 galaxies. Calibrated on deep COSMOS2020 photometric data, pop-cosmos specifies a physically motivated prior over the galaxy population up to $z \simeq 6$ in stellar population synthesis (SPS) parameter space. Using the Speculator SPS emulator with GPU (graphics processing unit)-accelerated Markov Chain Monte Carlo sampling, we perform full posterior inference at 8.2 GPU seconds per galaxy, obtaining joint constraints on galaxy redshifts and physical properties. We validate photometric redshifts against $\sim \!185\,\!000$ KiDS galaxies cross-matched to Dark Energy Spectroscopic Instrument Data Release 1 spectroscopic samples, achieving low bias ($2\times 10^{-3}$), scatter ($\sigma _{\mathrm{MAD}}=0.03$), and outlier fraction (3.2 per cent) for the Bright Galaxy Survey, with comparable performance (bias $3\times 10^{-2}$, $\sigma _{\mathrm{MAD}}=0.05$, 1.0 per cent outliers) for luminous red galaxies (LRGs). Within the LRG sample, we identify massive, dusty, star-forming contaminants at $z \simeq 0.4$ satisfying standard colour selections for quenched populations. We infer trends in stellar mass, star formation, metallicity, and dust across five tomographic redshift bins consistent with established scaling relations. Using specific star formation rate constraints, we identify $\sim$7 per cent of KiDS-1000 galaxies as quenched, versus 37 per cent implied by conservative colour cuts. This enables the construction of weak-lensing samples defined by physical properties while mitigating intrinsic alignment systematics and preserving statistical power. Our analysis validates pop-cosmos out of sample, establishing it as a scalable approach for galaxy evolution and cosmological analyses with photometric surveys.

Halder, Anik [Institute of Astronomy and Kavli Ins

An Advanced Machine Learning and Artificial Intelligence System for Demonstrating Radiation Regulatory Compliance in DOE Accelerator Facilities

In this Phase II proposal, Applied Research LLC (ARLLC), Thomas Jefferson National Accelerator Facility (Jefferson Lab), and Old Dominion University (ODU) propose the combination of domain knowledge (beam characteristics, fixed structural shielding, earthen burden (the soil and foliage added to the dome of the experimental halls as additional shielding), etc.), machine learning (ML) and/or artificial intelligence (AI) to correlate a variety of multi-modal onsite signals and the radiation fields seen in accessible areas of the accelerator site and the site boundary. The ML/AI will consider the complex influence of environmental parameters affecting the radon contribution of the measurements, focusing on actual data obtained from Jefferson Lab. In Phase I, the coded beam and location data were fed into a deep learning model to predict doses at several designated locations in Jefferson Lab’s facility. Moreover, a dense radiation map was generated using only a sparse collection of the samples in a facility. In Phase II, we will develop a software prototype containing a radiation prediction algorithm, dense radiation map algorithms, and background noise prediction algorithms, with actual data used to evaluate the prototype. This work will provide a framework for evaluation of radiation measurement results around the site based on learned responses. In addition, the proposed approach allows more granular mapping of radiation levels. Better understanding and communication of these levels is related to the overall approach in keeping doses to personnel ALARA.

43 PARTICLE ACCELERATORS

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY