Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning for Data Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Electronic Nose Development and Preliminary Human Breath Testing for Rapid, Non-Invasive COVID-19 Detection

We adapted an existing, spaceflight-proven, robust “electronic nose” (E-Nose) that uses an array of electrical resistivity-based nanosensors mimicking aspects of mammalian olfaction to conduct on-site, rapid screening for COVID-19 infection by measuring the pattern of sensor responses to volatile organic compounds (VOCs) in exhaled human breath. We built and tested multiple copies of a hand-held prototype E-Nose sensor system, composed of 64 chemically sensitive nanomaterial sensing elements tailored to COVID-19 VOC detection; data acquisition electronics; a smart tablet with software (App) for sensor control, data acquisition and display; and a sampling fixture to capture exhaled breath samples and deliver them to the sensor array inside the E-Nose. The sensing elements detect the combination of VOCs typical in breath at parts-per-billion (ppb) levels, with repeatability of 0.02% and reproducibility of 1.2%; the measurement electronics in the E-Nose provide measurement accuracy and signal-to-noise ratios comparable to benchtop instrumentation. Preliminary clinical testing at Stanford Medicine with 63 participants, their COVID-19-positive or COVID-19-negative status determined by concomitant RT-PCR, discriminated between these two categories of human breath with a 79% correct identification rate using “leave-one-out” training-and-analysis methods. Analyzing the E-Nose response in conjunction with body temperature and other non-invasive symptom screening using advanced machine learning methods, with a much larger database of responses from a wider swath of the population, is expected to provide more accurate on-the-spot answers. Additional clinical testing, design refinement, and a mass manufacturing approach are the main steps toward deploying this technology to rapidly screen for active infection in clinics and hospitals, public and commercial venues, or at home.

COVID-19↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Personnel Data Analysis and Retrieval of Phase 1 Move To LC-39 Area

As a technology major from Jackson State University (JSU) I was called in as a summer intern at Kennedy Space Center (KSC) to work in the NASA Engineering, Control and Data Systems (NE-C) Division supporting the Spaceport Command and Control System (SCCS) at the Space Station Processing Facility (SSPF). I was given a two-part project; the first consisted of lending support relocating SCCS Computer Equipment and Project Personnel to the Launch Control Center (LCC). This task involved me using a Microsoft Office data processing tool to assist with the analysis and information management of logistics worth millions of dollars. With the assistance of two other interns, I was responsible for collecting data on equipment used, on a daily basis, by over 200 KSC employees. The many network servers, enterprise switches, desktop computers, and fiber optics had to be handled in an equally prompt and precise manner in order to ensure a minimal amount of equipment down time; which is critical in ensuring a properly secured networking environment. The second part of my project was to assist KSC in developing a more cost effective way of maintaining and taking full advantage of the functionality of some new kiosk units. Since KSC currently has no expert on the servicing and maintenance of the units, I, as a computer technology major, was given the opportunity to assess the hardware and software of the machines. The goal was to learn to establish a secure and remote environment for the kiosks; a goal highly valuing convenience by preserving valuable man-hours saved by not having to travel to each individual kiosk location. In addition, I was to leave a clear and precise plan for future users and administrators of the devices to follow.

Davis, Derrick D.↗

The INSTEP Monitoring Network: Merging High-and-Low Cost Measurements to Characterize California Wildfires

Despite challenges with data quality and scope, low-cost sensor networks have skyrocketed in popularity over the last 15 years, making air quality data available on refined spatial scales. More recently, studies have leveraged both high and low-quality instruments to create stronger “hybrid” models, with most studies focusing on particulate matter. Low-cost measurements typically represent ground-level emissions only, providing context for human health issues from climate change-driven events such as wildfires. Since low-cost sensors’ capabilities are localized, daily events and microclimates tend to dominate the data rather than larger regional or atmospheric trends. Likewise, their low cost explains their high uncertainty. In contrast, some regulatory-grade instruments produce column measurements as well, providing reliable information on a broader scope. To bridge this gap while expanding into gas-phase measurements, we deployed 12 air quality sensor packages in California, USA during the 2022 wildfire season. These INSTEP (Inexpensive Network Sensor Technology Exploring Pollution) monitors measure carbon monoxide (CO), carbon dioxide (CO2), ozone (O3), nitrogen dioxide (NO2), and several hydrocarbons including methane (CH4) and formaldehyde (HCHO). Half of the monitors were co-located with remote sensing spectrometers: NASA Pandora and Total Column Carbon Observing Network (TCCON). The overlap in pollutants includes NO2, O3, and HCHO between the INSTEP monitors and the Pandora column measurements. TCCON covers column CO, CO2, and CH4, rounding out our comparison. Most of the monitors were distributed throughout the San Francisco Bay area, and an additional three were located within 100 km of Los Angeles. The sites ranged in geographic and population characteristics, including desert, mountainous, coastal, and urban locations. Since varying environmental conditions such as temperature and pressure are known to challenge sensor performance, we will apply newer sensor “calibration” techniques meant to combat this. We will normalize our sensor signals by z-scoring them prior to applying a single calibration model in the form of multivariate linear regression or an artificial neural network. While this technique has been validated for the hydrocarbon and ozone sensor types (metal oxide), it has not yet been tested on electrochemical and non-dispersive infrared sensors, which are also used in the INSTEP monitors. This will serve as a test to see if this normalization technique – or another – is most effective in accounting for environmental differences among sensors. Related data analysis efforts have found success with a variety of geospatial analysis techniques, including weighted network models in which high-quality instruments are given higher weights than their low-cost counterparts. Our preliminary analysis will focus on kriging, which uses a Gaussian algorithm to assign weights, providing estimated pollution levels at locations between monitors. Smoke trajectory and evolution will also be considered using both measurement types. We also aim to baseline subtract our emission estimates from each region to determine which portion of emissions are regional and local, further characterizing burn differences in northern and southern California fires. Future directions include using INSTEP jointly with TEMPO satellite data, and mobile deployments on aircraft and uncrewed aerial vehicles (UAV).

Low-cost sensors↗

Exploring the Capabilities of a Machine Learning Algorithm to Detect Space Weather-Significant Emerging Active Regions

Active regions are a source of various phenomena responsible for Space Weather disturbances; therefore, developing a technology for early warning about upcoming magnetic activity is crucial to mitigate its impact. However, observational limitations and the high nonlinearity of processes associated with the accumulation of magnetic flux and its interaction with the surrounding plasma during the emergence through the convection zone make early activity detection a challenging problem. To address these challenges, we developed a physics-driven machine learning model that allows us to detect active regions (ARs) before they become visible on the solar surface by analyzing the power spectra of acoustic oscillations observed by the SDO/HMI instrument. This study is based on a time series of Doppler shift maps of 31x31-degree areas tracked with the Carrington rotation rate for four days before and after the emergence. The Doppler shift time series are processed into the oscillation power maps for four frequency ranges and accompanied by line-of-sight magnetograms and the continuum intensity maps from SDO/HMI. The resulting data are converted into a 1D time series representing the mean temporal variations of these quantities. The redacted time series are used as input to predict AR emergence using the Long Short Term Memory (LSTM) method. The training of the LSTM model is based on 40 ARs, which includes an independent analysis for each sub region that exhibits AR emergence or remains quiet. The emergence of magnetic flux (defined as a decrease of the continuum intensity) was detected with the developed LSTM algorithm from 5 to 48 hours before the reported time by NOAA. The developed model is capable of pointing to the time and location of active region formation. In this presentation, we discuss reasons that impact how early in advance the model can identify the upcoming activity and the possibility of improving the current predictive skills and steps to transition to the operational forecast.

Heliophysics↗

Digital Lunar Exploration Sites (DLES) Terrain Crafting

Humans will soon be returning to the surface of the Moon with NASA’s Artemis program. The Artemis program is an international collaboration that will consist of a complex series of space systems and missions to explore the lunar surface and pave the way for the future exploration of Mars. NASA and its partners rely heavily on simulation for lighting and navigation studies as well as training astronauts, flight controllers, and mission support staff. The NASA Exploration Systems Simulations (NExSyS) team in the Simulation and Graphics Branch (ER7) in the Engineering Directorate at NASA’s Johnson Space Center has built up many simulation products to support this effort, one of which is the Digital Lunar Exploration Sites (DLES). DLES is a collection of products used to simulate and render the lunar surface in a digital environment. We discussed and presented an overview of the DLES products at the 2022 IEEE Aerospace Conference in Big Sky, MT with a paper titled "Digital Lunar Exploration Sites". This “DLES Terrain Crafting” paper will expand on the information previously provided in “DLES” paper and dive deeper into the details of the terrain crafting process and the toolsets used to support this task. The best digital data currently available of the lunar surface is provided by the Lunar Reconnaissance Orbiter (LRO). Its Lunar Orbiter Laser Altimeter (LOLA) achieves an impressive resolution of 5m per pixel at the Lunar South Pole (LSP) and can generate datasets covering a large continuous region near the LSP. There are a few additional methods, such as Shape from Shading which can infer higher resolution data (up to 1m per pixel) from the LRO Narrow Angle Camera (NAC) images. However, surface-based simulations require higher-resolution data, and this paper will discuss the process of enhancing the terrain to meet that need. The process begins with capturing statistical data of craters in the regions of interest using images provided by the LRO NAC. This data is then used to scatter artificial features which are not captured in the truth data, resulting in an enhanced DEM with a much higher resolution of 20cm per pixel. Many tools were built up to assist in the creation of these artificial Digital Elevation Models (DEM), which this paper will discuss in detail. DEMs themselves are a very powerful representation of a planetary surface, and many operations and tools can utilize the data they contain. This paper includes a description of the rendering of the lunar surface in a graphics engine, generation of contact patches to simulate tire to ground interaction, and ray tracing utilities to model Line of Sight (LOS) interactions with the terrain. This paper will also explore some new tool sets currently under development which aim to utilize Machine Learning (ML) to assist in the identification of craters from LRO NAC imagery. While this is not a novel idea, the NExSyS team is developing a unique approach which may result in more robust identification of crater characteristics.

Artemis↗

Coevolution of Machine Learning and Process-Based Modelling to Revolutionize Earth and Environmental Sciences: A Perspective

Machine learning (ML) applications in Earth and environmental sciences (EES) have gained incredible momentum in recent years. However, these ML applications have largely evolved in ‘isolation’ from the mechanistic, process-based modelling (PBM) paradigms, which have historically been the cornerstone of scientific discovery and policy support. In this perspective, we assert that the cultural barriers between the ML and PBM communities limit the potential of ML, and even its ‘hybridization’ with PBM, for EES applications. Fundamental, but often ignored, differences between ML and PBM are discussed as well as their strengths and weaknesses in light of three overarching modelling objectives in EES, (1) nowcasting and prediction, (2) scenario analysis, and (3) diagnostic learning. The paper ponders over a ‘coevolutionary’ approach to model building, shifting away from a borrowing to a co-creation culture, to develop a generation of models that leverage the unique strengths of ML such as scalability to big data and high-dimensional mapping, while remaining faithful to process-based knowledge base and principles of model explainability and interpretability, and therefore, falsifiability.

Saman Razavi↗

Review and Analysis of Algorithmic Approaches Developed for Prognostics on CMAPSS Dataset

Benchmarking of prognostic algorithms has been challenging due to limited availability of common datasets suitable for prognostics. In an attempt to alleviate this problem several benchmarking datasets have been collected by NASA's prognostic center of excellence and made available to the Prognostics and Health Management (PHM) community to allow evaluation and comparison of prognostics algorithms. Among those datasets are five C-MAPSS datasets that have been extremely popular due to their unique characteristics making them suitable for prognostics. The C-MAPSS datasets pose several challenges that have been tackled by different methods in the PHM literature. In particular, management of high variability due to sensor noise, effects of operating conditions, and presence of multiple simultaneous fault modes are some factors that have great impact on the generalization capabilities of prognostics algorithms. More than 70 publications have used the C-MAPSS datasets for developing data-driven prognostic algorithms. The C-MAPSS datasets are also shown to be well-suited for development of new machine learning and pattern recognition tools for several key preprocessing steps such as feature extraction and selection, failure mode assessment, operating conditions assessment, health status estimation, uncertainty management, and prognostics performance evaluation. This paper summarizes a comprehensive literature review of publications using C-MAPSS datasets and provides guidelines and references to further usage of these datasets in a manner that allows clear and consistent comparison between different approaches.

Uncertainty↗

Landslide Hazard is Projected to Increase Across High Mountain Asia

High Mountain Asia has long been known as a hotspot for landslide risk, and studies have suggested that landslide hazard is likely to increase in this region over the coming decades. Extreme precipitation may become more frequent, with a nonlinear response relative to increasing global temperatures. However, these changes are geographically varied. This article maps probable changes to landslide hazard, as shown by a landslide hazard indicator (LHI) derived from downscaled precipitation and temperature. In order to capture the nonlinear response of slopes to extreme precipitation, a simple machine-learning model was trained on a database of landslides across High Mountain Asia to develop a regional LHI. This model was applied to statistically downscaled data from the 30 members of the Seamless System for Prediction and Earth System Research large ensembles to produce a range of possible outcomes under the Shared Socioeconomic Pathways 2-4.5 and 5-8.5. The LHI reveals that landslide hazard will increase in most parts of High Mountain Asia. Absolute increases will be highest in already hazardous areas such as the Central Himalaya, but relative change is greatest on the Tibetan Plateau. Even in regions where landslide hazard declines by year 2100, it will increase prior to the mid-century mark. However, the seasonal cycle of landslide occurrence will not change greatly across High Mountain Asia. Although substantial uncertainty remains in these projections, the overall direction of change seems reliable. These findings highlight the importance of continued analysis to inform disaster risk reduction strategies for stakeholders across High Mountain Asia.

Thomas A Stanley↗

Session on High Speed Civil Transport Design Capability Using MDO and High Performance Computing

Since the inception of CAS in 1992, NASA Langley has been conducting research into applying multidisciplinary optimization (MDO) and high performance computing toward reducing aircraft design cycle time. The focus of this research has been the development of a series of computational frameworks and associated applications that increased in capability, complexity, and performance over time. The culmination of this effort is an automated high-fidelity analysis capability for a high speed civil transport (HSCT) vehicle installed on a network of heterogeneous computers with a computational framework built using Common Object Request Broker Architecture (CORBA) and Java. The main focus of the research in the early years was the development of the Framework for Interdisciplinary Design Optimization (FIDO) and associated HSCT applications. While the FIDO effort was eventually halted, work continued on HSCT applications of ever increasing complexity. The current application, HSCT4.0, employs high fidelity CFD and FEM analysis codes. For each analysis cycle, the vehicle geometry and computational grids are updated using new values for design variables. Processes for aeroelastic trim, loads convergence, displacement transfer, stress and buckling, and performance have been developed. In all, a total of 70 processes are integrated in the analysis framework. Many of the key processes include automatic differentiation capabilities to provide sensitivity information that can be used in optimization. A software engineering process was developed to manage this large project. Defining the interactions among 70 processes turned out to be an enormous, but essential, task. A formal requirements document was prepared that defined data flow among processes and subprocesses. A design document was then developed that translated the requirements into actual software design. A validation program was defined and implemented to ensure that codes integrated into the framework produced the same results as their standalone counterparts. Finally, a Commercial Off the Shelf (COTS) configuration management system was used to organize the software development. A computational environment, CJOPT, based on the Common Object Request Broker Architecture, CORBA, and the Java programming language has been developed as a framework for multidisciplinary analysis and Optimization. The environment exploits the parallelisms inherent in the application and distributes the constituent disciplines on machines best suited to their needs. In CJOpt, a discipline code is "wrapped" as an object. An interface to the object identifies the functionality (services) provided by the discipline, defined in Interface Definition Language (IDL) and implemented using Java. The results of using the HSCT4.0 capability are described. A summary of lessons learned is also presented. The use of some of the processes, codes, and techniques by industry are highlighted. The application of the methodology developed in this research to other aircraft are described. Finally, we show how the experience gained is being applied to entirely new vehicles, such as the Reusable Space Transportation System. Additional information is contained in the original.

Rehder, Joe↗

Exploring the Utility of Machine Learning-Based Passive Microwave Brightness Temperature Data Assimilation over Terrestrial Snow in High Mountain Asia

This study explores the use of a support vector machine (SVM) as the observation operator within a passive microwave brightness temperature data assimilation framework (herein SVM-DA) to enhance the characterization of snow water equivalent (SWE) over High Mountain Asia (HMA). A series of synthetic twin experiments were conducted with the NASA Land Information System (LIS) at a number of locations across HMA. Overall, the SVM-DA framework is effective at improving SWE estimates (~70% reduction in RMSE relative to the Open Loop) for SWE depths less than 200 mm during dry snowpack conditions. The SVM-DA framework also improves SWE estimates in deep, wet snow (~45% reduction in RMSE) when snow liquid water is well estimated by the land surface model, but can lead to model degradation when snow liquid water estimates diverge from values used during SVM training. In particular, two key challenges of using the SVM-DA framework were observed over deep, wet snowpacks. First, variations in snow liquid water content dominate the brightness temperature spectral difference (TB) signal associated with emission from a wet snowpack, which can lead to abrupt changes in SWE during the analysis update. Second, the ensemble of SVM-based predictions can collapse (i.e., yield a near-zero standard deviation across the ensemble) when prior estimates of snow are outside the range of snow inputs used during the SVM training procedure. Such a scenario can lead to the presence of spurious error correlations between SWE and TB, and as a consequence, can result in degraded SWE estimates from the analysis update. These degraded analysis updates can be largely mitigated by applying rule-based approaches. For example, restricting the SWE update when the standard deviation of the predicted TB is greater than 0.05 K helps prevent the occurrence of filter divergence. Similarly, adding a thin layer (i.e., 5 mm) of SWE when the synthetic TB is larger than 5 K can improve SVM-DA performance in the presence of a precipitation dry bias. The study demonstrates that a carefully constructed SVM-DA framework cognizant of the inherent limitations of passive microwave-based SWE estimation holds promise for snow mass data assimilation.

Kwon, Yonghwan↗

Future Model-Based Systems Engineering Vision and Strategy Bridge for NASA

A vision for the future of model-based systems engineering (MBSE) at NASA in 2029 and a strategy bridge towards that future are presented. Strategic thinking and leading change concepts were used to analyze reports and presentations on global trends and visionary thinking about the future of systems and digital engineering. The context, strategic time horizon, stakeholders, strategic challenges, strategic advantages, driving forces, and opportunities were considered. The analysis resulted in a future vision of MBSE that shows what NASA systems engineers and digital machines will do to perform rapid, extraordinary, and unprecedented missions. The NASA systems engineer, in this future vision, works with a global project team in a virtual and collaborative environment, engineers the system, and uses digital approaches as the routine and default way of working. The digital machines provide data-driven and automated mission designs; have a backbone of program and project management, systems engineering, and product life-cycle management; and are a knowledge-sharing infrastructure. The NASA systems engineer and the systems engineering team are envisioned to use digital machines to plan and perform rapid exploration missions, develop a digital twin that lasts across the life cycle, and develop enduring and adaptable systems. NASA has an engineering enterprise and a life-cycle management framework that endure, adapt, and respond. A strategy bridge based on the Baldrige Criteria for Performance Excellence Framework and lessons learned from a recent MBSE initiative illuminates a way forward from today to this desired future. The bridge lays out a strategy for leaders and recommends investments of today for immediate benefits and for benefits in 2029.

model-based systems engineering, digital engineeri↗

Regression Analysis of Top of Descent Location for Idle-thrust Descents

In this paper, multiple regression analysis is used to model the top of descent (TOD) location of user-preferred descent trajectories computed by the flight management system (FMS) on over 1000 commercial flights into Melbourne, Australia. The independent variables cruise altitude, final altitude, cruise Mach, descent speed, wind, and engine type were also recorded or computed post-operations. Both first-order and second-order models are considered, where cross-validation, hypothesis testing, and additional analysis are used to compare models. This identifies the models that should give the smallest errors if used to predict TOD location for new data in the future. A model that is linear in TOD altitude, final altitude, descent speed, and wind gives an estimated standard deviation of 3.9 nmi for TOD location given the trajec- tory parameters, which means about 80% of predictions would have error less than 5 nmi in absolute value. This accuracy is better than demonstrated by other ground automation predictions using kinetic models. Furthermore, this approach would enable online learning of the model. Additional data or further knowl- edge of algorithms is necessary to conclude definitively that no second-order terms are appropriate. Possible applications of the linear model are described, including enabling arriving aircraft to fly optimized descents computed by the FMS even in congested airspace. In particular, a model for TOD location that is linear in the independent variables would enable decision support tool human-machine interfaces for which a kinetic approach would be computationally too slow.

trajectory prediction↗

Enhancing NASA Earth Science Data Discovery from Scientific Publications

Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.

Irina V Gerasimov↗

Data-driven landslide nowcasting at the global scale

Landslides affect nearly every country in the world each year. To better understand this global hazard, the Landslide Hazard Assessment for Situational Awareness (LHASA) model was developed previously. LHASA version 1 combines satellite precipitation estimates with a global landslide susceptibility map to produce a gridded map of potentially hazardous areas from 60° North-South every 3 h. LHASA version 1 categorizes the world’s land surface into three ratings: high, moderate, and low hazard with a single decision tree that first determines if the last seven days of rainfall were intense, then evaluates landslide susceptibility. LHASA version 2 has been developed with a data-driven approach. The global susceptibility map was replaced with a collection of explanatory variables, and two new dynamically varying quantities were added: snow and soil moisture. Along with antecedent rainfall, these variables modulated the response to current daily rainfall. In addition, the Global Landslide Catalog (GLC) was supplemented with several inventories of rainfall-triggered landslide events. These factors were incorporated into the machine-learning framework XGBoost, which was trained to predict the presence or absence of landslides over the period 2015–2018, with the years 2019–2020 reserved for model evaluation. As a result of these improvements, the new global landslide nowcast was twice as likely to predict the occurrence of historical landslides as LHASA version 1, given the same global false positive rate. Furthermore, the shift to probabilistic outputs allows users to directly manage the trade-off between false negatives and false positives, which should make the nowcast useful for a greater variety of geographic settings and applications. In a retrospective analysis, the trained model ran over a global domain for 5 years, and results for LHASA version 1 and version 2 were compared. Due to the importance of rainfall and faults in LHASA version 2, nowcasts would be issued more frequently in some tropical countries, such as Colombia and Papua New Guinea; at the same time, the new version placed less emphasis on arid regions and areas far from the Pacific Rim. LHASA version 2 provides a nearly real-time view of global landslide hazard for a variety of stakeholders.

XGBoos↗

Detection of CH 4 Hotspots from NASA’s GEOS Composition Analysis System

Recent research indicates that 8 to 12% of the global oil and gas production methane emissions could be attributed to ultra-emitters, which result in high concentration ‘hotspots’ near point sources. Identifying these emissions in near real time provides useful information to the policy makers and private industry, who are working to reduce their impact. To meet this need, scientists are increasingly analyzing satellite data from the TROPOspheric Monitoring Instrument (TROPOMI) instrument aboard ESA’s Sentinel 5-Precursor mission. While direct analysis of TROPOMI level 2 swath data has been successful in identifying some large emission events, identifying hotpots is challenging because of the imaging noise due to a variety of artifacts and limits in daily coverage. Here we explore possible methodologies to detect methane hotspots using a new, gap-filled, and temporally continuous methane product from NASA’s Goddard Earth Observing System (GEOS) Constituent Data Assimilation System (CoDAS), which assimilates column averaged methane mole fractions from the TROPOMI with capabilities to assimilate other remote sensing measurements. The CoDAS has been expanded from a heritage of stratospheric composition and carbon dioxide assimilation allowing for the support of regional modeling, validation with non-coincident operations, and merging variety of datasets. The current work mainly explores Observing System Simulation Experiments with methane GEOS simulations without assimilation to prepare the groundwork for further experiments with assimilated TROPOMI. First, known hotspots based on the known inventory are identified to demonstrate the capability of the system to point out emission hotspots. In the next step, a variety of machine learning techniques such as Self-Organizing Maps and Deep Learning are explored to automate detection of the plumes. Finally, a few approaches to quantify emissions from the identified hotspots are presented and are evaluated against the inventory. The effort is directed toward a future evaluation of the CoDAS based methane monitoring system’s ability to successfully detect and quantify hotspots.

Nikolay Balashov↗

A Neural Network Parametrization of Volumetric Cloud Fraction Profiles Using Satellite Observations and MERRA-2 Reanalysis Meteorological Data

Clouds play a crucial role in regulating the hydrologic cycle and Earth's radiative energy budget, yet they are often poorly represented in global climate models (GCMs). This study applies deep machine learning techniques to develop a physical parameterization of volumetric cloud fraction (VCF), the fraction of a 3-D grid volume occupied by clouds using satellite lidar-radar measurements. The neural network (NN) captures the complicated relationships between observed VCF profiles and collocated meteorological variables from MERRA-2 reanalysis data. Our results show that the NN model, particularly a sequence-to-sequence long short-term memory (LSTM) network with a sixfactor loss function, effectively learns the underlying cloud physical processes. The NN model outperforms MERRA-2 reanalysis in representing low-level clouds in tropical and subtropical regions and low- and middle-level clouds over midlatitude storm-track regions, and also improves VCF histograms. These improvements are reflected in the vertical distributions of zonally, meridionally, and globally averaged VCFs, geographic distributions of low-, middle-, and high-level clouds, and seasonal variations in monthly-mean VCF. Furthermore, the NN predictions effectively capture the El Niño-Southern Oscillation (ENSO) effects and other interannual variations. The NN parameterization is further evaluated through a sensitivity analysis, in which a single predictor is perturbed at a time. This reveals that relative humidity (RH) is the dominant factor influencing variations in globally averaged VCF at low and middle altitudes, followed by temperature. At higher altitudes, temperature becomes the primary driver of VCF through its effect on RH. Changes in wind components had minimal impact on globally averaged VCF.

Shan Zeng↗

Global SO 2 Data Record from OMPS Instruments on the JPSS Constellation

NASA’s Earth Observing System (EOS) SO 2 climate data record (CDR) started in 2004, with the launch of the Aura/Ozone Monitoring Instrument (OMI) and is now being continued with the SNPP/Ozone Mapping and Profiler Suite (OMPS) launched in 2011. Both OMI and SNPP/OMPS SO 2 CDRs are produced with the Goddard principal component analysis (PCA) spectral fitting algorithm. An advantage of the data-driven PCA retrieval technique is that it enables highly consistent retrievals from different instruments, by inherently accounting for various instrumental factors. To further extend the EOS SO 2 CDR, we are implementing the PCA SO 2 retrieval algorithm with the L1B measurements from OMPS instruments flying on the Joint Polar Satellite System (JPSS) constellation. In this presentation, we will provide an update on our progress in NOAA-20 (launched in 2017) and NOAA-21 (launched in 2022) PCA SO2 retrievals. We will focus on our new NOAA-20/OMPS PCA SO 2 EOS continuity product, to be publicly released in fall of 2023. We will present statistical analyses on the quality of NOAA-20 PCA SO 2 product, including retrieval noise, biases over background areas, and long-term stability. We will compare our PCA SO 2 retrievals from NOAA-20 with those from OMI, SNPP/OMPS, and S5P/TROPOMI (TROPOspheric Monitoring Instrument) for anthropogenic sources as well as large volcanic plumes. We will also discuss the application of a new machine learning technique that helps to further reduce the noise of NOAA-20 SO 2 retrievals. In addition, we will present preliminary PCA SO 2 retrievals from NOAA-21/OMPS, including those from direct readout implementation for aviation disaster avoidance. Finally, we will share some first results applying the PCA algorithm to NASA’s geostationary TEMPO (Tropospheric Emissions: Monitoring of Pollution) instrument to obtain hourly, high resolution SO 2 data over North America.

SO2↗