Search NASASearch

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Approximate CFTs and random tensor models

Abstract A key issue in both the field of quantum chaos and quantum gravity is an effective description of chaotic conformal field theories (CFTs), that is CFTs that have a quantum ergodic limit. We develop a framework incorporating the constraints of conformal symmetry and locality, allowing the definition of ensembles of ‘CFT data’. These ensembles take on the same role as the ensembles of random Hamiltonians in more conventional quantum ergodic phases of many-body quantum systems. To describe individual members of the ensembles, we introduce the notion of approximate CFT, defined as a collection of ‘CFT data’ satisfying the usual CFT constraints approximately, i.e. up to small deviations. We show that they generically exist by providing concrete examples. Ensembles of approximate CFTs are very natural in holography, as every member of the ensemble is indistinguishable from a true CFT for low-energy probes that only have access to information from semi-classical gravity. To specify these ensembles, we impose successively higher moments of the CFT constraints. Lastly, we propose a theory of pure gravity in AdS 3 as a random matrix/tensor model implementing approximate CFT constraints. This tensor model is the maximum ignorance ensemble compatible with conformal symmetry, crossing invariance, and a primary gap to the black-hole threshold. The resulting theory is a random matrix/tensor model governed by the Virasoro 6j-symbol.

Physics

Hybrid Data Assimilation without Ensemble Filtering

The Global Modeling and Assimilation Office is preparing to upgrade its three-dimensional variational system to a hybrid approach in which the ensemble is generated using a square-root ensemble Kalman filter (EnKF) and the variational problem is solved using the Grid-point Statistical Interpolation system. As in most EnKF applications, we found it necessary to employ a combination of multiplicative and additive inflations, to compensate for sampling and modeling errors, respectively and, to maintain the small-member ensemble solution close to the variational solution; we also found it necessary to re-center the members of the ensemble about the variational analysis. During tuning of the filter we have found re-centering and additive inflation to play a considerably larger role than expected, particularly in a dual-resolution context when the variational analysis is ran at larger resolution than the ensemble. This led us to consider a hybrid strategy in which the members of the ensemble are generated by simply converting the variational analysis to the resolution of the ensemble and applying additive inflation, thus bypassing the EnKF. Comparisons of this, so-called, filter-free hybrid procedure with an EnKF-based hybrid procedure and a control non-hybrid, traditional, scheme show both hybrid strategies to provide equally significant improvement over the control; more interestingly, the filter-free procedure was found to give qualitatively similar results to the EnKF-based procedure.

Kalman Filter

Challenges and alternatives to empirical orthogonal functions for earth system data

Empirical orthogonal functions (EOFs) applied to gridded Earth system data enables users to diagnose modes of variability with relative ease. Yet, many challenges to interpretation exist such that they must be used with awareness and intention when applied to gridded climate data, especially with large ensembles. Utilizing data from two different Earth system modelling large ensemble frameworks, the Energy Exoscale Earth System Model and the Community Earth System Model, as well as reanalysis data, common EOF pitfalls are summarized and discussed. Challenges include erroneous mode swapping, sign flipping, and the temporal variability of the centers of action. For modes of variability with similar contribution to variance, mode swapping is not uncommon. Sign flipping can occur with almost any mode where the pattern is correct, but the sign is arbitrary. Although the variability of the center of action is not necessarily problematic, it potentially complicates interpretation over multi-century timescales. A wide variety of alternative methods to EOFs exist, but fitness-for-purpose must be evaluated. Additionally, illustrations of alternative methods and examples of proper use are provided. Alternative methods fit into three categories: EOF variants, linear methods, and multilinear methods.

54 ENVIRONMENTAL SCIENCES

Application of an Ensemble Smoother to Precipitation Assimilation

Assimilation of precipitation in a global modeling system poses a special challenge in that the observation operators for precipitation processes are highly nonlinear. In the variational approach, substantial development work and model simplifications are required to include precipitation-related physical processes in the tangent linear model and its adjoint. An ensemble based data assimilation algorithm "Maximum Likelihood Ensemble Smoother (MLES)" has been developed to explore the ensemble representation of the precipitation observation operator with nonlinear convection and large-scale moist physics. An ensemble assimilation system based on the NASA GEOS-5 GCM has been constructed to assimilate satellite precipitation data within the MLES framework. The configuration of the smoother takes the time dimension into account for the relationship between state variables and observable rainfall. The full nonlinear forward model ensembles are used to represent components involving the observation operator and its transpose. Several assimilation experiments using satellite precipitation observations have been carried out to investigate the effectiveness of the ensemble representation of the nonlinear observation operator and the data impact of assimilating rain retrievals from the TMI and SSM/I sensors. Preliminary results show that this ensemble assimilation approach is capable of extracting information from nonlinear observations to improve the analysis and forecast if ensemble size is adequate, and a suitable localization scheme is applied. In addition to a dynamically consistent precipitation analysis, the assimilation system produces a statistical estimate of the analysis uncertainty.

Zhang, Sara

Geology of Southern Guinevere Planitia, Venus, based on analyses of Goldstone radar data

The ensemble of 41 backscatter images of Venus acquired by the S Band (12.6 cm) Goldstone radar system covers approx. 35 million km and includes the equatorial portion of Guinevere Planitia, Navka Planitia, Heng-O Chasma, and Tinatin Planitia, and parts of Devana Chasma and Phoebe Regio. The images and associated altimetry data combine relatively high spatial resolution (1 to 10 km) with small incidence angles (less than 10 deg) for regions not covered by either Venera Orbiter or Arecibo radar data. Systematic analyses of the Goldstone data show that: (1) Volcanic plains dominate, including groups of small volcanic constructs, radar bright flows on a NW-SE arm of Phoebe Regio and on Ushas Mons and circular volcano-tectonic depressions; (2) Some of the regions imaged by Goldstone have high radar cross sections, including the flows on Ushas Mons and the NW-SE arm of Phoebe Regio, and several other unnamed hills, ridged terrains, and plains areas; (3) A 1000 km diameter multiringed structure is observed and appears to have a morphology not observed in Venera data (The northern section corresponds to Heng-O Chasma); (4) A 150 km wide, 2 km deep, 1400 km long rift valley with upturned flanks is located on the western flank of Phoebe Regio and extends into Devana Chasma; (5) A number of structures can be discerned in the Goldstone data, mainly trending NW-SE and NE-SW, directions similar to those discerned in Pioneer-Venus topography throughout the equatorial region; and (6) The abundance of circular and impact features is similar to the plains global average defined from Venera and Arecibo data, implying that the terrain imaged by Goldstone has typical crater retention ages, measured in hundreds of millions of years. The rate of resurfacing is less than or equal to 4 km/Ga.

Arvidson, R. E.

Dimensionality Reduction Through Classifier Ensembles

In data mining, one often needs to analyze datasets with a very large number of attributes. Performing machine learning directly on such data sets is often impractical because of extensive run times, excessive complexity of the fitted model (often leading to overfitting), and the well-known "curse of dimensionality." In practice, to avoid such problems, feature selection and/or extraction are often used to reduce data dimensionality prior to the learning step. However, existing feature selection/extraction algorithms either evaluate features by their effectiveness across the entire data set or simply disregard class information altogether (e.g., principal component analysis). Furthermore, feature extraction algorithms such as principal components analysis create new features that are often meaningless to human users. In this article, we present input decimation, a method that provides "feature subsets" that are selected for their ability to discriminate among the classes. These features are subsequently used in ensembles of classifiers, yielding results superior to single classifiers, ensembles that use the full set of features, and ensembles based on principal component analysis on both real and synthetic datasets.

Oza, Nikunj C.

Using feature importance as an exploratory data analysis tool on Earth system models

Abstract. Machine learning (ML) models are commonly used to generate predictions, but these models can also support the discovery of new science. Generating accurate predictions necessitates that a model captures the structure of the underlying data. If the structure is properly extracted, ML could be a useful exploratory and evidential tool. In this paper, we present a case study that demonstrates the use of ML for exploratory data analysis (EDA) in the climate space. We apply the ML explainability method of spatiotemporal zeroed feature importance (stZFI) to understand how climate-variable associations evolve over space and time. Our analyses focus on data from ensembles of Earth system models (ESMs) which provide data on different climate states and conditions. We elect to work with ESM ensembles since they allow us to compare feature importance across alternative scenarios not available with observed data. The ensembles also account for natural variability so that we can distinguish between signal and noise due to natural climate variability when computing feature importance. The use of perturbed initial condition ensembles introduces variability mimicking the natural variability in the atmosphere; thus the signals emerging using feature importance (FI) can be evaluated against the natural variability in the climate system. For our analyses, we consider the 1991 volcanic eruption of Mount Pinatubo, which was a large stratospheric aerosol injection. We explore the climate pathway associated with the eruption from aerosols to radiation to temperature at both the near-surface and stratospheric levels. In addition to applying the method to data generated from two different ESMs, we apply stZFI to reanalysis data to compare the associations identified by stZFI. We show how stZFI tracks the importance of aerosol optical depth over time on forecasting temperatures. This case study illustrates usefulness of an ML tool (stZFI) for EDA on a well-studied climate exemplar.

Ries, Daniel (ORCID:0000000250294647)

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment

Role of Forcing Uncertainty and Background Model Error Characterization in Snow Data Assimilation

Accurate specification of the model error covariances in data assimilation systems is a challenging issue. Ensemble land data assimilation methods rely on stochastic perturbations of input forcing and model prognostic fields for developing representations of input model error covariances. This article examines the limitations of using a single forcing dataset for specifying forcing uncertainty inputs for assimilating snow depth retrievals. Using an idealized data assimilation experiment, the article demonstrates that the use of hybrid forcing input strategies (either through the use of an ensemble of forcing products or through the added use of the forcing climatology) provide a better characterization of the background model error, which leads to improved data assimilation results, especially during the snow accumulation and melt-time periods. The use of hybrid forcing ensembles is then employed for assimilating snow depth retrievals from the AMSR2 (Advanced Microwave Scanning Radiometer 2) instrument over two domains in the continental USA with different snow evolution characteristics. Over a region near the Great Lakes, where the snow evolution tends to be ephemeral, the use of hybrid forcing ensembles provides significant improvements relative to the use of a single forcing dataset. Over the Colorado headwaters characterized by large snow accumulation, the impact of using the forcing ensemble is less prominent and is largely limited to the snow transition time periods. The results of the article demonstrate that improving the background model error through the use of a forcing ensemble enables the assimilation system to better incorporate the observational information.

assimilation

Ensemble Kalman filter for data assimilation coupled with low-resolution computations techniques applied in fluid dynamics

This paper presents an innovative Reduced-order model (ROM) for merging experimental and simulation data using data assimilation (DA) to estimate the "True" state of a fluid dynamics system, leading to more accurate predictions. Our methodology introduces a novel approach by implementing the ensemble Kalman filter (EnKF) within a reduced-dimensional framework, grounded in a robust theoretical foundation and applied to fluid dynamics. To address the substantial computational demands of DA, the proposed ROM employs low-resolution (LR) techniques to drastically reduce computational costs. This innovative approach involves downsampling datasets for DA computations, followed by an advanced reconstruction technique based on low-cost singular value decomposition (lcSVD). The lcSVD method, a key innovation in this paper, has never been applied to DA before and offers a highly efficient way to enhance resolution with minimal computational resources. Our results demonstrate significant reductions in both computation time and RAM usage through these LR techniques without compromising the accuracy of the estimations. For instance, in a turbulent test case, for a data compression rate of 15.9, the LR approach can achieve a speed-up of 13.7 and a RAM compression of 90.9% while maintaining a low relative root mean square error (RRMSE) of 2.6%, compared to 0.8% in the high-resolution (HR) reference. Furthermore, we highlight the effectiveness of the EnKF in estimating and predicting the state of fluid flow systems based on limited observations and given low-fidelity numerical data. This paper highlights the potential of the proposed DA method in fluid dynamics applications, particularly for improving computational efficiency in CFD and related fields. Its ability to balance accuracy with low computational and memory costs makes it especially suitable for large-scale and real-time applications, such as environmental monitoring or engineering design. This method will be incorporated into ModelFLOWs-app.

Data Assimilation

A test of a cumulus parameterization model using the GATE data

Two parametric, ensemble cloud models were tested with data obtained during GATE (GARP Atlantic Tropical Experiment) (1974). The first model is an adaption of the Arakawa-Schubert scheme which uses an entraining jet to represent individual cumulus cloud types. The second model consists of an ensemble of cylindrical cells to represent the convective cloud field.

Rodenhuis, D.

AeroCom Phase III Multi-Model Evaluation of the Aerosol Life Cycle and Optical Properties Using Ground and Space-Based Remote Sensing as Well as Surface In Situ Observations

Within the framework of the AeroCom (Aerosol Comparisons between Observations and Models) initiative, the state-of-the-art modelling of aerosol optical properties is assessed from 14 global models participating in the phase III control experiment (AP3). The models are similar to CMIP6/AerChemMIP Earth System Models (ESMs) and provide a robust multi-model ensemble. Inter-model spread of aerosol species lifetimes and emissions appears to be similar to that of mass extinction coefficients (MECs), suggesting that aerosol optical depth (AOD) uncertainties are associated with a broad spectrum of parameterised aerosol processes. Total AOD is approximately the same as in AeroCom phase I (AP1) simulations. However, we find a 50 % decrease in the optical depth (OD) of black carbon (BC), attributable to a combination of decreased emissions and lifetimes. Relative contributions from sea salt (SS) and dust (DU) have shifted from being approximately equal in AP1 to SS contributing about 2∕3 of the natural AOD in AP3. This shift is linked with a decrease in DU mass burden, a lower DU MEC, and a slight decrease in DU lifetime, suggesting coarser DU particle sizes in AP3 compared to AP1. Relative to observations, the AP3 ensemble median and most of the participating models underestimate all aerosol optical properties investigated, that is, total AOD as well as fine and coarse AOD (AODf, AODc), Ångström exponent (AE), dry surface scattering (SCdry), and absorption (ACdry) coefficients. Compared to AERONET, the models underestimate total AOD by ca. 21 % ± 20 % (as inferred from the ensemble median and interquartile range). Against satellite data, the ensemble AOD biases range from −37 % (MODIS-Terra) to −16 % (MERGED-FMI, a multi-satellite AOD product), which we explain by differences between individual satellites and AERONET measurements themselves. Correlation coefficients (R) between model and observation AOD records are generally high (R>0.75), suggesting that the models are capable of capturing spatio-temporal variations in AOD. We find a much larger underestimate in coarse AODc (∼ −45 % ± 25 %) than in fine AODf (∼ −15 % ± 25 %) with slightly increased inter-model spread compared to total AOD. These results indicate problems in the modelling of DU and SS. The AODc bias is likely due to missing DU over continental land masses (particularly over the United States, SE Asia, and S. America), while marine AERONET sites and the AATSR SU satellite data suggest more moderate oceanic biases in AODc. Column AEs are underestimated by about 10 % ± 16 %. For situations in which measurements show AE > 2, models underestimate AERONET AE by ca. 35 %. In contrast, all models (but one) exhibit large overestimates in AE when coarse aerosol dominates (bias ca. +140 % if observed AE < 0.5). Simulated AE does not span the observed AE variability. These results indicate that models overestimate particle size (or underestimate the fine-mode fraction) for fine-dominated aerosol and underestimate size (or overestimate the fine-mode fraction) for coarse-dominated aerosol. This must have implications for lifetime, water uptake, scattering enhancement, and the aerosol radiative effect, which we can not quantify at this moment. Comparison against Global Atmosphere Watch (GAW) in situ data results in mean bias and inter-model variations of −35 % ± 25 % and −20 % ± 18 % for SCdry and ACdry, respectively. The larger underestimate of SCdry than ACdry suggests the models will simulate an aerosol single scattering albedo that is too low. The larger underestimate of SCdry than ambient air AOD is consistent with recent findings that models overestimate scattering enhancement due to hygroscopic growth. The broadly consistent negative bias in AOD and surface scattering suggests an underestimate of aerosol radiative effects in current global aerosol models. Considerable inter-model diversity in the simulated optical properties is often found in regions that are, unfortunately, not or only sparsely covered by ground-based observations. This includes, for instance, the Sahara, Amazonia, central Australia, and the South Pacific. This highlights the need for a better site coverage in the observations, which would enable us to better assess the models, but also the performance of satellite products in these regions. Using fine-mode AOD as a proxy for present-day aerosol forcing estimates, our results suggest that models underestimate aerosol forcing by ca. −15 %, however, with a considerably large interquartile range, suggesting a spread between −35 % and +10 %.

space-based remote sensing

Flying-hot-wire study of two-dimensional mean flow past an NACA 4412 airfoil at maximum lift

Hot-wire measurements have been made in the boundary layer, the separated region, and the near wake for flow past an NACA 4412 airfoil at maximum lift. The Reynolds number based on chord was about 1,500,000. The main instrumentation was a hot-wire probe mounted on the end of a rotating arm. A digital computer was used to control synchronized sampling of hot-wire data at closely spaced points along the probe arc. Ensembles of data were obtained at several thousand locations in the flow field. The data include intermittency, two components of mean velocity, and twelve mean values for double, triple, and quadruple products of two velocity fluctuations. The data are available on punched cards in raw form and also after use of smoothing and interpolation routines to obtain values on a fine rectangular grid aligned with the airfoil chord. The data are displayed in the paper as contour plots.

Coles, D.

Parameterization-Induced Uncertainties and Impacts of Crop Management Harmonization in a Global Gridded Crop Model Ensemble

Global gridded crop models (GGCMs) combine agronomic or plant growth models with gridded spatial input data to estimate spatially explicit crop yields and agricultural externalities at the global scale. Differences in GGCM outputs arise from the use of different biophysical models, setups, and input data. GGCM ensembles are frequently employed to bracket uncertainties in impact studies without investigating the causes of divergence in outputs. This study explores differences in maize yield estimates from five GGCMs based on the public domain field-scale model Environmental Policy Integrated Climate (EPIC) that participate in the AgMIP Global Gridded Crop Model Intercomparison initiative. Albeit using the same crop model, the GGCMs differ in model version, input data, management assumptions, parameterization, and selection of subroutines affecting crop yield estimates via cultivar distributions, soil attributes, and hydrology among others. The analyses reveal inter-annual yield variability and absolute yield levels in the EPIC-based GGCMs to be highly sensitive to soil parameterization and crop management. All GGCMs show an intermediate performance in reproducing reported yields with a higher skill if a static soil profile is assumed or sufficient plant nutrients are supplied. An in-depth comparison of setup domains for two EPIC-based GGCMs shows that GGCM performance and plant stress responses depend substantially on soil parameters and soil process parameterization, i.e. hydrology and nutrient turnover, indicating that these often neglected domains deserve more scrutiny. For agricultural impact assessments, employing a GGCM ensemble with its widely varying assumptions in setups appears the best solution for coping with uncertainties from lack of comprehensive global data on crop management, cultivar distributions and coefficients for agro-environmental processes. However, the underlying assumptions require systematic specifications to cover representative agricultural systems and environmental conditions. Furthermore, the interlinkage of parameter sensitivity from various domains such as soil parameters, nutrient turnover coefficients, and cultivar specifications highlights that global sensitivity analyses and calibration need to be performed in an integrated manner to avoid bias resulting from disregarded core model domains. Finally, relating evaluations of the EPIC-based GGCMs to a wider ensemble based on individual core models shows that structural differences outweigh in general differences in configurations of GGCMs based on the same model, and that the ensemble mean gains higher skill from the inclusion of structurally different GGCMs. Although the members of the wider ensemble herein do not consider crop-soil-management interactions, their sensitivity to nutrient supply indicates that findings for the EPIC-based sub-ensemble will likely become relevant for other GGCMs with the progressing inclusion of such processes.

Folberth, Christian

Principle Component Analysis of AIRS and CrIS Data

Synthetic Eigen Vectors (EV) used for the statistical analysis of the PC reconstruction residual of large ensembles of data are a novel tool for the analysis of data from hyperspectral infrared sounders like the Atmospheric Infrared Sounder (AIRS) on the EOS Aqua and the Cross-track Infrared Sounder (CrIS) on the SUOMI polar orbiting satellites. Unlike empirical EV, which are derived from the observed spectra, the synthetic EV are derived from a large ensemble of spectra which are calculated assuming that, given a state of the atmosphere, the spectra created by the instrument can be accurately calculated. The synthetic EV are then used to reconstruct the observed spectra. The analysis of the differences between the observed spectra and the reconstructed spectra for Simultaneous Nadir Overpasses of tropical oceans reveals unexpected differences at the more than 200 mK level under relatively clear conditions, particularly in the mid-wave water vapor channels of CrIS. The repeatability of these differences using independently trained SEV and results from different years appears to rule out inconsistencies in the radiative transfer algorithm or the data simulation. The reasons for these discrepancies are under evaluation.

infrared

High resolution wind measurements for offshore wind energy development

A method, apparatus, system, article of manufacture, and computer readable storage medium provide the ability to measure wind. Data at a first resolution (i.e., low resolution data) is collected by a satellite scatterometer. Thin slices of the data are determined. A collocation of the data slices are determined at each grid cell center to obtain ensembles of collocated data slices. Each ensemble of collocated data slices is decomposed into a mean part and a fluctuating part. The data is reconstructed at a second resolution from the mean part and a residue of the fluctuating part. A wind measurement is determined from the data at the second resolution using a wind model function. A description of the wind measurement is output.

Nghiem, Son Van

Ocean Surface Vector Wind: Research Challenges and Operational Opportunities

The atmosphere and ocean are joined together over seventy percent of Earth, with ocean surface vector wind (OSVW) stress one of the linkages. Satellite OSVW measurements provide estimates of wind divergence at the bottom of the atmosphere and wind stress curl at the top of the ocean; both variables are critical for weather and climate applications. As is common with satellite measurements, a multitude of OSVW data products exist for each currently operating satellite instrument. In 2012 the Joint Technical Commission on Oceanography and Marine Meteorology (JCOMM) launched an initiative to coordinate production of OSVW data products to maximize the impact and benefit of existing and future OSVW measurements in atmospheric and oceanic applications. This paper describes meteorological and oceanographic requirements for OSVW data products; provides an inventory of unique data products to illustrate that the challenge is not the production of individual data products, but the generation of harmonized datasets for analysis and synthesis of the ensemble of data products; and outlines a vision for JCOMM, in partnership with other international groups, to assemble an international network to share ideas, data, tools, strategies, and deliverables to improve utilization of satellite OSVW data products for research and operational applications.

Joint Technical Commission on Oceanography and Mar