Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Analyzing Tropical Waves Using the Parallel Ensemble Empirical Model Decomposition Method: Preliminary Results from Hurricane Sandy

In this study, we discuss the performance of the parallel ensemble empirical mode decomposition (EMD) in the analysis of tropical waves that are associated with tropical cyclone (TC) formation. To efficiently analyze high-resolution, global, multiple-dimensional data sets, we first implement multilevel parallelism into the ensemble EMD (EEMD) and obtain a parallel speedup of 720 using 200 eight-core processors. We then apply the parallel EEMD (PEEMD) to extract the intrinsic mode functions (IMFs) from preselected data sets that represent (1) idealized tropical waves and (2) large-scale environmental flows associated with Hurricane Sandy (2012). Results indicate that the PEEMD is efficient and effective in revealing the major wave characteristics of the data, such as wavelengths and periods, by sifting out the dominant (wave) components. This approach has a potential for hurricane climate study by examining the statistical relationship between tropical waves and TC formation.

PEEMD↗

Program for narrow-band analysis of aircraft flyover noise using ensemble averaging techniques

A package of computer programs was developed for analyzing acoustic data from an aircraft flyover. The package assumes the aircraft is flying at constant altitude and constant velocity in a fixed attitude over a linear array of ground microphones. Aircraft position is provided by radar and an option exists for including the effects of the aircraft's rigid-body attitude relative to the flight path. Time synchronization between radar and acoustic recording stations permits ensemble averaging techniques to be applied to the acoustic data thereby increasing the statistical accuracy of the acoustic results. Measured layered meteorological data obtained during the flyovers are used to compute propagation effects through the atmosphere. Final results are narrow-band spectra and directivities corrected for the flight environment to an equivalent static condition at a specified radius.

Gridley, D.↗

Improving the Representation of Land Surface Processes using the Data Assimilation Research Testbed (DART)

The land surface is a critical part of the earth system as processes related to water, carbon, energy and nitrogen cycling have important implications for climate forcing, air quality, water availability and seasonal atmospheric forecasting. Despite advances in land surface modeling, land surface model performance is often limited because of errors related to initial and boundary conditions, model structure, and parameters. Data assimilation (DA) techniques combined with an expanding network of earth system observations present an opportunity to reduce these errors and improve simulations. Here we apply an Ensemble Kalman Filter DA system as part of the Data Assimilation Research Testbed to a variety of land surface simulations. First, we describe the use of remotely sensed biomass observations to provide improved simulations of plant phenology, carbon and water cycling for regions highly sensitive to climate change (Western US, China, and Arctic). We discuss approaches to account for systemic biases between models and observations, including the use of spatially-varying adaptive ensemble inflation as an alternative approach to re-scaling soil moisture observations. Finally, we discuss a strategy to incorporate complementary observations (snow water equivalent, solar-induced fluorescence) to better constrain the representation of carbon and water cycling across complex terrain.

DART↗

Analyzing Non Stationary Processes in Radiometers

The lack of well-developed techniques for modeling changing statistical moments in our observations has stymied the application of stochastic process theory for many scientific and engineering applications. Non linear effects of the observation methodology is one of the most perplexing aspects to modeling non stationary processes. This perplexing problem was encountered when modeling the effect of non stationary receiver fluctuations on the performance of radiometer calibration architectures. Existing modeling approaches were found not applicable; particularly problematic is modeling processes across scales over which they begin to exhibit non stationary behavior within the time interval of the calibration algorithm. Alternatively, the radiometer output is modeled as samples from a sequence random variables; the random variables are treated using a conditional probability distribution function conditioned on the use of the variable in the calibration algorithm. This approach of treating a process as a sequence of random variables with non stationary stochastic moments produce sensible predictions of temporal effects of calibration algorithms. To test these model predictions, an experiment using the Millimeter wave Imaging Radiometer (MIR) was conducted. The MIR with its two black body calibration references was configured in a laboratory setting to observe a third ultra-stable reference (CryoTarget). The MIR was programmed to sequentially sample each of the three references in approximately a 1 second cycle. Data were collected over a six-hour interval. The sequence of reference measurements form an ensemble sample set comprised of a series of three reference measurements. Two references are required to estimate the receiver response. A third reference is used to estimate the uncertainty in the estimate. Typically, calibration algorithms are designed to suppress the non stationary effects of receiver fluctuations. By treating the data sequence as an ensemble collection, it is possible to apply temporal algorithms which exacerbate the non stationary effects. By varying the algorithm, information about the properties of the non stationary receiver fluctuations is obtained. Comparisons of analytical calculations and statistical analysis of data demonstrate impressive agreement.

Racette, Paul↗

The structure of the vorticity field in turbulent channel flow. Part 2: Study of ensemble-averaged fields

Several conditional sampling techniques are applied to a data base generated by large-eddy simulation of turbulent channel flow. It is shown that the bursting process is associated with well-organized horseshoe vortices inclined at about 45 deg. to the wall. These vortical structures are identified by examining the vortex lines of three-dimensional, ensemble averaged vorticity fields. Two distinct horseshoe-shaped vortices corresponding to the sweep and ejection events are detected. These vortices are associated with high Reynolds shear stress and hence make a significant contribution to turbulent energy production. The dependency of the ensemble averaged vortical structures on the detection criteria, and the question of whether this ensemble-averaged structure is an artifact of the ensemble averaging process are examined. The ensemble-averaged pattern of these vortical structures that emerge from the analysis could provide the basis for a hypothetical model of the organized structures of wall-bounded shear flows.

Kim, J.↗

The structure of the vorticity field in turbulent channel flow. II - Study of ensemble-averaged fields

Several conditional sampling techniques are applied to a data base generated by large-eddy simulation of turbulent channel flow. It is shown that the bursting process is associated with well-organized horseshoe vortices inclined at about 45 deg to the wall. These vortical structures are identified by examining the vortex lines of three-dimensional, ensemble averaged vorticity fields. Two distinct horseshoe-shaped vortices corresponding to the sweep and ejection events are detected. These vortices are associated with high Reynolds shear stress and hence make a significant contribution to turbulent energy production. The dependency of the ensemble averaged vortical structures on the detection criteria, and the question of whether this ensemble-averaged structure is an artifact of the ensemble averaging process are examined. The ensemble-averaged pattern of these vortical structures that emerge from the analysis could provide the basis for a hypothetical model of the organized structures of wall-bounded shear flows.

Kim, J.↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Martian Polar Vortices: Comparison of Reanalyses

The structure and evolution of the Martian polar vortices is examined using two recently available reanalysis systems: version 1.0 of the Mars Analysis Correction Data Assimilation (MACDA) and a preliminary version of the Ensemble Mars Atmosphere Reanalysis System (EMARS). There is quantitative agreement between the reanalyses in the lower atmosphere, where Mars Global Surveyor (MGS) Thermal Emission Spectrometer (TES) data are assimilated, but there are differences at higher altitudes reflecting differences in the free-running general circulation model simulations used in the two reanalyses. The reanalyses show similar potential vorticity (PV) structure of the vortices: There is near-uniform small PV equatorward of the core of the westerly jet, steep meridional PV gradients on the polar side of the jet core, and a maximum of PV located off of the pole. In maps of 30 sol mean PV, there is a near-continuous elliptical ring of high PV with roughly constant shape and longitudinal orientation from fall to spring. However, the shape and orientation of the vortex varies on daily time scales, and there is not a continuous ring of PV but rather a series of smaller scale coherent regions of high PV. The PV structure of the Martian polar vortices is, as has been reported before, very different from that of Earth's stratospheric polar vortices, but there are similarities with Earth's tropospheric vortices which also occur at the edge of the Hadley Cell, and have near-uniform small PV equatorward of the jet, and a large increase of PV poleward of the jet due to increased stratification.

Waugh, D. W.↗

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto↗

Popnet : computer vision based deep learning model for forecasting gridded population

Here, this study introduces Popnet, a deep learning model for forecasting 1 km-gridded populations, integrating U-Net, ConvLSTM, a Spatial Autocorrelation module and deep ensemble methods. Using spatial variables and population data from 2000 to 2020, Popnet predicts South Korea’s population trends by age groups (under 14, 15-64 and over 65) up to 2040. In validation, it outperforms traditional machine learning and state-of-the-art computer vision models. The output of this model discovered significant polarisation: population growth in urban areas, especially the capital region, and severe depopulation in rural areas. Popnet is a robust tool for offering significant insights to policymakers and related stakeholders about the detailed future population, which allows them to establish detailed, localised planning and resource allocations.

computer vision↗

Error Estimation of An Ensemble Statistical Seasonal Precipitation Prediction Model

This NASA Technical Memorandum describes an optimal ensemble canonical correlation forecasting model for seasonal precipitation. Each individual forecast is based on the canonical correlation analysis (CCA) in the spectral spaces whose bases are empirical orthogonal functions (EOF). The optimal weights in the ensemble forecasting crucially depend on the mean square error of each individual forecast. An estimate of the mean square error of a CCA prediction is made also using the spectral method. The error is decomposed onto EOFs of the predictand and decreases linearly according to the correlation between the predictor and predictand. Since new CCA scheme is derived for continuous fields of predictor and predictand, an area-factor is automatically included. Thus our model is an improvement of the spectral CCA scheme of Barnett and Preisendorfer. The improvements include (1) the use of area-factor, (2) the estimation of prediction error, and (3) the optimal ensemble of multiple forecasts. The new CCA model is applied to the seasonal forecasting of the United States (US) precipitation field. The predictor is the sea surface temperature (SST). The US Climate Prediction Center's reconstructed SST is used as the predictor's historical data. The US National Center for Environmental Prediction's optimally interpolated precipitation (1951-2000) is used as the predictand's historical data. Our forecast experiments show that the new ensemble canonical correlation scheme renders a reasonable forecasting skill. For example, when using September-October-November SST to predict the next season December-January-February precipitation, the spatial pattern correlation between the observed and predicted are positive in 46 years among the 50 years of experiments. The positive correlations are close to or greater than 0.4 in 29 years, which indicates excellent performance of the forecasting model. The forecasting skill can be further enhanced when several predictors are used.

Shen, Samuel S. P.↗

Development and Implementation of a Comprehensive Radiometric Validation Protocol for the CERES Earth Radiation Budget Climate Record Sensors

The CERES Flight Models 1 through 4 instruments were launched aboard NASA's Earth Observing System (EOS) Terra and Aqua Spacecraft into 705 Km sun-synchronous orbits with 10:30 a.m. and 1:30 p.m. equatorial crossing times. These instruments supplement measurements made by the CERES Proto Flight Model (PFM) instrument launched aboard NASA's Tropical Rainfall Measuring Mission (TRMM) into a 350 Km, 38-degree mid-inclined orbit. CERES Climate Data Records consist of geolocated and calibrated instantaneous filtered and unfiltered radiances through temporally and spatially averaged TOA, Surface and Atmospheric fluxes. CERES filtered radiance measurements cover three spectral bands including shortwave (0.3 to 5 microns), total (0.3 to 100 microns) and an atmospheric window channel (8 to 12 microns). The CERES Earth Radiation Budget measurements represent a new era in radiation climate data, realizing a factor of 2 to 4 improvement in calibration accuracy and stability over the previous ERBE climate records, while striving for the next goal of 0.3-percent per decade absolute stability. The current improvement is derived from two sources: the incorporation of lessons learned from the ERBE mission in the design of the CERES instruments and the development of a rigorous and comprehensive radiometric validation protocol consisting of individual studies covering different spatial, spectral and temporal time scales on data collected both pre and post launch. Once this ensemble of individual perspectives is collected and organized, a cohesive and highly rigorous picture of the overall end-to-end performance of the CERES instrument's and data processing algorithms may be clearly established. This approach has resulted in unprecedented levels of accuracy for radiation budget instruments and data products with calibration stability of better than 0.2-percent and calibration traceability from ground to flight of 0.25-percent. The current work summarizes the development, philosophy and implementation of the protocol designed to rigorously quantify the quality of the data products as well as the level of agreement between the CERES TRMM, Terra and Aqua climate data records.

Priestley, K. J.↗

Testing Classical Properties from Quantum Data

Many properties of Boolean functions can be tested far more efficiently than the function itself can be learned. However, this dramatic advantage often disappears when testers are limited to random samples of ƒ instead of adaptively chosen queries to f. In this work we investigate the quantum version of this restriction: quantum algorithms that test properties of a Boolean function f solely from copies of either the function state |ƒ⟩ ∝ ∑ x |x, ƒ(x)⟩ or the phase state |(-1) ƒ ⟩ ∝ ∑ x (-1) ƒ(x) |x⟩. For monotonicity, symmetry, and triangle-freeness, we show passive quantum testers are unboundedly or super-polynomially better than their classical passive testing counterparts. They are competitive with classic query -based testers in each case. Our new testers use techniques beyond quantum Fourier sampling, and it turns out this is necessary: we show a certain class of bent functions can be tested from 𝒪(1) function states but has a sample complexity lower bound of 2 Ω(n) for any tester relying exclusively on Fourier and classical samples. Our passive quantum testers are competitive with classical query -based testers, but this isn't universal: we exhibit a testing problem that can be solved from 𝒪(1) classical queries but requires Ω(2 n/2 ) function state copies. The Forrelation problem provides a separation of the same magnitude in the opposite direction, so we conclude that quantum data and classical queries are "maximally incomparable" resources for testing. We also begin the study of lower bounds for testing from quantum data. For quantum monotonicity testing, we prove that the ensembles of [Goldreich et al., 2000; Black, 2024], which give exponential lower bounds for classical sample-based testing, do not yield any nontrivial lower bounds for testing from quantum data. New insights specific to quantum data will be required for proving copy complexity lower bounds for testing in this model.

Boolean Functions↗

Investigation of Particle Sampling Bias in the Shear Flow Field Downstream of a Backward Facing Step

The flow field about a backward facing step was investigated to determine the characteristics of particle sampling bias in the various flow phenomena. The investigation used the calculation of the velocity:data rate correlation coefficient as a measure of statistical dependence and thus the degree of velocity bias. While the investigation found negligible dependence within the free stream region, increased dependence was found within the boundary and shear layers. Full classic correction techniques over-compensated the data since the dependence was weak, even in the boundary layer and shear regions. The paper emphasizes the necessity to determine the degree of particle sampling bias for each measurement ensemble and not use generalized assumptions to correct the data. Further, it recommends the calculation of the velocity:data rate correlation coefficient become a standard statistical calculation in the analysis of all laser velocimeter data.

Meyers, James F.↗

SAGE III/ISS Rapid Data Analysis Through Dashboarding with Jupyter Notebooks

Spaceborne remote sensing observations of Earth’s atmosphere produce significant quantities of data over the life of each mission. In the case of the Stratospheric Aerosol and Gas Experiment III on the International Space Station (SAGE III/ISS) nearly four years of vertical profiles of atmospheric ozone, water vapor, and nitrogen dioxide concentrations as well as aerosol extinction coefficients have been released. The dichotomy of the desire for both long-term trends in the atmospheric state alongside the assessment of short-term impacts of major disruptive events such as volcanic eruptions and pyrocumulus injections requires agile tools to handle these cases in near real-time as new data are produced. The analysis landscape is further complicated by the desire to compare results between the numerous contemporary observations available for a given dataset. The SAGE III/ISS team has developed a suite of tools leveraging modern web-based frameworks allowing members to interact with a dashboard-style interface to load the data record, assess new profiles as they are generated and in ensemble, compare between species, and additionally add in measurements observed by other platforms as necessary. Leveraging a commonly packaged data format of NetCDF alongside the Python Jupyter Notebook framework, the data can be served to interested parties from an analysis server while still runnable on personal systems if required. This presentation illustrates the ecosystem developed by the SAGE III/ISS team, the applicability to measurements made by any limb-observing platform, and the benefit to transforming routine analyses into readily accessible dynamic plots. Frameworks currently exist at larger scales with projects such as GIOVANNI, and this illustration seeks to show that similar frameworks are accessible and possible within the local research environment while simultaneously unloading human processing cycles for more specialized analysis tasks.

Dashboarding↗

Economic Impact Assessments (EIA) of application of GEOGLOWS in Ecuador: Data Gaps, Limitations and Recommendations

In 2022, the United Nations launched the Early Warnings for All (EW4ALL) Program to establish global early warning systems by 2027. To assess the impact of the substantial $3.1 billion annual investment over five years, EW4ALL will consider factors that will require national coordination for the data needed for these assessments. In 2023, Ecuador was identified as one of the world's most climate-vulnerable countries, emphasizing the need to enhance its early warning systems. In 2020, the SERVIR Amazonia hub implemented the GEOGLOWS streamflow forecast service in collaboration with Ecuador's national meteorological agency (INAMHI). GEOGloWS provides 15-day ensemble forecasts and 80 years of historical streamflow data for every river worldwide through a free web service. The World Meteorological Organization has recognized this initiative as essential in contributing to the UN's call to ensure an 'Early Warning for All' by 2027. In 2023, as part of NASA's continuous efforts to fund research for Policy-Relevant Implementations, an economic impact assessment (EIA) was performed to understand the potential socioeconomic benefits of Early streamflow predictions in Ecuador using the GEOGLOWS service. Preliminary findings highlighted that gaps remain in effectively integrating socioeconomic and Earth observation (EO) data to capture the total value of these predictions. Implementing GEOGLOWS has led to valuable hydrological forecasts; however, the total economic benefits have yet to be documented. This study addresses the gaps and makes recommendations for future work that should focus on capturing the socio-economic benefits and costs associated with these forecasts, including their impact on decision-making at national and local levels. Despite the daily use of GEOGLOWS by key figures, including the President of Ecuador, the need for comprehensive recommendations and assessments is urgent.

Reetwika Basu↗

Comparison of CNN-Based Image Classification Approaches for Implementation of Low-Cost Multispectral Arcing Detection

Camera-based sensing has benefited in recent years from developments in machine learning data processing methods, as well as improved data collection options such as Unmanned Aerial Vehicles (UAV) mounted sensors. However, cost considerations, both for the initial purchase of sensors as well as updates, maintenance, or potential replacement if damaged, can limit adoption of more expensive sensing options for some applications. To evaluate more affordable options with less expensive, more available, and more easily replaceable hardware, we examine the use of machine learning-based image classification with custom datasets, utilizing deep learning based-image classification and the use of ensemble models for sensor fusion. Utilizing the same models for each camera to reduce technical overhead, we showed that for a very representative training dataset, camera-based detection can be successful for detection of electrical arcing. We also use multiple validation datasets, based on conditions expected to be of varying difficulty, to evaluate custom data. These results show that ensemble models of different data sources can mitigate risks from gaps in training data, though the system will be less redundant for those cases unless other precautions are taken. We found that with good quality custom datasets, data fusion models can be utilized without specialization in design to the specific cameras utilized, allowing for less specialized, more accessible equipment to be utilized as multispectral camera components. This approach can provide an alternative to expensive sensing equipment for applications in which lower-cost or more easily replaceable sensing equipment is desirable.

convolutional neural networks↗

Resolving Mesoscale Convective Systems: Grid Spacing Sensitivity in the Tropics and Midlatitudes

Abstract Mesoscale convective systems (MCSs) are a critical global water cycle component and drive extreme precipitation events in tropical and midlatitude regions. However, simulating deep convection remains challenging for modern numerical weather and climate models due to the complex interactions of processes from microscales to synoptic scales. Recent models with kilometer‐scale horizontal grid spacings offer notable improvements in simulating deep convection compared to coarser‐resolution models. Still, deficiencies in representing key physical processes, such as entrainment, lead to systematic biases. Additionally, evaluating model outputs using process‐oriented observational data remain difficult. This study presents an ensemble of MCS simulations with spanning the deep convective gray zone ( from 12 km to 125 m) in the Southern Great Plains of the U.S. and the Amazon Basin. Comparing these simulations with Atmospheric Radiation Measurement (ARM) wind profiler observations, we find greater sensitivity in the Amazon Basin compared to the Great Plains. Convective drafts converge structurally at sub‐kilometer scales, but some deficiencies remain. In both regions, simulated up and downdrafts are too deep and extreme downdrafts are not strong enough. Furthermore, Amazonian updrafts are too strong. Overall, we observe higher sensitivity in the tropics, including an artificial buildup in vertical kinetic energy at scales of , suggesting a need for 250 m in this region. Nevertheless, bulk convergence—agreement of storm‐average statistics—is achievable with kilometer‐scale simulations within a 10% error margin with 1 km providing a good balance between accuracy and computational cost.

54 ENVIRONMENTAL SCIENCES↗