Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

EFIT-Prime: Probabilistic and physics-constrained reduced-order neural network model for equilibrium reconstruction in DIII-D

We introduce EFIT-Prime, a novel machine learning surrogate model for EFIT (Equilibrium FIT) that integrates probabilistic and physics-informed methodologies to overcome typical limitations associated with deterministic and ad hoc neural network architectures. EFIT-Prime utilizes a neural architecture search-based deep ensemble for robust uncertainty quantification, providing scalable and efficient neural architectures that comprehensively quantify both data and model uncertainties. Physically informed by the Grad–Shafranov equation, EFIT-Prime applies a constraint on the current density J tor and a smoothness constraint on the first derivative of the poloidal flux, ensuring physically plausible solutions. Furthermore, the spatial location of the diagnostics is explicitly incorporated in the inputs to account for their spatial correlation. Extensive evaluations demonstrate EFIT-Prime's accuracy and robustness across diverse scenarios, most notably showing good generalization on negative-triangularity discharges that were excluded from training. Timing studies indicate an ensemble inference time of 15 ms for predicting a new equilibrium, offering the possibility of plasma control in real-time, if the model is optimized for speed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A New Ensemble Canonical Correlation Prediction Scheme for Seasonal Precipitation

Department of Mathematical Sciences, University of Alberta, Edmonton, Canada This paper describes the fundamental theory of the ensemble canonical correlation (ECC) algorithm for the seasonal climate forecasting. The algorithm is a statistical regression sch eme based on maximal correlation between the predictor and predictand. The prediction error is estimated by a spectral method using the basis of empirical orthogonal functions. The ECC algorithm treats the predictors and predictands as continuous fields and is an improvement from the traditional canonical correlation prediction. The improvements include the use of area-factor, estimation of prediction error, and the optimal ensemble of multiple forecasts. The ECC is applied to the seasonal forecasting over various parts of the world. The example presented here is for the North America precipitation. The predictor is the sea surface temperature (SST) from different ocean basins. The Climate Prediction Center's reconstructed SST (1951-1999) is used as the predictor's historical data. The optimally interpolated global monthly precipitation is used as the predictand?s historical data. Our forecast experiments show that the ECC algorithm renders very high skill and the optimal ensemble is very important to the high value.

Kim, Kyu-Myong↗

Large-Scale Groundwater Monitoring in Brazil Assisted With Satellite-Based Artificial Intelligence Techniques

Here, we develop and test an artificial intelligence (AI)-based approach to monitor major Brazilian aquifers. The approach combines Gravity Recovery and Climate Experiment (GRACE) data and ground-based hydrogeological measurements from Brazil’s Integrated Groundwater Monitoring Network at hundreds of wells distributed in twelve aquifers across the country. We tested model ensembles based on three AI approaches: Extreme Gradient Boost, Light Gradient Boosting Model and CatBoost, followed by a Linear Regression (LR) step. The approach is further boosted with wavelet and seasonal decomposition processes applied to GRACE data. To determine the AI-based model’s sensitivity to data availability, we propose four experiments combining hydrogeological measurements from different aquifers. Groundwater storage estimates from the Global Land Data Assimilation System (GLDAS) are used as benchmark. A sensitivity analysis shows that the LR-based model ensemble is the best suited and to reproduce groundwater storage change in all studied Brazilian aquifers. Results show that the proposed approach outperforms GLDAS in all experiments, with an RMSE value of 2.68cm for the experiment that covers all monitored wells in Brazil. GLDAS resulted in RMSE=6.76cm. Using our AI model outputs, we quantified the groundwater storage change of two major aquifers, Urucuia and Bauru-Caiuá, over the past two decades: -31km 3 and -6km 3 , respectively. Water loss is driven by a prolonged drought across most of the country and intensification of groundwater pumping for irrigation. This study demonstrates that combining satellite data and AI can be a cost-effective alternative to monitor poorly equipped aquifers at the continental scale, with possible global replicability.

GRACE↗

Improved Soil Moisture Estimation and Detection of Irrigation Signal By Incorporating SMAP Soil Moisture Into the Indian Land Data Assimilation System (ILDAS)

Land surface models have facilitated the estimation of soil moisture over a range of spatiotemporal scales. However, limitations in model parameterization and under-representation of anthropogenic processes restrict their ability to estimate local-scale soil moisture variability, especially over irrigated areas. Assimilation of satellite-based soil moisture retrievals into land surface models can be a viable approach to overcome these constraints, specially over highly irrigated countries such as India, where such applications are rare. Additionally, large-scale validation of modeled soil moisture has been limited over India till now due to lack of a representative station network. By assimilating Soil Moisture Active Passive (SMAP)-based estimates into the state-of-the-art Indian Land Data Assimilation System (ILDAS) and combining with a new soil moisture station network of more than 200 stations, this study demonstrates improved soil moisture estimations and capture of irrigation signals over the region. The Noah-MP land surface model is forced by multiple local and global meteorological datasets and Ensemble Kalman Filter (EnKF) is used for assimilation of soil moisture. Comparison of open-loop and data assimilated soil moisture against station soil moisture data shows relative spatial mean improvement of 0.0178 in correlation and 0.0029 m3/m3 in RMSE. Further statistical comparison with in-situ data has also shown better results over most of the stations, as evident from improved correlations and reduced unbiased RMSE after assimilation. Finally, the climatology of soil moisture over the different irrigation fractions reveals that data assimilated outputs over irrigated grid cells tend to have higher soil moisture during dry winter season, demonstrating the ability to capture irrigation signals. These findings quantify the value of data assimilation in improving soil moisture estimates and the ability to capture unmodeled processes such as irrigation, which lays the science groundwork for upcoming space missions such as NASA ISRO Synthetic Aperture Radar (NISAR).

Soil Moisture↗

A Model-Model and Data-Model Comparison for the Early Eocene Hydrological Cycle

A range of proxy observations have recently provided constraints on how Earth's hydrological cycle responded to early Eocene climatic changes. However, comparisons of proxy data to general circulation model (GCM) simulated hydrology are limited and inter-model variability remains poorly characterised. In this work, we undertake an intercomparison of GCM-derived precipitation and P - E distributions within the extended EoMIP ensemble (Eocene Modelling Intercomparison Project; Lunt et al., 2012), which includes previously published early Eocene simulations performed using five GCMs differing in boundary conditions, model structure, and precipitation-relevant parameterisation schemes. We show that an intensified hydrological cycle, manifested in enhanced global precipitation and evaporation rates, is simulated for all Eocene simulations relative to the preindustrial conditions. This is primarily due to elevated atmospheric paleo-CO2, resulting in elevated temperatures, although the effects of differences in paleogeography and ice sheets are also important in some models. For a given CO2 level, globally averaged precipitation rates vary widely between models, largely arising from different simulated surface air temperatures. Models with a similar global sensitivity of precipitation rate to temperature (dP=dT ) display different regional precipitation responses for a given temperature change. Regions that are particularly sensitive to model choice include the South Pacific, tropical Africa, and the Peri-Tethys, which may represent targets for future proxy acquisition. A comparison of early and middle Eocene leaf-fossil-derived precipitation estimates with the GCM output illustrates that GCMs generally underestimate precipitation rates at high latitudes, although a possible seasonal bias of the proxies cannot be excluded. Models which warm these regions, either via elevated CO2 or by varying poorly constrained model parameter values, are most successful in simulating a match with geologic data. Further data from low-latitude regions and better constraints on early Eocene CO2 are now required to discriminate between these model simulations given the large error bars on paleoprecipitation estimates. Given the clear differences between simulated precipitation distributions within the ensemble, our results suggest that paleohydrological data offer an independent means by which to evaluate model skill for warm climates.

Boundary conditions↗

Weather Impact on Airport Arrival Meter Fix Throughput

Time-based flow management provides arrival aircraft schedules based on arrival airport conditions, airport capacity, required spacing, and weather conditions. In order to meet a scheduled time at which arrival aircraft can cross an airport arrival meter fix prior to entering the airport terminal airspace, air traffic controllers make regulations on air traffic. Severe weather may create an airport arrival bottleneck if one or more of airport arrival meter fixes are partially or completely blocked by the weather and the arrival demand has not been reduced accordingly. Under these conditions, aircraft are frequently being put in holding patterns until they can be rerouted. A model that predicts the weather impacted meter fix throughput may help air traffic controllers direct arrival flows into the airport more efficiently, minimizing arrival meter fix congestion. This paper presents an analysis of air traffic flows across arrival meter fixes at the Newark Liberty International Airport (EWR). Several scenarios of weather impacted EWR arrival fix flows are described. Furthermore, multiple linear regression and regression tree ensemble learning approaches for translating multiple sector Weather Impacted Traffic Indexes (WITI) to EWR arrival meter fix throughputs are examined. These weather translation models are developed and validated using the EWR arrival flight and weather data for the period of April-September in 2014. This study also compares the performance of the regression tree ensemble with traditional multiple linear regression models for estimating the weather impacted throughputs at each of the EWR arrival meter fixes. For all meter fixes investigated, the results from the regression tree ensemble weather translation models show a stronger correlation between model outputs and observed meter fix throughputs than that produced from multiple linear regression method.

Machine Learning Model.↗

Using GPS and VLBI technology to maintain 14 digit synchronization

To facilitate the navigation of spacecraft to the outer planets, Jupiter and beyond, the JPL-NASA Deep Space Network (DSN) has implemented three ensembles of atomic clocks at widely separated locations. These clocks must be maintained, synchronized, to with a few parts in 10 to the 13th power of each other and, the entire group must be maintained, to a lesser degree, in synchronism with Coordinated Universal Time (UTC)NBS/USNO. Over the last 1 1/2 years the DSN has been using Global Positioning Satellites (GPS) and Very Long Baseline Interferometry (VLBI) technology to perform these critical Frequency and Time (F&T) synchronization tasks. A year of F&T synchronization data collected from the intercomparison of 3 sets of cesium and hydrogen maser driven clock ensembles through the use of GPS and VLBI techniques are covered. Also covered, are some of the problems met and limitations of these two techniques at their present level of technology.

Ward, S. C.↗

Exploring Data Set Bias and Decision Support with Predictive Uncertainty Through Bayesian Approximations and Convolutional Neural Networks

Individual seismic catalogs can contain multiscale observations from fault level to global scales and associated waveforms from discrete events reflect crustal structure across many different scales and locations. Seismic network aperture, geographic location, and observation distance may not provide informative guidance or intuition on how different catalogs will behave across models trained under different conditions. We rely on uncertainty to provide guardrails for when to trust model decisions, but understanding when our uncertainty is trustworthy is an open challenge. Here, in this work, we explore Bayesian approximation methods for assigning predictive uncertainty in seismic event classification problems. We find that computationally expensive Bayesian approximations do not outperform simple ensemble methods. We also find that when exploiting multiple seismic event catalogs, joint training with data from all the catalogs combined with Bayesian approximations and supervised training for classification can obscure bias and result in less robust uncertainty while also not providing substantial performance benefits compared to training individual models for each catalog.

58 GEOSCIENCES↗

Incremental Learning for Passive Microwave Precipitation Retrievals using Advanced Technology Microwave Sounder

Spaceborne passive microwave (PMW) radiometry is central to global precipitation monitoring, yet retrieval uncertainties remain substantial, particularly for cross-track sounders whose variable footprints and channel configurations are optimized for atmospheric temperature and moisture profiling rather than precipitation. Consequently, existing operational products often exhibit angular-dependent biases, limited effective swath utilization, unrealistic rainfall probability distributions, and systematic misclassification of precipitation phase. These limitations are further compounded by the scarcity of globally accurate and representative precipitation observations, as training data from the Dual-frequency Precipitation Radar (DPR) and the Cloud Profiling Radar (CPR) are spatially sparse, lack uniform global coverage, and exhibit heterogeneous error characteristics across precipitation regimes. To address these challenges, this study presents a supervised retrieval algorithm that incrementally trains an ensemble of extreme gradient-boosted decision trees by augmenting base learners with pre-training on reanalysis data and post-training on coincident DPR and CPR observations matched with the Advanced Technology Microwave Sounder (ATMS). By transferring prior information from reanalysis to posterior constraints from radar observations and adopting a sequential detection–estimation strategy for precipitation phase and rate retrieval, the proposed approach yields retrievals across the full ATMS swath that are largely free from persistent deficiencies in current Global Precipitation Measurement (GPM) passive microwave operational products. In particular, the method resolves bimodal artifacts in rainfall retrievals and mitigates systematic high-latitude snowfall biases, including overestimation across the Arctic and underestimation across the Antarctic. Validation against independent Multi-Radar Multi-Sensor (MRMS) data over the Contiguous United States (CONUS) further demonstrates improved performance in precipitation phase detection and rate estimation relative to both reanalysis and current GPM PMW products.

Mahyar Garshasbi↗

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations

Abstract Motivation Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. Results We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (×105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. Availability and implementation It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Du, Lin (ORCID:0009000289869812)↗

Estimating Future Changes of Energy Demand for Heating and Cooling Buildings at NASA Centers GC23J-1198

With its unique and trusted earth observations, NASA is a critical source in informing decisions that will help achieve the U.S. goal of Net-Zero Greenhouse Gas (GHG) Emissions by 2050. NASA’s Prediction of Worldwide Energy Resource (POWER) project facilitates the use of NASA Earth Science data holdings within the energy, agricultural, and building heating/cooling design industries. POWER packages solar and meteorological data from several NASA projects in a user friendly GIS-enabled web services system (https://power.larc.nasa.gov). As part of the development of new data products to support the energy and building heating/cooling design communities, we estimate the changes in energy required to heat and cool buildings in the future climate at 14 different NASA site locations spread throughout the continental United States, as projected by CMIP6 climate models under different emissions scenarios. Bias-corrected downscaled time series of meteorological variables are taken from NASA Earth Exchange (NEX) Global Daily Downscaled Projections (GDDP-CMIP6) downscaled climate model data. The spread between the different model projections is accounted for by analyzing both the ensemble average of 22 CMIP6 models and 6 representative models with different climate sensitivities and different interannual variability. Changes in energy use are estimated in two ways. First, changes in the total annual heating and cooling degree days (HDD and CDD, respectively) are calculated relative to the current climate. This is done at all 14 sites. Second, the downscaled time series are used as inputs into RETScreen(R), a clean energy management decision tool, to give an estimate of heating/cooling energy use for a typical office building. We use this estimation method with model data at Langley Research Center. In the next 50 years, the annual total of HDD (CDD) is projected to decrease by 8-38% (increase by 5-28%) at all sites, with the increase in CDD typically a larger magnitude the decrease in HDD. For a typical small office building at Langley Research Center, the amount of energy needed to cool increases by 33-54% and the amount of energy to heat decreases by 29-40%. POWER is working to develop long term climate data services based on these results to include in future data products to provide to users.

Bradley M. Hegyi↗

Decimated Input Ensembles for Improved Generalization

Recently, many researchers have demonstrated that using classifier ensembles (e.g., averaging the outputs of multiple classifiers before reaching a classification decision) leads to improved performance for many difficult generalization problems. However, in many domains there are serious impediments to such "turnkey" classification accuracy improvements. Most notable among these is the deleterious effect of highly correlated classifiers on the ensemble performance. One particular solution to this problem is generating "new" training sets by sampling the original one. However, with finite number of patterns, this causes a reduction in the training patterns each classifier sees, often resulting in considerably worsened generalization performance (particularly for high dimensional data domains) for each individual classifier. Generally, this drop in the accuracy of the individual classifier performance more than offsets any potential gains due to combining, unless diversity among classifiers is actively promoted. In this work, we introduce a method that: (1) reduces the correlation among the classifiers; (2) reduces the dimensionality of the data, thus lessening the impact of the 'curse of dimensionality'; and (3) improves the classification performance of the ensemble.

Tumer, Kagan↗

Simulations of the Mid-Pliocene Warm Period Using Two Versions of the NASA-GISS ModelE2-R Coupled Model

The mid-Pliocene Warm Period (mPWP) bears many similarities to aspects of future global warming as projected by the Intergovernmental Panel on Climate Change (IPCC, 2007). Both marine and terrestrial data point to high-latitude temperature amplification, including large decreases in sea ice and land ice, as well as expansion of warmer climate biomes into higher latitudes. Here we present our most recent simulations of the mid-Pliocene climate using the CMIP5 version of the NASAGISS Earth System Model (ModelE2-R). We describe the substantial impact associated with a recent correction made in the implementation of the Gent-McWilliams ocean mixing scheme (GM), which has a large effect on the simulation of ocean surface temperatures, particularly in the North Atlantic Ocean. The effect of this correction on the Pliocene climate results would not have been easily determined from examining its impact on the preindustrial runs alone, a useful demonstration of how the consequences of code improvements as seen in modern climate control runs do not necessarily portend the impacts in extreme climates.Both the GM-corrected and GM-uncorrected simulations were contributed to the Pliocene Model Intercomparison Project (PlioMIP) Experiment 2. Many findings presented here corroborate results from other PlioMIP multi-model ensemble papers, but we also emphasize features in the ModelE2-R simulations that are unlike the ensemble means. The corrected version yields results that more closely resemble the ocean core data as well as the PRISM3D reconstructions of the mid-Pliocene, especially the dramatic warming in the North Atlantic and Greenland-Iceland-Norwegian Sea, which in the new simulation appears to be far more realistic than previously found with older versions of the GISS model. Our belief is that continued development of key physical routines in the atmospheric model, along with higher resolution and recent corrections to mixing parameterisations in the ocean model, have led to an Earth System Model that will produce more accurate projections of future climate.

climate change↗

Are Atmospheric Models Too Cold in the Mountains? The State of Science and Insights from the SAIL Field Campaign

Mountains play an outsized role in water resource availability, and the amount and timing of water they provide depend strongly on temperature. To that end, we ask the question: How well are atmospheric models capturing mountain temperatures? We synthesize results showing that high-resolution, regionally relevant climate models produce 2-m air temperature (T2m) measurements colder than what is observed (a “cold bias”), particularly in snow-covered midlatitude mountain ranges during winter. We find common cold biases in 44 studies across global mountain ranges, including single-model and multimodel ensembles. We explore the factors driving these biases and examine the physical mechanisms, data limitations, and observational uncertainties behind T2m. Our analysis suggests that the biases are genuine and not due to observation sparsity or resolution mismatches. Cold biases occur primarily on mountain peaks and ridges, whereas valleys are often warm biased. Our literature review suggests that increasing model resolution does not clearly mitigate the bias. By analyzing data from the Surface Atmosphere Integrated Field Laboratory (SAIL) field campaign in the Colorado Rocky Mountains, we test various hypotheses related to cold biases and find that local wind circulations, longwave (LW) radiation, and surface-layer parameterizations contribute to the T2m biases in this particular location. We conclude by emphasizing the value of coordinated model evaluation and development efforts in heavily instrumented mountain locations for addressing the root cause(s) of T2m biases and improving predictive understanding of mountain climates.

54 ENVIRONMENTAL SCIENCES↗

Are Atmospheric Models Too Cold in the Mountains? The State of Science and Insights from the SAIL Field Campaign

Mountains play an outsized role in water resource availability, and the amount and timing of water they provide depend strongly on temperature. To that end, we ask the question: How well are atmospheric models capturing mountain temperatures? We synthesize results showing that high-resolution, regionally relevant climate models produce 2-m air temperature (T2m) measurements colder than what is observed (a “cold bias”), particularly in snow-covered midlatitude mountain ranges during winter. We find common cold biases in 44 studies across global mountain ranges, including single-model and multimodel ensembles. We explore the factors driving these biases and examine the physical mechanisms, data limitations, and observational uncertainties behind T2m. Our analysis suggests that the biases are genuine and not due to observation sparsity or resolution mismatches. Cold biases occur primarily on mountain peaks and ridges, whereas valleys are often warm biased. Our literature review suggests that increasing model resolution does not clearly mitigate the bias. By analyzing data from the Surface Atmosphere Integrated Field Laboratory (SAIL) field campaign in the Colorado Rocky Mountains, we test various hypotheses related to cold biases and find that local wind circulations, longwave (LW) radiation, and surface-layer parameterizations contribute to the T2m biases in this particular location. We conclude by emphasizing the value of coordinated model evaluation and development efforts in heavily instrumented mountain locations for addressing the root cause(s) of T2m biases and improving predictive understanding of mountain climates.

54 ENVIRONMENTAL SCIENCES↗

Particle phase function measurements by a new Fiber Array Nephelometer: FAN 1

A fiber array polar nephelometer of advanced design, the FAN I is capable of in-situ phase function measurements of scattered light from man-made or natural atmospheric particles. The scattered light is measured at 100 different angles throughout 360 degrees, thus providing a potential measurement of the asymmetry of irregularly shaped particles. Phase functions can be measured at 10 to 100 Hz rates and the range of measurable single particle sizes is from 5 micron m to as large as 8mm. For particles smaller than 5 micro m the ensemble average can be measured. The FAN I is microprocessor controlled and the data may be stored on floppy disk or printed out in tabular and/or graphical form. The optical head may be separated from the computer system for operation in field or adverse conditions. Examples of laboratory measured scattering phase functions obtained with the FAN I for spherical particles is given to illustrate its measurement capabilities.

Farmer, W. M.↗