Search NASA⌕ Search

SEARCH · Search NASA

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Machine Learning the COSMO Model for Predicting Thermodynamics of Electrolyte Mixtures

Bottom-up design of electrolyte mixtures for battery systems requires predicting macro thermodynamic properties from molecular constituents. For instance, molten salt electrolyte batteries require conditions far above room temperature to operate. Therefore, discovering mixtures with increasingly lower eutectic melting points is desirable. A model that can approximate chemical activity is a valuable tool to search through the vast compositional design space. Machine learning can predict properties of materials such as vibrational free energies, electronic energy gaps, and thermal conductivities. Moreover, they can learn physical models such as interatomic potentials. The COSMO-SAC model uses theory and empirical parameterization to predict liquid-vapor and liquid-solid properties using first-principles calculations. However, obtaining activity coefficients required for parameterizing the COSMO-SAC model is costly and limited to a select chemical space. In this work, we explored if machine learning methods could improve the COSMO-SAC model and bridge density functional theory calculations to liquid phase thermodynamic properties. Our data-driven approach uses existing databases for sigma-profiles of organic solvents and reconciles their methodological differences via ensemble averaging. First, an optimal machine learning model is constructed for each dataset. Our machine learning algorithms use the sigma-profile as an input feature to predict binary mixtures' activity coefficients using multi-output regression. Each dataset uses different choices of functionals, methods, and basis sets. Therefore, our ensemble model attempts to predict corrected activity coefficients given the combination of all the model outputs. The activity coefficients used for training are generated using the COSMO-SAC model. This approach enables the extraction of meaningful information from the existing datasets to improve the COSMO-SAC model for obtaining thermodynamic properties of electrolyte mixtures. With the liquid phase activities, we can identify electrolyte mixtures that meet desired phase equilibria conditions.

Thermodynamics↗

A time-accurate finite volume method valid at all flow velocities

A finite volume method to solve the Navier-Stokes equations at all flow velocities (e.g., incompressible, subsonic, transonic, supersonic and hypersonic flows) is presented. The numerical method is based on a finite volume method that incorporates a pressure-staggered mesh and an incremental pressure equation for the conservation of mass. Comparison of three generally accepted time-advancing schemes, i.e., Simplified Marker-and-Cell (SMAC), Pressure-Implicit-Splitting of Operators (PISO), and Iterative-Time-Advancing (ITA) scheme, are made by solving a lid-driven polar cavity flow and self-sustained oscillatory flows over circular and square cylinders. Calculated results show that the ITA is the most stable numerically and yields the most accurate results. The SMAC is the most efficient computationally and is as stable as the ITA. It is shown that the PISO is the most weakly convergent and it exhibits an undesirable strong dependence on the time-step size. The degenerated numerical results obtained using the PISO are attributed to its second corrector step that cause the numerical results to deviate further from a divergence free velocity field. The accurate numerical results obtained using the ITA is attributed to its capability to resolve the nonlinearity of the Navier-Stokes equations. The present numerical method that incorporates the ITA is used to solve an unsteady transitional flow over an oscillating airfoil and a chemically reacting flow of hydrogen in a vitiated supersonic airstream. The turbulence fields in these flow cases are described using multiple-time-scale turbulence equations. For the unsteady transitional over an oscillating airfoil, the fluid flow is described using ensemble-averaged Navier-Stokes equations defined on the Lagrangian-Eulerian coordinates. It is shown that the numerical method successfully predicts the large dynamic stall vortex (DSV) and the trailing edge vortex (TEV) that are periodically generated by the oscillating airfoil. The calculated streaklines are in very good comparison with the experimentally obtained smoke picture. The calculated turbulent viscosity contours show that the transition from laminar to turbulent state and the relaminarization occur widely in space as well as in time. The ensemble-averaged velocity profiles are also in good agreement with the measured data and the good comparison indicates that the numerical method as well as the multipletime-scale turbulence equations successfully predict the unsteady transitional turbulence field. The chemical reactions for the hydrogen in the vitiated supersonic airstream are described using 9 chemical species and 48 reaction-steps. Consider that a fast chemistry can not be used to describe the fine details (such as the instability) of chemically reacting flows while a reduced chemical kinetics can not be used confidently due to the uncertainty contained in the reaction mechanisms. However, the use of a detailed finite rate chemistry may make it difficult to obtain a fully converged solution due to the coupling between the large number of flow, turbulence, and chemical equations. The numerical results obtained in the present study are in good agreement with the measured data. The good comparison is attributed to the numerical method that can yield strongly converged results for the reacting flow and to the use of the multiple-time-scale turbulence equations that can accurately describe the mixing of the fuel and the oxidant.

Kim, S.-W.↗

Methodology of Blade Unsteady Pressure Measurement in the NASA Transonic Flutter Cascade

In this report the methodology adopted to measure unsteady pressures on blade surfaces in the NASA Transonic Flutter Cascade under conditions of simulated blade flutter is described. The previous work done in this cascade reported that the oscillating cascade produced waves, which for some interblade phase angles reflected off the wind tunnel walls back into the cascade, interfered with the cascade unsteady aerodynamics, and contaminated the acquired data. To alleviate the problems with data contamination due to the back wall interference, a method of influence coefficients was selected for the future unsteady work in this cascade. In this approach only one blade in the cascade is oscillated at a time. The majority of the report is concerned with the experimental technique used and the experimental data generated in the facility. The report presents a list of all test conditions for the small amplitude of blade oscillations, and shows examples of some of the results achieved. The report does not discuss data analysis procedures like ensemble averaging, frequency analysis, and unsteady blade loading diagrams reconstructed using the influence coefficient method. Finally, the report presents the lessons learned from this phase of the experimental effort, and suggests the improvements and directions of the experimental work for tests to be carried out for large oscillation amplitudes.

Lepicovsky, J.↗

Project FIRES. Volume 1: Program Overview and Summary, Phase 1B

Overall performance requirements and evaluation methods for firefighters protective equipment were established and published as the Protective Ensemble Performance Standards (PEPS). Current firefighters protective equipment was tested and evaluated against the PEPS requirements, and the preliminary design of a prototype protective ensemble was performed. In phase 1B, the design of the prototype ensemble was finalized. Prototype ensembles were fabricated and then subjected to a series of qualification tests which were based upon the PEPS requirements. Engineering drawings and purchase specifications were prepared for the new protective ensemble.

Abeles, F. J.↗

An Ensemble of Bayesian Neural Networks for Exoplanetary Atmospheric Retrieval

Machine learning (ML) is now used in many areas of astrophysics, from detecting exoplanets in Kepler transit signals to removing telescope systematics. Recent work demonstrated the potential of using ML algorithms for atmospheric retrieval by implementing a random forest (RF) to perform retrievals in seconds that are consistent with the traditional, computationally expensive nested-sampling retrieval method. We expand upon their approach by presenting a new ML model, plan-net, based on an ensemble of Bayesian neural networks (BNNs) that yields more accurate inferences than the RF for the same data set of synthetic transmission spectra. We demonstrate that an ensemble provides greater accuracy and more robust uncertainties than a single model. In addition to being the first to use BNNs for atmospheric retrieval, we also introduce a new loss function for BNNs that learns correlations between the model outputs. Importantly, we show that designing ML models to explicitly incorporate domain-specific knowledge both improves performance and provides additional insight by inferring the covariance of the retrieved atmospheric parameters. We apply plan-net to the Hubble Space Telescope Wide Field Camera 3 transmission spectrum for WASP-12b and retrieve an isothermal temperature and water abundance consistent with the literature. We highlight that our method is flexible and can be expanded to higher resolution spectra and a larger number of atmospheric parameters.

Adam D. Cobb↗

An investigation of turbulent transport in the extreme lower atmosphere

A model in which the Lagrangian autocorrelation is expressed by a domain integral over a set of usual Eulerian autocorrelations acquired concurrently at all points within a turbulence box is proposed along with a method for ascertaining the statistical stationarity of turbulent velocity by creating an equivalent ensemble to investigate the flow in the extreme lower atmosphere. Simultaneous measurements of turbulent velocity on a turbulence line along the wake axis were carried out utilizing a longitudinal array of five hot-wire anemometers remotely operated. The stationarity test revealed that the turbulent velocity is approximated as a realization of a weakly self-stationary random process. Based on the Lagrangian autocorrelation it is found that: (1) large diffusion time predominated; (2) ratios of Lagrangian to Eulerian time and spatial scales were smaller than unity; and, (3) short and long diffusion time scales and diffusion spatial scales were constrained within their Eulerian counterparts.

Koper, C. A., Jr.↗

Feature Selection in High-Dimensional Space with Applications to Gene Expression Data

Recent years have seen rapid growth in high-dimensional datasets. Most existing machine learning (ML) algorithms fail in high-dimensional settings where many features could be redundant. A critical process of feature selection is thus applied in such a setting that helps in identifying the most relevant features while removing redundant ones. With the increase in high dimensionality, one is also faced with problems of efficiency and interpretation in performing such selection methods. Therefore, this paper proposes a “novel” feature selection framework that uses an ensemble of interpretable ML algorithms to perform feature selection and the ranking of final features. Finally, this framework is applied to a gene expression dataset obtained through collaboration with the National Aeronautics and Space Administration (NASA)’s Biological and Physical Sciences (BPS) team and helps identify important and relevant genes contributing to specific target attributes through classification tasks.

Nishan Pantha↗

Multiple-Beam Detection of Fast Transient Radio Sources

A method has been designed for using multiple independent stations to discriminate fast transient radio sources from local anomalies, such as antenna noise or radio frequency interference (RFI). This can improve the sensitivity of incoherent detection for geographically separated stations such as the very long baseline array (VLBA), the future square kilometer array (SKA), or any other coincident observations by multiple separated receivers. The transients are short, broadband pulses of radio energy, often just a few milliseconds long, emitted by a variety of exotic astronomical phenomena. They generally represent rare, high-energy events making them of great scientific value. For RFI-robust adaptive detection of transients, using multiple stations, a family of algorithms has been developed. The technique exploits the fact that the separated stations constitute statistically independent samples of the target. This can be used to adaptively ignore RFI events for superior sensitivity. If the antenna signals are independent and identically distributed (IID), then RFI events are simply outlier data points that can be removed through robust estimation such as a trimmed or Winsorized estimator. The alternative "trimmed" estimator is considered, which excises the strongest n signals from the list of short-beamed intensities. Because local RFI is independent at each antenna, this interference is unlikely to occur at many antennas on the same step. Trimming the strongest signals provides robustness to RFI that can theoretically outperform even the detection performance of the same number of antennas at a single site. This algorithm requires sorting the signals at each time step and dispersion measure, an operation that is computationally tractable for existing array sizes. An alternative uses the various stations to form an ensemble estimate of the conditional density function (CDF) evaluated at each time step. Both methods outperform standard detection strategies on a test sequence of VLBA data, and both are efficient enough for deployment in real-time, online transient detection applications.

Thompson, David R.↗

Effects of bleed-hole geometry and plenum pressure on three-dimensional shock-wave/boundary-layer/bleed interactions

A numerical study was performed to investigate 3D shock-wave/boundary-layer interactions on a flat plate with bleed through one or more circular holes that vent into a plenum. This study was focused on how bleed-hole geometry and pressure ratio across bleed holes affect the bleed rate and the physics of the flow in the vicinity of the holes. The aspects of the bleed-hole geometry investigated include angle of bleed hole and the number of bleed holes. The plenum/freestream pressure ratios investigated range from 0.3 to 1.7. This study is based on the ensemble-averaged, 'full compressible' Navier-Stokes (N-S) equations closed by the Baldwin-Lomax algebraic turbulence model. Solutions to the ensemble-averaged N-S equations were obtained by an implicit finite-volume method using the partially-split, two-factored algorithm of Steger on an overlapping Chimera grid.

Chyu, Wei J.↗

Statistical Analysis of Large Simulated Yield Datasets for Studying Climate Effects

Many studies have been carried out during the last decade to study the effect of climate change on crop yields and other key crop characteristics. In these studies, one or several crop models were used to simulate crop growth and development for different climate scenarios that correspond to different projections of atmospheric CO2 concentration, temperature, and rainfall changes (Semenov et al., 1996; Tubiello and Ewert, 2002; White et al., 2011). The Agricultural Model Intercomparison and Improvement Project (AgMIP; Rosenzweig et al., 2013) builds on these studies with the goal of using an ensemble of multiple crop models in order to assess effects of climate change scenarios for several crops in contrasting environments. These studies generate large datasets, including thousands of simulated crop yield data. They include series of yield values obtained by combining several crop models with different climate scenarios that are defined by several climatic variables (temperature, CO2, rainfall, etc.). Such datasets potentially provide useful information on the possible effects of different climate change scenarios on crop yields. However, it is sometimes difficult to analyze these datasets and to summarize them in a useful way due to their structural complexity; simulated yield data can differ among contrasting climate scenarios, sites, and crop models. Another issue is that it is not straightforward to extrapolate the results obtained for the scenarios to alternative climate change scenarios not initially included in the simulation protocols. Additional dynamic crop model simulations for new climate change scenarios are an option but this approach is costly, especially when a large number of crop models are used to generate the simulated data, as in AgMIP. Statistical models have been used to analyze responses of measured yield data to climate variables in past studies (Lobell et al., 2011), but the use of a statistical model to analyze yields simulated by complex process-based crop models is a rather new idea. We demonstrate herewith that statistical methods can play an important role in analyzing simulated yield data sets obtained from the ensembles of process-based crop models. Formal statistical analysis is helpful to estimate the effects of different climatic variables on yield, and to describe the between-model variability of these effects.

climate↗

Rational Design of Nanoplasmonic Array Geometries for Biosensing

Background: Molecular diagnostics provide early and accurate diagnosis, which is essential for the prevention and treatment of infectious as well as chronic diseases. These tests are designed to detect disease-specific bioanalytes such as nucleic acid (DNA or RNA) or protein (antigens, antibodies) biomarkers. In the context of infectious disease diagnosis, nucleic acid-based detection methods are known to provide more specific and sensitive results. Here, the presence of a unique sequence belonging to the pathogenic genomic material is targeted to identify species, organism, genera and/or antimicrobial resistant gene markers. The majority of the common nucleic acid based diagnostic techniques require amplification (polymerase chain reaction, isothermal amplification etc.) of the pathogenic genetic material prior to detection impacting diagnostic speed, complexity, and cost thereby limiting ease of use. Thus, the development of simplified nucleic acid-based diagnostics that can be even used in resource-poor settings may hugely benefit patients across the globe. Nanopath is a molecular diagnostics company utilizing a solid-state nanosensor to enable sequence-specific detection of target nucleic acids without the need of amplification. These nanostructures enable ultra-sensitive biomarker detection using geometric, feature-dependent properties highly dependent on the local dielectric environment, allowing them to be sensitive to low concentration binding events. This paper describes an application of this approach to provide highly relevant clinical information within a single doctor’s office visit. Intro: The Nanopath team is in collaboration with NASA (National Aeronautics and Space Administration) and NIST (National Institute of Standards and Technology) to push the bounds of the fundamental physics associated with their biosensing platform. The ability of metals to support electromagnetic surface waves gives rise to surface plasmons when optically illuminated. This property, and its strong sensitivity to changes in the local refractive index, allows for the use of metal nanoparticles as ultra-sensitive transducers. In prior work by members of this team, ensembles of randomly oriented nanoparticles (i.e., colloidal nanorods dispersed on chip) were employed for sequence-specific nucleic acid sensing (1-3). While these particle sensors have the advantage of rapid fabrication, they suffer from low sensitivity and quality factor due to the random particle dispersity. In contrast, in this study we employ ordered array nanoparticle ensembles which can be used to improve sensor sensitivity and figure-of-merit. Study Methods Overview: In this talk, we detail the results of sensing experiments and computational simulations to outline a rational design of the structure of these plasmonic nanoparticle arrays for biomolecular sensing. Through simulation and experiment, we iteratively tailor nanostructure dimension to provide high quality signal and large resonance shifts upon modeled nucleic acid binding. In particular, full-wave electromagnetic simulations were conducted using Lumerical photonic simulation software in which periodic boundary conditions were applied in the x- and y- dimensions for each of the nanoplasmonic sensor geometries. To simulate the resonance response to changes in the bulk solution in contact with the sensor surface, the refractive index of the surrounding media was changed appropriately. Nucleic acid hybridization events were modeled using either using spherical structures approximating the relevant radius of genomic material as estimated by polymer models, or as conformal layers with the known refractive indices for nucleic acids. On the basis of initial simulations, nanosensors were fabricated using traditional electron-beam lithography protocols at NIST. To evaluate consensus between simulations and experiments, bulk sensing experiments were carried out in which the resonance peaks were obtained by submerging the sensors in refractive index standards. Key nanosensor characteristics including resonance peak locations, resonance peak shifts as a function of refractive index, and figure of merit (FOM) of extinction curves were examined between the experimental and simulation results prior to proceeding with simulations on additional geometries and more complex solution conditions, and further device fabrication. This iterative process is repeated toward a rational design of nanoplasmonic array geometries for biosensing optimizing response for targeted disease detection. In summary, this study puts forth a methodology for rational design and characterization of regularly spaced nanoparticle arrays for optics-based biosensing. The results of this study will allow for more informed design of nanostructure geometries towards sequence-specific nucleic acid detection. These improved designs have the potential to improve clinical sensitivity and limit-of-detection across disease indication.

sensor↗

Adaptive Fault Detection on Liquid Propulsion Systems with Virtual Sensors: Algorithms and Architectures

Prior to the launch of STS-119 NASA had completed a study of an issue in the flow control valve (FCV) in the Main Propulsion System of the Space Shuttle using an adaptive learning method known as Virtual Sensors. Virtual Sensors are a class of algorithms that estimate the value of a time series given other potentially nonlinearly correlated sensor readings. In the case presented here, the Virtual Sensors algorithm is based on an ensemble learning approach and takes sensor readings and control signals as input to estimate the pressure in a subsystem of the Main Propulsion System. Our results indicate that this method can detect faults in the FCV at the time when they occur. We use the standard deviation of the predictions of the ensemble as a measure of uncertainty in the estimate. This uncertainty estimate was crucial to understanding the nature and magnitude of transient characteristics during startup of the engine. This paper overviews the Virtual Sensors algorithm and discusses results on a comprehensive set of Shuttle missions and also discusses the architecture necessary for deploying such algorithms in a real-time, closed-loop system or a human-in-the-loop monitoring system. These results were presented at a Flight Readiness Review of the Space Shuttle in early 2009.

Matthews, Bryan L.↗

Investigating seasonal ENSO forecast amplitude calibration for GEOS-S2S-2

The GEOS-S2S-2 is a global coupled model and assimilation system, encompassing many aspects of the Earth climate system. This project looked at seasonal (nine-month) ensemble forecasts, over the period from 1982 through the present. We explored and validated a forecast amplitude correction technique used by the North American Multi-Model Ensemble (NMME), of which GEOS is a member, for Niño3.4 sea-surface temperature (SST) anomaly predictions. The method relies on deriving a set of correction factors based on the standard deviations of hindcast and observed SST anomalies. This algorithm was implemented at the Global Modeling and Assimilation Office (GMAO) and applied to the GEOS-S2S-2 Niño 3.4 hindcasts. In general, the correction reduced the forecast amplitude. This improved the quality of the forecasts when measured by the root mean square error (RMSE) of the ensemble mean forecast. RMSE decreased for most initialization months and forecast leads. The correction had a larger impact on seasons with strong El Niño-Southern Oscillation (ENSO) events than neutral seasons. The effects of the correction on the ensemble characteristics were also examined, and a variation of the algorithm was proposed that preserves the ensemble spread. The outcome of this project is a tool that can be used to calibrate GEOS-S2S-2 ENSO forecasts and potentially improve the precision of multi-model El Niño outlooks.

ENSO↗

Using the Bootstrap Method for a Statistical Significance Test of Differences between Summary Histograms

A new method is proposed to compare statistical differences between summary histograms, which are the histograms summed over a large ensemble of individual histograms. It consists of choosing a distance statistic for measuring the difference between summary histograms and using a bootstrap procedure to calculate the statistical significance level. Bootstrapping is an approach to statistical inference that makes few assumptions about the underlying probability distribution that describes the data. Three distance statistics are compared in this study. They are the Euclidean distance, the Jeffries-Matusita distance and the Kuiper distance. The data used in testing the bootstrap method are satellite measurements of cloud systems called cloud objects. Each cloud object is defined as a contiguous region/patch composed of individual footprints or fields of view. A histogram of measured values over footprints is generated for each parameter of each cloud object and then summary histograms are accumulated over all individual histograms in a given cloud-object size category. The results of statistical hypothesis tests using all three distances as test statistics are generally similar, indicating the validity of the proposed method. The Euclidean distance is determined to be most suitable after comparing the statistical tests of several parameters with distinct probability distributions among three cloud-object size categories. Impacts on the statistical significance levels resulting from differences in the total lengths of satellite footprint data between two size categories are also discussed.

Xu, Kuan-Man↗

Performance of Trajectory Models with Wind Uncertainty

Typical aircraft trajectory predictors use wind forecasts but do not account for the forecast uncertainty. A method for generating estimates of wind prediction uncertainty is described and its effect on aircraft trajectory prediction uncertainty is investigated. The procedure for estimating the wind prediction uncertainty relies uses a time-lagged ensemble of weather model forecasts from the hourly updated Rapid Update Cycle (RUC) weather prediction system. Forecast uncertainty is estimated using measures of the spread amongst various RUC time-lagged ensemble forecasts. This proof of concept study illustrates the estimated uncertainty and the actual wind errors, and documents the validity of the assumed ensemble-forecast accuracy relationship. Aircraft trajectory predictions are made using RUC winds with provision for the estimated uncertainty. Results for a set of simulated flights indicate this simple approach effectively translates the wind uncertainty estimate into an aircraft trajectory uncertainty. A key strength of the method is the ability to relate uncertainty to specific weather phenomena (contained in the various ensemble members) allowing identification of regional variations in uncertainty.

Lee, Alan G.↗

Simulations of Spray Reacting Flows in a Single Element LDI Injector With and Without Invoking an Eulerian Scalar PDF Method

This paper presents the numerical simulations of the Jet-A spray reacting flow in a single element lean direct injection (LDI) injector by using the National Combustion Code (NCC) with and without invoking the Eulerian scalar probability density function (PDF) method. The flow field is calculated by using the Reynolds averaged Navier-Stokes equations (RANS and URANS) with nonlinear turbulence models, and when the scalar PDF method is invoked, the energy and compositions or species mass fractions are calculated by solving the equation of an ensemble averaged density-weighted fine-grained probability density function that is referred to here as the averaged probability density function (APDF). A nonlinear model for closing the convection term of the scalar APDF equation is used in the presented simulations and will be briefly described. Detailed comparisons between the results and available experimental data are carried out. Some positive findings of invoking the Eulerian scalar PDF method in both improving the simulation quality and reducing the computing cost are observed.

Shih, Tsan-Hsing↗