Search NASASearch

SEARCH · Search NASA

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Microstructure Quantification and Random Forest Regression Models for Li4Ti5O12–Ni Property Prediction

All-solid-state structural lithium-ion batteries are sought to enable all-electric propulsion in next generation aerospace concepts through improved safety and systems level weight savings. In this work, the influence of processing conditions on microstructural evolution was evaluated for anode composites of strain-free Li4Ti5O12 and metallic nickel current collector. Beyond size distributions, this study explored methods of quantifying microstructural features that describe changes in the spatial distribution and coalescence of nickel particles as a function of sample composition and sintering conditions. Processing-microstructure-property relationships were described by microstructure quantifiers including nickel particle count per area, nearest neighbor distance distribution, and edge-to-edge distance distribution. Machine learning methods were applied to compare the relative influence of processing conditions and microstructural features on electrical conductivity and mechanical strength to optimize for simultaneous energy storage and load bearing performance. Insights gained from this work inform future evaluation of alternative energy storage materials and microstructures for multifunctional performance, and generation of microstructural descriptors strengthens modeling across length scales.

anode

Atmospheric Chemistry Modeling Using a Regression Forest Model

Atmospheric chemistry is central to many environmental issues such as air pollution, climate change, and stratospheric ozone loss. Chemistry Transport Models (CTM) are a central tool for understanding these issues, whether for research or for forecasting. These models split the atmosphere in a large number of grid-boxes and consider the emission of compounds into these boxes and their subsequent transport, deposition, and chemical processing. The chemistry is represented through a series of simultaneous ordinary differential equations, one for each compound. Given the difference in life-times between the chemical compounds (milli-seconds for O1D to years for CH4) these equations are numerically stiff and solving them consists of a significant fraction of the computational burden of a CTM. We have investigated a machine learning approach to solving the differential equations instead of solving them numerically. From an annual simulation of the GEOS-Chem model we have produced a training dataset consisting of the concentration of compounds before and after the differential equations are solved, together with some key physical parameters for every grid-box and time-step. From this dataset we have trained a machine learning algorithm (regression forest) to be able to predict the concentration of the compounds after the integration step based on the concentrations and physical state at the beginning of the time step. We have then included this algorithm back into the GEOS-Chem model, bypassing the need to integrate the chemistry. This machine learning approach shows many of the characteristics of the full simulation and has the potential to be substantially faster. There are a wide range of application for such an approach - generating boundary conditions, for use in air quality forecasts, chemical data assimilation systems, centennial scale climate simulations etc. We discuss our approches' speed and accuracy, and highlight some potential future directions for improving this approach.

Keller, Christoph A.

Radar Altimetry as a Proxy for Determining Terrestrial Water Storage Variability in Tropical Basins

The Gravity Recovery and Climate Experiment (GRACE) mission has provided us with unforeseen information on terrestrial water-storage (TWS) variability, contributing to our understanding of global hydrological processes, including hydrological extreme events and anthropogenic impacts on water storage. Attempts to decompose GRACE-based TWS signals into its different water storage layers, i.e., surface water storage (SWS), soil moisture, groundwater and snow, have shown that SWS is a principal component, particularly in the tropics, where major rivers flow over arid regions at high latitudes. Here, we demonstrate that water levels, measured with radar altimeters at a limited number of locations, can be used to reconstruct gridded GRACE-based TWS signals in the Amazon basin, at spatial resolutions ranging from 0.5 to 3°, with mean absolute errors (MAE) as low as 2.5 cm and correlations as high as 0.98. We show that, at 3° spatial resolution, spatially-distributed TWS time series can be precisely reconstructed with as few as 41 water-level time series located within the basin. The proposed approach is competitive when compared to existing TWS estimates derived from physically based and computationally expensive methods. Also, a validation experiment indicates that TWS estimates can be extrapolated to periods beyond that of the model regression with low errors. The approach is robust, based on regression models and interpolation techniques, and offers a new possibility to reproduce spatially and temporally distributed TWS that could be used to fill inter-mission gaps and to extend GRACE-based TWS time series beyond its timespan.

Terrestrial water storage

Relationship of physiography and snow area to stream discharge

The author has identified the following significant results. A comparison of snowmelt runoff models shows that the accuracy of the Tangborn model and regression models is greater if the test data falls within the range of calibration than if the test data lies outside the range of calibration data. The regression models are significantly more accurate for forecasts of 60 days or more than for shorter prediction periods. The Tangborn model is more accurate for forecasts of 90 days or more than for shorter prediction periods. The Martinec model is more accurate for forecasts of one or two days than for periods of 3,5,10, or 15 days. Accuracy of the long-term models seems to be independent of forecast data. The sufficiency of the calibration data base is a function not only of the number of years of record but also of the accuracy with which the calibration years represent the total population of data years. Twelve years appears to be a sufficient length of record for each of the models considered, as long as the twelve years are representative of the population.

Mccuen, R. H.

Comparison of Iterative and Non-Iterative Strain-Gage Balance Load Calculation Methods

The accuracy of iterative and non-iterative strain-gage balance load calculation methods was compared using data from the calibration of a force balance. Two iterative and one non-iterative method were investigated. In addition, transformations were applied to balance loads in order to process the calibration data in both direct read and force balance format. NASA's regression model optimization tool BALFIT was used to generate optimized regression models of the calibration data for each of the three load calculation methods. This approach made sure that the selected regression models met strict statistical quality requirements. The comparison of the standard deviation of the load residuals showed that the first iterative method may be applied to data in both the direct read and force balance format. The second iterative method, on the other hand, implicitly assumes that the primary gage sensitivities of all balance gages exist. Therefore, the second iterative method only works if the given balance data is processed in force balance format. The calibration data set was also processed using the non-iterative method. Standard deviations of the load residuals for the three load calculation methods were compared. Overall, the standard deviations show very good agreement. The load prediction accuracies of the three methods appear to be compatible as long as regression models used to analyze the calibration data meet strict statistical quality requirements. Recent improvements of the regression model optimization tool BALFIT are also discussed in the paper.

Ulbrich, N.

Neural Network and Regression Soft Model Extended for PAX-300 Aircraft Engine

In fiscal year 2001, the neural network and regression capabilities of NASA Glenn Research Center's COMETBOARDS design optimization testbed were extended to generate approximate models for the PAX-300 aircraft engine. The analytical model of the engine is defined through nine variables: the fan efficiency factor, the low pressure of the compressor, the high pressure of the compressor, the high pressure of the turbine, the low pressure of the turbine, the operating pressure, and three critical temperatures (T(sub 4), T(sub vane), and T(sub metal)). Numerical Propulsion System Simulation (NPSS) calculations of the specific fuel consumption (TSFC), as a function of the variables can become time consuming, and numerical instabilities can occur during these design calculations. "Soft" models can alleviate both deficiencies. These approximate models are generated from a set of high-fidelity input-output pairs obtained from the NPSS code and a design of the experiment strategy. A neural network and a regression model with 45 weight factors were trained for the input/output pairs. Then, the trained models were validated through a comparison with the original NPSS code. Comparisons of TSFC versus the operating pressure and of TSFC versus the three temperatures (T(sub 4), T(sub vane), and T(sub metal)) are depicted in the figures. The overall performance was satisfactory for both the regression and the neural network model. The regression model required fewer calculations than the neural network model, and it produced marginally superior results. Training the approximate methods is time consuming. Once trained, the approximate methods generated the solution with only a trivial computational effort, reducing the solution time from hours to less than a minute.

Patnaik, Surya N.

Winter Wheat Yield Assessment from Landsat 8 and Sentinel-2 Data: Incorporating Surface Reflectance, Through Phenological Fitting, into Regression Yield Models

A combination of Landsat 8 and Sentinel-2 offers a high frequency of observations (3–5 days) at moderate spatial resolution (10–30 m), which is essential for crop yield studies. Existing methods traditionally apply vegetation indices (VIs) that incorporate surface reflectances (SRs) in two or more spectral bands into a single variable, and rarely address the incorporation of SRs into empirical regression models of crop yield. In this work, we address these issues by normalizing satellite data (both VIs and SRs) derived from NASA’s Harmonized Landsat Sentinel-2 (HLS) product, through a phenological fitting. We apply a quadratic function to fit VIs or SRs against accumulated growing degree days (AGDDs), which affects the rate of crop development. The derived phenological metrics for VIs and SRs, namely peak, area under curve (AUC), and fitting coefficients from a quadratic function, were used to build empirical regression winter wheat models at a regional scale in Ukraine for three years, 2016–2018. The best results were achieved for the model with near infrared (NIR) and red spectral bands and derived AUC, constant, linear, and quadratic coefficients of the quadratic model. The best model yielded a root mean square error (RMSE) of 0.201 t/ha (5.4%) and coefficient of determination R2 = 0.73 on cross-validation.

phenological fitting

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences

Development of a Non-Iterative Balance Load Prediction Algorithm for the NASA Ames Unitary Plan Wind Tunnel

A non-iterative load prediction algorithm for strain-gage balances was developed for the NASA Ames Unitary Plan Wind Tunnels that computes balance loads from the electrical outputs of the balance bridges and a set of state variables. A state variable could be, for example, a balance temperature difference or the bellows pressure of a flow-through balance. The algorithm directly uses regression models of the balance loads for the load prediction that were obtained by applying global regression analysis to balance calibration data. This choice greatly simplifies both implementation and use of the load prediction process for complex balance configurations as no load iteration needs to be performed. The regression model of a balance load is constructed by using terms from a total of nine term groups. Four term groups are derived from a Taylor Series expansion of the relationship between the load, gage outputs, and state variables. The remaining five term groups are defined by using absolute values of the gage outputs and state variables. Terms from these groups should only be included in the regression model if calibration data from a balance with known bi-directional outputs is analyzed. It is illustrated in detail how global regression analysis may be applied to obtain the coefficients of the chosen regression model of a load component assuming that no linear or massive near-linear dependencies between the regression model terms exist. Data from the machine calibration of a six-component force balance is used to illustrate both application and accuracy of the non-iterative load prediction process.

Ulbrich, Norbert M.

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Reflectance of vegetation, soil, and water

The author has identified the following significant results. The Kubelka-Munk model, a regression model, and a combination of these models were used to extract plant, soil, and shadow reflectance components of vegetated surfaces. The combination model was superior to the others; it explained 86% of the variation in band 5 reflectance of corn and sorghum, and 90% of the variation in band 6 reflectance of cotton. A fractional shadow term substantially increased the proportion of the digital count sum of squares explained when plant parameters alone explained 85% or less of the variation. Overall recognition of 94 agricultural fields using simultaneously acquired aircraft and spacecraft MSS data was 61.8 and 62.8%, respectively; recognition of vegetable fields larger than 10 acres and taller than 25 cm, rose to 88.9 and 100% for aircraft and spacecraft, respectively. Agriculture and rangeland, were well discriminated for the entire county but level 2 categories of vegetables, citrus, and idle cropland, except for citrus, were not.

Wiegand, C. L.

Accelerated test modeling

Cycle life regression model, cycle life prediction model, and acceleration factors are discussed. A method was presented to: (1) select a mathematical model; (2) determine model coefficients using accelerated test data; (3) test model fit of the accelerated test data; and (4) predict normal packs.

Schwartz, D.

External Tank Liquid Hydrogen (LH2) Prepress Regression Analysis Independent Review Technical Consultation Report

The request to conduct an independent review of regression models, developed for determining the expected Launch Commit Criteria (LCC) External Tank (ET)-04 cycle count for the Space Shuttle ET tanking process, was submitted to the NASA Engineering and Safety Center NESC on September 20, 2005. The NESC team performed an independent review of regression models documented in Prepress Regression Analysis, Tom Clark and Angela Krenn, 10/27/05. This consultation consisted of a peer review by statistical experts of the proposed regression models provided in the Prepress Regression Analysis. This document is the consultation's final report.

Parsons, Vickie s.