Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

Application of Temperature Sensitivities During Iterative Strain-Gage Balance Calibration Analysis

A new method is discussed that may be used to correct wind tunnel strain-gage balance load predictions for the influence of residual temperature effects at the location of the strain-gages. The method was designed for the iterative analysis technique that is used in the aerospace testing community to predict balance loads from strain-gage outputs during a wind tunnel test. The new method implicitly applies temperature corrections to the gage outputs during the load iteration process. Therefore, it can use uncorrected gage outputs directly as input for the load calculations. The new method is applied in several steps. First, balance calibration data is analyzed in the usual manner assuming that the balance temperature was kept constant during the calibration. Then, the temperature difference relative to the calibration temperature is introduced as a new independent variable for each strain--gage output. Therefore, sensors must exist near the strain--gages so that the required temperature differences can be measured during the wind tunnel test. In addition, the format of the regression coefficient matrix needs to be extended so that it can support the new independent variables. In the next step, the extended regression coefficient matrix of the original calibration data is modified by using the manufacturer specified temperature sensitivity of each strain--gage as the regression coefficient of the corresponding temperature difference variable. Finally, the modified regression coefficient matrix is converted to a data reduction matrix that the iterative analysis technique needs for the calculation of balance loads. Original calibration data and modified check load data of NASA's MC60D balance are used to illustrate the new method.

Ulbrich, N.↗

Applications of cluster analysis to satellite soundings

The advantages of the use of cluster analysis in the improvement of satellite temperature retrievals were evaluated since the use of natural clusters, which are associated with atmospheric temperature soundings characteristic of different types of air masses, has the potential for improving stratified regression schemes in comparison with currently used methods which stratify soundings based on latitude, season, and land/ocean. The method of discriminatory analysis was used. The correct cluster of temperature profiles from satellite measurements was located in 85% of the cases. Considerable improvement was observed at all mandatory levels using regression retrievals derived in the clusters of temperature (weighted and nonweighted) in comparison with the control experiment and with the regression retrievals derived in the clusters of brightness temperatures of 3 MSU and 5 IR channels.

Munteanu, M. J.↗

Analysis of oscillatory motion of a light airplane at high values of lift coefficient

A modified stepwise regression is applied to flight data from a light research air-plane operating at high angles at attack. The well-known phenomenon referred to as buckling or porpoising is analyzed and modeled using both power series and spline expansions of the aerodynamic force and moment coefficients associated with the longitudinal equations of motion.

Batterson, J. G.↗

Analysis of upper stratospheric Umkehr ozone profile data for trends and the effects of stratospheric aerosols

The effect of stratospheric aerosols on Umkehr estimates of long-term ozone depletion associated with chlorofluoromethanes (CFMs) is considered in a statistical time series trend analysis. Time series models are estimated using monthly averages of Umkehr measurements made over the last 15 to 20 years. The time series regression models incorporate seasonal, trend and noise factors and an additional factor to account for the effects of atmospheric aerosols on the Umkehr measurements. The analysis indicates a statistically significant relation with atmospheric aerosol transmission in the Umkehr layers and implies that the relation is an important factor in any time series trend analysis of Umkehr data. Taking this relation into account, statistically significant negative trends were found in the Upper Umkehr layers. It is pointed out that upper stratospheric ozone could be sensitive to long-term solar variability as well as other possible influences in addition to CFM-induced effects, and therefore the cause or causes of the estimated trend cannot be unambiguously estimated using current Umkehr data.

Reinsel, G. C.↗

Evaluation of the annoyance due to helicopter rotor noise

A program was conducted in which 25 test subjects adjusted the levels of various helicopter rotor spectra until the combination of the harmonic noise and a broadband background noise was judged equally annoying as a higher level of the same broadband noise spectrum. The subjective measure of added harmonic noise was equated to the difference in the two levels of broadband noise. The test participants also made subjective evaluations of the rotor noise signatures which they created. The test stimuli consisted of three degrees of rotor impulsiveness, each presented at four blade passage rates. Each of these 12 harmonic sounds was combined with three broadband spectra and was adjusted to match the annoyance of three different sound pressure levels of broadband noise. Analysis of variance indicated that the important variables were level and impulsiveness. Regression analyses indicated that inclusion of crest factor improved correlation between the subjective measures and various objective or physical measures.

Sternfeld, H., Jr.↗

A Method for Aircraft Concept Selection Using Multicriteria Interactive Genetic Algorithms

The problem of aircraft concept selection has become increasingly difficult in recent years as a result of a change from performance as the primary evaluation criteria of aircraft concepts to the current situation in which environmental effects, economics, and aesthetics must also be evaluated and considered in the earliest stages of the decision-making process. This has prompted a shift from design using historical data regression techniques for metric prediction to the use of physics-based analysis tools that are capable of analyzing designs outside of the historical database. The use of optimization methods with these physics-based tools, however, has proven difficult because of the tendency of optimizers to exploit assumptions present in the models and drive the design towards a solution which, while promising to the computer, may be infeasible due to factors not considered by the computer codes. In addition to this difficulty, the number of discrete options available at this stage may be unmanageable due to the combinatorial nature of the concept selection problem, leading the analyst to arbitrarily choose a sub-optimum baseline vehicle. These concept decisions such as the type of control surface scheme to use, though extremely important, are frequently made without sufficient understanding of their impact on the important system metrics because of a lack of computational resources or analysis tools. This paper describes a hybrid subjective/quantitative optimization method and its application to the concept selection of a Small Supersonic Transport. The method uses Genetic Algorithms to operate on a population of designs and promote improvement by varying more than sixty parameters governing the vehicle geometry, mission, and requirements. In addition to using computer codes for evaluation of quantitative criteria such as gross weight, expert input is also considered to account for criteria such as aeroelasticity or manufacturability which may be impossible or too computationally expensive to consider explicitly in the analysis. Results indicate that concepts resulting from the use of this method represent designs which are promising to both the computer and the analyst, and that a mapping between concepts and requirements that would not otherwise be apparent is revealed.

Buonanno, Michael↗

Determinants of Time to Fatigue during Non-Motorized Treadmill Exercise

Treadmill exercise is commonly used for aerobic and anaerobic conditioning. During non-motorized treadmill exercise, the subject must provide the power necessary to drive the treadmill belt. The purpose of this study was to determine what factors affected the time to fatigue on a pair of non-motorized treadmills. Twenty subjects (10 males/10 females) attempted to complete five minutes of locomotion during separate trials at 3.22, 4.83, 6.44, 8.05, 9.66, and 11.27 km (raised dot) h(sup -1). Total exercise time (less than or equal to 5 min) was recorded. Exercise time was converted to the amount of 15 second intervals completed. Peak oxygen uptake (VO2) was measured using a graded exercise test on a standard treadmill, and anthropometric measures were collected from each subject before entering into the study. A Cox proportional hazards regression model was used to determine significant predictive factors in a multivariate analysis. Non-motorized treadmill speed and absolute peak VO2 were found to be significant predictors of exercise time, but there was no effect of anthropometric characteristics. Gender was found to be a predictor of treadmill time, but this was likely due to a higher peak VO2 in males than in females. These results were not affected by the type of treadmill tested in this study. Coaches and therapists should consider the cardiovascular fitness of an athlete or client when prescribing target speed since these factors are related to the total exercise time than can be achieved on a non-motorized treadmill.

DeWitt, John K.↗

Significant Tropospheric O3 Production from Extratropical Forest Fires:When? When Not?

There is significant controversy on whether extratropical fires contribute importantly to widespread ozone productionabsent the addition of urban pollutant nitrogen oxides (Jaffe and Widger, 2012, Singh et al., 2012,2013). Wereport a significant range of O3 production early in the fire plume history and report controls on which notable O3production occurs. The current data set is the airborne observations made on board NASAs highly instrumentedDC-8 aircraft during the ARCTAS (2008) and SEAC4RS (2013) campaigns within the Western United States. Aclear analysis of fire emissions was aided considerably by MERET (a Mixed Effects Regressions Emissions Technique).This technique allows consistent emission factors and enhancement ratios for O3 for individual aircraftsamples largely free from uncertainties resulting from mixing and entrainment which plague many published estimates.(Yokelson et al, 2013, Chatfield and Andreae, 2016). We find support for the reasonable idea that the ratioof nitrogen oxides (NOx) to volatile organic carbon (VOC) emissions controls new O3 production. The evidence isfor significant O3 production from high-fuel-nitrogen fuels (as evidenced by acetonitrile) and extremely hot largefires (like the Yosemite Rim Fire of 2013 which we analyze), while others do not. VOC emissions factors varysignificantly from fire to fire (and from different samples in the Rim Fire). Relative CO production (aka modifiedcombustion efficiency) is only one factor describing VOC emissions factors.

Chatfield, R.↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗

Pulsed Phase Lock Loop Device for Monitoring Intracranial Pressure During Space Flight

We have developed an ultrasonic device to monitor ICP waveforms non-invasively from cranial diameter oscillations using a NASA-developed pulsed phase lock loop (PPLL) technique. The purpose of this study was to attempt to validate the PPLL device for reliable recordings of ICP waveforms and analysis of ICP dynamics in vivo. METHODS: PPLL outputs were recorded in patients during invasive ICP monitoring at UCSD Medical Center (n=10). RESULTS: An averaged linear regression coefficient between ICP and PPLL waveform data during one cardiac cycle in all patients is 0.88 +/- 0.02 (mean +/- SE). Coherence function analysis indicated that ICP and PPLL waveforms have high correlation in the lst, 2nd, and 3rd harmonic waves associated with a cardiac cycle. CONCLUSIONS: PPLL outputs represent ICP waveforms in both frequency and time domains. PPLL technology enables in vivo evaluation of ICP dynamics non-invasively, and can acquire continuous ICP waveforms during spaceflight because of compactness and non-invasive nature.

Ueno, Toshiaki↗

Partial Least Squares and Neural Networks for Quantitative Calibration of Laser-induced Breakdown Spectroscopy (LIBs) of Geologic Samples

The ChemCam instrument [1] on the Mars Science Laboratory (MSL) rover will be used to obtain the chemical composition of surface targets within 7 m of the rover using Laser Induced Breakdown Spectroscopy (LIBS). ChemCam analyzes atomic emission spectra (240-800 nm) from a plasma created by a pulsed Nd:KGW 1067 nm laser. The LIBS spectra can be used in a semiquantitative way to rapidly classify targets (e.g., basalt, andesite, carbonate, sulfate, etc.) and in a quantitative way to estimate their major and minor element chemical compositions. Quantitative chemical analysis from LIBS spectra is complicated by a number of factors, including chemical matrix effects [2]. Recent work has shown promising results using multivariate techniques such as partial least squares (PLS) regression and artificial neural networks (ANN) to predict elemental abundances in samples [e.g. 2-6]. To develop, refine, and evaluate analysis schemes for LIBS spectra of geologic materials, we collected spectra of a diverse set of well-characterized natural geologic samples and are comparing the predictive abilities of PLS, cascade correlation ANN (CC-ANN) and multilayer perceptron ANN (MLP-ANN) analysis procedures.

Anderson, R. B.↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Relationship Between Column-Density and Surface Mixing Ratio: Statistical Analysis of O3 and NO2 Data from the July 2011 Maryland DISCOVER-AQ Mission

To investigate the ability of column (or partial column) information to represent surface air quality, results of linear regression analyses between surface mixing ratio data and column abundances for O3 and NO2 are presented for the July 2011 Maryland deployment of the DISCOVER-AQ mission. Data collected by the P-3B aircraft, ground-based Pandora spectrometers, Aura/OMI satellite instrument, and simulations for July 2011 from the CMAQ air quality model during this deployment provide a large and varied data set, allowing this problem to be approached from multiple perspectives. O3 columns typically exhibited a statistically significant and high degree of correlation with surface data (R(sup 2) > 0.64) in the P- 3B data set, a moderate degree of correlation (0.16 < R(sup 2) < 0.64) in the CMAQ data set, and a low degree of correlation (R(sup 2) < 0.16) in the Pandora and OMI data sets. NO2 columns typically exhibited a low to moderate degree of correlation with surface data in each data set. The results of linear regression analyses for O3 exhibited smaller errors relative to the observations than NO2 regressions. These results suggest that O3 partial column observations from future satellite instruments with sufficient sensitivity to the lower troposphere can be meaningful for surface air quality analysis.

nitrogen dioxide↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

Acousto-ultrasonic verification of the strength of filament wound composite material

The concept of acousto-ultrasonic (AU) waveform partitioning was applied to nondestructive evaluation of mechanical properties in filament wound composites (FWC). A series of FWC test specimens were subjected to AU analysis and the results were compared with destructively measured interlaminar shear strengths (ISS). AU stress-wave factor (SWF) measurements gave greater than 90 percent correlation coefficient upon regression against the ISS. This high correlation was achieved by employing the appropriate time and frequency domain partitioning as dictated by wave propagation path analysis. There is indication that different SWF frequency partitions are sensitive to ISS at different depths below the surface.

Kautz, H. E.↗

Experiments with Test Case Generation and Runtime Analysis

Software testing is typically an ad hoc process where human testers manually write many test inputs and expected test results, perhaps automating their execution in a regression suite. This process is cumbersome and costly. This paper reports preliminary results on an approach to further automate this process. The approach consists of combining automated test case generation based on systematically exploring the program's input domain, with runtime analysis, where execution traces are monitored and verified against temporal logic specifications, or analyzed using advanced algorithms for detecting concurrency errors such as data races and deadlocks. The approach suggests to generate specifications dynamically per input instance rather than statically once-and-for-all. The paper describes experiments with variants of this approach in the context of two examples, a planetary rover controller and a space craft fault protection system.

Artho, Cyrille↗

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Meteoric 10Be Flux Calibration Data for the East River Watershed, Colorado, USA

This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.

East River↗