Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

A Universal Threshold for the Assessment of Load and Output Residuals of Strain-Gage Balance Data

A new universal residual threshold for the detection of load and gage output residual outliers of wind tunnel strain{gage balance data was developed. The threshold works with both the Iterative and Non{Iterative Methods that are used in the aerospace testing community to analyze and process balance data. It also supports all known load and gage output formats that are traditionally used to describe balance data. The threshold's definition is based on an empirical electrical constant. First, the constant is used to construct a threshold for the assessment of gage output residuals. Then, the related threshold for the assessment of load residuals is obtained by multiplying the empirical electrical constant with the sum of the absolute values of all first partial derivatives of a given load component. The empirical constant equals 2.5 microV/V for the assessment of balance calibration or check load data residuals. A value of 0.5 microV/V is recommended for the evaluation of repeat point residuals because, by design, the calculation of these residuals removes errors that are associated with the regression analysis of the data itself. Data from a calibration of a six-component force balance is used to illustrate the application of the new threshold definitions to real{world balance calibration data.

wind tunnel balance↗

Changes in the Structure and Propagation of the MJO with Increasing CO2

Changes in the Madden-Julian Oscillation (MJO) with increasing CO2 concentrations are examined using the Goddard Institute for Space Studies Global Climate Model (GCM). Four simulations performed with fixed CO2 concentrations of 0.5, 1, 2 and 4 times pre-industrial levels using the GCM coupled with a mixed layer ocean model are analyzed in terms of the basic state, rainfall and moisture variability, and the structure and propagation of the MJO.The GCM simulates basic state changes associated with increasing CO2 that are consistent with results from earlier studies: column water vapor increases at approximately 7.1% K(exp -1), precipitation also increases but at a lower rate (approximately 3% K(exp -1)), and column relative humidity shows little change. Moisture and rainfall variability intensify with warming. Total moisture and rainfall variability increases at a rate that is similar to that of the mean state change. The intensification is faster in the MJO-related anomalies than in the total anomalies, though the ratio of the MJO band variability to its westward counterpart increases at a much slower rate. On the basis of linear regression analysis and space-time spectral analysis, it is found that the MJO exhibits faster eastward propagation, faster westward energy dispersion, a larger zonal scale and deeper vertical structure in warmer climates.

Adames, Angel F.↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

Assessing Flooding Vulnerability to Assist High Water Intervention and Urban Planning Programs in the Charles River Watershed with NASA DEVELOP

The Charles River watershed intersects 35 municipalities within the Boston Metropolitan Area and has a total population of 1.2 million, making it one of the most densely populated watersheds in New England. In recent years, the watershed has observed higher rates of flood inundation, mainly due to increased development, extreme precipitation events, and increased surface runoff. As the frequency of flooding events increases and a changing climate poses an ongoing threat to local communities, governments and organizations in Massachusetts are in need of accurate flood risk assessments. NASA DEVELOP partnered with the Charles River Watershed Association, the Town of Natick’s Office of Sustainability, and the Massachusetts Audubon Society to assess flood vulnerability and susceptibility in the watershed. The team used Landsat 5 Thematic Mapper, Landsat 8 Operational Land Imager, Sentinel-1 C-Band Synthetic Aperture Radar, and Sentinel-2 MultiSpectral Instrument to assess the feasibility of identifying the extent of past flood events using remote sensing. After identifying images that overlapped with reported flood events, the team concluded that it was not feasible to use Earth observation data to detect localized flooding in the time available for this study. Instead, the Federal Emergency Management Agency (FEMA) 100-year floodplain was used as a proxy for areas where flooding may occur. The team used statistical regression analysis and validation and supervised classification to develop a flood susceptibility map, incorporating several flood conditioning factors. The susceptibility maps were calibrated to various thresholds, including two that highlight hypothetical flooding under more liberal and more conservative planning scenarios. These were overlaid with demographic and socioeconomic data to create flood vulnerability maps. The team’s flood susceptibility maps showed an improvement in capturing known flood events over the FEMA 100-year and 500-year floodplain maps. These results will be improved with the addition of stormwater drainage mapping and precipitation data. Results can be used to fill in the gaps to help the stakeholders understand their communities’ vulnerability and susceptibility to flooding and improve their preparedness plans.

Trista Brophy↗

Application of Support Vector Regression to Derive Crater Depth/Diameter From Satellite Images

Through the study of impact crater shapes, one can draw important conclusions about the nature and evolution of planetary surfaces [e.g., 1-4].In particular, studying the depth (d) to diameter (D)ratio (d/D) of a population of impact craters, in combination with crater count statistics, can yield valuable insights regarding rates of erosion and burial[5]. Motivated by the great abundance of available planetary surface image data, the goal of this project is to develop an efficient way to estimate d/D from satellite images of impact craters for which stereo information is not available [6]. We set out to develop and train a machine learning algorithm to extract d/D from a dataset of synthetic impact crater images for which model d/D is known. The applications of machine learning to planetary science are numerous and diverse [7], including automatic planetary surface mapping [8] and the detection of impact craters [9]. Our algorithm makes use of Support Vector Regression (SVR), which is a type of Support Vector Machine (SVM) [10, 11].SVMs are a branch of supervised machine learning valued for their straightforward implementation and versatility in solving both classification and regression problems. In regression analysis, an SVR algorithm produces a hyperplane function to fit the training data points, as well as an ε-tube that surrounds the hyperplane. Tunable hyperparameters include the width of the ε-tube (ε) and the amount an algorithm is penalized for points which fall outside the ε-tube.

L R Chin↗

Assessment of aerosol burden over Ghana

Although air pollution in Ghana is ranked number one in environmental health threats to public health and sixth to cause of deaths, routine monitoring is rare. This paper presents fourteen years (2005-2018) assessment of aerosol optical depth (AOD) at 3 km resolution from MODIS Aqua and Terra satellites to ascertain the Spatio-temporal and seasonal distribution of aerosols over Ghana and its major cities. The MODIS AOD at 3 km were validated against ground-based Aerosol Robotic Network (AERONET) AODs to ascertain the suitability of the MODIS 3 km data for air quality application in the region. The contribution of distant aerosols to city aerosol loadings was also assessed with Hybrid Single-Particle Lagrangian Integrated Trajectory (HYSPLIT) backscatter model. A moderate-high aerosol burden (AODs ~ 0.50) was observed over Ghana with a significant contribution from the pre-monsoon season. City centres of Takoradi and Kumasi showed higher aerosol loads (AODs ~ 0.80) than Accra and Tamale. The HYSPLIT model showed that distant or transported aerosol sources to the city centres were of both marine and land generated origins. Linear regression analysis between MODIS AOD and AERONET AOD showed a reasonably good correlation of ~ 0.60 for Aqua and Terra. From the validation analysis, both Aqua and Terra satellites can be used for air quality monitoring over Ghana; however, more ground research must be conducted to ascertain better aerosol model assumptions for the region.

Aerosols↗

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Central Park Ecological Conservation: Assessing Tree Health Conditions in New York City’s Central Park with NASA Earth Observation Data

The Central Park Conservancy stewards New York City’s iconic Central Park with a mission to preserve the park for all. This mission is complicated by the spread of Dutch elm disease (DED) which has threatened the culturally and ecologically significant American elm tree (Ulmus americana). Central Park is home to one of the largest and last remaining urban concentrations of American elm and the Conservancy currently protects them through integrated pest management. This paper discusses an interdisciplinary feasibility study that assessed the application of NASA Earth observations from 2014 to 2023 to detect changes in forest phenology possibly related to DED. Landsat 8 and 9 imagery was used to calculate multiyear time series of the Normalized Difference Vegetation Index (NDVI) and quantify changes in land surface phenology for a given year. A pixel-based logistic regression analysis was performed using changes in NDVI, tree site locations, and recorded occurrences of trees infected with DED as inputs. The results of this analysis show that changes in NDVI derived from Landsat data are capable of detecting unhealthy tree canopies with 71% precision and healthy tree canopies with 41% precision. The study had uncertainties and limitations due to the spatial and temporal resolutions of Landsat, the natural variability in land surface phenology and NDVI, and the attempt to detect disease impacts while disease prevention and mitigation is occurring. As is, the findings of this study and its methods provide managers with an approach for integrating Earth observations to make more informed decisions in the application and timing of urban forest management activities.

Central Park↗

Assessing Tree Health Conditions in New York City’s Central Park with Earth Observation Data

The Central Park Conservancy stewards New York City’s iconic Central Park with a mission to preserve the park for all. This mission is complicated by the spread of Dutch elm disease (DED) which has threatened the culturally and ecologically significant American elm tree (Ulmus americana). Central Park is home to one of the largest and last remaining urban concentrations of American elm and the Conservancy currently protects them through integrated pest management. This project is an interdisciplinary feasibility study that assessed the application of NASA Earth observations from 2014 to 2023 to detect changes in forest phenology possibly related to DED. Landsat 8 and 9 imagery was used to calculate a multiyear time series of the Normalized Difference Vegetation Index (NDVI) and quantify changes in land surface phenology. A pixel-based logistic regression analysis was performed using changes in NDVI, tree site locations, and recorded occurrences of trees infected with DED as inputs. The results of this analysis show that changes in NDVI derived from Landsat data are capable of detecting unhealthy tree canopies with 71% precision and healthy tree canopies with 41% precision. The study had uncertainties and limitations due to the spatial and temporal resolutions of Landsat, the natural variability in land surface phenology and NDVI, and the attempt to detect disease impacts while disease prevention and mitigation are occurring. As is, the findings of this study and its methods provide managers with an approach for integrating Earth observations to make more informed decisions in the application and timing of urban forest management activities.

John Hocknell↗

Understanding and Modeling Pooled Rideshare Acceptance: Influential Factors, Preferred User Experiences, and Implications

This dissertation explores factors influencing pooled rideshare (PR) adoption to provide actionable insights for transportation network companies (TNCs) and policymakers. PR allows travelers to share rides with unknown passengers, offering benefits such as cost reduction and congestion relief. However, adoption remains limited due to safety concerns, privacy issues, and trust in rideshare platforms. A national U.S. survey with 5,385 respondents examined transportation preferences and barriers to PR adoption. Exploratory and confirmatory factor analyses identified five key factors influencing PR consideration—safety, service experience, privacy, traffic/environment, and time/cost. Second factor analyses examined ways to optimize PR experiences, revealing four factors—comfort/ease of use, convenience, vehicle technology/accessibility, and passenger safety. Privacy concerns, for instance, using regression analysis, were found to reduce the likelihood of PR adoption by 77%, and convenience had the potential to increase it by 156%. The Pooled Rideshare Acceptance Model (PRAM), based on the Technology Acceptance Model, assessed the impact of these factors using the Structural Equation Model (SEM). Privacy, safety, trust, and convenience had a large effect (Cohen's f2 > 0.35) on PR acceptance, while multigroup analyses (PRAMMA) explored 16 demographic variables such as gender, generation, and income, emphasizing the need for tailored strategies. Based on all the statistical analysis and workshops using descriptive statistics, 95 actionable recommendations were made from the riders' perspective. Findings highlight the importance of customized services, user experience improvements, and policy interventions to enhance PR adoption. This dissertation provides a roadmap for future research and policy development, ensuring evidence-based, practical strategies to improve PR services in the U.S. and beyond.

Gangadharaiah, Rakesh↗

Understanding Peelle’s Pertinent Puzzle bias in generalized least squares regression through eigenspectrum analysis

Certain correlation structures in the data covariance matrix (DCM) used for generalized least squares (GLS) regression can result in biased estimates, commonly known in the field of nuclear data evaluation as Peele’s Pertinent Puzzle (PPP). This article introduces a generative, forward modeling framework within which the PPP bias is characterized through an eigenspectrum analysis of the DCM. This analysis highlights the root cause of the bias, generalizes the problem beyond the nuclear data field, and provides insight to the problem regimes where it can occur. What follows is an understanding that the bias can show up for any experimental neutron time-of-flight data for which systematic uncertainties have been quantified. Lastly, a discussion of the adaptation of cross validation approaches that require pre-whitening to incorporate the known ‘fix’ to the PPP bias in the GLS estimator.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed in 2011 that may be used to analyze both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The new algorithm is a simplified version of a more complex search algorithm that was originally developed at the NASA Ames Balance Calibration Laboratory. The new algorithm has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression models. Therefore, the simplified search algorithm is not intended to replace the original search algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm either fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new regression model search algorithm.

multivariate experimental data↗

Developing Empirical Lightning Cessation Forecast Guidance for the Cape Canaveral Air Force Station and Kennedy Space Center

This research addresses the 45th Weather Squadron's (45WS) need for improved guidance regarding lightning cessation at Cape Canaveral Air Force Station and Kennedy Space Center (KSC). KSC's Lightning Detection and Ranging (LDAR) network was the primary observational tool to investigate both cloud-to-ground and intracloud lightning. Five statistical and empirical schemes were created from LDAR, sounding, and radar parameters derived from 116 storms. Four of the five schemes were unsuitable for operational use since lightning advisories would be canceled prematurely, leading to safety risks to personnel. These include a correlation and regression tree analysis, three variants of multiple linear regression, event time trending, and the time delay between the greatest height of the maximum dBZ value to the last flash. These schemes failed to adequately forecast the maximum interval, the greatest time between any two flashes in the storm. The majority of storms had a maximum interval less than 10 min, which biased the schemes toward small values. Success was achieved with the percentile method (PM) by separating the maximum interval into percentiles for the 100 dependent storms.

LDAR (LIGHTNING DETECTION AND RANGING)↗

Elongated particles in flow: commentary on small-angle scattering investigations

Here, this work thoroughly examines several analytical tools, each possessing a different level of mathematical intricacy, for the purpose of characterizing the orientation distribution function of elongated objects under flow. Our investigation places an emphasis on connecting the orientation distribution to the small-angle scattering spectra measured experimentally. The diverse range of mathematical approaches investigated herein provide insights into the flow behavior of elongated particles from different perspectives and serve as powerful tools for elucidating the complex interplay between flow dynamics and the orientation distribution function.

36 MATERIALS SCIENCE↗

Serum bile acid and unsaturated fatty acid profiles of non-alcoholic fatty liver disease in type 2 diabetic patients

The understanding of bile acid (BA) and unsaturated fatty acid (UFA) profiles, as well as their dysregulation, remains elusive in individuals with type 2 diabetes mellitus (T2DM) coexisting with non-alcoholic fatty liver disease (NAFLD). Investigating these metabolites could offer valuable insights into the pathophy-siology of NAFLD in T2DM. Our aim is to identify potential metabolite biomarkers capable of distinguishing between NAFLD and T2DM. A training model was developed involving 399 participants, comprising 113 healthy controls (HCs), 134 individuals with T2DM without NAFLD, and 152 individuals with T2DM and NAFLD. External validation encompassed 172 participants. NAFLD patients were divided based on liver fibrosis scores. The analytical approach employed univariate testing, orthogonal partial least squares-discriminant analysis, logistic regression, receiver operating characteristic curve analysis, and decision curve analysis to pinpoint and assess the diagnostic value of serum biomarkers. Compared to HCs, both T2DM and NAFLD groups exhibited diminished levels of specific BAs. In UFAs, particular acids exhibited a positive correlation with NAFLD risk in T2DM, while the ω-6:ω-3 UFA ratio demonstrated a negative correlation. Levels of α-linolenic acid and γ-linolenic acid were linked to significant liver fibrosis in NAFLD. The validation cohort substantiated the predictive efficacy of these biomarkers for assessing NAFLD risk in T2DM patients. This study underscores the connection between altered BA and UFA profiles and the presence of NAFLD in individuals with T2DM, proposing their potential as biomarkers in the pathogenesis of NAFLD.

60 APPLIED LIFE SCIENCES↗

Cascade Optimization for Aircraft Engines With Regression and Neural Network Analysis - Approximators

The NASA Engine Performance Program (NEPP) can configure and analyze almost any type of gas turbine engine that can be generated through the interconnection of a set of standard physical components. In addition, the code can optimize engine performance by changing adjustable variables under a set of constraints. However, for engine cycle problems at certain operating points, the NEPP code can encounter difficulties: nonconvergence in the currently implemented Powell's optimization algorithm and deficiencies in the Newton-Raphson solver during engine balancing. A project was undertaken to correct these deficiencies. Nonconvergence was avoided through a cascade optimization strategy, and deficiencies associated with engine balancing were eliminated through neural network and linear regression methods. An approximation-interspersed cascade strategy was used to optimize the engine's operation over its flight envelope. Replacement of Powell's algorithm by the cascade strategy improved the optimization segment of the NEPP code. The performance of the linear regression and neural network methods as alternative engine analyzers was found to be satisfactory. This report considers two examples-a supersonic mixed-flow turbofan engine and a subsonic waverotor-topped engine-to illustrate the results, and it discusses insights gained from the improved version of the NEPP code.

Patnaik, Surya N.↗