Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sparse regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Sparse Regression as a Sparse Eigenvalue Problem

We extend the l0-norm "subspectral" algorithms for sparse-LDA [5] and sparse-PCA [6] to general quadratic costs such as MSE in linear (kernel) regression. The resulting "Sparse Least Squares" (SLS) problem is also NP-hard, by way of its equivalence to a rank-1 sparse eigenvalue problem (e.g., binary sparse-LDA [7]). Specifically, for a general quadratic cost we use a highly-efficient technique for direct eigenvalue computation using partitioned matrix inverses which leads to dramatic x103 speed-ups over standard eigenvalue decomposition. This increased efficiency mitigates the O(n4) scaling behaviour that up to now has limited the previous algorithms' utility for high-dimensional learning problems. Moreover, the new computation prioritizes the role of the less-myopic backward elimination stage which becomes more efficient than forward selection. Similarly, branch-and-bound search for Exact Sparse Least Squares (ESLS) also benefits from partitioned matrix inverse techniques. Our Greedy Sparse Least Squares (GSLS) generalizes Natarajan's algorithm [9] also known as Order-Recursive Matching Pursuit (ORMP). Specifically, the forward half of GSLS is exactly equivalent to ORMP but more efficient. By including the backward pass, which only doubles the computation, we can achieve lower MSE than ORMP. Experimental comparisons to the state-of-the-art LARS algorithm [3] show forward-GSLS is faster, more accurate and more flexible in terms of choice of regularization

Exact Sparse Least Squares (ESLS)↗

Reevaluation of Stratospheric Ozone Trends From SAGE II Data Using a Simultaneous Temporal and Spatial Analysis

This paper details a new method of regression for sparsely sampled data sets for use with time-series analysis, in particular the Stratospheric Aerosol and Gas Experiment (SAGE) II ozone data set. Non-uniform spatial, temporal, and diurnal sampling present in the data set result in biased values for the long-term trend if not accounted for. This new method is performed close to the native resolution of measurements and is a simultaneous temporal and spatial analysis that accounts for potential diurnal ozone variation. Results show biases, introduced by the way data is prepared for use with traditional methods, can be as high as 10%. Derived long-term changes show declines in ozone similar to other studies but very different trends in the presumed recovery period, with differences up to 2% per decade. The regression model allows for a variable turnaround time and reveals a hemispheric asymmetry in derived trends in the middle to upper stratosphere. Similar methodology is also applied to SAGE II aerosol optical depth data to create a new volcanic proxy that covers the SAGE II mission period. Ultimately this technique may be extensible towards the inclusion of multiple data sets without the need for homogenization.

Damadeo, R. P.↗

Application of Sparse Identification of Nonlinear Dynamics for Physics-Informed Learning

Advances in machine learning and deep neural networks has enabled complex engineering tasks like image recognition, anomaly detection, regression, and multi-objective optimization, to name but a few. The complexity of the algorithm architecture, e.g., the number of hidden layers in a deep neural network, typically grows with the complexity of the problems they are required to solve, leaving little room for interpreting (or explaining) the path that results in a specific solution. This drawback is particularly relevant for autonomous aerospace and aviation systems, where certifications require a complete understanding of the algorithm behavior in all possible scenarios. Including physics knowledge in such data-driven tools may improve the interpretability of the algorithms, thus enhancing model validation against events with low probability but relevant for system certification. Such events include, for example, spacecraft or aircraft sub-system failures, for which data may not be available in the training phase. This paper investigates a recent physics-informed learning algorithm for identification of system dynamics, and shows how the governing equations of a system can be extracted from data using sparse regression. The learned relationships can be utilized as a surrogate model which, unlike typical data-driven surrogate models, relies on the learned underlying dynamics of the system rather than large number of fitting parameters. The work shows that the algorithm can reconstruct the differential equations underlying the observed dynamics using a single trajectory when no uncertainty is involved. However, the training set size must increase when dealing with stochastic systems, e.g., nonlinear dynamics with random initial conditions.

Corbetta, Matteo↗

Efficient Parametric Uncertainty Analysis of an Earth Entry Vehicle Concept Using Least Angle Regression

The objective of this work was to outline and apply an efficient and accurate parametric un-certainty propagation approach to the analysis of convective heating on an Earth entry vehicle concept. The described approach was based on Least Angle Regression used to solve a sparse and underdetermined linear system in the point-collocation non-intrusive polynomial chaos surrogate method. This approach involved an iterative process to computing the non-zero terms of the underlying polynomial chaos model using only enough samples to converge uncertainty interval predictions and Sobol index values based global nonlinear sensitivity estimates. The Earth entry vehicle was analyzed at three points along a representative trajectory for a Mars return mission. 329 sources of uncertainty were identified in the computational fluid dynamics model used to predict the forebody convective heating. These included uncertainty in flow field chemical rates, collision integrals, heats of formation, surface finite rate char model reaction rates, wall roughness height, and the turbulent Schmidt number. Results from this study showed that convective heating uncertainty as high as 50% of the nominal was predicted with only about 50 evaluations of the computational model. This was far fewer than would be required for a sampling-based approach or a full basis polynomial chaos model, which would have required over 50,000 samples. Additionally, results showed that over 90% of the total convective heating uncertainty was due to uncertainty in the N2catalytic rate on the surface, while the remainder of the uncertainty was attributed to the turbulent Schmidt number and the wall roughness uncertainties.

Thomas K West IV↗

Review: Strategies for Using Satellite-Based Products in Modeling PM2.5 and Short-Term Pollution Episodes

Short-term air pollution episodes motivate improved understanding of the association between air pollution and acute morbidity and mortality episodes, and triggers required mitigation plans. A variety of methods have been employed to estimate exposure to air pollution episodes, including GIS-based dispersion models, interpolation between sparse monitoring sites, land-use regression models, optimization models, line- or area-dispersion plume models, and models using information from imaging satellites, often including land-use and meteorological variables. There has been increasing use of satellite-borne aerosol products for assessing short-term air quality events. They provide better spatial coverage, but currently at the price of low temporal coverage and rather crude spatial resolution. This brief review of using satellite data for modeling short-term air quality and pollution events. The review can be pursued as a practical guide for modeling air quality with satellite-based products, as it includes important questions that should be considered in both the study design as well as the model development stages. Progress in this field is detailed and includes published models and their use in environmental and health studies. Both current and future satellite-borne capabilities are covered. It also provides links to access and download relevant datasets and some R code for data processing and modeling.

Meytar Sorek-Hamer↗

Estimation of Surface Air Temperature Over Central and Eastern Eurasia from MODIS Land Surface Temperature

Surface air temperature (T(sub a)) is a critical variable in the energy and water cycle of the Earth.atmosphere system and is a key input element for hydrology and land surface models. This is a preliminary study to evaluate estimation of T(sub a) from satellite remotely sensed land surface temperature (T(sub s)) by using MODIS-Terra data over two Eurasia regions: northern China and fUSSR. High correlations are observed in both regions between station-measured T(sub a) and MODIS T(sub s). The relationships between the maximum T(sub a) and daytime T(sub s) depend significantly on land cover types, but the minimum T(sub a) and nighttime T(sub s) have little dependence on the land cover types. The largest difference between maximum T(sub a) and daytime T(sub s) appears over the barren and sparsely vegetated area during the summer time. Using a linear regression method, the daily maximum T(sub a) were estimated from 1 km resolution MODIS T(sub s) under clear-sky conditions with coefficients calculated based on land cover types, while the minimum T(sub a) were estimated without considering land cover types. The uncertainty, mean absolute error (MAE), of the estimated maximum T(sub a) varies from 2.4 C over closed shrublands to 3.2 C over grasslands, and the MAE of the estimated minimum Ta is about 3.0 C.

Shen, Suhung↗

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Interhemispheric comparison of atmospheric circulation features as evaluated from NIMBUS satellite data

Findings are presented for IRIS data from NIMBUS 3 in mapping the global ozone distribution. The seasonal and regional variations of ozone, especially in the Southern Hemisphere, reveal features that were not evident from the sparse ground-based ozone observation network in this hemisphere. A regression analysis was undertaken for temperature and height fields on radiance data. Spectrum analyses of upper wind data from the North American section and Australia were completed.

Reiter, E. R.↗

LAI inversion from optical reflectance using a neural network trained with a multiple scattering model

The inversion of the leaf area index (LAI) canopy parameter from optical spectral reflectance measurements is obtained using a backpropagation artificial neural network trained using input-output pairs generated by a multiple scattering reflectance model. The problem of LAI estimation over sparse canopies (LAI < 1.0) with varying soil reflectance backgrounds is particularly difficult. Standard multiple regression methods applied to canopies within a single homogeneous soil type yield good results but perform unacceptably when applied across soil boundaries, resulting in absolute percentage errors of >1000 percent for low LAI. Minimization methods applied to merit functions constructed from differences between measured reflectances and predicted reflectances using multiple-scattering models are unacceptably sensitive to a good initial guess for the desired parameter. In contrast, the neural network reported generally yields absolute percentage errors of <30 percent when weighting coefficients trained on one soil type were applied to predicted canopy reflectance at a different soil background.

Smith, James A.↗

Dynamic Ensemble Prediction of Cognitive Performance in Space

Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.

Danni Tu↗

A Study on the Potential Applications of Satellite Data in Air Quality Monitoring and Forecasting

In this study we explore the potential applications of MODIS (Moderate Resolution Imaging Spectroradiometer) -like satellite sensors in air quality research for some Asian regions. The MODIS aerosol optical thickness (AOT), NCEP global reanalysis meteorological data, and daily surface PM(sub 10) concentrations over China and Thailand from 2001 to 2009 were analyzed using simple and multiple regression models. The AOT-PM(sub 10) correlation demonstrates substantial seasonal and regional difference, likely reflecting variations in aerosol composition and atmospheric conditions, Meteorological factors, particularly relative humidity, were found to influence the AOT-PM(sub 10) relationship. Their inclusion in regression models leads to more accurate assessment of PM(sub 10) from space borne observations. We further introduced a simple method for employing the satellite data to empirically forecast surface particulate pollution, In general, AOT from the previous day (day 0) is used as a predicator variable, along with the forecasted meteorology for the following day (day 1), to predict the PM(sub 10) level for day 1. The contribution of regional transport is represented by backward trajectories combined with AOT. This method was evaluated through PM(sub 10) hindcasts for 2008-2009, using ohservations from 2005 to 2007 as a training data set to obtain model coefficients. For five big Chinese cities, over 50% of the hindcasts have percentage error less than or equal to 30%. Similar performance was achieved for cities in northern Thailand. The MODIS AOT data are responsible for at least part of the demonstrated forecasting skill. This method can be easily adapted for other regions, but is probably most useful for those having sparse ground monitoring networks or no access to sophisticated deterministic models. We also highlight several existing issues, including some inherent to a regression-based approach as exemplified by a case study for Beijing, Further studies will be necessa1Y before satellite data can see more extensive applications in the operational air quality monitoring and forecasting.

Li, Can↗

Mitigating the Impacts of Measurement Error in the Quesst Mission Community Noise Study

Beginning in 2025, the NASA Quesst mission will conduct a series of community response tests involving flyovers of the X-59 aircraft at select localities across the United States. Several waves of a longitudinal survey will be administered over approximately one month of testing in order to capture perceptual responses to low-amplitude sonic booms, or “sonic thumps”. Simultaneously, noise exposure levels will be estimated by fusing model-based predictions with measurements taken from a sparse network of monitors in the region. As one of the aims of the study is to produce a dose-response curve, a regression model relating perceptual response to noise exposure levels, it is important to acknowledge the potential attenuation bias that results from measurement error in the estimated noise exposure levels. In this presentation we review and compare several methods for dealing with measurement error in generalized linear mixed models. The methods are demonstrated on simulated data and real data collected during past NASA risk reduction studies.

measurement error↗

Some methods of computing platform transmitter terminal location estimates

A position estimation algorithm was developed to track a humpback whale tagged with an ARGOS platform after a transmitter deployment failure and the whale's diving behavior precluded standard methods. The algorithm is especially useful where a transmitter location program exists; it determines the classical keplarian elements from the ARGOS spacecraft position vectors included with the probationary file messages. A minimum of three distinct messages are required. Once the spacecraft orbit is determined, the whale is located using standard least squares regression techniques. Experience suggests that in instances where circumstances inherent in the experiment yield message data unsuitable for the standard ARGOS reduction, (message data may be too sparse, span an insufficient period, or include variable-length messages). System ARGOS can still provide much valuable location information if the user is willing to accept the increased location uncertainties.

Hoisington, C. M.↗

A Comparison of Soil Moisture Retrieval Models Using SIR-C Measurements over the Little Washita River Watershed

Six SIR-C L-band measurements over the Little Washita River watershed in Chickasha, Oklahoma during 11-17 April 1994 have been analyzed for studying the change of soil moisture in the region. Two algorithms developed recently for estimation of moisture content in bare soil were applied to these measurements and the results were compared with those sampled on the ground. There is a good agreement between the values of soil moisture estimated by either one of the algorithms and those measured from ground sampling for bare or sparsely vegetated fields. The standard error from this comparison is on the order of 0.05-0.06 cu cm/cu cm, which is comparable to that expected from a regression between backscattering coefficients and measured soil moisture. Both algorithms provide a poor estimation of soil moisture or fail to give solutions to areas covered with moderate or dense vegetation. Even for bare soils the number of pixels that bear no numerical solution from the application of either one of the two algorithms to the data is not negligible. Results from using one of these algorithms indicate that the fraction of these pixels becomes larger as the bare soils become drier. The other algorithm generally gives a larger fraction of these pixels when the fields are vegetation-covered. The implication and impact of these features are discussed in this article.

Wang, J. R.↗

The effect of modeling dose uncertainty on low-boom community noise dose-response curves

In logistic dose-response modeling, failing to account for uncertainty in estimated doses can cause an artificial flattening or attenuation of the slope of the summary curve. In Lee et al. [J. Acoust. Soc. Am. 147(4), pp. 2222-2234 (2020)], data from two NASA low-amplitude sonic boom community noise survey tests were modeled using a Bayesian multilevel logistic regression (MLR) statistical model that assumed there was no uncertainty in the noise dose estimates. However, in these community tests, the noise dose uncertainty was estimated by Page et al. [NASA/CR-2014-218180 and NASA/CR-2020-220589/Volume I] using a leave-one-out method. In the current work, a term was added to extend the Bayesian MLR model to account for the estimated noise dose uncertainty quantified in the Page et al. analyses. This uncertainty term was included in two ways, either as classical or as Berkson uncertainty, and yield similar results. When the uncertainty is accounted for in the Bayesian MLR model, the dose-response curves become 5-10% steeper, but the difference in the noise dose that elicits a 5% highly annoyed response is small (less than 1 dB). This result is encouraging for future X-59 community tests whose survey area will be sparsely populated with noise monitors.

X-59↗

Results of a statistical approach to rainfall estimation using Nimbus 5 6.7 micrometers and 11.5 micrometers THIR data

Nimbus 5 6.7 mm and 11.5 mm temperature humidity infrared radiometer (THIR) data were used in a simple multiple regression scheme to test the feasibility of using these data to estimate hourly rainfall. Throughout the test area (85 W to 105 W and 45 N to 30 N) subareas (8 deg x 6 deg) were chosen from which point to point and areal statistics were obtained. Four subsets of data were used. The first consisted of only those surface stations indicating precipitation whose latitude and longitude coincided with the THIR grid points. A second used surface stations 0.1 degree from the THIR grid points. The third was a combination of subsets one and two. A reciprocal distance weighting scheme was used to derive precipitation values in data sparse areas. A fourth subset was made using these data combined with the data from subsets one and two. Point estimates resulted in negative correlations between estimated and grid derived "surface" precipitation. One degree areal estimates showed a slight improvement with a correlation coefficient of approximately 0.11. Single regression areal estimates resulted in correlations of approximately 0.11 and 0.20 for the 6.7 mm and 11.5 mm data respectively. These poor results were attributed to problems which are inherent in the satellite data (location errors, short temporal span of data, wavelength of sensors, etc.) and the lack of sufficient surface data to better verify the satellite estimate.

Ormsby, J. P.↗