Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Machine Learning Application to Atmospheric Chemistry Modeling

Atmospheric chemistry is a high-dimensionality, large-data problem and thus may be suited to machine-learning algorithms. We show here the potential of a random forest regression algorithm to replace the gas-phase chemistry solver in the GEOS-Chem chemistry model. In this proof-of-concept study, we used one month of model output to train random forest regression models to predict the concentrations of each long-lived chemical species after integration based upon the physical and chemical conditions before the chemical integration. The choice of prediction type has a strong impact on the skill of the regression model. We find best results from predicting the change in concentration for very long-lived species and the absolute concentration for shorter lived species. The skill of the machine learning algorithm is further improved by using a family approach for NO and NO2 rather than treating them independently.By replacing the numerical integrator with the random forest algorithm and running this model for one month, we find that the model is able to reproduce many of the features of the reference chemistry simulation. Replacing the integration methodology with a machine learning algorithm has the potential to be substantially faster. There are a wide range of applications for such an approach, e.g. to generate boundary conditions, for use in air quality forecasts or chemical data assimilation systems, etc.

Keller, Christoph A.↗

Challenges in integrating dissolved organic matter chemodiversity into kinetic models of soil respiration

The chemodiversity of dissolved organic matter (DOM) in soil has been proposed to influence the microbial metabolism and fate of belowground organic carbon (C). However, integrating DOM chemistry into soil C cycle models to improve predictions of C stocks and fluxes—beyond simply considering DOM pool size—remains a challenge. While recent research suggests that incorporating DOM chemodiversity into models can improve predictions of microbial respiration, there is still a lack of mechanistic understanding describing how DOM chemodiversity affects microbial metabolism and soil respiration. Here, we evaluated whether DOM chemodiversity was a determinant of soil respiration using paired measurements of high-resolution DOM chemistry, obtained from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), and potential soil respiration rates from across the United States (U.S.), all data provided by the Molecular Observation Network. Our objectives were to (1) assess statistical relationships between DOM chemodiversity and microbial respiration, and (2) evaluate the ability of kinetic models to leverage DOM chemistry to explain empirical relationships found in statistical models. Statistical regressions revealed that DOM chemodiversity (alpha diversity) was nonlinearly related to potential soil respiration rates, both independently and through its interactions with DOM and total C concentrations. In soils with relatively high DOM but low total C concentrations, potential soil respiration rates were negatively correlated with DOM alpha diversity, whereas in soils with relatively low DOM and high total C concentrations showed the opposite trend. However, when metabolic transition theory kinetic models were modified to include chemodiversity, their performance was comparable to traditional Monod kinetics approaches, which simulate respiration rates as a function of DOM concentration. The inability to account for nonlinearities in DOM chemodiversity–respiration relationships highlight an opportunity to advance substrate uptake kinetics by establishing causal links between DOM chemodiversity, microbial metabolism trade-offs, and potential interactions under varied environmental conditions.

Bioenergetic model↗

Modeling Forest Understory Fires in an Eastern Amazonian Landscape

Forest understory fires are an increasingly important cause of forest impoverishment in Ammonia, but little is known of the landscape characteristics and climatic phenomena that determine their occurrence. We developed empirical functions relating the occurrence of understory fires to landscape features near Paragominas, a 35- yr-old ranching and logging center in eastern Ammonia. An historical sequence of maps of forest understory fire was created based on field interviews With local farmers and Landsat TM images. Several landscape features that might explain spatial variations in the occurrence of understory fires were also mapped and co-registered for each of the sample dates, including: forest fragment size and shape, forest impoverishment through logging and understory fires, source of ignition (settlements and charcoal pits), roads, forest edges, and others. The spatial relationship between forest understory fire and each landscape characteristic was tested by regression analyses. Fire probability models were then developed for various combinations of landscape characteristics. The analyses were conducted separately for years of the El Nino Southern Oscillation (ENSO), which are associated with severe drought in eastern Amazonia, and non-ENS0 years. Most (91 %) of the forest area that burned during the 10-yr sequence caught fire during ENSO years, when severe drought may have increased both forest flammability and the escape of agricultural management fires. Forest understory fires were associated with forest edges, as reported in previous studies from Ammonia. But the strongest predictor of forest fire was the percentage of the forest fragment that had been previously logged or burned. Forest fragment size, distance to charcoal pits, distance to agricultural settlement, proximity to forest edge, and distance to roads were also correlated with forest understory fire. Logistic regression models using information on fragment degradation and distance to ignition sources accurately predicted the location of lss than 80% of the forest fires observed during the ENSO event of 1997- 1998. In this Amazon landscape, forest understory fire is a complex function of several variables that influence both the flammability and ignition exposure of the forest.

Alencar, A. A. C.↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Interpretation of Probabilistic Surface Ozone Forecasts: A Case Study for Philadelphia

The use of probabilistic forecasting has been growing in a variety of disciplines because of its potential to emphasize the degree of uncertainty inherent in a prediction. Interpretation of probabilistic forecasts, however, is oftentimes difficult, deterring users who may benefit from such forecasts. To encourage broader use of probabilistic forecasts in the field of air quality, a process for interpreting forecasts from a statistical probabilistic air quality surface ozone model [the Regression in Self Organizing Map (REGiS)] is demonstrated. Four procedures to convert probabilistic to deterministic forecasts are explored for the Philadelphia, Pennsylvania, metropolitan area. These procedures calibrate the predicted probability of daily maximum 8-h-average ozone exceeding a standard value by 1) estimating climatological relative frequency, 2) establishing a probability of an exceedance threshold as 50%, 3) maximizing the threat score, and 4) determining the unit bias ratio. REGiS is trained using 2000–11 ozone-season (1 May–30 September) data, calibrated using 2012–14 data, and evaluated using 2015–18 data. Assessment of the calibration data with the Pierce skill score suggests an exceedance threshold based on climatological relative frequency for the conversion from probabilistic to deterministic forecasts. Calibrated REGiS generally compares well to predictions from the U.S. national air quality model and operational “expert” forecasts over the evaluation period. For other probabilistic models and situations, different procedures of converting probabilistic to deterministic forecasts may be more beneficial. The methods presented in this paper represent an approach for operational air quality forecasters seeking to use probabilistic model output to support forecasts designed to protect public health.

Nikolay Balashov↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Convective Weather Forecast Accuracy Analysis at Center and Sector Levels

This paper presents a detailed convective forecast accuracy analysis at center and sector levels. The study is aimed to provide more meaningful forecast verification measures to aviation community, as well as to obtain useful information leading to the improvements in the weather translation capacity models. In general, the vast majority of forecast verification efforts over past decades have been on the calculation of traditional standard verification measure scores over forecast and observation data analyses onto grids. These verification measures based on the binary classification have been applied in quality assurance of weather forecast products at the national level for many years. Our research focuses on the forecast at the center and sector levels. We calculate the standard forecast verification measure scores for en-route air traffic centers and sectors first, followed by conducting the forecast validation analysis and related verification measures for weather intensities and locations at centers and sectors levels. An approach to improve the prediction of sector weather coverage by multiple sector forecasts is then developed. The weather severe intensity assessment was carried out by using the correlations between forecast and actual weather observation airspace coverage. The weather forecast accuracy on horizontal location was assessed by examining the forecast errors. The improvement in prediction of weather coverage was determined by the correlation between actual sector weather coverage and prediction. observed and forecasted Convective Weather Avoidance Model (CWAM) data collected from June to September in 2007. CWAM zero-minute forecast data with aircraft avoidance probability of 60% and 80% are used as the actual weather observation. All forecast measurements are based on 30-minute, 60- minute, 90-minute, and 120-minute forecasts with the same avoidance probabilities. The forecast accuracy analysis for times under one-hour showed that the errors in intensity and location for center forecast are relatively low. For example, 1-hour forecast intensity and horizontal location errors for ZDC center were about 0.12 and 0.13. However, the correlation between sector 1-hour forecast and actual weather coverage was weak, for sector ZDC32, about 32% of the total variation of observation weather intensity was unexplained by forecast; the sector horizontal location error was about 0.10. The paper also introduces an approach to estimate the sector three-dimensional actual weather coverage by using multiple sector forecasts, which turned out to produce better predictions. Using Multiple Linear Regression (MLR) model for this approach, the correlations between actual observation and the multiple sector forecast model prediction improved by several percents at 95% confidence level in comparison with single sector forecast.

Wang, Yao↗

Atmospheric Correction of Satellite Ocean-Color Imagery During the PACE Era

The Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission will carry into space the Ocean Color Instrument (OCI), a spectrometer measuring at 5 nm spectral resolution in the ultraviolet (UV) to near infrared (NIR) with additional spectral bands in the shortwave infrared (SWIR), and two multi-angle polarimeters that will overlap the OCI spectral range and spatial coverage, i. e., the Spectrometer for Planetary Exploration (SPEXone) and the Hyper-Angular Rainbow Polarimeter (HARP2). These instruments, especially when used in synergy, have great potential for improving estimates of water reflectance in the post Earth Observing System (EOS) era. Extending the top-of-atmosphere (TOA) observations to the UV, where aerosol absorption is effective, adding spectral bands in the SWIR, where even the most turbid waters are black and sensitivity to the aerosol coarse mode is higher than at shorter wavelengths, and measuring in the oxygen A-band to estimate aerosol altitude will enable greater accuracy in atmospheric correction for ocean color science. The multi-angular and polarized measurements, sensitive to aerosol properties (e.g., size distribution, index of refraction), can further help to identify or constrain the aerosol model, or to retrieve directly water reflectance. Algorithms that exploit the new capabilities are presented, and their ability to improve accuracy is discussed. They embrace a modern, adapted heritage two-step algorithm and alternative schemes (deterministic, statistical) that aim at inverting the TOA signal in a single step. These schemes, by the nature of their construction, their robustness, their generalization properties, and their ability to associate uncertainties, are expected to become the new standard in the future. A strategy for atmospheric correction is presented that ensures continuity and consistency with past and present ocean-color missions while enabling full exploitation of the new dimensions and possibilities. Despite the major improvements anticipated with the PACE instruments, gaps/issues remain to be filled/tackled. They include dealing properly with whitecaps, taking into account Earth-curvature effects, correcting for adjacency effects, accounting for the coupling between scattering and absorption, modeling accurately water reflectance, and acquiring a sufficiently representative dataset of water reflectance in the UV to SWIR. Dedicated efforts, experimental and theoretical, are in order to gather the necessary information and rectify inadequacies. Ideas and solutions are put forward to address the unresolved issues. Thanks to its design and characteristics, the PACE mission will mark the beginning of a new era of unprecedented accuracy in ocean-color radiometry from space.

ocean color↗

Analysis of fractal dimensions of rat bones from film and digital images

OBJECTIVES: (1) To compare the effect of two different intra-oral image receptors on estimates of fractal dimension; and (2) to determine the variations in fractal dimensions between the femur, tibia and humerus of the rat and between their proximal, middle and distal regions. METHODS: The left femur, tibia and humerus from 24 4-6-month-old Sprague-Dawley rats were radiographed using intra-oral film and a charge-coupled device (CCD). Films were digitized at a pixel density comparable to the CCD using a flat-bed scanner. Square regions of interest were selected from proximal, middle, and distal regions of each bone. Fractal dimensions were estimated from the slope of regression lines fitted to plots of log power against log spatial frequency. RESULTS: The fractal dimensions estimates from digitized films were significantly greater than those produced from the CCD (P=0.0008). Estimated fractal dimensions of three types of bone were not significantly different (P=0.0544); however, the three regions of bones were significantly different (P=0.0239). The fractal dimensions estimated from radiographs of the proximal and distal regions of the bones were lower than comparable estimates obtained from the middle region. CONCLUSIONS: Different types of image receptors significantly affect estimates of fractal dimension. There was no difference in the fractal dimensions of the different bones but the three regions differed significantly.

NASA Discipline Musculoskeletal↗

The effect of modeling dose uncertainty on low-boom community noise dose-response curves

In logistic dose-response modeling, failing to account for uncertainty in estimated doses can cause an artificial flattening or attenuation of the slope of the summary curve. In Lee et al. [J. Acoust. Soc. Am. 147(4), pp. 2222-2234 (2020)], data from two NASA low-amplitude sonic boom community noise survey tests were modeled using a Bayesian multilevel logistic regression (MLR) statistical model that assumed there was no uncertainty in the noise dose estimates. However, in these community tests, the noise dose uncertainty was estimated by Page et al. [NASA/CR-2014-218180 and NASA/CR-2020-220589/Volume I] using a leave-one-out method. In the current work, a term was added to extend the Bayesian MLR model to account for the estimated noise dose uncertainty quantified in the Page et al. analyses. This uncertainty term was included in two ways, either as classical or as Berkson uncertainty, and yield similar results. When the uncertainty is accounted for in the Bayesian MLR model, the dose-response curves become 5-10% steeper, but the difference in the noise dose that elicits a 5% highly annoyed response is small (less than 1 dB). This result is encouraging for future X-59 community tests whose survey area will be sparsely populated with noise monitors.

X-59↗

Atmospheric moisture fields derived by satellite observations over the tropical Pacific Ocean

Values of precipitable water are retrieved over the tropical and subtropical Pacific Ocean from TOVS infrared and microwave channel brightness temperature and OLR observations by means of stepwise linear regression. The most useful temperature and moisture sensing channels are pre-selected from sensitivity tests of a radiative transfer model. Numerous models are developed and tested against collocated radiosonde observations and Nimbus-7 SMMR precipitable water estimates. For RAOB comparisons, the best estimator used 15 TOVS predictors and captured 71.1 deg percent of the variance (+0.62 g/sq cm standard error) for column precipitable water; for precipitable water of 700-500 mb bulk layer, these values were 71.7 percent and +/- 0.17 g/sq cm. Little skill of estimated precipitable water was obtained for moisture above 500 mb. Regressions were less skillful against SMMR, unless collocation parameters were tightly controlled; SMMR was less acceptable than RAOB's because of observational drift and errors. Generally, the most skillful predictors were boundary layer brightness temperatures of TOVS channels and satellite estimated stability indices. 'Moisture channels' were hardly useful except for estimating middle and upper tropospheric moisture. Additional regression models were constructed testing the sensitivity to different observational and meteorological characteristics. Models which used some in situ observations surface observations or stabilities calculated from RAOB, were the most successful. The best of these explained 87.5 percent of the variance but the regression selected almost no TOVS channels, relying instead on conventional RAOB and surface observations. A set of four regression models were developed, stratifying atmospheric characteristics on the basis of collocated OLR values. These models improved the variance explained by 5.0 percent; the model associated with the highest OLR values (275 W/sq m less than or equal to OLR; that is, no cloud) showed only marginal skill. Precipitable water fields were generated from the best TOVS-only model for seven days in January 1983 and compared with SMMR-estimated fields. OLR fields and ECMWF precipitable water analysis. The TOVS regression model compared favorably to the SMMR analysis in amplitudes and features. It revealed more evolving synoptic signal than the ECMWF analysis. In synoptically active regions, it differed with respect to the OLR analysis, primarily because of actual differences in the vertical distribution of water vapor.

Chung, Hyosang↗

Search for the associated production of charm quarks and a Higgs boson decaying into a photon pair with the ATLAS detector

A search for the production of a Higgs boson and one or more charm quarks, in which the Higgs boson decays into a photon pair, is presented. This search uses proton-proton collision data with a centre-of-mass energy of $\sqrt{s}$ = 13 TeV and an integrated luminosity of 140 fb −1 recorded by the ATLAS detector at the Large Hadron Collider. The analysis relies on the identification of charm-quark-containing jets, and adopts an approach based on Gaussian process regression to model the non-resonant di-photon background. The observed (expected, assuming the Standard Model signal) upper limit at the 95% confidence level on the cross-section for producing a Higgs boson and at least one charm-quark-containing jet that passes a fiducial selection is found to be 10.6 pb (8.8 pb). The observed (expected) measured cross-section for this process is 5.3 ± 3.2 pb (2.9 ± 3.1 pb).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Temperature and Composition Dependence Modeling of Viscosity and Electrical Conductivity of Low-Activity Waste Glass Melts

The development of models that accurately relate the properties of a glass melt to its temperature and composition is important for glass formulation, melter control, and modeling the melt flow, refractory corrosion, and production rate. Using a database consisting of more than 4,000 data points measured between 900 °C and 1250 °C for over 600 unique low-activity waste glass compositions, we developed models for the melt viscosity and electrical conductivity. Models based on the Gaussian process regression approach outperformed models based on the Vogel–Fulcher–Tammann equation according to four standard metrics and yielded reliable prediction intervals. The models found primarily linear effects between properties and individual components, except for the effect of the Na 2 O mass fraction on the electrical conductivity. The effects were found to be consistent with current theories on physical processes involved with those properties.

36 MATERIALS SCIENCE↗

Protocol to detect dilution cycles in chemostat experiments and estimate growth rate slopes with linear modeling with R software chemostat_regression

Chemostat growth chambers measure optical density over time and require manual calculation of growth rates. Here, we present chemostat_regression, R software that enables users to automatically identify chemostat cycles and estimate growth rate using a linear regression approach. We describe steps for creating requisite software environment(s), formatting input data, executing the software via command line/RStudio/R-Shiny, interpreting results, assessing the validity of results, and modifying input parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Insights into Tetravalent Np Speciation in HNO 3 through Spectroelectrochemistry and Multivariate Analysis

In situ optical spectroscopy, spectropotentiometry, and multivariate analysis were applied to the Np(IV) nitrate system to better understand speciation and quantify HNO 3 concentration. Thin-layer spectropotentiometry, or spectroelectrochemistry, was leveraged to isolate and stabilize Np(IV) without compromising the solution conditions and generate representative Vis-NIR absorption spectra from 0.5 to 10 M HNO 3 and benchmark the corresponding Np(IV) molar absorptivity coefficients. Spectra were described with principal component analysis (PCA) to identify the purest Np(IV) absorbance spectra among other oxidation states [e.g., Np(V/VI)] at each acid concentration and then to identify the primary sources of variance within each Np(IV) spectrum with respect to Np(IV) nitrate complexes. Then, partial least-squares regression (PLSR) and support vector regression (SVR) models were built to predict HNO 3 concentration from the Np(IV) spectral data. The nonlinear SVR model outperformed the linear PLSR model for the HNO 3 concentration predictions. Finally, the inclusion of spectra collected in edge and center point HNO 3 concentrations in the calibration set was determined to be crucial for producing models with strong predictive capabilities. The multivariate approach used in this study makes it possible to quantify HNO 3 concentration solely based on Np(IV) absorption spectra, which is essential to quantifying processing streams in various online monitoring applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Near-Real-Time Material Tracking: Combining Vis–NIR Spectroscopy with Flow Sensing for Accurate Nd(III) Quantification

A fiber-optic visible–near-infrared (vis–NIR) absorption spectroscopy and flow sensor system has been developed for near-real-time tracking of Nd mass in the effluent stream from a column in a fume hood. The approach leverages two unique data streams and a partial least-squares regression (PLSR) model trained on vis–NIR absorption spectra of Nd(III) (0–1.5 M) in 1 M HNO 3 . In-line volumetric flow rate and vis–NIR spectra are measured in sequence after a chromatography column. The time stamps from each data stream are then synchronized, which allows integrated volumes to be combined with Nd(III) molarities predicted by a PLSR model to accurately calculate the Nd mass flowing through the column. This integrated measurement provides instantaneous mass flow and accumulates these data over time to obtain the total mass processed. The methodology developed in this study contributes critical technical infrastructure to improve monitoring capabilities to support chemical separations and the production of strategic materials and isotopes.

Irvine, Sawyer B. [Oak Ridge National Laboratory (↗

The Role of Data Filtering in Open Source Software Ranking and Selection

Faced with more than 100M open source projects, a more manageable small subset is needed for most empirical investigations. More than half of the research papers in leading venues investigated filtering projects by some measure of popularity with explicit or implicit arguments that unpopular projects are not of interest, may not even represent "real" software projects, or that less popular projects are not worthy of study. However, such filtering may have enormous effects on the results of the studies if and precisely because the sought-out response or prediction is in any way related to the filtering criteria.This paper exemplifies the impact of this common practice on research outcomes, specifically how filtering of software projects on GitHub based on inherent characteristics affects the assessment of their popularity. Using a dataset of over 100,000 repositories, we used multiple regression to model the number of stars -a commonly used proxy for popularity- based on factors such as the number of commits, the duration of the project, the number of authors and the number of core developers. Our control model included the entire dataset, while a second filtered model considered only projects with ten or more authors. The results indicated that while certain characteristics of the repository consistently predict popularity, the filtering process significantly alters the relationships between these characteristics and the response. We found that the number of commits exhibited a positive correlation with popularity in the control sample but showed a negative correlation in the filtered sample. These findings highlight the potential biases introduced by data filtering and emphasize the need for careful sample selection in empirical research of mining software repositories. We recommend that empirical work should either analyze complete datasets such as World of Code, or employ stratified random sampling from a complete dataset to ensure that filtering is not biasing the results.

Malviya Thakur, Addi↗

BASS

SAND2026-17001O BASS implements Bayesian Adaptive Spline Surfaces in MATLAB and serves as a surrogate model for regression applications. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Tucker, J. Derek [Sandia National Lab. (SNL-CA), L↗