Search NASASearch

SEARCH · Search NASA

Results for “feature importance analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Reverse Osmosis (RO) are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in (ultra-filtration) UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square error (RMSE) metric. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent covariates across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is studied for both direct and recursive RF modelling approaches across increasing forecast horizons. Accurate prediction of initial TMP is critical for optimizing RO operations, as it enables the development of robust modelling frameworks by accurately estimating membrane fouling trends, thereby enhancing process efficiency and long-term reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse

Textural signatures for wetland vegetation

This investigation indicates that unique textural signatures do exist for specific wetland communities at certain times in the growing season. When photographs with the proper resolution are obtained, the textural features can identify the spectral features of the vegetation community seen with lower resolution mapping data. The development of a matrix of optimum textural signatures is the goal of this research. Seasonal variations of spectral and textural features are particularly important when performing a vegetations analysis of fresh water marshes. This matrix will aid in flight planning, since expected seasonal variations and resolution requirements can be established prior to a given flight mission.

Whitman, R. I.

Navier-Stokes analyses of the redistribution of inlet temperature distortions in a turbine

The flow exiting the combustor and entering the turbine of a gas turbine engine is known to contain both spatial and temporal variations in total temperature. Although historically it has been presumed that the turbine rotor responded to the average temperature, recent experimental evidence has demonstrated that the rotor actually separated the hotter and cooler streams of fluid so that the hotter fluid migrated toward the pressure surface and the cooler fluid migrated toward the suction surface. In the present study a time-accurate, two-dimensional, thin-layer, Navier-Stokes analysis of a turbine stage was used to analyze this phenomenon. The rough qualitative agreement between the measured and the computed results indicated that the analysis had successfully captured many of the important features of the flow.

Rai, Man Mohan

Mesospheric KHI signatures in reconstructed broadband rocket probe data

In an accepted method of measuring D region electron density with a rocket-borne dc-mode Langmuir probe, the resulting fine-structure signal is modified by a high-pass filter. While this form of data has been successfully used for power spectral analysis, it fails to reveal important spatial domain features without further processing. It is demonstrated that the broadband signal can be recovered by use of a digital filter that corrects for the on-board high-pass filter effects. The D-region broadband structure so recovered for a daytime equatorial flight shows variations in electron density consistent with a model of Kelvin-Helmholtz billows, and not consistent with inertial range turbulence. The isolated regions of irregularity are about 100 m thick near an altitude of 75 km, imbedded in a region with a staircase-structure electron density profile with about 200-m steps. Suggestive similarities to a numerical model of KHI are to undersea observations of KHI-caused microstructure are shown.

Parker, Jay W.

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES

An Eigensystem Realization Algorithm in Frequency Domain for modal parameter identification

This paper demonstrates the close conceptual relationships between time domain and frequency domain approaches to identification of modal parameters for linear systems. A frequency domain eigensystem realization algorithm, via transfer functions, is developed using a known procedure formulated for a time domain eigensystem realization algorithm, via free decay measurement data. An important feature is the capability of windowing to concentrate analysis on the frequency range of interest. The procedure of overlap averaging is used to produce smoother spectra to reduce the effect of noise on identified modal parameters. Examples from simulation and experiments are given to illustrate the validity of formulations derived in the paper.

Juang, J.-N.

Detecting Important Drivers of Gridded Population Modeling With Machine Learning

High-resolution population datasets have been lever-aged across a broad swath of domains, such as climate change, public policy, humanitarian aid, and rescue operations, among others. Machine learning methods were adopted to generate high-resolution or gridded population estimates by using various geospatial input features such as buildings, roads, and nighttime lights. In this study, we evaluate the importance of population features using Random Forest models across three levels of analysis, utilizing permutation measures. Our research aims to address key questions to enhance our understanding of high-resolution population modeling, such as: Are certain features globally (10 countries collectively) more important than others? Do optimal features vary by country? Within each country, do feature importance differ across administrative units? What similarities exist in feature importance at the global, country, and administrative unit levels? To answer these questions, we leverage the Kneedle algorithm to automate the selection of optimum features. We find that there are patterns displayed by features across spatial boundaries, evidenced by the same feature being the most important indicator of population across 7 of the 10 countries modeled. Our findings indicate that while important features may vary across geographies, certain features consistently hold greater importance than others agnostic of geography.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914

Counter tube window and X-ray fluorescence analyzer study

A study was performed to determine the best design tube window and X-ray fluorescence analyzer for quantitative analysis of Venusian dust and condensates. The principal objective of the project was to develop the best counter tube window geometry for the sensing element of the instrument. This included formulation of a mathematical model of the window and optimization of its parameters. The proposed detector and instrument has several important features. The instrument will perform a near real-time analysis of dust in the Venusian atmosphere, and is capable of measuring dust layers less than 1 micron thick. In addition, wide dynamic measurement range will be provided to compensate for extreme variations in count rates. An integral pulse-height analyzer and memory accumulate data and read out spectra for detail computer analysis on the ground.

Hertel, R.

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS

Compressive strain in Lunae Planum-shortening across wrinkle ridges

Wrinkle ridges have long been considered to be structural or structurally controlled features. Most, but not all, recent studies have converged on a model in which wrinkle ridges are structural features formed under compressive stress; the deformation being accommodated by faulting and folding. Given that wrinkle ridges are compressive tectonic features, an analysis of the associated shortening and strain provides important quantitative information about local and regional deformation. Lunae Planum is dominated by north-south trending ridges extending from Kasei Valles in the north to Valles Marineris in the south. To quantify the morphometric character, a photoclinometric study was undertaken for ridges on Lunae Planum using the Davis and Soderblom. More than 25 ridges were examined between long. 57 and 80 deg, lat. 5 to 25 deg N. For each ridge, several profiles were obtained along its length. Ridge width, total relief, and elevation offset were measured for each ridge. Analyses are given.

Plescia, J. B.

Machine Learning-Assisted Recovery of Delicate Kinetic Information from Transient Reactor Experiments

Identifying active sites and their roles in chemical reaction steps remains a vital challenge in heterogeneous catalysis. Transient experiments offer a unique way to probe active sites and distinguish subtle kinetic features. Although physics-based analysis methods may be well-developed, they can be highly susceptible to experimental noise, and smoothing methods may erase or even distort important features; a smooth curve is not always the best curve. We demonstrate a new workflow for the direct interpretation of intrinsic kinetic information from exit flux curves measured in transient reactor experiments. This workflow contains three artificial neural networks (ANNs), including a noise reducer, a concentration predictor, and a rate predictor to analyze experimental data, followed by the virtual TAP (VTAP) physics-based reactor model and density functional theory (DFT) calculations of adsorption energies on specific sites. We use this workflow to analyze the data from experiments titrating Pt/Al 2 O 3 and Pt/SiO 2 catalysts with carbon monoxide (CO) in the temporal analysis of products (TAP) reactor. Our workflow separates the time-evolving chemical reaction and mass transfer information contained in the TAP pulse response. The existence of strong- and weak-binding sites on the Pt/Al 2 O 3 catalyst is observed in the catalyst titration experiment in the transient reactor. The structures of the strong- and weak-binding sites are then identified by using DFT calculations. We find that the Pt/SiO 2 catalyst has only strong-binding sites, which aligns with the inactive support effect of SiO 2 . We demonstrate how machine learning methods provide unique insights with high-resolution data analysis that cannot be achieved by using state-of-the-art physics-based methods.

Adsorption

The Effect of Point-spread Function Interaction with Radiance from Heterogeneous Scenes on Multitemporal Signature Analysis

The point-spread function is an important factor in determining the nature of feature types on the basis of multispectral recorded radiance, particularly from heterogeneous scenes and particularly from scenes which are imaged repetitively, in order to provide thematic characterization by means of multitemporal signature. To demonstrate the effect of the interaction of scene heterogeneity with the point spread function (PSF)1, a template was constructed from the line spread function (LSF) data for the thematic mapper photoflight model. The template was in 0.25 (nominal) pixel increments in the scan line direction across three scenes of different heterogeneity. The sensor output was calculated by considering the calculated scene radiance from each scene element occurring between the contours of the PSF template, plotted on a movable mylar sheet while it was located at a given position.

Duggin, M. J.

Atmospheric-water absorption features near 2.2 micrometers and their importance in high spectral resolution remote sensing

Selective absorption of electromagnetic radiation by atmospheric gases and water vapor is an accepted fact in terrestrial remote sensing. Until recently, only a general knowledge of atmospheric effects was required for analysis of remote sensing data; however, with the advent of high spectral resolution imaging devices, detailed knowledge of atmospheric absorption bands has become increasingly important for accurate analysis. Detailed study of high spectral resolution aircraft data at the U.S. Geological Survey has disclosed narrow absorption features centered at approximately 2.17 and 2.20 micrometers not caused by surface mineralogy. Published atmospheric transmission spectra and atmospheric spectra derived using the LOWTRAN-5 computer model indicate that these absorption features are probably water vapor. Spectral modeling indicates that the effects of atmospheric absorption in this region are most pronounced in spectrally flat materials with only weak absorption bands. Without correction and detailed knowledge of the atmospheric effects, accurate mapping of surface mineralogy (particularly at low mineral concentrations) is not possible.

Kruse, F. A.

In search of stratospheric bromine oxide

The Imaging Spectrometric Observatory (ISO) is capable of recording spectra in the wavelength range of 200 to 12000 Angstroms. Data from a recent Spacelab 1 ATLAS mission has imaged the terrestrial airglow at tangent ray heights of 90 and 150 km. These data contain information about trace atmospheric constituents such as bromine oxide (BrO), hydroxyl (OH), and chlorine dioxide (OClO). The abundances of these species are critical to stratospheric models of catalytic ozone destruction. Heretofore, very few observations were made especially for BrO. Software was developed to purge unwanted solar features from the airglow spectra. The next step is a measure of the strength of the emission features for BrO. The final analysis will yield the scale height of this important compound.

Lestrade, John Patrick

The application of encapsulation material stability data to photovoltaic module life assessment

For any piece of hardware that degrades when subject to environmental and application stresses, the route or sequence that describes the degradation process may be summarized in terms of six key words: LOADS, RESPONSE, CHANGE, DAMAGE, FAILURE, and PENALTY. Applied to photovoltaic modules, these six factors form the core outline of an expanded failure analysis matrix for unifying and integrating relevant material degradation data and analyses. An important feature of this approach is the deliberate differentiation between factors such as CHANGE, DAMAGE, and FAILURE. The application of this outline to materials degradation research facilitates the distinction between quantifying material property changes and quantifying module damage or power loss with their economic consequences. The approach recommended for relating material stability data to photovoltaic module life is to use the degree of DAMAGE to (1) optical coupling, (2) encapsulant package integrity, (3) PV circuit integrity or (4) electrical isolation as the quantitative criterion for assessing module potential service life rather than simply using module power loss.

Coulbert, C. D.

Experimental and theoretical aerodynamic characteristics of a high-lift semispan wing model

Experimental and theoretical aerodynamic characteristics were compared for a high-lift, semispan wing configuration that incorporated a slightly modified version of the NASA Advanced Laminar Flow Control airfoil section. The experimental investigation was conducted in the Langley 14- by 22-Foot Subsonic Tunnel at chord Reynolds numbers of 2.36 and 3.33 million. A two-dimensional airfoil code and a three-dimensional panel code were used to obtain aerodynamic predictions. Two-dimensional data were corrected for three-dimensional effects. Comparisons between predicted and measured values were made for the cruise configuration and for various high-lift configurations. Both codes predicted lift and pitching moment coefficients that agreed well with experiment for the cruise configuration. These parameters were overpredicted for all high-lift configurations. Drag coefficient was underpredicted for all cases. Corrected two-dimensional pressure distributions typically agreed well with experiment, while the panel code overpredicted the leading-edge suction peak on the wing. One important feature missing from both of these codes was a capability for separated flow analysis. The major cause of disparity between the measured data and predictions presented herein was attributed to separated flow conditions.

Applin, Zachary T.