Search NASA⌕ Search

SEARCH · Search NASA

Results for “categorical variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Advanced Method Optimization with Categorical and Constrained Continuous Parameters

Traditional approaches to analytical method optimization (e.g., univariate and “guess-and-check”) can be time-consuming, costly, and often fail to identify true optima within the parameter space. Previous work defined and implemented a generalized technique for method optimization for continuous method parameters, but a knowledge gap remains for the incorporation of categorical variables into these advanced method optimization schemes. This work presents and validates a generalized optimization approach that incorporates both continuous and categorical variables while also utilizing a multivariate, multiobjective optimization scheme with Karush–Kuhn–Tucker conditions to bound the optimization space to solutions within the physical limitations of the parameter space. Method optimization from a case study using GC–MS for the analysis of 11 analytical standards with objectives to minimize peak width and maximize peak height resulted in a 3 orders of magnitude improvement in the average peak height and a 2 orders of magnitude improvement in the average peak width compared to the least optimal (but reasonable) instrumental parameters utilized in this study. This approach to optimization allows for a customizable method optimization in which users can include both continuous and categorical variables to achieve objectives specific to their analytical goals. This approach significantly reduces the labor and cost associated with traditional method development approaches and can be applied in a variety of scientific fields across a range of laboratory techniques (e.g., instrument method development, sample preparation, and extraction techniques).

Amorphous materials↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Neural Network Burst Pressure Prediction in Composite Overwrapped Pressure Vessels

Acoustic emission data were collected during the hydroburst testing of eleven 15 inch diameter filament wound composite overwrapped pressure vessels. A neural network burst pressure prediction was generated from the resulting AE amplitude data. The bottles shared commonality of graphite fiber, epoxy resin, and cure time. Individual bottles varied by cure mode (rotisserie versus static oven curing), types of inflicted damage, temperature of the pressurant, and pressurization scheme. Three categorical variables were selected to represent undamaged bottles, impact damaged bottles, and bottles with lacerated hoop fibers. This categorization along with the removal of the AE data from the disbonding noise between the aluminum liner and the composite overwrap allowed the prediction of burst pressures in all three sets of bottles using a single backpropagation neural network. Here the worst case error was 3.38 percent.

Hill, Eric v. K.↗

Temporal studies of compact galactic X-ray sources

Advances in X-ray astronomy due to temporal studies are discussed. Attention is given to X-ray temporal variability categorized into periodic, quasi-periodic, and chaotic. Approximately 95% of all known X-ray sources exhibit chaotic variability, in which there is no apparent pattern or periodicities. The orbits of seven binary X-ray sources are presented, along with empirical findings of neutron star masses. In addition, the interior structure of a supergiant star, HD 77581, is investigated by attempting to measure its apsidal motion. Graphs are provided for the power density spectrum of 4U 1626-67, GX301-2 Doppler delay data, and empirically determined error contours for the mass and radius of four companions stars in X-ray binary systems.

Rappaport, S.↗

CTGAN-TVAE

SAND2026-18914O CTGAN-TVAE (Conditional Tabular Generative Adversarial Networks-Tabular Variational Autoencoders) generates extensive sets of variable generation data through a hybrid framework. It enhances latent space representation by combining TVAE's robust feature-embedding with CTGAN's ability to condition categorical variables such as time. CTGAN-TVAE employs a fully connected neural network within a conditional generative adversarial network framework to manage continuous and categorical data effectively, capturing complex feature interactions without needing sequential modeling. This was developed as part of NNSA-MSIPP: Minority Serving Institution Partnership Program, Grant Number DE-NA0004016. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Newlun, Cody [Sandia National Lab. (SNL-CA), Liver↗

Use of collateral information to improve LANDSAT classification accuracies

Methods to improve LANDSAT classification accuracies were investigated including: (1) the use of prior probabilities in maximum likelihood classification as a methodology to integrate discrete collateral data with continuously measured image density variables; (2) the use of the logit classifier as an alternative to multivariate normal classification that permits mixing both continuous and categorical variables in a single model and fits empirical distributions of observations more closely than the multivariate normal density function; and (3) the use of collateral data in a geographic information system as exercised to model a desired output information layer as a function of input layers of raster format collateral and image data base layers.

Strahler, A. H.↗

Contributions of Astronauts Aerobic Exercise Intensity and Time on Change in VO2peak during Spaceflight

There is considerable variability among astronauts with respect to changes in maximal aerobic capacity (VO2peak) during International Space Station (ISS) missions, ranging from a 5% increase to 30% decline. Individual differences may be due to in-flight aerobic exercise time and intensity. PURPOSE: To evaluate the effects of in-flight aerobic exercise time and intensity on change in VO2peak during ISS missions. METHODS: Astronauts (N=11) performed peak cycle tests approx 60 days before flight (L-60), on flight day (FD) approx 14, and every approx 30 days thereafter. Metabolic gas analysis and heart rate (HR) were measured continuously during the test using the portable pulmonary function system. HR and duration of each in-flight cycle ergometer and treadmill (TM) session were recorded and averaged in time segments corresponding to each peak test. Mixed effects linear regression with exercise mode (TM or cycle) as a categorical variable was used to assess the contributions of exercise intensity (%time >70% peak HR or %time >90% peak HR) and time (min/wk), adjusted for body weight, on %change in VO2peak during the mission, and incorporating the repeated-measures experimental design. RESULTS: 110 observations were included in the model (4-6 peak cycle tests per astronaut, 2 exercise devices). VO2peak was reduced from preflight throughout the mission (FD14: 13+/-13% and FD 105: 8+/-10%). Exercise intensity (%peak HR: FD14=66+/-14; FD105=75+/-8) and time (min/wk: FD14=82+/-46; FD105=158+/-40) increased during flight. The models showed main effects for exercise time and intensity with no interactions between time, intensity, and device (70% peak HR: time [z-score=2.39; P=0.017], intensity [z-score=3.51; P=0.000]; 90% peak HR: time [zscore= 3.31; P=0.001], intensity [z-score=2.24; P=0.025]). CONCLUSION: Exercise time and intensity independently contribute to %change in VO2peak during ISS missions, indicating that there are minimal values for exercise time and intensity required to maintain VO2peak. As the FD105 average exercise intensity and time did not prevent a decline in VO2peak from preflight, astronauts' exercise prescriptions should target at least 160 min of weekly aerobic exercise at an average above 75% peak HR with increased time at intensities above 90% of peak HR starting early in the mission.

Downs, Meghan E.↗

Application of a data-mining method based on Bayesian networks to lesion-deficit analysis

Although lesion-deficit analysis (LDA) has provided extensive information about structure-function associations in the human brain, LDA has suffered from the difficulties inherent to the analysis of spatial data, i.e., there are many more variables than subjects, and data may be difficult to model using standard distributions, such as the normal distribution. We herein describe a Bayesian method for LDA; this method is based on data-mining techniques that employ Bayesian networks to represent structure-function associations. These methods are computationally tractable, and can represent complex, nonlinear structure-function associations. When applied to the evaluation of data obtained from a study of the psychiatric sequelae of traumatic brain injury in children, this method generates a Bayesian network that demonstrates complex, nonlinear associations among lesions in the left caudate, right globus pallidus, right side of the corpus callosum, right caudate, and left thalamus, and subsequent development of attention-deficit hyperactivity disorder, confirming and extending our previous statistical analysis of these data. Furthermore, analysis of simulated data indicates that methods based on Bayesian networks may be more sensitive and specific for detecting associations among categorical variables than methods based on chi-square and Fisher exact statistics.

NASA Discipline Neuroscience↗

The Effect of Habitual Smoking on VO2max

VO2max is associated with many factors, including age, gender, physical activity, and body composition. It is popularly believed that habitual smoking lowers aerobic fitness. PURPOSE: to determine the effect of habitual smoking on VO2max after controlling for age, gender, activity and BMI. METHODS: 2374 men and 375 women employed at the NASA/Johnson Space Center were measured for VO2max by indirect calorimetry (RER>=1.1), activity by the 11 point (0-10) NASA Physical Activity Status Scale (PASS), BMI and smoking pack-yrs (packs day*y of smoking). Age was recorded in years and gender was coded as M=1, W=0. Pack.y was made a categorical variable consisting of four levels as follows: Never Smoked (0), Light (1-10), Regular (11-20), Heavy (>20). Group differences were verified by ANOVA. A General Linear Models (GLM) was used to develop two models to examine the relationship of smoking behavior on VO2max. GLM #1(without smoking) determined the combined effects of age, gender, PASS and BMI on VO2max. GLM #2 (with smoking) determined the added effects of smoking (pack.y groupings) on VO2max after controlling for age, gender, PASS and BMI. Constant errors (CE) were calculated to compare the accuracy of the two models for estimating the VO2max of the smoking subgroups. RESULTS: ANOVA affirmed the mean VO2max of each pack.y grouping decreased significantly (p<0.01) as the level of smoking exposure increased. GLM #1 showed that age, gender, PASS and BMI were independently related with VO2max (R2 = 0.642, SEE = 4.90, p<0.001). The added pack.y variables in GLM #2 were statistically significant (R2 change = 0.7%, p<0.01). Post hoc analysis showed that compared to Never Smoked, the effects on VO2max from Light and Regular smoking habits were -0.83 and -0.85 ml.kg- 1.min-1 respectively (p<0.05). The effect of Heavy smoking on VO2max was -2.56 ml.kg- 1.min-1 (p<0.001). The CE s of each smoking group in GLM #2 was smaller than the CE s of the smoking group counterparts in GLM #1. CONCLUSIONS: After accounting for the effects of gender, age, PASS and BMI the effect of habitual smoking on reducing VO2max is minimal, about 0.85 ml/kg/min, until the habit exceeds 20 pack.y at which point an additional decrease of 1.71 ml/kg/min is noted. Adding pack.y data improves the accuracy of predicting the VO2max of smokers.

Wier, Larry T.↗

Predictive Models of Duration of Ground Delay Programs in New York Area Airports

Initially planned GDP duration often turns out to be an underestimate or an overestimate of the actual GDP duration. This, in turn, results in avoidable airborne or ground delays in the system. Therefore, better models of actual duration have the potential of reducing delays in the system. The overall objective of this study is to develop such models based on logs of GDPs. In a previous report, we described descriptive models of Ground Delay Programs. These models were defined in terms of initial planned duration and in terms of categorical variables. These descriptive models are good at characterizing the historical errors in planned GDP durations. This paper focuses on developing predictive models of GDP duration. Traffic Management Initiatives (TMI) are logged by Air Traffic Control facilities with The National Traffic Management Log (NTML) which is a single system for automated recoding, coordination, and distribution of relevant information about TMIs throughout the National Airspace System. (Brickman, 2004 Yuditsky, 2007) We use 2008-2009 GDP data from the NTML database for the study reported in this paper. NTML information about a GDP includes the initial specification, possibly one or more revisions, and the cancellation. In the next section, we describe general characteristics of Ground Delay Programs. In the third section, we develop models of actual duration. In the fourth section, we compare predictive performance of these models. The final section is a conclusion.

Kulkarni, Deepak↗

A Statistical Theory of Bidirectionality

Original concepts related to the quantification and assessment of bidirectionality in strain-gage balances were introduced by Ulbrich in 2012. These concepts are extended here in three ways: 1) the metric originally proposed by Ulbrich is normalized, 2) a categorical variable is introduced in the regression analysis to account for load polarity, and 3) the uncertainty in both normalized and non-normalized bidirectionality metrics is quantified. These extensions are applied to four representative balances to assess the bidirectionality characteristics of each. The paper is tutorial in nature, featuring reviews of certain elements of regression and formal inference. Principal findings are that bidirectionality appears to be a common characteristic of most balance outputs and that unless it is taken into account, it is likely to consume the entire error budget of a typical balance calibration experiment. Data volume and correlation among calibration loads are shown to have a significant impact on the precision with which bidirectionality metrics can be assessed.

DeLoach, Richard↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

A Framework for Categorizing Important Project Variables

While substantial research has led to theories concerning the variables that affect project success, no universal set of such variables has been acknowledged as the standard. The identification of a specific set of controllable variables is needed to minimize project failure. Much has been hypothesized about the need to match project controls and management processes to individual projects in order to increase the chance for success. However, an accepted taxonomy for facilitating this matching process does not exist. This paper surveyed existing literature on classification of project variables. After an analysis of those proposals, a simplified categorization is offered to encourage further research.

Parsons, Vickie S.↗

Neural Network Burst Pressure Prediction in Graphite/Epoxy Pressure Vessels from Acoustic Emission Amplitude Data

Acoustic emission (AE) data were taken during hydroproof for three sets of ASTM standard 5.75 inch diameter filament wound graphite/epoxy bottles. All three sets of bottles had the same design and were wound from the same graphite fiber; the only difference was in the epoxies used. Two of the epoxies had similar mechanical properties, and because the acoustic properties of materials are a function of their stiffnesses, it was thought that the AE data from the two sets might also be similar; however, this was not the case. Therefore, the three resin types were categorized using dummy variables, which allowed the prediction of burst pressures all three sets of bottles using a single neural network. Three bottles from each set were used to train the network. The resin category, the AE amplitude distribution data taken up to 25 % of the expected burst pressure, and the actual burst pressures were used as inputs. Architecturally, the network consisted of a forty-three neuron input layer (a single categorical variable defining the resin type plus forty-two continuous variables for the AE amplitude frequencies), a fifteen neuron hidden layer for mapping, and a single output neuron for burst pressure prediction. The network trained on all three bottle sets was able to predict burst pressures in the remaining bottles with a worst case error of + 6.59%, slightly greater than the desired goal of + 5%. This larger than desired error was due to poor resolution in the amplitude data for the third bottle set. When the third set of bottles was eliminated from consideration, only four hidden layer neurons were necessary to generate a worst case prediction error of - 3.43%, well within the desired goal.

Hill, Eric v. K.↗

Autonomous Control of Space Nuclear Reactors

Nuclear reactors to support future robotic and manned missions impose new and innovative technological requirements for their control and protection instrumentation. Long-duration surface missions necessitate reliable autonomous operation, and manned missions impose added requirements for failsafe reactor protection. There is a need for an advanced instrumentation and control system for space-nuclear reactors that addresses both aspects of autonomous operation and safety. The Reactor Instrumentation and Control System (RICS) consists of two functionally independent systems: the Reactor Protection System (RPS) and the Supervision and Control System (SCS). Through these two systems, the RICS both supervises and controls a nuclear reactor during normal operational states, as well as monitors the operation of the reactor and, upon sensing a system anomaly, automatically takes the appropriate actions to prevent an unsafe or potentially unsafe condition from occurring. The RPS encompasses all electrical and mechanical devices and circuitry, from sensors to actuation device output terminals. The SCS contains a comprehensive data acquisition system to measure continuously different groups of variables consisting of primary measurement elements, transmitters, or conditioning modules. These reactor control variables can be categorized into two groups: those directly related to the behavior of the core (known as nuclear variables) and those related to secondary systems (known as process variables). Reliable closed-loop reactor control is achieved by processing the acquired variables and actuating the appropriate device drivers to maintain the reactor in a safe operating state. The SCS must prevent a deviation from the reactor nominal conditions by managing limitation functions in order to avoid RPS actions. The RICS has four identical redundancies that comply with physical separation, electrical isolation, and functional independence. This architecture complies with the safety requirements of a nuclear reactor and provides high availability to the host system. The RICS is intended to interface with a host computer (the computer of the spacecraft where the reactor is mounted). The RICS leverages the safety features inherent in Earth-based reactors and also integrates the wide range neutron detector (WRND). A neutron detector provides the input that allows the RICS to do its job. The RICS is based on proven technology currently in use at a nuclear research facility. In its most basic form, the RICS is a ruggedized, compact data-acquisition and control system that could be adapted to support a wide variety of harsh environments. As such, the RICS could be a useful instrument outside the scope of a nuclear reactor, including military applications where failsafe data acquisition and control is required with stringent size, weight, and power constraints.

Merk, John↗

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics↗