Search NASASearch

SEARCH · Search NASA

Results for “multivariate experimental data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Pattern recognition systems and procedures

The objectives of the pattern recognition tasks are to develop (1) a man-machine interactive data processing system; and (2) procedures to determine effective features as a function of time for crops and soils. The signal analysis and dissemination equipment, SADE, is being developed as a man-machine interactive data processing system. SADE will provide imagery and multi-channel analog tape inputs for digitation and a color display of the data. SADE is an essential tool to aid in the investigation to determine useful features as a function of time for crops and soils. Four related studies are: (1) reliability of the multivariate Gaussian assumption; (2) usefulness of transforming features with regard to the classifier probability of error; (3) advantage of selecting quantizer parameters to minimize the classifier probability of error; and (4) advantage of using contextual data. The study of transformation of variables (features), especially those experimental studies which can be completed with the SADE system, will be done.

Nelson, G. D.

Investigating lab-scaled offshore wind aerodynamic testing failure and developing solutions for early anomaly detections

As offshore wind systems become more complex, the risk of human error or equipment malfunction increases during experimental testing. This study investigates a lab-scale incident involving a 1 : 50 scale 5 MW wind turbine, where a generator failure led to rotor overspeed and a blade–tower strike. To improve early fault detection, we propose a data-driven method based on multivariate long short-term memory (LSTM) models. High-frequency measurements are projected onto principal components, and anomalies are identified using reconstruction error and its time derivative. Two models are trained on different healthy datasets and tested using single- and multi-principal component (1PC and MPC) variations. Results show that combining both error and error derivative improves detection accuracy. The 1PC model detects faults faster, has a higher recall rate, and achieves a 43 % improvement in anomaly detection accuracy, while the MPC model yields higher precision. This approach provides a simple and effective tool for early anomaly detection in lab-scale experiments, helping to reduce the risk of future failures during the testing of new technologies.

17 WIND ENERGY

A modal analysis of flexible aircraft dynamics with handling qualities implications

A multivariable modal analysis technique is presented for evaluating flexible aircraft dynamics, focusing on meaningful vehicle responses to pilot inputs and atmospheric turbulence. Although modal analysis is the tool, vehicle time response is emphasized, and the analysis is performed on the linear, time-domain vehicle model. In evaluating previously obtained experimental pitch tracking data for a family of vehicle dynamic models, it is shown that flexible aeroelastic effects can significantly affect pitch attitude handling qualities. Consideration of the eigenvalues alone, of both rigid-body and aeroelastic modes, does not explain the simulation results. Modal analysis revealed, however, that although the lowest aeroelastic mode frequency was still three times greater than the short-period frequency, the rigid-body attitude response was dominated by this aeroelastic mode. This dominance was defined in terms of the relative magnitudes of the modal residues in selected vehicle responses.

Schmidt, D. K.

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry

Nonlinear aerodynamic modeling using multivariate orthogonal functions

A technique was developed for global modeling of nonlinear aerodynamic coefficients using multivariate orthogonal functions based on the data. Each orthogonal function retained in the model was decomposed into an expansion of ordinary polynomials in the independent variables, so that the final model could be interpreted as selectively retained terms from a multivariable power series expansion. A predicted squared-error metric was used to determine the orthogonal functions to be retained in the model; analytical derivatives were easily computed. The approach was demonstrated on the Z-body axis aerodynamic force coefficient (Cz) wind tunnel data for an F-18 research vehicle which came from a tabular wind tunnel and covered the entire subsonic flight envelope. For a realistic case, the analytical model predicted experimental values of Cz very well. The modeling technique is shown to be capable of generating a compact, global analytical representation of nonlinear aerodynamics. The polynomial model has good predictive capability, global validity, and analytical differentiability.

Morelli, Eugene A.

Uncertainty Analysis of Inertial Model Attitude Sensor Calibration and Application with a Recommended New Calibration Method

Statistical tools, previously developed for nonlinear least-squares estimation of multivariate sensor calibration parameters and the associated calibration uncertainty analysis, have been applied to single- and multiple-axis inertial model attitude sensors used in wind tunnel testing to measure angle of attack and roll angle. The analysis provides confidence and prediction intervals of calibrated sensor measurement uncertainty as functions of applied input pitch and roll angles. A comparative performance study of various experimental designs for inertial sensor calibration is presented along with corroborating experimental data. The importance of replicated calibrations over extended time periods has been emphasized; replication provides independent estimates of calibration precision and bias uncertainties, statistical tests for calibration or modeling bias uncertainty, and statistical tests for sensor parameter drift over time. A set of recommendations for a new standardized model attitude sensor calibration method and usage procedures is included. The statistical information provided by these procedures is necessary for the uncertainty analysis of aerospace test results now required by users of industrial wind tunnel test facilities.

Tripp, John S.

Multivariate analyses of crater parameters and the classification of craters

Multivariate analyses were performed on certain linear dimensions of six genetic types of craters. A total of 320 craters, consisting of laboratory fluidization craters, craters formed by chemical and nuclear explosives, terrestrial maars and other volcanic craters, and terrestrial meteorite impact craters, authenticated and probable, were analyzed in the first data set in terms of their mean rim crest diameter, mean interior relief, rim height, and mean exterior rim width. The second data set contained an additional 91 terrestrial craters of which 19 were of experimental percussive impact and 28 of volcanic collapse origin, and which was analyzed in terms of mean rim crest diameter, mean interior relief, and rim height. Principal component analyses were performed on the six genetic types of craters. Ninety per cent of the variation in the variables can be accounted for by two components. Ninety-nine per cent of the variation in the craters formed by chemical and nuclear explosives is explained by the first component alone.

Siegal, B. S.

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Wind Tunnel Database Development using Modern Experiment Design and Multivariate Orthogonal Functions

A wind tunnel experiment for characterizing the aerodynamic and propulsion forces and moments acting on a research model airplane is described. The model airplane called the Free-flying Airplane for Sub-scale Experimental Research (FASER), is a modified off-the-shelf radio-controlled model airplane, with 7 ft wingspan, a tractor propeller driven by an electric motor, and aerobatic capability. FASER was tested in the NASA Langley 12-foot Low-Speed Wind Tunnel, using a combination of traditional sweeps and modern experiment design. Power level was included as an independent variable in the wind tunnel test, to allow characterization of power effects on aerodynamic forces and moments. A modeling technique that employs multivariate orthogonal functions was used to develop accurate analytic models for the aerodynamic and propulsion force and moment coefficient dependencies from the wind tunnel data. Efficient methods for generating orthogonal modeling functions, expanding the orthogonal modeling functions in terms of ordinary polynomial functions, and analytical orthogonal blocking were developed and discussed. The resulting models comprise a set of smooth, differentiable functions for the non-dimensional aerodynamic force and moment coefficients in terms of ordinary polynomials in the independent variables, suitable for nonlinear aircraft simulation.

Morelli, Eugene A.

Diurnal Variability of Vertical Structure from a TRMM Passive Microwave "Virtual Radar" Retrieval

Robust description of the diurnal cycle from TRMM observations is complicated by the limitations of Low Earth Orbit (LEO) sampling; from a 'climatological' perspective, sufficient sampling must exist to control for both spatial and seasonal variability, before tackling an additional diurnal component (e.g., with 8 additional 3-hourly or 24 1-hourly bins). For documentation of vertical structure, the narrow sample swath of the TRMM Precipitation Radar limits the resolution of any of these components. A neural-network based 'virtual radar" retrieval has been trained and internally validated, using multifrequency / multipolarization passive microwave(TM1) brightness temperatures and textures parameters and lightning (LIS) observations, as inputs, and PR volumetric reflectivity as targets (outputs). By training the algorithms (essentially highly multivariate, nonlinear regressions) on a very large sample of high-quality co-located data from the center of the TRMM swath, 3D radar reflectivity and derived parameters (VIL, IWC, Echo Tops, etc.) can be retrieved across the entire TMI swath, good to 8-9% over the dynamic range of parameters. As a step in the retrieval (and as an output of the process), each TMI multifrequency pixel (at 85 GHz resolution) is classified into one of the 25 archetypal radar profile vertical structure "types", previously identified using cluster analysis. The dynamic range of retrieved vertical structure appears to have higher fidelity than the current (Version 6) experimental GPROF hydrometeor vertical structure retrievals. This is attributable to correct representation of the prior probabilities of vertical structure variability in the neural network training data, unlike the GPROF cloud-resolving model training dataset used in the V6 algorithms. The LIS lightning inputs are supplementary inputs, and a separate offline neural network has been trained to impute (predict) LIS lightning from passive-microwave-only data. The virtual radar retrieval is thus, in principle, extensible to Aqua/AMSR-E and NPOESS/CMIS passive microwave instruments. The virtual radar approach yields a threefold increase in effective sampling from the mission, albeit of lower-quality "retrieved" data, reducing the variance of local estimates by one third (or the standard deviation by-0.57). In this talk, the variance reduction is leveraged to more finely resolve global diurnal variability in both space and time (local hour).

Boccippio, Dennis J.

A case study demonstration of the soil temperature extrema recovery rates after precipitation cooling at 10-cm soil depth

Since the invention of maximum and minimum thermometers in the 18th century, diurnal temperature extrema have been taken for air worldwide. At some stations, these extrema temperatures were collected at various soil depths also, and the behavior of these temperatures at a 10-cm depth at the Tifton Experimental Station in Georgia is presented. After a precipitation cooling event, the diurnal temperature maxima drop to a minimum value and then start a recovery to higher values (similar to thermal inertia). This recovery represents a measure of response to heating as a function of soil moisture and soil property. Eight different curves were fitted to a wide variety of data sets for different stations and years, and both power and exponential curves were fitted to a wide variety of data sets for different stations and years. Both power and exponential curve fits were consistently found to be statistically accurate least-square fit representations of the raw data recovery values. The predictive procedures used here were multivariate regression analyses, which are applicable to soils at a variety of depths besides the 10-cm depth presented.

Welker, Jean Edward

A gradient model of vegetation and climate utilizing NOAA satellite imagery. Phase 1: Texas transect

A new experimental climatological model/variable termed the sponge, a measure of moisture availability based on daily temperature maxima and minima and precipitation, is tested for potential biogeographic, ecological, and agro-climatological applications. Results, depicted in tabular and graphic from, suggest that, as a generalized climatic index, sponge's simplicity and sensitivity make particularly appropriate for trans-regional biogeographic studies (e.g., large-area and global vegetation monitoring). The feasibility of utilizing NOAA/AVHRR data for vegetation classification was investigated and a vegetation gradient model that utilizes sponge, and AVHRR pixel data (channels 1 and 2) were obtained for 12 locations. The normalized difference values for the AVHRR data when plotted against vegetation characteristics (biomass, net productivity, leaf area) and sponge values suggest that a multivariate gradient model incorporating AVHRR and sponge data may indeed be useful in global vegetation stratification and monitoring.

Greegor, D. H.

Search for Light Pseudoscalar Bosons, Pair-Produced in Higgs Boson Decays in the Four-Electron Final State in Proton-Proton Collisions at $\sqrt{s}=13$ TeV

A search for pairs of light neutral pseudoscalar bosons (𝐴) resulting from the decay of a Higgs boson is performed. The search is conducted using LHC proton-proton collision data at $\sqrt{s}=13$ TeV, collected with the CMS detector in 2016–2018 and corresponding to an integrated luminosity of 138 fb −1 . The 𝐴 boson decays into a highly collimated electron-positron pair. A novel multivariate algorithm using tracks and calorimeter information is developed to identify these distinctive signatures, and events are selected with two such merged electron-positron pairs. No significant excess above the standard model background predictions is observed. Upper limits on the branching fraction for 𝐻 → 𝐴⁢𝐴 → 4⁢𝑒 are set at 95% confidence level, for masses between 10 and 100 MeV and proper decay lengths below 100 μ⁢m, reaching branching fraction sensitivities as low as 10 −5 . This is the first search for Higgs boson decays to four electrons via light pseudoscalars at the LHC. It significantly improves the experimental sensitivity to axionlike particles with masses below 100 MeV.

Hayrapetyan, A. [Yerevan Physics Institute]

Vibration attenuation of the NASA Langley evolutionary structure experiment using H(sub infinity) and structured singular value (micron) robust multivariable control techniques

The use is studied of active control to attenuate structural vibrations of the NASA Langley Phase Zero Evolutionary Structure due to external disturbance excitations. H sub infinity and structured singular value (mu) based control techniques are used to analyze and synthesize control laws for the NASA Langley Controls Structures Interaction (CSI) Evolutionary Model (CEM). The CEM structure experiment provides an excellent test bed to address control design issues for large space structures. Specifically, control design for structures with numerous lightly damped, coupled flexible modes, collocated and noncollocated sensors and actuators and stringent performance specifications. The performance objectives are to attenuate the vibration of the structure due to external disturbances, and minimize the actuator control force. The control design problem formulation for the CEM Structure uses a mathematical model developed with finite element techniques. A reduced order state space model for the control design is formulated from the finite element model. It is noted that there are significant variations between the design model and the experimentally derived transfer function data.

Balas, Gary J.

Remote sensing of earth terrain

Two monographs and 85 journal and conference papers on remote sensing of earth terrain have been published, sponsored by NASA Contract NAG5-270. A multivariate K-distribution is proposed to model the statistics of fully polarimetric data from earth terrain with polarizations HH, HV, VH, and VV. In this approach, correlated polarizations of radar signals, as characterized by a covariance matrix, are treated as the sum of N n-dimensional random vectors; N obeys the negative binomial distribution with a parameter alpha and mean bar N. Subsequently, and n-dimensional K-distribution, with either zero or non-zero mean, is developed in the limit of infinite bar N or illuminated area. The probability density function (PDF) of the K-distributed vector normalized by its Euclidean norm is independent of the parameter alpha and is the same as that derived from a zero-mean Gaussian-distributed random vector. The above model is well supported by experimental data provided by MIT Lincoln Laboratory and the Jet Propulsion Laboratory in the form of polarimetric measurements.

Kong, J. A.

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition

Statistical methods and neural network approaches for classification of data from multiple sources

Statistical methods for classification of data from multiple data sources are investigated and compared to neural network models. A problem with using conventional multivariate statistical approaches for classification of data of multiple types is in general that a multivariate distribution cannot be assumed for the classes in the data sources. Another common problem with statistical classification methods is that the data sources are not equally reliable. This means that the data sources need to be weighted according to their reliability but most statistical classification methods do not have a mechanism for this. This research focuses on statistical methods which can overcome these problems: a method of statistical multisource analysis and consensus theory. Reliability measures for weighting the data sources in these methods are suggested and investigated. Secondly, this research focuses on neural network models. The neural networks are distribution free since no prior knowledge of the statistical distribution of the data is needed. This is an obvious advantage over most statistical classification methods. The neural networks also automatically take care of the problem involving how much weight each data source should have. On the other hand, their training process is iterative and can take a very long time. Methods to speed up the training procedure are introduced and investigated. Experimental results of classification using both neural network models and statistical methods are given, and the approaches are compared based on these results.

Benediktsson, Jon Atli

Experimental analysis of computer system dependability

This paper reviews an area which has evolved over the past 15 years: experimental analysis of computer system dependability. Methodologies and advances are discussed for three basic approaches used in the area: simulated fault injection, physical fault injection, and measurement-based analysis. The three approaches are suited, respectively, to dependability evaluation in the three phases of a system's life: design phase, prototype phase, and operational phase. Before the discussion of these phases, several statistical techniques used in the area are introduced. For each phase, a classification of research methods or study topics is outlined, followed by discussion of these methods or topics as well as representative studies. The statistical techniques introduced include the estimation of parameters and confidence intervals, probability distribution characterization, and several multivariate analysis methods. Importance sampling, a statistical technique used to accelerate Monte Carlo simulation, is also introduced. The discussion of simulated fault injection covers electrical-level, logic-level, and function-level fault injection methods as well as representative simulation environments such as FOCUS and DEPEND. The discussion of physical fault injection covers hardware, software, and radiation fault injection methods as well as several software and hybrid tools including FIAT, FERARI, HYBRID, and FINE. The discussion of measurement-based analysis covers measurement and data processing techniques, basic error characterization, dependency analysis, Markov reward modeling, software-dependability, and fault diagnosis. The discussion involves several important issues studies in the area, including fault models, fast simulation techniques, workload/failure dependency, correlated failures, and software fault tolerance.

Iyer, Ravishankar, K.