Search NASASearch

SEARCH · Search NASA

Results for “high dimensional data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Physics-Informed Active Learning With Simultaneous Weak-Form Latent Space Dynamics Identification

The parametric greedy latent space dynamics identification (gLaSDI) framework has demonstrated promising potential for accurate and efficient modeling of high-dimensional nonlinear physical systems. However, it remains challenging to handle noisy data. Here, to enhance robustness against noise, we incorporate the weak-form estimation of nonlinear dynamics (WENDy) into gLaSDI. In the proposed weak-form gLaSDI (WgLaSDI) framework, an autoencoder and WENDy are trained simultaneously to discover intrinsic nonlinear latent-space dynamics of high-dimensional data. Compared with the standard sparse identification of nonlinear dynamics (SINDy) employed in gLaSDI, WENDy enables variance reduction and robust latent space discovery, therefore leading to more accurate and efficient reduced-order modeling. Furthermore, the greedy physics-informed active learning in WgLaSDI enables adaptive sampling of optimal training data on the fly for enhanced modeling accuracy. The effectiveness of the proposed framework is demonstrated by modeling various nonlinear dynamical problems, including viscous and inviscid Burgers' equations, time-dependent radial advection, and the Vlasov equation for plasma physics. With data that contains 5%–10% Gaussian white noise, WgLaSDI outperforms gLaSDI by orders of magnitude, achieving 1%–7% relative errors. Compared with the high-fidelity models, WgLaSDI achieves 121 to 1779x speed-up.

97 MATHEMATICS AND COMPUTING

Comparisons of a Three-Dimensional, Full Navier Stokes Computer Model with High Mach Number Combuster Test Data

Comparisons between scramjet combustor data and a three-dimensional full Navier-Stokes calculation have been made to verify and substantiate computational fluid dynamics (CFD) codes and application procedures. High Mach number scramjet combustor development will rely heavily on CFD applications to provide wind tunnel-equivalent data of quality sufficient to design, build and fly hypersonic aircraft. Therefore. detailed comparisons between CFD results and test data are imperative. An experimental case is presented, for which combustor wall static pressures were measured and flow-fieid interferograms were obtained. A computer model was done of the experiment, and counterpart parameters are compared with experiment. The experiment involved a subscale combustor designed and fabricated for the National Aero-Space Plane Program, and tested in the Calspan Corporation 96" hypersonic shock tunnel. The combustor inlet ramp was inclined at a 20 angle to the shock tunnel nozzle axis, and resulting combustor entrance flow conditions simulated freestream M=10. The combustor body and cowl walls were instrumented with static pressure transducers, and the combustor lateral walls contained windows through which flowfield holographic interferograms were obtained. The CFD calculation involved a three-dimensional time-averaged full Navier-Stokes code applied to the axial flow segment containing fuel injection and combustion. The full Navier-Stokes approach allowed for mixed supersonic and subsonic flow, downstream-upstream communication in subsonic flow regions, and effects of adverse pressure gradients. The code included hydrogen-air chemistry in the combustor segment which begins near fuel injection and continues through combustor exhaust. Combustor ramp and inlet segments on the combustor lateral centerline were modelled as two dimensional. Comparisons to be shown include calculated versus measured wall static pressures as functions of axial flow coordinate, and calculated path-averaged density contours versus an holographic Interferogram.

Watkins, William B.

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation

Data-driven projection pursuit adaptation of polynomial chaos expansions for dependent high-dimensional parameters

Uncertainty quantification (UQ) and inference involving a large number of parameters are valuable tools for problems associated with heterogeneous and non-stationary behaviors. The difficulty with these problems is exacerbated when these parameters are statistically dependent requiring statistical characterization over joint measures. Probabilistic modeling methodologies stand as effective tools in the realms of UQ and inference. Among these, polynomial chaos expansions (PCE), when adapted to low-dimensional quantities of interest (QoI), provide effective yet accurate approximations for these QoI in terms of an adapted orthogonal basis. These adaptation techniques have been cast as projection pursuits in Gaussian Hilbert space in what has been referred to as a projection pursuit adaptation (PPA) by Xiaoshu Zeng and Roger Ghanem (2023). The PPA method efficiently identifies an optimal low-dimensional space for representing the QoI and simultaneously evaluates an optimal PCE within that space. The quality of this approximation clearly depends on the size of the training dataset, which is typically a function of the adapted reduced dimension. Here, the complexity of the problem is thus mediated by the complexity of the low-dimensional quantity of interest and not the complexity of the high-dimensional parameter space.

Data-driven

Data driven discovery and quantification of hyperspectral leaf reflectance phenotypes across a maize diversity panel

Abstract Estimates of plant traits derived from hyperspectral reflectance data have the potential to efficiently substitute for traits, which are time or labor intensive to manually score. Typical workflows for estimating plant traits from hyperspectral reflectance data employ supervised classification models that can require substantial ground truth datasets for training. We explore the potential of an unsupervised approach, autoencoders, to extract meaningful traits from plant hyperspectral reflectance data using measurements of the reflectance of 2151 individual wavelengths of light from the leaves of maize ( Zea mays ) plants harvested from 1658 field plots in a replicated field trial. A subset of autoencoder‐derived variables exhibited significant repeatability, indicating that a substantial proportion of the total variance in these variables was explained by difference between maize genotypes, while other autoencoder variables appear to capture variation resulting from changes in leaf reflectance between different batches of data collection. Several of the repeatable latent variables were significantly correlated with other traits scored from the same maize field experiment, including one autoencoder‐derived latent variable (LV8) that predicted plant chlorophyll content modestly better than a supervised model trained on the same data. In at least one case, genome‐wide association study hits for variation in autoencoder‐derived variables were proximal to genes with known or plausible links to leaf phenotypes expected to alter hyperspectral reflectance. In aggregate, these results suggest that an unsupervised, autoencoder‐based approach can identify meaningful and genetically controlled variation in high‐dimensional, high‐throughput phenotyping data and link identified variables back to known plant traits of interest.

Tross, Michael C.

COOP 3D ARPA Experiment 109 National Center for Atmospheric Research

Coupled atmospheric and hydrodynamic forecast models were executed on the supercomputing resources of the National Center for Atmospheric Research (NCAR) in Boulder, Colorado and the Ohio Supercomputing Center (OSC)in Columbus, Ohio. respectively. The interoperation of the forecast models on these geographically diverse, high performance Cray platforms required the transfer of large three dimensional data sets at very high information rates. High capacity, terrestrial fiber optic transmission system technologies were integrated with those of an experimental high speed communications satellite in Geosynchronous Earth Orbit (GEO) to test the integration of the two systems. Operation over a spacecraft in GEO orbit required modification of the standard configuration of legacy data communications protocols to facilitate their ability to perform efficiently in the changing environment characteristic of a hybrid network. The success of this performance tuning enabled the use of such an architecture to facilitate high data rate, fiber optic quality data communications between high performance systems not accessible to standard terrestrial fiber transmission systems. Thus obviating the performance degradation often found in contemporary earth/satellite hybrids.

Source record

Selection of Hyperspectral Narrowbands (HNBs) and Composition of Hyperspectral Twoband Vegetation Indices (HVIs) for Biophysical Characterization and Discrimination of Crop Types Using Field Reflectance and Hyperion-EO-1 Data

The overarching goal of this study was to establish optimal hyperspectral vegetation indices (HVIs) and hyperspectral narrowbands (HNBs) that best characterize, classify, model, and map the world's main agricultural crops. The primary objectives were: (1) crop biophysical modeling through HNBs and HVIs, (2) accuracy assessment of crop type discrimination using Wilks' Lambda through a discriminant model, and (3) meta-analysis to select optimal HNBs and HVIs for applications related to agriculture. The study was conducted using two Earth Observing One (EO-1) Hyperion scenes and other surface hyperspectral data for the eight leading worldwide crops (wheat, corn, rice, barley, soybeans, pulses, cotton, and alfalfa) that occupy approx. 70% of all cropland areas globally. This study integrated data collected from multiple study areas in various agroecosystems of Africa, the Middle East, Central Asia, and India. Data were collected for the eight crop types in six distinct growth stages. These included (a) field spectroradiometer measurements (350-2500 nm) sampled at 1-nm discrete bandwidths, and (b) field biophysical variables (e.g., biomass, leaf area index) acquired to correspond with spectroradiometer measurements. The eight crops were described and classified using approx. 20 HNBs. The accuracy of classifying these 8 crops using HNBs was around 95%, which was approx. 25% better than the multi-spectral results possible from Landsat-7's Enhanced Thematic Mapper+ or EO-1's Advanced Land Imager. Further, based on this research and meta-analysis involving over 100 papers, the study established 33 optimal HNBs and an equal number of specific two-band normalized difference HVIs to best model and study specific biophysical and biochemical quantities of major agricultural crops of the world. Redundant bands identified in this study will help overcome the Hughes Phenomenon (or "the curse of high dimensionality") in hyperspectral data for a particular application (e.g., biophysical characterization of crops). The findings of this study will make a significant contribution to future hyperspectral missions such as NASA's HyspIRI. Index Terms-Hyperion, field reflectance, imaging spectroscopy, HyspIRI, biophysical parameters, hyperspectral vegetation indices, hyperspectral narrowbands, broadbands.

Vegetation

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network

Two-dimensional wind-tunnel tests of a NASA supercritical airfoil with various high-lift systems. Volume 1: Data analysis

High-lift systems for a NASA, 9.3%, method for calculating the viscous flow about two-dimensional multicomponent airfoils was evaluated by comparing its predictions with test data. High-lift systems derived from supercritical airfoils were compared in terms of performance to high-lift systems derived from conventional airfoils. The high-lift systems for the supercritical airfoil were designed to achieve maximum lift and consisted of: a single-slotted flap; a double-slotted flap and a leading-edge slat; and a triple-slotted flap and a leading-edge slat. Agreement between theoretical predictions and experimental results are also discussed.

Omar, E.

Toward particle accelerator machine state embeddings as a modality for large language models

Understanding and diagnosing the state of a particle accelerator requires navigating high-dimensional control system data, often involving hundreds of interdependent parameters. We propose a novel multimodal embedding framework that jointly learns representations of machine states from both numerical control system readouts and natural language descriptions. This enables the translation of complex machine conditions into human-readable summaries while maintaining fidelity to the underlying physical system. The obtained embeddings are subsequently adapted to an open-weights large language model via cross-attention conditioning. We demonstrate a first implementation trained on European XFEL machine state data. This work covers the embedding model architecture, training methodology, and presents initial examples demonstrating the model's capabilities in action. Due to the general concept of machine state, the model can be easily adapted to other facilities and control system environments.

Accelerator Physics

Machine Learning Application to Atmospheric Chemistry Modeling

Atmospheric chemistry is a high-dimensionality, large-data problem and thus may be suited to machine-learning algorithms. We show here the potential of a random forest regression algorithm to replace the gas-phase chemistry solver in the GEOS-Chem chemistry model. In this proof-of-concept study, we used one month of model output to train random forest regression models to predict the concentrations of each long-lived chemical species after integration based upon the physical and chemical conditions before the chemical integration. The choice of prediction type has a strong impact on the skill of the regression model. We find best results from predicting the change in concentration for very long-lived species and the absolute concentration for shorter lived species. The skill of the machine learning algorithm is further improved by using a family approach for NO and NO2 rather than treating them independently.By replacing the numerical integrator with the random forest algorithm and running this model for one month, we find that the model is able to reproduce many of the features of the reference chemistry simulation. Replacing the integration methodology with a machine learning algorithm has the potential to be substantially faster. There are a wide range of applications for such an approach, e.g. to generate boundary conditions, for use in air quality forecasts or chemical data assimilation systems, etc.

Keller, Christoph A.

Deep Learning Emulation of Atmospheric Correction for Geostationary Sensors

New generation geostationary satellites make reflectance observations available at a continental scale with unprecedented spatiotemporal resolution and spectral range. Generating Earth monitoring products from these observations requires retrieval of the basic parameter, surface reflectance (SR), by atmospheric correction (AC). Algorithms for atmospheric correction, including Multi-Angle Implementation of Atmospheric Correction (MAIAC), are adapted for each sensor and are too computationally complex to be run in real time, relying instead on look-up tables with precomputed values. Machine learning methods, including convolutional neural networks, have demonstrated performance in learning complex, nonlinear mappings and extracting insight from high-dimensional remote sensing data. In this work, we present a deep learning emulator of MAIAC to retrieve both SR and cloud products. Using this adaptation of deep learning-based emulation to remote sensing, we demonstrate stable SR retrieval over a variety of land covers and viewing conditions and accurate cloud detection. Further, a comparison of computation time suggests emulation as a compelling alternative for expensive physical simulation, especially for applications benefited by near-real time data, such as agricultural management and disaster response.

Duffy, Kate

Coevolution of Machine Learning and Process-Based Modelling to Revolutionize Earth and Environmental Sciences: A Perspective

Machine learning (ML) applications in Earth and environmental sciences (EES) have gained incredible momentum in recent years. However, these ML applications have largely evolved in ‘isolation’ from the mechanistic, process-based modelling (PBM) paradigms, which have historically been the cornerstone of scientific discovery and policy support. In this perspective, we assert that the cultural barriers between the ML and PBM communities limit the potential of ML, and even its ‘hybridization’ with PBM, for EES applications. Fundamental, but often ignored, differences between ML and PBM are discussed as well as their strengths and weaknesses in light of three overarching modelling objectives in EES, (1) nowcasting and prediction, (2) scenario analysis, and (3) diagnostic learning. The paper ponders over a ‘coevolutionary’ approach to model building, shifting away from a borrowing to a co-creation culture, to develop a generation of models that leverage the unique strengths of ML such as scalability to big data and high-dimensional mapping, while remaining faithful to process-based knowledge base and principles of model explainability and interpretability, and therefore, falsifiability.

Saman Razavi

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Visions of visualization aids - Design philosophy and observations

Aids for the visualization of high-dimensional scientific or other data must be designed. Simply casting multidimensional data into a two-dimensional or three-dimensional spatial metaphor does not guarantee that the presentation will provide insight or a parsimonious description of phenomena implicit in the data. Useful visualization, in contrast to glitzy, high-tech, computer-graphics imagery, is generally based on preexisting theoretical beliefs concerning the underlying phenomena. These beliefs guide selection and formatting of the plotted variables. Visualization tools are useful for understanding naturally three-dimensional data bases such as those used by pilots or astronauts. Two examples of such aids for spatial maneuvering illustrate that informative geometric distortion may be introduced to assist visualization and that visualization of complex dynamics alone may not be adequate to provide the necessary insight into the underlying processes.

Ellis, Stephen R.

Intrinsic Dimensionality as a Metric for the Impact of Mission Design Parameters

High-resolution space-based spectral imaging of the Earth's surface delivers critical information for monitoring changes in the Earth system as well as resource management and utilization. Orbiting spectrometers are built according to multiple design parameters, including ground sampling distance (GSD), spectral resolution, temporal resolution, and signal-to-noise ratio. Different applications drive divergent instrument designs, so optimization for wide-reaching missions is complex. The Surface Biology and Geology component of NASA's Earth System Observatory addresses science questions and meets applications needs across diverse fields, including terrestrial and aquatic ecosystems, natural disasters, and the cryosphere. The algorithms required to generate the geophysical variables from the observed spectral imagery each have their own inherent dependencies and sensitivities, and weighting these objectively is challenging. Here, we introduce intrinsic dimensionality (ID), a measure of information content, as an applications-agnostic, data-driven metric to quantify performance sensitivity to various design parameters. ID is computed through the analysis of the eigenvalues of the image covariance matrix, and can be thought of as the number of significant principal components. This metric is extremely powerful for quantifying the information content in high-dimensional data, such as spectrally resolved radiances and their changes over space and time. We find that the ID decreases for coarser GSD, decreased spectral resolution and range, less frequent acquisitions, and lower signal-to-noise levels. This decrease in information content has implications for all derived products. ID is simple to compute, providing a single quantitative standard to evaluate combinations of design parameters, irrespective of higher-level algorithms, products, applications, or disciplines.

Intrinsic dimensionality

SAR processing based on the exact two-dimensional transfer function

The two-dimensional transfer functions of several synthetic aperture radar (SAR) focusing algorithms are derived considering the spaceborne SAR environments. The formulation includes the factors of the earth rotation and the antenna squint angles. The resultant transfer functions are explicitly expressed in terms of Doppler centroid frequency and Doppler frequency rate, which can be accurately estimated from the SAR data. Point target simulation results show that the algorithm based on the two-dimensional Fourier transformation outperforms the one-dimensional one for processing data acquired from high squint angles. The two-dimensional Fourier transformation approach appears to be a viable and simple solution for the processor design of future spaceborne SAR systems.

Chang, C. Y.

Completed Tabulation in the United States of Tests of 24 Airfoils at High Mach Numbers (Derived from Interrupted Work at Guidonia, Italy in the 1.31- by 1.74-Foot High-Speed Tunnel)

Two-dimensional data were obtained in Mach range of from 0.40 to 0.94 and Reynolds Number range of (3.4 - 4.2) X 10 Degrees. Results indicate that thickness ratio is dominating shape parameter at high Mach numbers and that aerodynamic advantages are attainable by using thinnest possible sections. Effects of jet boundaries, Reynolds Number, and Data presented are free from jet-boundary and humidity effects.

AIR FLOW VELOCITY, HIGH - SUBSONIC