Search NASA⌕ Search

SEARCH · Search NASA

Results for “High dimensional data,”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Effective Uncertainty Quantification for Multi-Angle Polarimetric Aerosol Remote Sensing Over Ocean

Multi-angle polarimetric (MAP) measurements can enable detailed characterization of aerosol microphysical and optical properties and improve atmospheric correction in ocean color remote sensing. Advanced retrieval algorithms have been developed to obtain multiple geophysical parameters in the atmosphere–ocean system. Theoretical pixel-wise retrieval uncertainties based on error propagation have been used to quantify retrieval performance and determine the quality of data products. However, standard error propagation techniques in high-dimensional retrievals may not always represent true retrieval errors well due to issues such as local minima and the nonlinear dependence of the forward model on the retrieved parameters near the solution. In this work, we analyze these theoretical uncertainty estimates and validate them using a flexible Monte Carlo approach. The Fast Multi-Angular Polarimetric Ocean coLor (FastMAPOL) retrieval algorithm, based on efficient neural network forward models, is used to conduct the retrievals and uncertainty quantification on both synthetic HARP2 (Hyper-Angular Rainbow Polarimeter 2) and AirHARP (airborne version of HARP2) datasets. In addition, for practical application of the uncertainty evaluation technique in operational data processing, we use the automatic differentiation method to calculate derivatives analytically based on the neural network models. Both the speed and accuracy associated with uncertainty quantification for MAP retrievals are addressed in this study. Pixel-wise retrieval uncertainties are further evaluated for the real AirHARP field campaign data. The uncertainty quantification methods and results can be used to evaluate the quality of data products, as well as guide MAP algorithm development for current and future satellite systems such as NASA’s Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission.

PACE↗

Two-dimensional convolute integers for optical image data processing and surface fitting

An approach toward low-pass, high-pass and band-pass filtering is presented. Convolution coefficients possessing the filtering speed associated with a moving smoothing average without suffering a loss of resolution are discussed. Resolution was retained because the coefficients represented the equivalance of applying high order two-dimensional regression calculations to an image without considering the time-consuming summations associated with the usual normal equations. The smoothing (low-pass) and roughing (high-pass) aspects of the filters are a result of being derived from regression theory. The coefficients are universal integer valves completely described by filter size and surface order, and possess a number of symmetry properties. Double convolution lead to a single set of coefficients with an expanded mask which can yield band-pass filtering and the surface normal. For low order surfaces (0,1), the two-dimensional convolute integers were equivalent to a moving smoothing average.

Edwards, T. R.↗

Climatespark: an In-Memory Distributed Computing Framework for Big Climate Data Analytics

The unprecedented growth of climate data creates new opportunities for climate studies, and yet big climate data pose a grand challenge to climatologists to efficiently manage and analyze big data. The complexity of climate data content and analytical algorithms increases the difficulty of implementing algorithms on high performance computing systems. This paper proposes an in-memory, distributed computing framework, ClimateSpark, to facilitate complex big data analytics and time-consuming computational tasks. Chunking data structure improves parallel I/O efficiency, while a spatiotemporal index is built for the chunks to avoid unnecessary data reading and preprocessing. An integrated, multi-dimensional, array-based data model (ClimateRDD) and ETL operations are developed to address big climate data variety by integrating the processing components of the climate data lifecycle. ClimateSpark utilizes Spark SQL and Apache Zeppelin to develop a web portal to facilitate the interaction among climatologists, climate data, analytic operations and computing resources (e.g., using SQL query and Scala/Python notebook). Experimental results show that ClimateSpark conducts different spatiotemporal data queries/analytics with high efficiency and data locality. ClimateSpark is easily adaptable to other big multiple- dimensional, array-based datasets in various geoscience domains.

Hu, Fei↗

A score-based diffusion model approach for adaptive learning of stochastic partial differential equation solutions

In this paper, we propose a novel framework for adaptively learning the time-evolving solutions of stochastic partial differential equations (SPDEs) using score-based diffusion models within a recursive Bayesian inference setting. SPDEs play a central role in modeling complex physical systems under uncertainty, but their numerical solutions often suffer from model errors and reduced accuracy due to incomplete physical knowledge and environmental variability. To address these challenges, we encode the governing physics into the score function of a diffusion model using simulation data and incorporate observational information via a likelihood-based correction in a reverse-time stochastic differential equation. This enables adaptive learning through iterative refinement of the solution as new data becomes available. To improve computational efficiency in high-dimensional settings, we introduce the ensemble score filter, a training-free approximation of the score function designed for real-time inference. Numerical experiments on benchmark SPDEs demonstrate the accuracy and robustness of the proposed method under sparse and noisy observations.

97 MATHEMATICS AND COMPUTING↗

Two-Dimensional High-Lift Aerodynamic Optimization Using Neural Networks

The high-lift performance of a multi-element airfoil was optimized by using neural-net predictions that were trained using a computational data set. The numerical data was generated using a two-dimensional, incompressible, Navier-Stokes algorithm with the Spalart-Allmaras turbulence model. Because it is difficult to predict maximum lift for high-lift systems, an empirically-based maximum lift criteria was used in this study to determine both the maximum lift and the angle at which it occurs. The 'pressure difference rule,' which states that the maximum lift condition corresponds to a certain pressure difference between the peak suction pressure and the pressure at the trailing edge of the element, was applied and verified with experimental observations for this configuration. Multiple input, single output networks were trained using the NASA Ames variation of the Levenberg-Marquardt algorithm for each of the aerodynamic coefficients (lift, drag and moment). The artificial neural networks were integrated with a gradient-based optimizer. Using independent numerical simulations and experimental data for this high-lift configuration, it was shown that this design process successfully optimized flap deflection, gap, overlap, and angle of attack to maximize lift. Once the neural nets were trained and integrated with the optimizer, minimal additional computer resources were required to perform optimization runs with different initial conditions and parameters. Applying the neural networks within the high-lift rigging optimization process reduced the amount of computational time and resources by 44% compared with traditional gradient-based optimization procedures for multiple optimization runs.

Greenman, Roxana M.↗

MULTI-LEADER: MULTI-source LEarning-Accelerated Design of high-Efficiency multi-stage compRessor (Final Technical Report)

The objective of MULTI-LEADER is to cut design costs by 80% while generating more energy-efficient designs of multi-stage compressors by developing and implementing novel machine learning (ML) techniques, which enable faster and fewer design iterations, improved solver performance, and concurrent multi-disciplinary design. Current industrial practices for the design of multi-stage compressors involve simulation-based design optimization with successive levels of model fidelity, iteratively evaluated between distinct disciplines, one stage at a time to tackle the high dimensional design variations. This project addresses these key design challenges: (1) concurrent optimization of multiple stages under many non-linear constraints; (2) multitude of evaluation of high-fidelity and expensive solvers and their gradients during optimization convergence in high-dimensional design; (3) multi-disciplinary design to maximize aerodynamic performance while guaranteeing structural integrity and additive manufacturability; (4) utilization of multiple fidelity of solvers with disparate parameterization and modeling assumptions. MULTI-LEADER achieved more than 5x speed up in detailed design of more energy-efficient compressors via these machine learning (ML) innovations: (i) rapid design surrogates by multi-source learning from diverse fidelities across multiple disciplines, (ii) physics-constrained data-augmented modeling for improved empiricism, (iii) generative manifold embedding for high dimensional concurrent design without gradient information; (iv) budget-constrained fidelity-adaptive sampling towards fewer design iterations.

33 ADVANCED PROPULSION SYSTEMS↗

Tropospheric chemical models

The differences in atmospheric composition over the globe and the short- and long-term variations in this composition are the net effect of several atmospheric and biospheric processes: biospheric emissions, atmospheric circulation, atmospheric chemical transformations and finally deposition back to the surface. Accurate and realistic atmospheric chemistry and circulation models are essential to interpret the observed global distributions and trends of atmospheric species in terms of these underlying processes. Comparisons between model predictions and observations test current understanding of these processes and models used in conjunction with inverse methods allow deductions of the rates of these processes from the observations. With the planned inclusion of at least CO and CH4 observations on the Earth Observing System (EOS) satellites, together with the large global data set expected from in situ observations under the International Global Atmospheric Chemistry (IGAC) Project, the further development of global three-dimensional high-resolution atmospheric chemistry and circulation models in order to interpret this new data is a high-priority endeavor.

Prinn, R. G.↗

Two-dimensional VLA maps of solar bursts at 15 and 23 GHz with arcsec resolution

Three small, impulsive solar-microwave bursts (peak fluxes 4.9, 3.0, and 0.4 sfu) were observed during 1979 September 7-9, using the Very Large Array (VLA) at 15.05 or 22.5 GHz. The data from 10 antennas distributed on three arms of the array provide the first two-dimensional burst images with spatial resolution as high as 1.0 x 0.75 arcsec. Comparison with optical data showed in the impulsive phase of all three flares, the microwave emission was dominated by a compact source located between the H-alpha kernels. In the post-impulsive phase, the microwave source was larger and elongated in a direction consistent with the orientation of the magnetic field lines joining the H-alpha kernels. These results are interpreted to imply that the initial energy release occurs near the top of the magnetic arch joining the H-alpha kernels.

Marsh, K. A.↗

Personalized and uncertainty-aware coronary hemodynamics simulations: From Bayesian estimation to improved multi-fidelity uncertainty quantification

Non-invasive simulations of coronary hemodynamics have improved clinical risk stratification and treatment outcomes for coronary artery disease, compared to relying on anatomical imaging alone. However, simulations typically use empirical approaches to distribute total coronary flow amongst the arteries in the coronary tree, which ignores patient variability, the presence of disease, and other clinical factors. Further, uncertainty in the clinical data often remains unaccounted for in the modeling pipeline. We present an end-to-end uncertainty-aware pipeline to (1) personalize coronary flow simulations by incorporating vessel-specific coronary flows as well as cardiac function; and (2) predict clinical and biomechanical quantities of interest with improved precision, while accounting for uncertainty in the clinical data. We assimilate patient-specific measurements of myocardial blood flow from clinical CT myocardial perfusion imaging to estimate branch-specific coronary artery flows. Simulated noise in the clinical data is used to estimate the joint posterior distributions of the model parameters using adaptive Markov Chain Monte Carlo sampling. Additionally, the posterior predictive distribution for the relevant quantities of interest is determined using a new approach combining multi-fidelity Monte Carlo estimation with non-linear, data-driven dimensionality reduction. This leads to improved correlations between high- and low-fidelity model outputs. Our framework accurately recapitulates clinically measured cardiac function as well as branch-specific coronary flows under measurement noise uncertainty. We observe substantial reductions in confidence intervals for estimated quantities of interest compared to single-fidelity Monte Carlo estimation and state-of-the-art multi-fidelity Monte Carlo methods. This holds especially true for quantities of interest that showed limited correlation between the low- and high-fidelity model predictions. In addition, the proposed multi-fidelity Monte Carlo estimators are significantly cheaper to compute than traditional estimators, under a specified confidence level or variance. The proposed pipeline for personalized and uncertainty-aware predictions of coronary hemodynamics is based on routine clinical measurements and recently developed techniques for CT myocardial perfusion imaging. The proposed pipeline offers significant improvements in precision and reduction in computational cost.

Bayesian parameter estimation↗

Two-dimensional hydrodynamic viscous electron flow in annular Corbino rings

The concept of fluidic viscosity is ubiquitous in condensed-matter systems hosting a continuum where macroscopic properties can emerge. While an important property of liquids and some solids, only recently was the viscosity of an electron shown to play a role in electronic transport experiments. In this Letter, we present nonlocal electronic transport measurements in concentric annular rings formed in high-mobility two-dimensional electron gases, and the resulting data show that viscous hydrodynamic flow can occur far away from the source-drain current region. Our conclusion of viscous electronic transport is further corroborated by simulations of the Navier-Stokes equations that are found to be in agreement with our measurements below T = 1 K . Finally, this work emphasizes the key role played by viscosity via electron-electron ( e − e ) interaction even when the electronic transport is restricted radially, and for which it should have played no major role. Published by the American Physical Society 2025

Vijayakrishnan, Sujatha (ORCID:0009000080933182)↗

Data Analytics for Catalysis Predictions: Are We Ready Yet?

Catalysis informatics has received tremendous attention in recent years as a tool to design catalysts and discover unique descriptors that capture the relationships between chemical properties and catalytic performance. One of the stop-gaps in understanding catalytic effects, which is often ignored and limits the deployment of data science tools, relates to the lack of uniform data. The catalytic cleavage of C–X (X= H, C, N, and O) bonds is relevant to many fundamental catalytic processes. In this Perspective, we performed data analytics on four groups of C–X cleavage reactions that are common in production, upcycling, or reactive separation: the C–C cleavage in cyclopropyl alcohol, the C–H cleavage in hydroacylation reactions, the C–O cleavage in β-O-4 linkages, and the C–N cleavage in amides, using experimental data collected from the literature to understand their underlying correlations. Experimental variables of high impact are identified for each reaction by dimensionality reduction methods. We highlight the urgent need for experimental data sets that include full details on the reaction conditions, such as reagent concentration, reaction temperature, or time in machine-readable forms. We discuss the potential improvement of the data of these reactions and promising approaches such as autonomous experiments to fill the gaps in unbiased experimental data. Finally, we also address the early stage consideration of separation aspects in the experimental design of efficient catalytic systems for these fundamental examples of chemical reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ML-based Dimension Reduction Strategies

Deep learning (DL)--based surrogate models have achieved success in various applications in carbon capture and storage (CCS). However, the model training on high-dimensional spaces is computationally expensive and impractical for large-scale and complex geological models, because the models usually contain hundreds of thousands to millions of grid cells, each with a set of parameters. Furthermore, the high cost of generating training data with sufficient variation is another limitation of model training on high-dimensional spaces, which may result in overfitting and reduce the model efficiency and prediction performance. We proposed the workflow incorporating dimension reduction methods and deep learning models, which aim to extract the latent variables of input parameters and output state variables, and then build the mapping function at the latent spaces. The proposed workflow can significantly reduce the computational complexity in solving both forward and inverse problems compared to models trained on high-dimensional spaces. Dimensionality reduction models showed great potential in workflows for fast reservoir simulation, history matching, prior model generation, visualization, and more, ultimately enhancing DL model performance in related SMART Work Packages.

Hosseini, Seyyed↗

Simplified Micromechanics of Plain Weave Composites

A micromechanics based methodology to simulate the complete hygro-thermomechanical behavior of plain weave composites is developed. This methodology is based on micromechanics and the classical laminate theory. The methodology predicts a complete set of thermal, hygral and mechanical properties of plain woven composites, generates necessary data for use in a finite element structural analysis, and predicts stresses all the way from the laminate to the constituent level. This methodology is used in conjunction with a composite mechanics code to analyze and predict the properties/response of a generic graphite/epoxy woven textile composite and a plain weave ceramic composite. The fiber architecture, including the fiber waviness and fiber end distributions through the thickness, is properly accounted for. Predicted results compare reasonably well with those from detailed three-dimensional finite element analyses as well as available experimental data. However, the main advantage of the proposed methodology is its high computational efficiency as compared with three-dimensional finite element analyses.

Mital, Subodh K.↗

Endwall Heat Transfer Measurements in a Transonic Turbine Cascade

Turbine blade endwall heat transfer measurements are given for a range of Reynolds and Mach numbers. Data were obtained for Reynolds numbers based on inlet conditions of 0.5 and 1.0 x 106, for isentropic exit Mach numbers of 1.0 and 1.3, and for freestream turbulence intensities of 0.25% and 7.0%. Tests were conducted in a linear cascade at the NASA Lewis Transonic Turbine Blade Cascade Facility. The test article was a turbine rotor with 136' of turning and an axial chord of 12.7 cm. The large scale allowed for very detailed measurements of both flow field and surface phenomena. The intent of the work is to provide benchmark quality data for computational fluid dynamics (CFD) code and model verification. The flow field in the cascade is highly three-dimensional as a result of thick boundary layers at the test section inlet. Endwall heat transfer data were obtained using a steady-state liquid crystal technique.

Giel, P. W.↗

Chaos: Understanding and Controlling Laser Instability

In order to characterize the behavior of tunable diode lasers (TDL), the first step in the project involved the redesign of the TDL system here at the University of Tennessee Molecular Systems Laboratory (UTMSL). Having made these changes it was next necessary to optimize the new optical system. This involved the fine adjustments to the optical components, particularly in the monochromator, to minimize the aberrations of coma and astigmatism and to assure that the energy from the beam is focused properly on the detector element. The next step involved the taking of preliminary data. We were then ready for the analysis of the preliminary data. This required the development of computer programs that use mathematical techniques to look for signatures of chaos. Commercial programs were also employed. We discovered some indication of high dimensional chaos, but were hampered by the low sample rate of 200 KSPS (kilosamples/sec) and even more by our sample size of 1024 (1K) data points. These limitations were expected and we added a high speed data acquisition board. We incorporated into the system a computer with a 40 MSPS (million samples/sec) data acquisition board. This board can also capture 64K of data points so that were then able to perform the more accurate tests for chaos. The results were dramatic and compelling, we had demonstrated that the lead salt diode laser had a chaotic frequency output. Having identified the chaotic character in our TDL data, we proceeded to stage two as outlined in our original proposal. This required the use of an Occasional Proportional Feedback (OPF) controller to facilitate the control and stabilization of the TDL system output. The controller was designed and fabricated at GSFC and debugged in our laboratories. After some trial and error efforts, we achieved chaos control of the frequency emissions of the laser. The two publications appended to this introduction detail the entire project and its results.

Blass, William E.↗

Accelerating particle-in-cell kinetic plasma simulations via reduced-order modeling of space-charge dynamics using dynamic mode decomposition

We present a data-driven reduced-order modeling of the space-charge dynamics for electromagnetic particle-in-cell (EMPIC) plasma simulations based on dynamic mode decomposition (DMD). The dynamics of the charged particles in kinetic plasma simulations such as EMPIC is manifested through the plasma current density defined along the edges of the spatial mesh. We showcase the efficacy of DMD in modeling the time evolution of current density through a low-dimensional feature space. Not only do such DMD based predictive reduced-order models help accelerate EMPIC simulations, they also have the potential to facilitate investigative analysis and control applications. Here, we demonstrate the proposed DMD-EMPIC scheme for reduced-order modeling of current density and speedup in EMPIC simulations involving electron beam under the influence of magnetic field, virtual cathode oscillations, and backward wave oscillator.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics↗

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗