Search NASA⌕ Search

SEARCH · Search NASA

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A summary of research on the NASA-Global Atmospheric Sampling Program performed by the Atmospheric Sciences Research Center

The annual variations of ozone near the tropopause are derived from aircraft exhibit year-to-year differences which are not explicitly accounted for by the simple, classical ozone transport theory. Phenomena such as tropopause lifting, interannual variations in the rates of stratospheric-tropospheric exchange and meridional mixing, contribute differently to the distribution of ozone in this altitude region. Ozone encounter climatologies have been represented by global maps which show the probabilities of exceeding ambient ozone levels of 200, 300, and 400 ppbV along flight routes during the year. Continuous ozone records obtained from the GASP system revealed the presence of gravity waves whose wavelength is of the order 20 km. The GASP data cannot, however, be utilized for the evaluation of horizontal fluxes of such quantities as ozone, sensible heat, and zonal momentum; the data are too sparsely and irregularly distributed for the computation of stable correlations. Multiple species data from the unique circumglobal flight of a Pan American airliner on 28-30 October 1977 are discussed with particular regard to the apparent interhemispheric differences in tropospheric species concentrations, variation between the Arctic and Antarctic stratospheres, to possible covariations between species, and to potential source regions for various constituents.

Falconer, P. D.↗

Prediction of Aerodynamic Coefficients for Wind Tunnel Data using a Genetic Algorithm Optimized Neural Network

A fast, reliable way of predicting aerodynamic coefficients is produced using a neural network optimized by a genetic algorithm. Basic aerodynamic coefficients (e.g. lift, drag, pitching moment) are modelled as functions of angle of attack and Mach number. The neural network is first trained on a relatively rich set of data from wind tunnel tests of numerical simulations to learn an overall model. Most of the aerodynamic parameters can be well-fitted using polynomial functions. A new set of data, which can be relatively sparse, is then supplied to the network to produce a new model consistent with the previous model and the new data. Because the new model interpolates realistically between the sparse test data points, it is suitable for use in piloted simulations. The genetic algorithm is used to choose a neural network architecture to give best results, avoiding over-and under-fitting of the test data.

Rajkumar, T.↗

Evaporation over global oceans derived from satellite data and AGCM

Evaporation over ocean at 100 km has been estimated using operational satellite data. Recent satellite results are compared to data obtained through atmospheric general circulation models (AGCMs). A large systematic error is found in AGCM data which is considered to be due to overestimation of atmospheric moisture. It is suggested that the discrepancy may be caused by errors in operational infrared sounder data assimilated into the model. The error may also be caused by interpolation of very sparse in situ data in areas of sharp gradients.

Liu, W. T.↗

Incorporating TRMM and Other High-Quality Estimates into the One-Degree Daily (1DD) Global Precipitation Product

The One-Degree Daily (1DD) precipitation dataset was recently developed for the Global Precipitation Climatology Project (GPCP). The IDD provides a globally-complete, observation-only estimate of precipitation on a daily 1 deg x 1 deg grid for the period 1997 through late 1999 (by the time of the conference). In the latitude band 40 N - 40 S the IDD uses the Threshold-Matched Precipitation Index (TMPI), a GPI-like IR product with the T(sub b) threshold and (single) conditional rain rate determined locally for each month by the frequency of precipitation in the GPROF SSNU product and by the precipitation amount in the GPCP satellite-gauge (SG) combination. Outside 40 N - 40 S the 1DD uses a scaled TOVS precipitation estimate that has adjustments based on the TMPI and the SG. This first-generation 1DD has been in beta test preparatory to release as an official GPCP product. In this paper we discuss further development of the 1DD framework to allow the direct incorporation of TRMM and other high-quality precipitation estimates. First, these data are generally sparse (typically from low-orbit satellites), so a fair amount of work was devoted to data boundaries. Second, these data are not the same as the original 1DD estimates, so we had to give careful consideration to the best scheme for forcing the 1DD to sum to the SG for the month. Finally, the non-sun-synchronous, low-inclination orbit occupied by TRMM creates interesting variations against the sun-synchronous, high-inclination orbits occupied by the Defense Meteorological Satellite Program satellites that carry the SSM/I. Examples will be given of each of the development issues, then comparisons will be made to daily raingauge analyses.

Huffman, George J.↗

Development of Aerodynamic Loads Databases for the Space Launch System Booster Separation Event

Booster separation is a mission-critical event within the orbital ascent of the Space Launch System (SLS). The complexity of the engine plume-affected, multibody, supersonic aerodynamics is compounded by the large span of the likely trajectory space. Characterization of the multiple input, multiple output system requires a combination of wind tunnel testing and computational simulation, but additional data processing is also required before the sparse, high-fidelity data can be fused into a continuous database with acceptable uncertainty quantification. This paper outlines the state of this approach as it has been applied to the most recent SLS booster separation aerodynamic loads database: that of the Artemis II launch vehicle.

Michael W Lee↗

Development of Aerodynamic Loads Databases for the Space Launch System Booster Separation Event

Booster separation is a mission-critical event within the orbital ascent of the Space Launch System (SLS). The complexity of the engine plume-affected, multibody, supersonic aerodynamics is compounded by the large span of the likely trajectory space. Characterization of the multiple input, multiple output system requires a combination of wind tunnel testing and computational simulation, but additional data processing is also required before the sparse, high-fidelity data can be fused into a continuous database with acceptable uncertainty quantification. This paper outlines the state of this approach as it has been applied to the most recent SLS booster separation aerodynamic loads database: that of the Artemis II launch vehicle.

Michael Lee↗

Computational Fluid Dynamics Simulations of the Transonic Dynamics Tunnel Airstream Oscillator System

This paper presents a Computational Fluid Dynamics (CFD) model of the flow in the NASA Langley Research Center Transonic Dynamics Tunnel (TDT) with the Airstream Oscillator System (AOS) in operation. The TDT is a continuous-flow, closed circuit wind tunnel with a 16- by 16-foot slotted test section with cropped corners. The tunnel was originally built as the 19-ft Pressure Tunnel in 1938, but it was converted to the current transonic tunnel in the 1950s, with capabilities to use either air or heavy gas (R-134a) as the test medium. The TDT was also fitted with an AOS which can be used to create a gust field in the test section. To date, no computational analyses of the TDT involving the AOS have been performed. In this study, experimental data acquired of an oscillating airstream in the tunnel will be used to calibrate the computational analyses. A significant motivation of this work is to attempt to compare computational gust velocities to those recorded in the TDT literature. In addition to a validation of this model with experimental data, it may also be possible to supplement the experimental data with computational data. The experimental data is rather sparse and at only a few Mach numbers. Computational results may be able to expand the AOS data set. Another motivation for this work is that the Integrated Adaptive Wing Technology Maturation (IAWTM) semi-span model of the high-aspect-ratio CRM with 10 trailing edge control surfaces is currently being fabricated and will be delivered to the NASA Langley Transonic Dynamics Tunnel (TDT) for testing in late 2020. Among tests to be conducted, gust load alleviation (GLA) will be demonstrated at transonic conditions using the AOS.

Computational Fluid Dynamics↗

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana↗

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks↗

Prong Segmentation using Point Set Transformers in Multiple View Neutrino Detectors

NOvA is a long-baseline neutrino experiment studying neutrino oscillations by detecting neutrinos from the NuMI beam at Fermilab. Its physics analysis relies on accurate prong segmentation, which involves matching each hit to its source particle and identifying the particle type. This task has commonly been addressed using a combination of traditional clustering algorithms and convolutional neural networks (CNNs). However, NOvA’s detector design presents data as two sparse and decoupled 2D images (XZ and YZ views) rather than a native 3D representation, posing a significant challenge for traditional CNN-based models. In this talk, we propose a novel neural network based on the Point Set Transformer. By treating detector hits as sparse point clouds and implementing a cross-view attention mechanism, our model enables efficient information mixing between both views. Evaluated on NOvA simulated data, our model achieves superior accuracy while requiring significantly fewer computational resources compared to other models. Furthermore, the model demonstrates great performance when applied to Liquid Argon Time Projection Chamber (LArTPC) data, which shows its potential as a universal prong segmentation algorithm for multiple view neutrino detectors.

Liu, Jiaxi [UC, Irvine]↗

Validation and Development of the GPCP Experimental One-Degree Daily (1DD) Global Precipitation Product

The One-Degree Daily (1DD) precipitation dataset has been developed for the Global Precipitation Climatology Project (GPCP) and is currently in beta test preparatory to release as an official GPCP product. The 1DD provides a globally-complete, observation-only estimate of precipitation on a daily 1 deg. x 1 deg. grid for the period 1997 through early 2000 (by the time of the conference). In the latitude band 40N-40S the 1DD uses the Threshold-Matched Precipitation Index (TMPI), a GPI-like IR product with the pixel-level T(sub b) threshold and (single) conditional rain rate determined locally for each month by the frequency of precipitation in the GPROF SSM/I product and by, the precipitation amount in the GPCP monthly satellite-gauge (SG) combination. Outside 40N-40S the 1DD uses a scaled TOVS precipitation estimate that has month-by-month adjustments based on the TMPI and the SG. Early validation results are encouraging. The 1DD shows relatively large scatter about the daily validation values in individual grid boxes, as expected for a technique that depends on cloud-sensing schemes such as the TMPI and TOVS. On the other hand, the time series of 1DD shows good correlation with validation in individual boxes. For example, the 1997-1998 time series of 1DD and Oklahoma Mesonet values in a grid box in northeastern Oklahoma have the correlation coefficient = 0.73. Looking more carefully at these two time series, the number of raining days for the 1DD is within 7% of the Mesonet value, while the distribution of daily rain values is very similar. Other tests indicate that area- or time-averaging improve the error characteristics, making the data set highly attractive to users interested in stream flow, short-term regional climatology, and model comparisons. The second generation of the 1DD product is currently under development; it is designed to directly incorporate TRMM and other high-quality precipitation estimates. These data are generally sparse because they are observed by low-orbit satellites, so a fair amount of work must be devoted to analyzing the effect of data boundaries. This work is laying, the groundwork for effective use of the NASA Global Precipitation Mission, which will have full Global coverage by low-orbit passive microwave satellites every three hours.

Huffman, George J.↗

Calibration and evaluation of Skylab altimetry for geodetic determination of the geoid

The author has identified the following significant results. The Skylab altimeter experiment has proven the capability of the altimeter for measurement of sea surface topography. The geometric determination of the geoid/mean sea level from satellite altimetry is a new approach having significant applications in many disciplines including geodesy and oceanography. A generalized least squares collocation technique was developed for determination of the geoid from altimetry data. The technique solves for the altimetry geoid and determines one bias term for the combined effect of sea state, orbit, tides, geoid, and instrument error using sparse ground truth data. The influence of errors in orbit and a priori geoid values are discussed. Although the Skylab altimeter instrument accuracy is about plus or minus 1m, significant results were obtained in identification of large geoidal features such as over the Puerto Rico trench. Comparison of the results of several passes shows that good agreement exists between the general slopes of the altimeter geoid and the ground truth, and that the altimeter appears to be capable of providing more details than are now available with best known geoids.

Mourad, A. G.↗

The significance of the Skylab altimeter experiment results and potential applications

The Skylab Altimeter Experiment has proven the capability of the altimeter for measurement of sea surface topography. The geometric determination of the geoid/mean sea level from satellite altimetry is a new approach having significant applications in many disciplines including geodesy and oceanography. A Generalized Least Squares Collocation Technique was developed for determination of the geoid from altimetry data. The technique solves for the altimetry geoid and determines one bias term for the combined effect of sea state, orbit, tides, geoid, and instrument error using sparse ground truth data. The influence of errors in orbit and a priori geoid values are discussed. Although the Skylab altimeter instrument accuracy is about + or - 1 m, significant results were obtained in identification of large geoidal features such as over the Puerto Rico trench. Comparison of the results of several passes shows that good agreement exists between the general slopes of the altimeter geoid and the ground truth, and that the altimeter appears to be capable of providing more details than are now available with best known geoids. The altimetry geoidal profiles show excellent correlations with bathymetry and gravity. Potential applications of altimetry results to geodesy, oceanography, and geophysics are discussed.

Mourad, A. G.↗

Machine Learning Based Metamodel for Faster Life Cycle Assessment of Large Portfolio of Buildings

Managing a large portfolio of buildings involves decisions on reuse, retrofit, renovation, rehabilitation, and new construction, influenced by trade-offs between performance metrics such as cost, time, and operational flexibility over the building's life cycle. Traditional life cycle assessment tools for evaluating these metrics can be labor- and compute-intensive, requiring extensive data and modeling for each building. Metamodels (or surrogate models) using machine learning have been explored as faster alternatives, but training these models has been hindered by the limited availability of comprehensive data on key life cycle metrics. Recent advancements in machine learning, particularly deep learning techniques like zero-shot and few-shot learning, allow models to learn from sparse or limited data. We propose a machine learning-based metamodel that leverages these techniques for rapid estimation of key building life cycle metrics. This presentation will cover the model architecture, data collection, training, and validation processes, along with an ongoing case study applied to a large portfolio of buildings. We will discuss the model's performance in terms of accuracy, compute time, limitations, and its potential for expanding to additional life cycle metrics. This data-driven approach offers a promising direction for the rapid evaluation of large building portfolios.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

Open Science for Plants in Space: Improvements in NASA's Open Science Data Repository

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.

FAIR↗

Applications of Data Assimilation to Analysis of the Ocean on Large Scales

It is commonplace to begin talks on this topic by noting that oceanographic data are too scarce and sparse to provide complete initial and boundary conditions for large-scale ocean models. Even considering the availability of remotely-sensed data such as radar altimetry from the TOPEX and ERS-1 satellites, a glance at a map of available subsurface data should convince most observers that this is still the case. Data are still too sparse for comprehensive treatment of interannual to interdecadal climate change through the use of models, since the new data sets have not been around for very long. In view of the dearth of data, we must note that the overall picture is changing rapidly. Recently, there have been a number of large scale ocean analysis and prediction efforts, some of which now run on an operational or at least quasi-operational basis, most notably the model based analyses of the tropical oceans. These programs are modeled on numerical weather prediction. Aside from the success of the global tide models, assimilation of data in the tropics, in support of prediction and analysis of seasonal to interannual climate change, is probably the area of large scale ocean modeling and data assimilation in which the most progress has been made. Climate change is a problem which is particularly suited to advanced data assimilation methods. Linear models are useful, and the linear theory can be exploited. For the most part, the data are sufficiently sparse that implementation of advanced methods is worthwhile. As an example of a large scale data assimilation experiment with a recent extensive data set, we present results of a tropical ocean experiment in which the Kalman filter was used to assimilate three years of altimetric data from Geosat into a coarsely resolved linearized long wave shallow water model. Since nonlinear processes dominate the local dynamic signal outside the tropics, subsurface dynamical quantities cannot be reliably inferred from surface height anomalies. Because of its potential for large scale synoptic coverage of the deep ocean, acoustic travel time data should be a natural complement to satellite altimetry. Satellite data give us vertical integrals associated with thermodynamic and dynamic processes.

Miller, Robert N.↗

AEPF: Attention-Enabled Point Fusion for 3D Object Detection

Current state-of-the-art (SOTA) LiDAR-only detectors perform well for 3D object detection tasks, but point cloud data are typically sparse and lacks semantic information. Detailed semantic information obtained from camera images can be added with existing LiDAR-based detectors to create a robust 3D detection pipeline. With two different data types, a major challenge in developing multi-modal sensor fusion networks is to achieve effective data fusion while managing computational resources. With separate 2D and 3D feature extraction backbones, feature fusion can become more challenging as these modes generate different gradients, leading to gradient conflicts and suboptimal convergence during network optimization. To this end, we propose a 3D object detection method, Attention-Enabled Point Fusion (AEPF). AEPF uses images and voxelized point cloud data as inputs and estimates the 3D bounding boxes of object locations as outputs. An attention mechanism is introduced to an existing feature fusion strategy to improve 3D detection accuracy and two variants are proposed. These two variants, AEPF-Small and AEPF-Large, address different needs. AEPF-Small, with a lightweight attention module and fewer parameters, offers fast inference. AEPF-Large, with a more complex attention module and increased parameters, provides higher accuracy than baseline models. Experimental results on the KITTI validation set show that AEPF-Small maintains SOTA 3D detection accuracy while inferencing at higher speeds. AEPF-Large achieves mean average precision scores of 91.13, 79.06, and 76.15 for the car class’s easy, medium, and hard targets, respectively, in the KITTI validation set. Results from ablation experiments are also presented to support the choice of model architecture.

Chemistry↗