Search NASASearch

SEARCH · Search NASA

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

NASA's Science Discovery Engine: Enabling Interdisciplinary Open Science

NASA is currently implementing capabilities to enable open science access. The Science Discovery Engine (SDE) is a major endeavor in this effort. The SDE supports discovery and access to complex, heterogeneous science data and information across science topical areas.

Kaylin Mclendon Bugbee

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]

Method for Accessing Distributed Heterogeneous Databases

A scenario of relational, hierarchial, and network data bases is presented and a distributed access view integrated data base system (DAVID) is described for uniformly accessing data bases which are heterogeneous and physically distributed. The DAVID system is based on data base logic so that the relational approach is generalized to the heterogeneous approach. The global data manager is explained as are global data manipulation languages which can operate on all the data bases and can query the data dictionary and the data directory.

Jacobs, B. E.

Stewardship Best Practices for Improved Discovery and Reuse of Heterogeneous and Cross-Disciplinary Earth System Data

Some of the Earth system data products such as those from NASA airborne and field investigations (a.k.a. campaigns), are highly heterogeneous and cross-disciplinary, making the data extremely challenging to manage. For example, airborne and field campaign measurements tend to be sporadic over a period of time, with large gaps. Data products generated are of various processing levels and utilized for a wide range of inter- and cross-disciplinary research and applications. Data and derived products have been historically stored in a variety of domain-specific standard (and some non-standard) formats and in various locations such as NASA Distributed Active Archive Centers (DAACs), NASA airborne science facilities, field archives, or even individual scientists’ computer hard drives. As a result, airborne and field campaign data products have often been managed and represented differently, making it onerous for data users to find, access, and utilize campaign data. Some difficulties in discovering and accessing the campaign data originate from the incomplete data product and contextual metadata that may contain details relevant to the campaign (e.g. campaign acronym and instrument deployment locations), but tend to lack other significant information needed to understand conditions surrounding the data. Such details can be burdensome to locate after the conclusion of a campaign. Utilizing consistent terminology, essential for improved discovery and reuse, is also challenging due to the variety of involved disciplines. To help address the aforementioned challenges faced by many repositories and data managers handling airborne and field data, this presentation will describe stewardship practices developed by the Airborne Data Management Group (ADMG) within the Interagency Implementation and Advanced Concepts Team (IMPACT) under the NASA’s Earth Science Data systems (ESDS) Program.

best practices

Space-Time Data Fusion

Space-time Data Fusion (STDF) is a methodology for combing heterogeneous remote sensing data to optimally estimate the true values of a geophysical field of interest, and obtain uncertainties for those estimates. The input data sets may have different observing characteristics including different footprints, spatial resolutions and fields of view, orbit cycles, biases, and noise characteristics. Despite these differences all observed data can be linked to the underlying field, and therefore the each other, by a statistical model. Differences in footprints and other geometric characteristics are accounted for by parameterizing pixel-level remote sensing observations as spatial integrals of true field values lying within pixel boundaries, plus measurement error. Both spatial and temporal correlations in the true field and in the observations are estimated and incorporated through the use of a space-time random effects (STRE) model. Once the models parameters are estimated, we use it to derive expressions for optimal (minimum mean squared error and unbiased) estimates of the true field at any arbitrary location of interest, computed from the observations. Standard errors of these estimates are also produced, allowing confidence intervals to be constructed. The procedure is carried out on a fine spatial grid to approximate a continuous field. We demonstrate STDF by applying it to the problem of estimating CO2 concentration in the lower-atmosphere using data from the Atmospheric Infrared Sounder (AIRS) and the Japanese Greenhouse Gasses Observing Satellite (GOSAT) over one year for the continental US.

Greenhouse Gases Observing Satellite (GOSAT)

IPC-Fusion (Infrastructure Perception and Control (IPC): Multisensor Data Fusion Software) [SWR-25-153]

As part of the National Laboratory of the Rockies' (NLR’s) Infrastructure Perception and Control Laboratory, the IPC-Fusion toolkit provides a probabilistic, scalable, multi-sensor fusion framework that integrates (late-stage fusion) heterogeneous object detection data from traffic sensors to enable robust, real-time tracking of roadway occupants. The algorithmic design of the toolkit is motivated by the need for creating a digital twin of traffic at the edge in a scalable and affordable manner. The software operates by combining object-level measurements (such as position and velocity) from a suite of sensors (such as radar, lidar, camera) using Kalman filtering and probabilistic data association techniques to overcome individual sensor limitations and achieve superior tracking performance in complex traffic zones. The framework addresses key challenges including heterogeneous measurement uncertainties, asynchronous data streams, varying spatiotemporal data resolutions, robust data association, and adaptive object lifecycle management. Validated on real-world traffic intersection data including vehicles and pedestrians, IPC-Fusion demonstrates enhanced tracking reliability across scenarios involving occlusions, sensor failures, and varying traffic densities, supporting the broader IPC initiative's goal of transforming transportation infrastructure through advanced perception capabilities for intelligent transportation systems, traffic safety applications, and autonomous vehicle support.

Sandhu, Rimple [National Laboratory of the Rockies

A distributed data base management facility for the CAD/CAM environment

Current/PAD research in the area of distributed data base management considers facilities for supporting CAD/CAM data management in a heterogeneous network of computers encompassing multiple data base managers supporting a variety of data models. These facilities include coordinated execution of multiple DBMSs to provide for administration of and access to data distributed across them.

Balza, R. M.

The role of data management in discipline-independent data visualization

The common data format (CDF) is described in terms of its support applications for the database management of visualization systems. The CDF is a self-describing data abstraction technique for the storage and manipulation of multidimensional data that are based on block structures. The discipline-independent approach is designed to manage, manipulate, archive, display, and analyze data, and can be applied to heterogeneous equipment communicating different data structures over networks. An improved CDF version incorporates a hyperplane access allowing random aggregate access to subdimensional blocks within a multidimensional variable. The visualization pipeline is also discussed, which controls the flow of data and permits the visualization of different classes of data representation techniques. The system is found to accommodate a large variety of scientific data structures and large disk-based data sets.

Treinish, Lloyd A.

Data Grid Management Systems

The "Grid" is an emerging infrastructure for coordinating access across autonomous organizations to distributed, heterogeneous computation and data resources. Data grids are being built around the world as the next generation data handling systems for sharing, publishing, and preserving data residing on storage systems located in multiple administrative domains. A data grid provides logical namespaces for users, digital entities and storage resources to create persistent identifiers for controlling access, enabling discovery, and managing wide area latencies. This paper introduces data grids and describes data grid use cases. The relevance of data grids to digital libraries and persistent archives is demonstrated, and research issues in data grids and grid dataflow management systems are discussed.

Moore, Reagan W.

Variability and Extremes of Precipitation in the Global Climate as Determined by the 25-Year GEWEX/GPCP Data Set

The Global Precipitation Climatology Project (GPCP) 25-year precipitation data set is used to evaluate the variability and extremes on global and regional scales. The variability of precipitation year-to-year is evaluated in relation to the overall lack of a significant global trend and to climate events such as ENSO and volcanic eruptions. The validity of conclusions and limitations of the data set are checked by comparison with independent data sets (e.g., TRMM). The GPCP data set necessarily has a heterogeneous time series of input data sources, so part of the assessment described above is to test the initial results for potential influence by major data boundaries in the record. Regional trends, or inter-decadal changes, are also analyzed to determine validity and correlation with other long-term data sets related to the hydrological cycle (e.g., clouds and ocean surface fluxes). Statistics of extremes (both wet and dry) are analyzed at the monthly time scale for the 25 years. A preliminary result of increasing frequency of extreme monthly values will be a focus to determine validity. Daily values for an eight-year are also examined for variation in extremes and compared to the longer monthly-based study.

Adler, R. F.

Means, Variability and Trends of Precipitation in the Global Climate as Determined by the 25-year GEWEWGPCP Data Set

The Global Precipitation Climatology Project (GPCP) 25-year precipitation data set is used as a basis to evaluate the mean state, variability and trends (or inter-decadal changes) of global and regional scales of precipitation. The uncertainties of these characteristics of the data set are evaluated by examination of other, parallel data sets and examination of shorter periods with higher quality data (e.g., TRMM). The global and regional means are assessed for uncertainty by comparing with other satellite and gauge data sets, both globally and regionally. The GPCP global mean of 2.6 mdday is divided into values of ocean and land and major latitude bands (Tropics, mid-latitudes, etc.). Seasonal variations globally and by region are shown and uncertainties estimated. The variability of precipitation year-to-year is shown to be related to ENS0 variations and volcanoes and is evaluated in relation to the overall lack of a significant global trend. The GPCP data set necessarily has a heterogeneous time series of input data sources, so part of the assessment described above is to test the initial results for potential influence by major data boundaries in the record.

Adler, R. F.

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT

An Extensible, Interchangeable and Sharable Database Model for Improving Multidisciplinary Aircraft Design

Crucial to an efficient aircraft simulation-based design is a robust data modeling methodology for both recording the information and providing data transfer readily and reliably. To meet this goal, data modeling issues involved in the aircraft multidisciplinary design are first analyzed in this study. Next, an XML-based. extensible data object model for multidisciplinary aircraft design is constructed and implemented. The implementation of the model through aircraft databinding allows the design applications to access and manipulate any disciplinary data with a lightweight and easy-to-use API. In addition, language independent representation of aircraft disciplinary data in the model fosters interoperability amongst heterogeneous systems thereby facilitating data sharing and exchange between various design tools and systems.

Lin, Risheng

Flammability of Heterogeneously Combusting Metals

Most engineering materials, including some metals, most notably aluminum, burn in homogeneous combustion. 'Homogeneous' refers to both the fuel and the oxidizer being in the same phase, which is usually gaseous. The fuel and oxidizer are well mixed in the combustion reaction zone, and heat is released according to some relation like q(sub c) = delta H(sub c)c[((rho/rho(sub 0))]exp a)(exp -E(sub c)/RT), Eq. (1) where the pressure exponent a is usually close to unity. As long as there is enough heat released, combustion is sustained. It is useful to conceive of a threshold pressure beyond which there is sufficient heat to keep the temperature high enough to sustain combustion, and beneath which the heat is so low that temperature drains away and the combustion is extinguished. Some materials burn in heterogeneous combustion, in which the fuel and oxidizer are in different phases. These include iron and nickel based alloys, which burn in the liquid phase with gaseous oxygen. Heterogeneous combustion takes place on the surface of the material (fuel). Products of combustion may appear as a solid slag (oxide) which progressively covers the fuel. Propagation of the combustion melts and exposes fresh fuel. Heterogeneous combustion heat release also follows the general form of Eq.(1), except that the pressure exponent a tends to be much less than 1. Therefore, the increase in heat release with increasing pressure is not as dramatic as it is in homogeneous combustion. Although the concept of a threshold pressure still holds in heterogeneous combustion, the threshold is more difficult to identify experimentally, and pressure itself becomes less important relative to the heat transfer paths extant in any specific application. However, the constants C, a, and E(sub c) may still be identified by suitable data reduction from heterogeneous combustion experiments, and may be applied in a heat transfer model to judge the flammability of a material in any particular actual-use situation. In order to support the above assertions, two investigations are undertaken: 1) PCT data are examined in detail to discover the pressure dependence of heterogeneous combustion experiment results; and 2) heterogeneous combustion in a PCT situation is described by a heat transfer model, which is solved first in simplified form for a simple actual-use situation, and then extended to apply to PCT data reduction (combustion constant identification).

Jones, Peter D.

AXAF user interfaces for heterogeneous analysis environments

The AXAF Science Center (ASC) will develop software to support all facets of data center activities and user research for the AXAF X-ray Observatory, scheduled for launch in 1999. The goal is to provide astronomers with the ability to utilize heterogeneous data analysis packages, that is, to allow astronomers to pick the best packages for doing their scientific analysis. For example, ASC software will be based on IRAF, but non-IRAF programs will be incorporated into the data system where appropriate. Additionally, it is desired to allow AXAF users to mix ASC software with their own local software. The need to support heterogeneous analysis environments is not special to the AXAF project, and therefore finding mechanisms for coordinating heterogeneous programs is an important problem for astronomical software today. The approach to solving this problem has been to develop two interfaces that allow the scientific user to run heterogeneous programs together. The first is an IRAF-compatible parameter interface that provides non-IRAF programs with IRAF's parameter handling capabilities. Included in the interface is an application programming interface to manipulate parameters from within programs, and also a set of host programs to manipulate parameters at the command line or from within scripts. The parameter interface has been implemented to support parameter storage formats other than IRAF parameter files, allowing one, for example, to access parameters that are stored in data bases. An X Windows graphical user interface called 'agcl' has been developed, layered on top of the IRAF-compatible parameter interface, that provides a standard graphical mechanism for interacting with IRAF and non-IRAF programs. Users can edit parameters and run programs for both non-IRAF programs and IRAF tasks. The agcl interface allows one to communicate with any command line environment in a transparent manner and without any changes to the original environment. For example, the authors routinely layer the GUI on top of IRAF, ksh, SMongo, and IDL. The agcl, based on the facilities of a system called Answer Garden, also has sophisticated support for examining documentation and help files, asking questions of experts, and developing a knowledge base of frequently required information. Thus, the GUI becomes a total environment for running programs, accessing information, examining documents, and finding human assistance. Because the agcl can communicate with any command-line environment, most projects can make use of it easily. New applications are continually being found for these interfaces. It is the authors' intention to evolve the GUI and its underlying parameter interface in response to these needs - from users as well as developers - throughout the astronomy community. This presentation describes the capabilities and technology of the above user interface mechanisms and tools. It also discusses the design philosophies guiding the work, as well as hopes for the future.

Mandel, Eric

An overview of the EOSDIS V0 information management system: Lessons learned from the implementation of a distributed data system

The EOSDIS Version 0 system, released in July, 1994, is a working prototype of a distributed data system. One of the purposes of the V0 project is to take several existing data systems and coordinate them into one system while maintaining the independent nature of the original systems. The project is a learning experience and the lessons are being passed on to the architects of the system which will distribute the data received from the planned EOS satellites. In the V0 system, the data resides on heterogeneous systems across the globe but users are presented with a single, integrated interface. This interface allows users to query the participating data centers based on a wide set of criteria. Because this system is a prototype, we used many novel approaches in trying to connect a diverse group of users with the huge amount of available data. Some of these methods worked and others did not. Now that V0 has been released to the public, we can look back at the design and implementation of the system and also consider some possible future directions for the next generation of EOSDIS.

Ryan, Patrick M.