Search NASASearch

SEARCH · Search NASA

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Automated Classification of ROSAT Sources Using Heterogeneous Multiwavelength Source Catalogs

We describe an on-line system for automated classification of X-ray sources, ClassX, and present preliminary results of classification of the three major catalogs of ROSAT sources, RASS BSC, RASS FSC, and WGACAT, into six class categories: stars, white dwarfs, X-ray binaries, galaxies, AGNs, and clusters of galaxies. ClassX is based on a machine learning technology. It represents a system of classifiers, each classifier consisting of a considerable number of oblique decision trees. These trees are built as the classifier is 'trained' to recognize various classes of objects using a training sample of sources of known object types. Each source is characterized by a preselected set of parameters, or attributes; the same set is then used as the classifier conducts classification of sources of unknown identity. The ClassX pipeline features an automatic search for X-ray source counterparts among heterogeneous data sets in on-line data archives using Virtual Observatory protocols; it retrieves from those archives all the attributes required by the selected classifier and inputs them to the classifier. The user input to ClassX is typically a file with target coordinates, optionally complemented with target IDs. The output contains the class name, attributes, and class probabilities for all classified targets. We discuss ways to characterize and assess the classifier quality and performance and present the respective validation procedures. Based on both internal and external validation, we conclude that the ClassX classifiers yield reasonable and reliable classifications for ROSAT sources and have the potential to broaden class representation significantly for rare object types.

McGlynn, Thomas

An Ecological Forecasting Agent

The project goals are: Make data analysis faster and cheaper. Increase use of NASA data by removing barriers to data access. Cope with data heterogeneity. Support code reuse and rapid application development. Support multiple applications, users. Including fire and health domains. Improve QOS. Always provide an answer. Tell user how good it is, where it come from.

Golden, Keith

Adding Hierarchical Objects to Relational Database General-Purpose XML-Based Information Managements

NETMARK is a flexible, high-throughput software system for managing, storing, and rapid searching of unstructured and semi-structured documents. NETMARK transforms such documents from their original highly complex, constantly changing, heterogeneous data formats into well-structured, common data formats in using Hypertext Markup Language (HTML) and/or Extensible Markup Language (XML). The software implements an object-relational database system that combines the best practices of the relational model utilizing Structured Query Language (SQL) with those of the object-oriented, semantic database model for creating complex data. In particular, NETMARK takes advantage of the Oracle 8i object-relational database model using physical-address data types for very efficient keyword searches of records across both context and content. NETMARK also supports multiple international standards such as WEBDAV for drag-and-drop file management and SOAP for integrated information management using Web services. The document-organization and -searching capabilities afforded by NETMARK are likely to make this software attractive for use in disciplines as diverse as science, auditing, and law enforcement.

Lin, Shu-Chun

Bridging the Last Mile with Open-Source Advancements: Empowering Communities through Fusion of Aerosol Optical Depth (AOD) Products from Multi-Satellite Sensors

Aerosol Optical Depth (AOD) is a crucial parameter for understanding atmospheric aerosol distribution and their impact on climate and air quality. With the growing number of Earth observation satellites, there is an abundance of AOD products derived from various sensors onboard both geostationary and low-orbit satellites. The availability of multiple datasets provides an opportunity to harness the strengths of each sensor and create comprehensive and accurate AOD datasets for climate and air quality studies at different temporal and spatial scales. Our NASA aerosol MEaSURES project has made significant strides in recent years by undertaking the ambitious task of developing an open-source package tailored for fusing AOD products from different sources. The package is based on OOP (Object-Oriented Programming) design and is implemented in Python modules. Generic interfaces enable easy inclusion of large and heterogeneous data. The package may be utilized to produce harmonized AOD datasets with enhanced spatial and temporal coverage. The latest version of the package is able to process and integrate the dark-target AOD data from six different sensors: AHI Himawari-8, ABI GOES-West, ABI GOES-East, MODIS AQUA, MODIS TERRA, and VIIRS SNPP. Rigorous validation and intercomparison studies have been performed to assess the accuracy and reliability of the fused AOD product against ground-based measurements and reference datasets. The open-source nature of the developed package ensures transparency, reproducibility, and community engagement. The research community and stakeholders can access, contribute to, and further improve the fusion methodology, making it adaptable to other studies, or expanding it to include new satellite data as they become available. In this poster presentation, we will introduce the accomplishments and challenges faced during the development of the open-source package for AOD data fusion, and demonstrate the advantages of combining AOD products from the six aforementioned satellite sensors. The presentation aims to foster discussions, collaborations, and future directions in integrating Earth observation and remote sensing data, which may contribute to a better understanding of atmospheric aerosols and their impacts on our environment.

Zhaohui Zhang

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]

Issues and Solutions for Bringing Heterogeneous Water Cycle Data Sets Together

The water cycle research community has generated many regional to global scale products using data from individual NASA missions or sensors (e.g., TRMM, AMSR-E); multiple ground- and space-based data sources (e.g., Global Precipitation Climatology Project [GPCP] products); and sophisticated data assimilation systems (e.g., Land Data Assimilation Systems [LDAS]). However, it is often difficult to access, explore, merge, analyze, and inter-compare these data in a coherent manner due to issues of data resolution, format, and structure. These difficulties were substantiated at the recent Collaborative Energy and Water Cycle Information Services (CEWIS) Workshop, where members of the NASA Energy and Water cycle Study (NEWS) community gave presentations, provided feedback, and developed scenarios which illustrated the difficulties and techniques for bringing together heterogeneous datasets. This presentation reports on the findings of the workshop, thus defining the problems and challenges of multi-dataset research. In addition, the CEWIS prototype shown at the workshop will be presented to illustrate new technologies that can mitigate data access roadblocks encountered in multi-dataset research, including: (1) Quick and easy search and access of selected NEWS data sets. (2) Multi-parameter data subsetting, manipulation, analysis, and display tools. (3) Access to input and derived water cycle data (data lineage). It is hoped that this presentation will encourage community discussion and feedback on heterogeneous data analysis scenarios, issues, and remedies.

Acker, James

A Comparison of FIFE Observation with GEOS Assimilated Data Including a Heterogeneous LSM

Several recent studies have shown that much can be learned by comparing grid-point data from a data assimilation system with in-situ observations from field experiments. While the surface heterogeneity is acknowledged in these studies, they lack quantitative representations of the influence of heterogeneity on the near-surface meteorology and surface hydrologic and energy balance. Here, we use the Betts and Ball FIFE site-averaged data. Standard deviations of the site-average will provide an estimate of the FIFE site heterogeneity. Recently, the Mosaic Land-Surface Model (LSM) has been incorporated into the Goddard Earth Observing System (GEOS) Data Assimilation System (DAS). The Mosaic LSM computes the surface energy and hydrologic balance for nine distinct surface types at each grid-point. Each surface type is proportionally weighted to determine the mean grid point properties. Hence, we can compare modeled and observed grid-point variability in addition to the mean properties. Also, assimilated data sets created with and without the LSM are compared. The results indicate the importance of including quantitative estimates of heterogeneity in the analysis of the land surface hydrology and energy balances in assimilation systems.

Bosilovich, M.

Problems related to the determination of land surface parameters and fluxes over heterogeneous media from satellite data

Problems encountered in the definition of relevant parameters which can describe heterogeneous media as a whole are discussed. A procedure to extend the definition of parameters from local to regional scales via inversion of appropriate models is tentatively proposed. The adequacy of these models for describing physical processes at the earth/atmosphere interface as observed with satellite systems is addressed.

Becker, F.

A Study on Co-existing Heterogeneous Wireless Networks for Data Transmission within a Nuclear Facility

Deployment of wireless technologies is a salient need for modernization, automation and improved operation of nuclear power plants (NPPs). As a single technology cannot support the ever-changing needs, it is required to have a heterogeneous wireless network architecture to address the different technical and economic challenges. However, the coexistence of these multiband heterogeneous wireless networks brings numerous challenges due to the factors including dissimilarity in their channel access mechanism, distance between nodes, transmit power level and many more. This paper develops real-world experiments and simulations of wireless coexistence for Wi-Fi, Fifth generation cellular (5G) and Zigbee in the unlicensed band to understand the challenges and opportunities. The experiments were conducted over the Platform for Open Wireless Data-driven Experimental Research (POWDER) testbed at the university of Utah. In addition, this paper is the first to propose a novel packet rate control technique at the network layer to create temporary opportunities for 5G or Zigbee signal transmissions focusing its application in a nuclear facility while using the shared band. The performance of the proposed coexistence solution is validated with experimental results and simulation.

5G

Development of a Heterogeneous sUAS High-Accuracy Positional Flight Data Acquisition System

Recently, a heterogeneous FDAS, consisting of a diverse range of instruments was developed to support acoustic flight research programs at NASA Langley Research Center. In addition to a conventional GPS to measure latitude, longitude and altitude, the FDAS also utilizes a small, light-weight, low-cost DGPS system to obtain centimeter accuracy to measure the distance traveled by sound from a sUAS vehicle to a microphone on the ground. Acoustic flight testing using the FDAS installed on several different sUAS platforms has been conducted in support of the NASA CAS DELIVER and ERA ITD projects (Reference 1). The first FDAS prototype was assembled and implemented in the acoustic/flight measurement system in December 2014 to support DELIVER acoustic flight tests. Evaluation of the system performance and results from the data analyses were used to further test, develop and enhance the FDAS over a six-month period to support acoustic flight research for the ERA.

McSwain, Robert G.

Automated Traffic Management System and Method

A data management system and method that enables acquisition, integration, and management of real-time data generated at different rates, by multiple heterogeneous incompatible data sources. The system achieves this functionality by using an expert system to fuse data from a variety of airline, airport operations, ramp control, and air traffic control tower sources, to establish and update reference data values for every aircraft surface operation. The system may be configured as a real-time airport surface traffic management system (TMS) that electronically interconnects air traffic control, airline data, and airport operations data to facilitate information sharing and improve taxi queuing. In the TMS operational mode, empirical data shows substantial benefits in ramp operations for airlines, reducing departure taxi times by about one minute per aircraft in operational use, translating as $12 to $15 million per year savings to airlines at the Atlanta, Georgia airport. The data management system and method may also be used for scheduling the movement of multiple vehicles in other applications, such as marine vessels in harbors and ports, trucks or railroad cars in ports or shipping yards, and railroad cars in switching yards. Finally, the data management system and method may be used for managing containers at a shipping dock, stock on a factory floor or in a warehouse, or as a training tool for improving situational awareness of FAA tower controllers, ramp and airport operators, or commercial airline personnel in airfield surface operations.

Glass, Brian J.

Real-Time Surface Traffic Adviser

A real-time data management system which uses data generated at different rates by multiple heterogeneous incompatible data sources are presented. In one embodiment, the invention is as an airport surface traffic data management system (traffic adviser) that electronically interconnects air traffic control, airline, and airport operations user communities to facilitate information sharing and improve taxi queuing. The system uses an expert system to fuse dam from a variety of airline, airport operations, ramp control, and air traffic control sources, in order to establish, predict, and update reference data values for every aircraft surface operation.

Brian J Glass

Electrical Load Forecasting Over Multihop Smart Metering Networks With Federated Learning

Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) record household energy data. Traditional machine learning (ML) methods are often employed for load forecasting, but require data sharing, which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. Here, this article presents a novel personalized FL (PFL) method for high-quality load forecasting in metering networks. A meta-learning-based strategy is developed to address data heterogeneity at local SMs in the collaborative training of local load forecasting models. Moreover, to minimize the load forecasting delays in our PFL model, we study a new latency optimization problem based on optimal resource allocation at SMs. A theoretical convergence analysis is also conducted to provide insights into FL design for federated load forecasting. Extensive simulations from real-world datasets show that our method outperforms existing approaches regarding better load forecasting and reduced operational latency costs.

Rahman, Ratun [Univ. of Alabama, Huntsville, AL (U

Deep Domain Adaptation based Cloud Type Detection using Active and Passive Satellite Data

Domain adaptation techniques have been developed to handle data from multiple sources or domains. Most existing domain adaptation models assume that source and target domains are homogeneous, i.e., they have the same feature space. Nevertheless, many real world applications often deal with data from heterogeneous domains that come from completely different feature spaces. In our remote sensing application, data in source domain (from an active spaceborne Lidar sensor CALIOP onboard CALIPSO satellite) contain 25 attributes, while data in target domain (from a passive spectroradiometer sensor VIIRS onboard Suomi-NPP satellite) contain 20 different attributes. CALIOP has better representation capability and sensitivity to aerosol types and cloud phase, while VIIRS has wide swaths and better spatial coverage but has inherent weakness in differentiating atmospheric objects on different vertical levels. To address this mismatch of features across the domains/sensors, we propose a novel end-to-end deep domain adaptation with domain mapping and correlation alignment (DAMA) to align the heterogeneous source and target domains in active and passive satellite remote sensing data. It can learn domain invariant representation from source and target domains by transferring knowledge across these domains, and achieve additional performance improvement by incorporating weak label information into the model (DAMA-WL). Our experiments on a collocated CALIOP and VIIRS dataset show that DAMA and DAMA-WL can achieve higher classification accuracy in predicting cloud types.

domain adaptation

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning

Version 0 EOSDIS - An overview

Attention is given to NASA's Earth Observing System Data and Information System (EOSDIS), which is to be a single, distributed but internally consistent evolutionary system to support the planning and execution of EOS data acquisitions and to process, archive, and distribute EOS data products and selected non-EOS data to enable interdisciplinary studies of the earth. V0 EOSDIS, a logical step in this evolutionary process, is to address both technical and managerial challenges. Technical challenges include developing a multidiscipline, distributed system for searching and ordering data in a heterogeneous environment, and standardizing data formats and distribution techniques among differing communities and organizations. Managerial challenges include establishing and maintaining a structure consisting of geographically distributed entities such that cooperative development is carried out effectively despite organizational differences, maintaining interactions with the scientific community to ensure its close involvement despite its size and diversity, and keeping the expectations for V0 consistent with its schedules and resources.

Ramapriyan, H. K.