Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

Integration of socioeconomic data and remotely sensed imagery for land use applications

The Image Based Information System (IBIS) uses techniques of digital image processing to interface geocoded data sets, information management systems, thematic maps, and remotely sensed imagery. For two cases IBIS has been used to integrate remotely sensed imagery and socioeconomic data. In the first case thematically classified Landsat imagery covering Orange County, California is combined with census tracts and municipalities. In the second case, thematically classified Landsat imagery over Los Angeles, California is combined with census tracts.

Bryant, N. A.↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Semantic Representation and Scale-Up of Integrated Air Traffic Management Data

Each day, the global air transportation industry generates a vast amount of heterogeneous data from air carriers, air traffic control providers, and secondary aviation entities handling baggage, ticketing, catering, fuel delivery, and other services. Generally, these data are stored in isolated data systems, separated from each other by significant political, regulatory, economic, and technological divides. These realities aside, integrating aviation data into a single, queryable, big data store could enable insights leading to major efficiency, safety, and cost advantages. In this paper, we describe an implemented system for combining heterogeneous air traffic management data using semantic integration techniques. The system transforms data from its original disparate source formats into a unified semantic representation within an ontology-based triple store. Our initial prototype stores only a small sliver of air traffic data covering one day of operations at a major airport. The paper also describes our analysis of difficulties ahead as we prepare to scale up data storage to accommodate successively larger quantities of data -- eventually covering all US commercial domestic flights over an extended multi-year timeframe. We review several approaches to mitigating scale-up related query performance concerns.

data management↗

Semantic Representation and Scale-Up of Integrated Air Traffic Management Data

Each day, the global air transportation industry generates a vast amount of heterogeneous data from air carriers, air traffic control providers, and secondary aviation entities handling baggage, ticketing, catering, fuel delivery, and other services. Generally, these data are stored in isolated data systems, separated from each other by significant political, regulatory, economic, and technological divides. These realities aside, integrating aviation data into a single, queryable, big data store could enable insights leading to major efficiency, safety, and cost advantages. In this paper, we describe an implemented system for combining heterogeneous air traffic management data using semantic integration techniques. The system transforms data from its original disparate source formats into a unified semantic representation within an ontology-based triple store. Our initial prototype stores only a small sliver of air traffic data covering one day of operations at a major airport. The paper also describes our analysis of difficulties ahead as we prepare to scale up data storage to accommodate successively larger quantities of data -- eventually covering all US commercial domestic flights over an extended multi-year timeframe. We review several approaches to mitigating scale-up related query performance concerns.

air traffic management↗

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

he exponential growth of data centers—driven by artificial intelligence and cloud computing—is reshaping the U.S. energy landscape, presenting urgent challenges and transformative opportunities for data center developers, nuclear energy providers, and utility operators. As data centers are projected to consume up to 12% of U.S. electricity by 2028, stakeholders must address rapid deployment needs, grid congestion, and the demand for reliable, high-quality power. This presentation explores the multifaceted barriers to integrating data centers with nuclear and utility infrastructure, including land use constraints, public perception, regulatory complexity, and workforce alignment. It highlights the distinct priorities and operational cultures of each sector, and the friction that arises from misaligned planning horizons and risk tolerances. We examine collaborative strategies such as co-siting, hybrid power-purchase agreements, unified community engagement, and innovative financing models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Integrating Gridded NASA Hydrological Data into CUAHSI HIS

The amount of hydrological data available from NASA remote sensing and modeling systems is vast and ever-increasing;but, one challenge persists:increasing the usefulness of these data for, and thus their use by, end user communities. The Hydrology Data and Information Services Center (HDISC), part of the Goddard Earth Sciences DISC, has continually worked to better understand the hydrological data needs of different end users, to thus better able to bridge the gap between NASA data and end user communities. One effective strategy is integrating the data in to end user community tools and environments. There is an ongoing collaborative effort between NASA HDISC, NASA Hydrological Sciences Branch, and CUAHSI to integrate NASA gridded hydrology data in to the CUAHSI Hydrologic Information System (HIS).

Rui, Hualan↗

Coupled Data Assimilation for Integrated Earth System Analysis and Prediction: Goals, Challenges, and Recommendations

The purpose of this report is to identify fundamental issues for coupled data assimilation (CDA), such as gaps in science and limitations in forecasting systems, in order to provide guidance to the World Meteorological Organization (WMO) on how to facilitate more rapid progress internationally. Coupled Earth system modeling provides the opportunity to extend skillful atmospheric forecasts beyond the traditional two-week barrier by extracting skill from low-frequency state components such as the land, ocean, and sea ice. More generally, coupled models are needed to support seamless prediction systems that span timescales from weather, subseasonal to seasonal (S2S), multiyear, and decadal. Therefore, initialization methods are needed for coupled Earth system models, either applied to each individual component (called Weakly Coupled Data Assimilation - WCDA) or applied the coupled Earth system model as a whole (called Strongly Coupled Data Assimilation - SCDA). Using CDA, in which model forecasts and potentially the state estimation are performed jointly, each model domain benefits from observations in other domains either directly using error covariance information known at the time of the analysis (SCDA), or indirectly through flux interactions at the model boundaries (WCDA). Because the non-atmospheric domains are generally under-observed compared to the atmosphere, CDA provides a significant advantage over single-domain analyses. Next, we provide a synopsis of goals, challenges, and recommendations to advance CDA: Goals: (a) Extend predictive skill beyond the current capability of NWP (e.g. as demonstrated by improving forecast skill scores), (b) produce physically consistent initial conditions for coupled numerical prediction systems and reanalyses (including consistent fluxes at the domain interfaces), (c) make best use of existing observations by allowing observations from each domain to influence and improve the full earth system analysis, (d) develop a robust observation-based identification and understanding of mechanisms that determine the variability of weather and climate, (e) identify critical weaknesses in coupled models and the earth observing system, (f) generate full-field estimates of unobserved or sparsely observed variables, (g) improve the estimation of the external forcings causing changes to climate, (h) transition successes from idealized CDA experiments to real-world applications. Challenges: (a) Modeling at the interfaces between interacting components of coupled Earth system models may be inadequate for estimating uncertainty or error covariances between domains, (b) current data assimilation methods may be insufficient to simultaneously analyze domains containing multiple spatiotemporal scales of interest, (c) there is no standardization of observation data or their delivery systems across domains, (d) the size and complexity of many large-scale coupled Earth system models makes it is difficult to accurately represent uncertainty due to model parameters and coupling parameters, (e) model errors lead to local biases that can transfer between the different Earth system components and lead to coupled model biases and long-term model drift, (e) information propagation across model components with different spatiotemporal scales is extremely complicated, and must be improved in current coupled modeling frameworks, (h) there is insufficient knowledge on how to represent evolving errors in non-atmospheric model components (e.g. as sea ice, land and ocean) on the timescales of NWP.

CDA↗

A Comprehensive Search for Gamma-Ray Lines in the First Year of Data from the INTEGRAL Spectrometer

Gamma-ray lines are produced in nature by a variety of different physical processes. They can be valuable astrophysical diagnostics providing information the may be unobtainable by other means. We have carried out an extensive search for gamma-ray lines in the first year of public data from the Spectrometer (SPI) on the INTEGRAL mission. INTEGRAL has spent a large fraction of its observing time in the Galactic Plane with particular concentration in the Galactic Center (GC) region (approximately 3 Msec in the first year). Hence the most sensitive search regions are in the Galactic Plane and Center. The phase space of the search spans the energy range 20-8000 keV, and line widths from 0-1000 keV (FWHM) and includes both diffuse and point-like emission. We have searched for variable emission on time scales down to approximately 1000 sec. Diffuse emission has been searched for on a range of different spatial scales from approximately 20 degrees (the approximate field-of-view of the spectrometer) up to the entire Galactic Plane. Our search procedures were verified by the recovery of the known gamma-ray lines at 511 keV and 1809 keV at the appropriate intensities and significances. We find no evidence for any previously unknown gamma-ray lines. The upper limits range from a few x10(exp -5) per square centimeter per second to a few x10(exp -3) per square centimeter per second depending on line width, energy and exposure. Comparison is made between our results and various prior predictions of astrophysical lines

Teegarden, B. J.↗

The Space Station Data Management System - Avionics that integrate

The Space Station Data Management System (DMS) comprises the networked computers, mass storage, workstations, and instrumentation interfaces required to support onboard systems and payload operations. This paper gives an overview of the current DMS architecture and discusses its role as onboard integrator in four of its major functional areas: (1) data communication; (2) data processing; (3) data administration, storage and retrieval; and (4) data presentation at the human-computer interface.

Whitelaw, Virginia↗

Integrated Task And Data Parallel Programming: Language Design

his research investigates the combination of task and data parallel language constructs within a single programming language. There are an number of applications that exhibit properties which would be well served by such an integrated language. Examples include global climate models, aircraft design problems, and multidisciplinary design optimization problems. Our approach incorporates data parallel language constructs into an existing, object oriented, task parallel language. The language will support creation and manipulation of parallel classes and objects of both types (task parallel and data parallel). Ultimately, the language will allow data parallel and task parallel classes to be used either as building blocks or managers of parallel objects of either type, thus allowing the development of single and multi-paradigm parallel applications. 1995 Research Accomplishments In February I presented a paper at Frontiers '95 describing the design of the data parallel language subset. During the spring I wrote and defended my dissertation proposal. Since that time I have developed a runtime model for the language subset. I have begun implementing the model and hand-coding simple examples which demonstrate the language subset. I have identified an astrophysical fluid flow application which will validate the data parallel language subset. 1996 Research Agenda Milestones for the coming year include implementing a significant portion of the data parallel language subset over the Legion system. Using simple hand-coded methods, I plan to demonstrate (1) concurrent task and data parallel objects and (2) task parallel objects managing both task and data parallel objects. My next steps will focus on constructing a compiler and implementing the fluid flow application with the language. Concurrently, I will conduct a search for a real-world application exhibiting both task and data parallelism within the same program m. Additional 1995 Activities During the fall I collaborated with Andrew Grimshaw and Adam Ferrari to write a book chapter which will be included in Parallel Processing in C++ edited by Gregory Wilson. I also finished two courses, Compilers and Advanced Compilers, in 1995. These courses complete my class requirements at the University of Virginia. I have only my dissertation research and defense to complete.

Grimshaw, Andrew S.↗

Integrated Task and Data Parallel Programming

This research investigates the combination of task and data parallel language constructs within a single programming language. There are an number of applications that exhibit properties which would be well served by such an integrated language. Examples include global climate models, aircraft design problems, and multidisciplinary design optimization problems. Our approach incorporates data parallel language constructs into an existing, object oriented, task parallel language. The language will support creation and manipulation of parallel classes and objects of both types (task parallel and data parallel). Ultimately, the language will allow data parallel and task parallel classes to be used either as building blocks or managers of parallel objects of either type, thus allowing the development of single and multi-paradigm parallel applications. 1995 Research Accomplishments In February I presented a paper at Frontiers 1995 describing the design of the data parallel language subset. During the spring I wrote and defended my dissertation proposal. Since that time I have developed a runtime model for the language subset. I have begun implementing the model and hand-coding simple examples which demonstrate the language subset. I have identified an astrophysical fluid flow application which will validate the data parallel language subset. 1996 Research Agenda Milestones for the coming year include implementing a significant portion of the data parallel language subset over the Legion system. Using simple hand-coded methods, I plan to demonstrate (1) concurrent task and data parallel objects and (2) task parallel objects managing both task and data parallel objects. My next steps will focus on constructing a compiler and implementing the fluid flow application with the language. Concurrently, I will conduct a search for a real-world application exhibiting both task and data parallelism within the same program. Additional 1995 Activities During the fall I collaborated with Andrew Grimshaw and Adam Ferrari to write a book chapter which will be included in Parallel Processing in C++ edited by Gregory Wilson. I also finished two courses, Compilers and Advanced Compilers, in 1995. These courses complete my class requirements at the University of Virginia. I have only my dissertation research and defense to complete.

Grimshaw, A. S.↗

The Surface Longwave Downward Fluxes of the NASA GEWEX SRB Release 4.0 IP Products: Validation Against the Surface-Based BSRN and PMEL Observed Data

Since the NASA Global Energy Water Exchanges (GEWEX) Surface Radiation Budget (SRB) project released its 3rd version of products in 2010, the GEWEX Data Assessments Panel (GDAP) has been working on integrating various data products to address issues in the closing of the global energy and water cycles. The 4th version of the SRB products, Rel. 4.0-IP, has integrated data products from the cloud, aerosol, atmosphere, ocean surface, and land surface projects, coordinating within GDAP, to produce a long-term time series of TOA and surface radiative estimates. The Rel. 4-IP shortwave products span 34 years continuously from July 1983 to June 2017 on a quasi-equal-area 1degree longitude by 1degree latitude grid system. The longwave products are for land only from 1983-07 to 1987-12; for both land and ocean from 1988-01 to 2009-12; and ocean only from 2010-01 to 2017-06. The data are provided at 3 hourly, 3-hourly-monthly, daily and monthly means. The ISCCP HXS clouds and radiances are the key cloud input of the current GEWEX SRB algorithms. In addition, the longwave algorithm has also made changes in cloud microphysical property, surface skin temperature input, surface emissivity, atmospheric profile, adding longwave aerosol optical properties, revising cloud overlap procedure, and so on. Details of changes in both inputs and algorithms are documented in a NASA Algorithm Theoretical Basis Document (ATBD). We have validated the surface longwave downward fluxes against the surface-based Baseline Surface Radiation Network (BSRN) and the Pacific Marine Environmental Laboratory (PMEL) buoy data. As of 2020, the BSRN archive has 12,116 site-months of observed records from 73 stations on all seven continents, and as of 2017, PMEL archive has 4389 buoy months of observed records from 64 buoys deployed in the tropics of Pacific, Atlantic and Indian Oceans. This paper presents how the SRB Rel. 4.0-IP surface longwave downward fluxes compare with these surface-based measurements and how the comparison statistics differ from that of Rel. 3.0.

GEWEX SRB↗

Materials Informatics at NASA GRC: Machine Learning Surrogate Modeling, Data Management, and Integrated Toolsets for Establishing/Maintaining the Digital Thread

Integrated Computational Materials Engineering (ICME) has recently received widespread attention due to its promises in reducing dependence on physical testing for engineering design by relying on simulation, reducing both time and cost to market for various applications. ICME however requires validated multiscale material models, which heavily depend on available test data with full material and test pedigree, including material processing, test and measurement equipment, raw data collection, and analysis methodology and results that is findable and usable, along with integrated, efficient toolsets for effectively passing information across various length and time scales across such models. At the NASA Glenn Research Center under the Transformational Tools and Technologies Project, significant recent efforts have been directed towards establishing the required cyberinfrastructure to enable optimized ICME processes and the design of “fit-for-purpose” materials to achieve the goals outlined in the NASA Vision 2040 report. Such efforts include development of multiscale physics-based material models, which can be used to train highly efficient surrogate machine learning models, development of best practices and infrastructure for effective, traceable materials information management, and development of toolsets that integrate with physics-based codes, machine learning models, and an information management system to enable high throughput of materials data collection and analysis, establishment of digital twins and the digital thread, and automation of the ICME design process for material optimization.

Machine Learning↗