Search NASA⌕ Search

SEARCH · Search NASA

Results for “data aggregation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multidimensional Data Aggregation in the Cloud with Application to Geostationary Satellite-based Air Quality Monitoring

Scientists use satellite data for studying Earth's systems, and the remote sensing data that these satellites collect are typically separated into files of a size small enough for efficient network transfer and storage. However, researchers usually prefer to analyze the data based on real-world dimensions like time, space, or elevation. To help with this, NASA's Atmospheric Science Data Center (ASDC) developed a new cloud-based tool that combines these smaller data chunks into larger, more useful datasets. The tool works on Network Common Data Form (netCDF4) and some HDF5 formatted files, and it is available as a service in NASA's Earthdata Cloud. In this presentation, we showcase this service using data from the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. By combining TEMPO's continuous observations over time, we create longer and more informative analysis-ready time series to facilitate the study of air quality patterns. Insights gained will provide a more comprehensive understanding of pollution sources, transport patterns, and their effects on the environment and human health.

Daniel Kaufman↗

Integrating Engineering Data Systems for NASA Spaceflight Projects

NASA has a large range of custom-built and commercial data systems to support spaceflight programs. Some of the systems are re-used by many programs and projects over time. Management and systems engineering processes require integration of data across many of these systems, a difficult problem given the widely diverse nature of system interfaces and data models. This paper describes an ongoing project to use a central data model with a web services architecture to support the integration and access of linked data across engineering functions for multiple NASA programs. The work involves the implementation of a web service-based middleware system called Data Aggregator to bring together data from a variety of systems to support space exploration. Data Aggregator includes a central data model registry for storing and managing links between the data in disparate systems. Initially developed for NASA's Constellation Program needs, Data Aggregator is currently being repurposed to support the International Space Station Program and new NASA projects with processes that involve significant aggregating and linking of data. This change in user needs led to development of a more streamlined data model registry for Data Aggregator in order to simplify adding new project application data as well as standardization of the Data Aggregator query syntax to facilitate cross-application querying by client applications. This paper documents the approach from a set of stand-alone engineering systems from which data are manually retrieved and integrated, to a web of engineering data systems from which the latest data are automatically retrieved and more quickly and accurately integrated. This paper includes the lessons learned through these efforts, including the design and development of a service-oriented architecture and the evolution of the data model registry approaches as the effort continues to evolve and adapt to support multiple NASA programs and priorities.

Carvalho, Robert E.↗

Relationship Between Ecosystem Productivity and Photosynthetically Active Radiation for Northern Peatlands

We analyzed the relationship between net ecosystem exchange of carbon dioxide (NEE) and irradiance (as photosynthetic photon flux density or PPFD), using published and unpublished data that have been collected during midgrowing season for carbon balance studies at seven peatlands in North America and Europe, NEE measurements included both eddy-correlation tower and clear, static chamber methods, which gave very similar results. Data were analyzed by site, as aggregated data sets by peatland type (bog, poor fen, rich fen, and all fens) and as a single aggregated data set for all peatlands. In all cases, a fit with a rectangular hyperbola (NEE = alpha PPFD P(sub max)/(alpha PPFD + P(sub max) + R) better described the NEE-PPFD relationship than did a linear fit (NEE = beta PPFD + R). Poor and rich fens generally had similar NEE-PPFD relationships, while bogs had lower respiration rates (R = -2.0 micro mol m(exp -2) s(exp -1) for bogs and -2.7 micro mol m(exp -2) s(exp -1)) for fens) and lower NEE at moderate and high light levels (P(sub max)= 5.2 micro mol m(exp -2) s(exp -1) for bogs and 10.8 micro mol m(exp -2) s(exp -1) for fens). As a single class, northern peatlands had much smaller ecosystem respiration (R = -2.4 micro mol m(exp -2) s(exp -1)) and NEE rates (alpha = 0.020 and P(sub max)= 9.2 micro mol m(exp -2) s(exp -1)) than the upland ecosystems (closed canopy forest, grassland, and cropland). Despite this low productivity, northern peatland soil carbon pools are generally 5-50 times larger than upland ecosystems because of slow rates of decomposition caused by litter quality and anaerobic, cold soils.

Frolking, S. E.↗

Forming Aggregations using Virtual Sharding: Lessons Learned from Simple Scalable Storage (S3)

Data aggregation is the ability to combine separate datasets to form a single new logical dataset provides users with a powerful abstraction. The advantage of an aggregate dataset is that the users are freed from having to understand, and incorporate into their workflow, knowledge about the (ad hoc) organization of the constituent datasets. However, aggregating large numbers of files can be computationally complex with data server systems performing many repetitive operations. As part of the authors work on subsetting data stored on Amazon Web Service (AWS) Simple Storage Service (S3), we developed technology to read portions of otherwise monolithic data files. This enables the formation of virtual shards for user in subsetting data stored in HDF5 (hierarchical data format, version 5) files. This same tool can be used to form aggregations that combine data stored in many HDF5 files when those files are stored on S3. The nature of the virtual sharding and the algorithm that exploits it for subsetting is such that it can also be used for aggregation with the need for many of the repetitive operations required by the per file aggregation techniques. We will present timing information that demonstrates the flexibility of this approach. However, the lessons learned is that while this is a useful result in and of itself, these very same techniques can be applied in other contexts where data are stored in services and on media other than S3. For example, this same technique can be applied to data stored on spinning disk. Pushing the envelope for S3 forced a reexamination of our data access techniques which lead to unexpected positive benefits.

Gallagher, James↗

On-board processing for future satellite communications systems: Satellite-Routed FDMA

A frequency division multiple access (FDMA) 30/20 GHz satellite communications architecture without on-board baseband processing is investigated. Conceptual system designs are suggested for domestic traffic models totaling 4 Gb/s of customer premises service (CPS) traffic and 6 Gb/s of trunking traffic. Emphasis is given to the CPS portion of the system which includes thousands of earth terminals with digital traffic ranging from a single 64 kb/s voice channel to hundreds of channels of voice, data, and video with an aggregate data rate of 33 Mb/s. A unique regional design concept that effectively smooths the non-uniform traffic distribution and greatly simplifies the satellite design is employed. The satellite antenna system forms thirty-two 0.33 deg beam on both the uplinks and the downlinks in one design. In another design matched to a traffic model with more dispersed users, there are twenty-four 0.33 deg beams and twenty-one 0.7 deg beams. Detailed system design techniques show that a single satellite producing approximately 5 kW of dc power is capable of handling at least 75% of the postulated traffic. A detailed cost model of the ground segment and estimated system costs based on current information from manufacturers are presented.

Berk, G.↗

Streak camera based SLR receiver for two color atmospheric measurements

To realize accurate two-color differential measurements, an image digitizing system with variable spatial resolution was designed, built, and integrated to a photon-counting picosecond streak camera, yielding a temporal scan resolution better than 300 femtosecond/pixel. The streak camera is configured to operate with 3 spatial channels; two of these support green (532 nm) and uv (355 nm) while the third accommodates reference pulses (764 nm) for real-time calibration. Critical parameters affecting differential timing accuracy such as pulse width and shape, number of received photons, streak camera/imaging system nonlinearities, dynamic range, and noise characteristics were investigated to optimize the system for accurate differential delay measurements. The streak camera output image consists of three image fields, each field is 1024 pixels along the time axis and 16 pixels across the spatial axis. Each of the image fields may be independently positioned across the spatial axis. Two of the image fields are used for the two wavelengths used in the experiment; the third window measures the temporal separation of a pair of diode laser pulses which verify the streak camera sweep speed for each data frame. The sum of the 16 pixel intensities across each of the 1024 temporal positions for the three data windows is used to extract the three waveforms. The waveform data is processed using an iterative three-point running average filter (10 to 30 iterations are used) to remove high-frequency structure. The pulse pair separations are determined using the half-max and centroid type analysis. Rigorous experimental verification has demonstrated that this simplified process provides the best measurement accuracy. To calibrate the receiver system sweep, two laser pulses with precisely known temporal separation are scanned along the full length of the sweep axis. The experimental measurements are then modeled using polynomial regression to obtain a best fit to the data. Data aggregation using normal point approach has provided accurate data fitting techniques and is found to be much more convenient than using the full rate single shot data. The systematic errors from this model have been found to be less than 3 ps for normal points.

Varghese, Thomas K.↗

Issues in knowledge representation to support maintainability: A case study in scientific data preparation

Scientific data preparation is the process of extracting usable scientific data from raw instrument data. This task involves noise detection (and subsequent noise classification and flagging or removal), extracting data from compressed forms, and construction of derivative or aggregate data (e.g. spectral densities or running averages). A software system called PIPE provides intelligent assistance to users developing scientific data preparation plans using a programming language called Master Plumber. PIPE provides this assistance capability by using a process description to create a dependency model of the scientific data preparation plan. This dependency model can then be used to verify syntactic and semantic constraints on processing steps to perform limited plan validation. PIPE also provides capabilities for using this model to assist in debugging faulty data preparation plans. In this case, the process model is used to focus the developer's attention upon those processing steps and data elements that were used in computing the faulty output values. Finally, the dependency model of a plan can be used to perform plan optimization and runtime estimation. These capabilities allow scientists to spend less time developing data preparation procedures and more time on scientific analysis tasks. Because the scientific data processing modules (called fittings) evolve to match scientists' needs, issues regarding maintainability are of prime importance in PIPE. This paper describes the PIPE system and describes how issues in maintainability affected the knowledge representation used in PIPE to capture knowledge about the behavior of fittings.

Chien, Steve↗

Intelligent assistance in scientific data preparation

Scientific data preparation is the process of extracting usable scientific data from raw instrument data. This task involves noise detection (and subsequent noise classification and flagging or removal), extracting data from compressed forms, and construction of derivative or aggregate data (e.g. spectral densities or running averages). A software system called PIPE provides intelligent assistance to users developing scientific data preparation plans using a programming language called Master Plumber. PIPE provides this assistance capability by using a process description to create a dependency model of the scientific data preparation plan. This dependency model can then be used to verify syntactic and semantic constraints on processing steps to perform limited plan validation. PIPE also provides capabilities for using this model to assist in debugging faulty data preparation plans. In this case, the process model is used to focus the developer's attention upon those processing steps and data elements that were used in computing the faulty output values. Finally, the dependency model of a plan can be used to perform plan optimization and run time estimation. These capabilities allow scientists to spend less time developing data preparation procedures and more time on scientific analysis tasks.

Chien, Steve↗

Cost effective data system design approach for EOS AM-1

The question of how to design a cost-effective data system for a space-based application with a high aggregate data rate is addressed. The steps used by the EOS AM-1 design team are defined. A summary of the design as of January 1993 is outlined. The AM-1 Science Data system is comprised of one subsystem that ingests, formats, and routes data - the Science Formatting Equipment (SFE), and another that records the data. Key interfaces are the interfaces to the high rate science instruments, the real-time transmitters, and the internal interface between the SFR and the Solid State Recorder. A cost-effective design approach for the AM-1 data system must take into account impacts on ground systems, data processing facilities, availability of data for real-time users, and future spacecraft designs.

Westmeyer, Paul A.↗

Hyperspectral Microwave Atmospheric Sounder (HyMAS) Architecture and Design Accommodations

The Hyperspectral Microwave Atmospheric Sounder (HyMAS) is being developed at Lincoln Laboratories and accommodated by the Goddard Space Flight Center for a flight opportunity on a NASA research aircraft. The term "hyperspectral microwave" is used to indicate an all-weather sounding that performs equivalent to hyperspectral infrared sounders in clear air with vertical resolution of approximately 1 km. Deploying the HyMAS equipped scanhead with the existing Conical Scanning Microwave Imaging Radiometer (CoSMIR) shortens the path to a flight demonstration. Hyperspectral microwave is achieved through the use of independent RF antennas that sample the volume of the Earth s atmosphere through various levels of frequencies, thereby producing a set of dense, spaced vertical weighting functions. The simulations proposed for HyMAS 118/183-GHz system should yield surface precipitation rate and water path retrievals for small hail, soft hail, or snow pellets, snow, rainwater, etc. with accuracies comparable to those of the Advanced Technology Microwave Sounder. Further improvements in retrieval methodology (for example, polarization exploitation) are expected. The CoSMIR instrument is a packaging concept re-used on HyMAS to ease the integration features of the scanhead. The HyMAS scanhead will include an ultra-compact Intermediate Frequency Processor (IFP) module that is mounted inside the door to improve thermal management. The IFP is fabricated with materials made of Low-Temperature Co-fired Ceramic (LTCC) technology integrated with detectors, amplifiers, A/D conversion and data aggregation. The IFP will put out 52 channels of 16 bit data comprised of 4-9 channel data streams for temperature profiles and 2-8 channel streams for water vapor. With the limited volume of the existing CoSMIR scanhead and new HyMAS front end components, the HyMAS team at Goddard began preliminary layout work inside the new drum. Importing and re-using models of the shell, the scan head computer, and the slip rings developed for CoSMIR was the starting point. The next step was to modify the antenna faceplate to accommodate the dimensions of the three dual polarization Gaussian Optics Antenna (GOA) assemblies. Two mechanical concepts for the core technology, the hyperspectral IFP, were captured in a design tradeoff. Connector models considered minimum bend radii for the IFP analog connectors. Hyperspectral imaging is accomplished by strategically using a short wavelength intermediate frequency of 18-29 GHz, and thus reducing the size of components in the connection of the front end to the IFP. The SMK (2.92mm) Series connector will lay near the hinge line to minimize its flexing. The digital output of the IFP will use a Serial Peripheral Interface (SPI) that must be accommodated by the scan head computer. To make that computer more reliable, maintainable, and forward compatible with the 52 HyMAS channels, a testbed of the scan head, calibration, and archive computers and the PIC24 microprocessor that resides on the IFP is in development. The computers will be programmed using a new framework application called Interoperable Remote Component (IRC). This software allows flexibility to program computers that communicate with each other and can adapt easily to the emerging HyMAS requirements for data format, algorithms, and graphical user interface (GUI). It is expected that the CoSMIR instrument will cut over to the IRC after it is adapted on an updated CoSMIR testbed.

Hilliard, Lawrence↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

Mobile and replicated alignment of arrays in data-parallel programs

When a data-parallel language like FORTRAN 90 is compiled for a distributed-memory machine, aggregate data objects (such as arrays) are distributed across the processor memories. The mapping determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. A common approach is to break the mapping into two stages: first, an alignment that maps all the objects to an abstract template, and then a distribution that maps the template to the processors. We solve two facets of the problem of finding alignments that reduce residual communication: we determine alignments that vary in loops, and objects that should have replicated alignments. We show that loop-dependent mobile alignment is sometimes necessary for optimum performance, and we provide algorithms with which a compiler can determine good mobile alignments for objects within do loops. We also identify situations in which replicated alignment is either required by the program itself (via spread operations) or can be used to improve performance. We propose an algorithm based on network flow that determines which objects to replicate so as to minimize the total amount of broadcast communication in replication. This work on mobile and replicated alignment extends our earlier work on determining static alignment.

Chatterjee, Siddhartha↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment-distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

Selection of a map grid for data analysis and archival

Arguments for selection of a map grid are reiterated by illustrating the quantitative effects on data quality caused by using different grids. It is shown that the use of the rectangular latitude-longitude grid actually degrades data quality, increases stored data volume, and increases the complexity of data manipulation. Use of an equal-area grid is shown to be a proper way to aggregate data, but such grids are thought to be inconvenient. A solution to this dilemma is proposed by showing that the analysis and archival map grids need not be the same. The results of a proper analysis of the data on an equal-area grid can be remapped to the more 'convenient' rectangular latitude-longitude grid without loss of quality.

Rossow, W. B.↗

NASA Earth Sciences Data Support System and Services for the Northern Eurasia Earth Science Partnership Initiative

The presentation describes data management of NASA remote sensing data for Northern Eurasia Earth Science Partnership Initiative (NEESPI). Many types of ground and integrative (e.g., satellite, GIs) data will be needed and many models must be applied, adapted or developed for properly understanding the functioning of Northern Eurasia cold and diverse regional system. Mechanisms for obtaining the requisite data sets and models and sharing them among the participating scientists are essential. The proposed project targets integration of remote sensing data from AVHRR, MODIS, and other NASA instruments on board US- satellites (with potential expansion to data from non-US satellites), customized data products from climatology data sets (e.g., ISCCP, ISLSCP) and model data (e.g., NCEPNCAR) into a single, well-architected data management system. It will utilize two existing components developed by the Goddard Earth Sciences Data & Information Services Center (GES DISC) at the NASA Goddard Space Flight Center: (1) online archiving and distribution system, that allows collection, processing and ingest of data from various sources into the online archive, and (2) user-friendly intelligent web-based online visualization and analysis system, also known as Giovanni. The former includes various kinds of data preparation for seamless interoperability between measurements by different instruments. The latter provides convenient access to various geophysical parameters measured in the Northern Eurasia region without any need to learn complicated remote sensing data formats, or retrieve and process large volumes of NASA data. Initial implementation of this data management system will concentrate on atmospheric data and surface data aggregated to coarse resolution to support collaborative environment and climate change studies and modeling, while at later stages, data from NASA and non-NASA satellites at higher resolution will be integrated into the system.

Leptoukh, Gregory↗

NASA Earth Sciences Data Support System and Services for the Northern Eurasia Earth Science Partnership Initiative

The presentation describes the recently awarded ACCESS project to provide data management of NASA remote sensing data for the Northern Eurasia Earth Science Partnership Initiative (NEESPI). The project targets integration of remote sensing data from MODIS, and other NASA instruments on board US-satellites (with potential expansion to data from non-US satellites), customized data products from climatology data sets (e.g., ISCCP, ISLSCP) and model data (e.g., NCEP/NCAR) into a single, well-architected data management system. It will utilize two existing components developed by the Goddard Earth Sciences Data & Information Services Center (GES DISC) at the NASA Goddard Space Flight Center: (1) online archiving and distribution system, that allows collection, processing and ingest of data from various sources into the online archive, and (2) user-friendly intelligent web-based online visualization and analysis system, also known as Giovanni. The former includes various kinds of data preparation for seamless interoperability between measurements by different instruments. The latter provides convenient access to various geophysical parameters measured in the Northern Eurasia region without any need to learn complicated remote sensing data formats, or retrieve and process large volumes of NASA data. Initial implementation of this data management system will concentrate on atmospheric data and surface data aggregated to coarse resolution to support collaborative environment and climate change studies and modeling, while at later stages, data from NASA and non-NASA satellites at higher resolution will be integrated into the system.

Leptoukh, Gregory↗

NASA Wrangler: Automated Cloud-Based Data Assembly in the RECOVER Wildfire Decision Support System

NASA Wrangler is a loosely-coupled, event driven, highly parallel data aggregation service designed to take advantageof the elastic resource capabilities of cloud computing. Wrangler automatically collects Earth observational data, climate model outputs, derived remote sensing data products, and historic biophysical data for pre-, active-, and post-wildfire decision making. It is a core service of the RECOVER decision support system, which is providing rapid-response GIS analytic capabilities to state and local government agencies. Wrangler reduces to minutes the time needed to assemble and deliver crucial wildfire-related data.

decision support↗