Search NASA⌕ Search

SEARCH · Search NASA

Results for “small files”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

Internet Protocol Enhanced over Satellite Networks

Extensive research conducted by the Satellite Networks and Architectures Branch of the NASA Lewis Research Center led to an experimental change to the Internet's Transmission Control Protocol (TCP) that will increase performance over satellite channels. The change raises the size of the initial burst of data TCP can send from 1 packet to 4 packets or roughly 4 kilobytes (kB), whichever is less. TCP is used daily by everyone on the Internet for e-mail and World Wide Web access, as well as other services. TCP is one of the feature protocols used in computer communications for reliable data delivery and file transfer. Increasing TCP's initial data burst from the previously specified single segment to approximately 4 kB may improve data transfer rates by up to 27 percent for very small files. This is significant because most file transfers in wide-area networks today are small files, 4 kilobytes or less. In addition, because data transfers over geostationary satellites can take 5 to 20 times longer than over typical terrestrial connections, increasing the initial burst of data that can be sent is extremely important. This research along with research from other institutions has led to the release of two new Request for Comments from the Internet Engineering Task Force (IETF, the international body that sets Internet standards). In addition, two studies of the implications of this mechanism were also funded by NASA Lewis.

Ivancic, William D.↗

Zebra: A striped network file system

The design of Zebra, a striped network file system, is presented. Zebra applies ideas from log-structured file system (LFS) and RAID research to network file systems, resulting in a network file system that has scalable performance, uses its servers efficiently even when its applications are using small files, and provides high availability. Zebra stripes file data across multiple servers, so that the file transfer rate is not limited by the performance of a single server. High availability is achieved by maintaining parity information for the file system. If a server fails its contents can be reconstructed using the contents of the remaining servers and the parity information. Zebra differs from existing striped file systems in the way it stripes file data: Zebra does not stripe on a per-file basis; instead it stripes the stream of bytes written by each client. Clients write to the servers in units called stripe fragments, which are analogous to segments in an LFS. Stripe fragments contain file blocks that were written recently, without regard to which file they belong. This method of striping has numerous advantages over per-file striping, including increased server efficiency, efficient parity computation, and elimination of parity update.

Hartman, John H.↗

Using Transparent Informed Prefetching (TIP) to reduce file read latency

As processor performance gains continue to outstrip Input/Output gains, I/O performance is becoming critical to overall system performance. File read latency is the most significant bottleneck for high performance I/O. Other aspects of I/O performance benefit from recent advances in disk bandwidth and throughput resulting from disk arrays, and in write performance derived from buffered write behind and the Log-structured File System. The access gap problem limiting improvements in read latency is exacerbated by distributed file systems operating over networks with diverse bandwidth. Focus is on extending the power of caching and prefetching to reduce file read latencies by exploiting hints from high-levels of a system. Such Transparent Informed Prefetching, TIP, and its benefits are described. It is argued that hints that disclose high level knowledge are a means for transferring optimization information across, without violating, module boundaries. How TIP can be used to convert the high throughput of new technologies such as disk arrays and log-structured file systems into low latency for applications is discussed. Our preliminary experiments show reductions in wall - clock execution time of 13 percent and 20 percent for a multiple module compilation tool (make) accessing data on a local disk and remote Coda file server, respectively, and a reduction of 30 percent for a text search (grep) remotely accessing many small files.

Patterson, R. H.↗

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics↗

Acquisition of and Access to Research Omics Data

Omics data are essential for understanding the myriad and complex effects of space environments on humans. To assure maximum benefit from these kinds of data, the NASA Human Research Program Data Management Plan stipulates that human omics data should be archived within and accessed through the NASA Life Sciences Portal (NLSP). The NLSP has the capability to acquire and provision access to omics (and other kinds of) research results for individual and ad-hoc groups of subjects at the direction of institutional review boards, or other authorizing bodies or individuals, per institutional, program and investigation-specific policies and procedures. However, because some single-subject omics data, like CT scans and other kinds of large, complex biomedical data, could be used to identify heretofore unknown risks to the subject’s health, or, in certain cases, be used to identify a subject, NASA Policy Directive 7170.1 describes various policies regarding the management of and access to “research genetic testing” data, which includes many kinds of omics data. For example, NPD 7170.1 prohibits access to human research genetic data by NASA personnel who make employment decisions for the subjects from whom the data were obtained. To meet the objective of acquiring research omics data for NLSP in compliance with the policies in NPD 7170.1 and other applicable NASA policies, we designed NOMADS (the NLSP Omics Multimodal Acquisition of Data System), a new component that supports the transfer of large research data files, including research genetic testing data, using one of several different transfer mechanisms. The choice of mechanism is made by the submitter of the data, with guiding information from the system, and is likely to often be determined in large part by the nature and source location of the data. For example, for small files where the source data files are not already stored in a cloud storage system, users are likely to prefer to transfer their data to the NLSP via a web browser. Conversely, for large sets of files already organized and stored in a cloud storage system, users may opt for NOMAD’s cloud-to-cloud transfer method. All omics datasets targeted for the NASA Life Sciences Data Archive must pass a variety of quality checks to ensure data integrity and adherence to the standards defined by the LSDA Data Submission Guidelines (DSG) (see https://nlsp.nasa.gov/explore/lsdahome/datasubmit). These include requirements that data are consistent with open standards established by the omics community. Non-compliant data will not be accepted however archivists are available to advise submitters on how to revise data submissions and re-submit until compliance is achieved. Following compliance with the LSDA DSG, omics data next undergo a variety of additional quality checks to ensure the data meet omics community standards. Domain specific Omics data quality control tools and techniques are continually evolving and linked to the advancements in omics assays utilized and thus, the tools and techniques utilized by the LSDA for data quality control and validation will need to be sustained accordingly. All human omics data will be access controlled according to the policies described above, and requiring IRB approval for any additional access grants once the data are acquired (including access for analysis using the NLSP workspace tools).

Omics↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X↗

Rucio at LSST/Rubin

In this presentation, we will explore the Rucio experience with the Rubin Observatory experiment. Our discussion will cover several key areas: Scalability Tests: Insights into the performance and scalability evaluations of Rucio in the context of Rubin's data needs and what we have learned, especially with many small files. Role in Rubin's Data Curation: Rubin's Data Butler: An overview of how Rucio, along with with Rubin's Data Butler using Hermes-K, which involves message passing through Kafka, is integrated in the Rubin's data curation system. Monitoring and Support: Current status of Rucio and PostgreSQL monitoring and Rucio deployment and support within the Rubin environment. Tape RSE Implementation: Deal with the order of magnitude more files going to tape than HEP. Future Needs: An examination of Rubin's evolving requirements for Rucio services and how we plan to address them.

Lee, Dennis [Fermilab]↗

Human Flight to Lunar and Beyond - Re-Learning Operations Paradigms

For the first time since the Apollo era, NASA is planning on sending astronauts on flights beyond LEO. The Human Space Flight (HSF) program started with a successful initial flight in Earth orbit, in December 2014. The program will continue with two Exploration Missions (EM): EM-1 will be unmanned and EM-2, carrying astronauts, will follow. NASA established a multi-center team to address the communications, and related tacking/navigation needs. This paper will focus on the lessons learned by the team designing the architecture and operations for the missions. Many of these Beyond Earth Orbit lessons had to be re-learned, as the HSF program has operated for many years in Earth orbit. Unlike the Apollo missions that were largely tracked by a dedicated ground network, the HSF planned missions will be tracked (at distances beyond GEO) by the DSN, a network that mostly serves robotic missions. There have been surprising challenges to the DSN as unique modern human spaceflight needs stretch the experience base beyond that of tracking robotic missions in deep space. Close interaction between the DSN and the HSF community to understand the unique needs (e.g. 2-way voice) resulted in a Concept of Operations (ConOps) that leverages both the deep space robotic and the Human LEO experiences. Several examples will be used to highlight the unique challenges the team faced in establishing the communications and tracking capabilities for HSF missions beyond Earth Orbit, including: Navigation. At LEO, HSF missions can rely on GPS devices for orbit determination. For Lunar-and-beyond HSF missions, techniques such as precision 2-way and 3-way Doppler and ranging, Delta-Difference-of-range, and eventually possibly on-board navigation will be used. At the same time, HSF presents a challenge to navigators, beyond those presented by robotic missions - navigating a dynamic/"noisy" spacecraft. Impact of latency - the delay associated with Round-Trip-Light-Time (RTLT). Imagine trying to have a 2-way discussion (audio or video) with an astronaut, with a 2-3 sec or more delay inserted (for lunar distances) or 20 minutes delay (for Mars distances). Balanced communications link. For robotic missions, there has been a heavy emphasis on higher downlink data rates, e.g. bringing back science data. Higher uplink data rates were of secondary importance, as uplink was used only to send commands (and occasionally small files) to the spacecraft. The ratio of downlink-to-uplink data rates was often 10:1 or more. For HSF, a continuous forward link is established and rates for uplink and downlink are more similar.

Kenny, Edward (Ted)↗

Usage analysis of user files in UNIX

Presented is a user-oriented analysis of short term file usage in a 4.2 BSD UNIX environment. The key aspect of this analysis is a characterization of users and files, which is a departure from the traditional approach of analyzing file references. Two characterization measures are employed: accesses-per-byte (combining fraction of a file referenced and number of references) and file size. This new approach is shown to distinguish differences in files as well as users, which cam be used in efficient file system design, and in creating realistic test workloads for simulations. A multi-stage gamma distribution is shown to closely model the file usage measures. Even though overall file sharing is small, some files belonging to a bulletin board system are accessed by many users, simultaneously and otherwise. Over 50% of users referenced files owned by other users, and over 80% of all files were involved in such references. Based on the differences in files and users, suggestions to improve the system performance were also made.

Devarakonda, Murthy V.↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Software to model AXAF image quality

This draft final report describes the work performed under this delivery order from May 1992 through June 1993. The purpose of this contract was to enhance and develop an integrated optical performance modeling software for complex x-ray optical systems such as AXAF. The GRAZTRACE program developed by the MSFC Optical Systems Branch for modeling VETA-I was used as the starting baseline program. The original program was a large single file program and, therefore, could not be modified very efficiently. The original source code has been reorganized, and a 'Make Utility' has been written to update the original program. The new version of the source code consists of 36 small source files to make it easier for the code developer to manage and modify the program. A user library has also been built and a 'Makelib' utility has been furnished to update the library. With the user library, the users can easily access the GRAZTRACE source files and build a custom library. A user manual for the new version of GRAZTRACE has been compiled. The plotting capability for the 3-D point spread functions and contour plots has been provided in the GRAZTRACE using the graphics package DISPLAY. The Graphics emulator over the network has been set up for programming the graphics routine. The point spread function and the contour plot routines have also been modified to display the plot centroid, and to allow the user to specify the plot range, and the viewing angle options. A Command Mode version of GRAZTRACE has also been developed. More than 60 commands have been implemented in a Code-V like format. The functions covered in this version include data manipulation, performance evaluation, and inquiry and setting of internal parameters. The user manual for these commands has been formatted as in Code-V, showing the command syntax, synopsis, and options. An interactive on-line help system for the command mode has also been accomplished to allow the user to find valid commands, command syntax, and command function. A translation program has been written to convert FEA output from structural analysis to GRAZTRACE surface deformation file (.dfm file). The program can accept standard output files and list files from COSMOS/M and NASTRAN finite analysis programs. Some interactive options are also provided, such as Cartesian or cylindrical coordinate transformation, coordinate shift and scale, and axial length change. A computerized database for technical documents relating to the AXAF project has been established. Over 5000 technical documents have been entered into the master database. A user can now rapidly retrieve the desired documents relating to the AXAF project. The summary of the work performed under this contract is shown.

Ahmad, Anees↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

Solar Astronomy Data Base: Packaged Information on Diskette

In its role as a library, the National Geophysical Data Center has transferred to diskette a collection of small, digital files of routinely measured solar indices for use on an IBM-compatible desktop computer. Recording these observations on diskette allows the distribution of specialized information to researchers with a wide range of expertise in computer science and solar astronomy. Every data set was made self-contained by including formats, extraction utilities, and plain-language descriptive text. Moreover, for several archives, two versions of the observations are provided - one suitable for display, the other for analysis with popular software packages. Since the files contain no control characters, each one can be modified with any text editor.

Mckinnon, John A.↗

An area model for on-chip memories and its application

An area model suitable for comparing data buffers of different organizations and arbitrary sizes is described. The area model considers the supplied bandwidth of a memory cell and includes such buffer overhead as control logic, driver logic, and tag storage. The model gave less than 10 percent error when verified against real caches and register files. It is shown that, comparing caches and register files in terms of area for the same storage capacity, caches generally occupy more area per bit than register files for small caches because the overhead dominates the cache area at these sizes. For larger caches, the smaller storage cells in the cache provide a smaller total cache area per bit than the register set. Studying cache performance (traffic ratio) as a function of area, it is shown that, for small caches, direct-mapped caches perform significantly better than four-way set-associative caches and, for caches of medium areas, both direct-mapped and set-associative caches perform better than fully associative caches.

Mulder, Johannes M.↗

Commercialization of LARC (TradeMark) -SI Polyimide Technology

LARC(TradeMark)-SI, Langley Research Center- Soluble Imide, was developed in 1992, with the first patent issuing in 1997, and then subsequent patents issued in 1998 and 2000. Currently, this polymer has been successfully licensed by NASA, and has generated revenues, at the time of this reporting, in excess of $1.4 million. The success of this particular polymer has been due to many factors and many lessons learned to the point that the invention, while important, is the least significant part in the commercialization of this material. Commercial LARC(TradeMark)-SI is a polyimide composed of two molar equivalents of dianhydrides: 4,4 -oxydiphthalic anhydride (ODPA), and 3,3 ,4,4 -biphenyltetracarboxylic dianhydride (BPDA) and 3,4 -oxydianiline (3,4 -ODA) as the diamine. The unique feature of this aromatic polyimide is that it remains soluble after solution imidization in high-boiling, polar aprotic solvents, even at solids contents of 50-percent by weight. However, once isolated and heated above its T(sub g) of 240 C, it becomes insoluble and exhibits high-temperature thermoplastic melt-flow behavior. With these unique structure property characteristics, it was thought this would be an advantage to have an aromatic polyimide that is both solution and melt processable in the imide form. This could potentially lead to lower cost production as it was not as equipment- or labor-intensive as other high-performance polyimide materials that either precipitate or are intractable. This unique combination of properties allowed patents with broad claim coverage and potential commercialization. After the U.S. Patent applications were filed, a Small Business Innovation Research (SBIR) contract was awarded to Imtec, Inc. to develop and supply the polyimide to NASA and the general public. Some examples of demonstration parts made with LARC(TradeMark)-SI ranged from aircraft wire and multilayer printed-circuit boards, to gears, composite panels, supported adhesive tape, composite coatings, cookware, and polyimide foam. Even with its unique processing characteristics, the thermal and mechanical properties were not drastically different from other solution or meltprocessable polyimides developed by NASA. LARC(TradeMark)-SI risked becoming another interesting, but costly, high-performance material.

Bryant, Robert G.↗

Providing Data Access and Analysis Capabilities to SERVIR’s Data-Sparse Regions

In developing regions of the world, the communications infrastructure pose enormous challenges for using Earth observation data. Limited internet bandwidth along with the high costs make it almost impossible to process and extract zonal statistics over large periods of time for even small geographic areas. In such cases, downloading daily rainfall data or dekadal series of NDVI data would take days and consume all the bandwidth allocated to an organization (for reference, internet connections in Niger would cost thousands of dollars per month at a maximum - and unreliable - bandwidth of just 10 Mbps). Running crop models or hydrological models typically require several years of historic data over the area of interest (AOI). In some cases, these AOIs are relatively small compared to the footprint of individual earth observation granules. Hence, systems that let the stakeholders subset the data to download to a user specified area, or even submit processing requests that let them download small result files for the AOI become critical. The SERVIR program has developed a tool to provide this type of access to help decision makers in developing regions use long time series of adjusted rainfall data (CHIRPS), NDVI values, seasonal weather forecasts, evaporative stress indices and others in a very efficient manner. This system, named ClimateSERV (https://climateserv.servirglobal.net) ingests the datasets in an automated fashion and allows interactive access (through a web application), or automated access through a simple API that developers can quickly incorporate in independent applications. This way, the extraction of daily averages of rainfall over a 50 square Km area through 30 years of archived data takes only a few seconds to process, and the results can be presented on an online chart or downloaded in a comma separated file that's only a few Kb.

Ashmall, William↗