Search NASA⌕ Search

SEARCH · Search NASA

Results for “pipeline data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

AstroAmpSeq: Microbial Bioinformatics Education with NASA GeneLab’s Amplicon Pipeline

The prevalence and importance of large sequencing datasets in microbiology has led to a movement to share microbial ecology experimental data through open-access databases. This is particularly true of experiments that are difficult to replicate, such as those conducted in the spaceflight environment and shared via NASA GeneLab. It is now possible and indeed valuable for students to access and re-analyze these shared datasets for educational and research purposes. To provide students with experience utilizing microbial bioinformatics tools, GeneLab for Colleges and Universities (GL4U) has designed AstroAmpSeq, a week-long, virtually implemented project-based learning (PBL) minicourse to instruct undergraduate students on 16S amplicon sequencing. AstroAmpSeq was created to be accessible to students without prior bioinformatics or microbial ecology experience. During the minicourse students work in teams to process, analyze, and visualize a subsample of GeneLab dataset GLDS-280 using GeneLab’s standard amplicon processing pipeline, which is based in R. Students develop a hypothesis related to the dataset then generate and analyze figures to evaluate their hypothesis. Formative assessment of student learning is determined via pre- and post-evaluations, peer feedback, and self-reflection. Project and presentation rubrics serve as a summative assessment of student learning. GL4U AstroAmpSeq not only meets American Society for Microbiology Curriculum Guidelines, but also incites student interest in research by an inquiry-based approach and can be made part of a larger semester-long curriculum. GL4U AstroAmpSeq raises awareness of space microbiology and bioinformatics as a field and career path among undergraduates. Further, by using a GeneLab dataset and nesting microbiology techniques into the real-world application of space biology, AstroAmpSeq enforces deeper and longer-lasting student learning.

microbiology↗

Status of the TESS Science Processing Operations Center

The Transiting Exoplanet Survey Satellite (TESS) science pipeline is being developed by the Science Processing Operations Center (SPOC) at NASA Ames Research Center based on the highly successful Kepler Mission science pipeline. Like the Kepler pipeline, the TESS science pipeline will provide calibrated pixels, simple and systematic error-corrected aperture photometry, and centroid locations for all 200,000+ target stars, observed over the 2-year mission, along with associated uncertainties. The pixel and light curve products are modeled on the Kepler archive products and will be archived to the Mikulski Archive for Space Telescopes (MAST). In addition to the nominal science data, the 30-minute Full Frame Images (FFIs) simultaneously collected by TESS will also be calibrated by the SPOC and archived at MAST. The TESS pipeline will search through all light curves for evidence of transits that occur when a planet crosses the disk of its host star. The Data Validation pipeline will generate a suite of diagnostic metrics for each transit-like signature discovered, and extract planetary parameters by fitting a limb-darkened transit model to each potential planetary signature. The results of the transit search will be modeled on the Kepler transit search products (tabulated numerical results, time series products, and pdf reports) all of which will be archived to MAST.

high performance computing↗

Quantitative Comparison of Proprietary and Open-Source Georeferencing Tools for Use with Astronaut Photography

The Crew Earth Observations (CEO) Facility within the Earth Science and Remote Sensing Unit at NASA’s Johnson Space Center supports the acquisition, analysis, and curation of astronaut photography of Earth’s surface and atmosphere. Astronauts on the International Space Station (ISS) respond to requests from CEO to acquire imagery of scientific and education targets, to include high profile targets in response to activations from the International Charter for Space & Major Disasters (also known as the International Disaster Charter, or IDC) and NASA’s Disasters Program. CEO facilitates the acquisition of astronaut photography in response to IDC events and delivers georeferenced data products to the United States Geological Survey (USGS) for distribution to the disaster community. Using GeoRef, an internal web-based tool developed in collaboration with NASA’s Ames Research Center, CEO generates data packages of georeferenced imagery, uncertainty images for assessing control and tie point accuracy, and metadata documenting raw and processed data. Operational experience with the Georef software identified vulnerabilities to internal code and server errors that can significantly increase time of data production. As such, CEO developed a backup procedure in case the GeoRef software experiences front-end or back-end errors. A system using OSGEO’s open-source QGIS software combined with a semi-automated pipeline using the object-oriented Python language and the Geospatial Abstract Library for generating metadata is quantitatively compared to GeoRef’s data package for quality and productivity. Root Mean Square Error (RMSE) provides a standard measurement of data quality as it relates to ground error. Assessing RMSE measurements generated from georeferenced astronaut photographs acquired with different obliquity and focal length offers a comprehensive accuracy assessment of the software’s transformation algorithms. This assessment will indicate the software's ability to produce data products with the least ground-error or highest data quality regarding ground accuracy. In addition, a comparison of the software’s efficiency in generating a data package that includes georeferenced images, metadata, and uncertainty images for measuring tie/ground point error was performed. Initial results, based on the comparison of three nadir-facing astronaut photographs acquired with a 95mm focal length, reveal the QGIS-based system's average RMSE is 2.36 (pixels) suggesting its georectification system produces data products that meet and perhaps improve upon Georef solution's average RMSE of 32.99 (pixels). However, the QGIS system was unable to reproduce two unique Georef data products, uncertainty images for measuring tie and control point errors and a translated unwrapped image. In addition, the Georef software is designed to accept handheld camera pose information from a hardware component (Geosens) scheduled for deployment on the ISS in late 2018; this information is intended to provide increased accuracy and auto-registration capability for astronaut photographs. Future work is expected to determine the QGIS-based georectification system’s potential as an open-source alternative (and operational backup) to Georef for georeferencing the full range of resolutions and viewing angles unique to handheld digital camera imagery in support of ISS disaster response activities.

Jagge, Amy M.↗

The Use of Field Programmable Gate Arrays (FPGA) in Small Satellite Communication Systems

This paper will describe the use of digital Field Programmable Gate Arrays (FPGA) to contribute to advancing the state-of-the-art in software defined radio (SDR) transponder design for the emerging SmallSat and CubeSat industry and to provide advances for NASA as described in the TAO5 Communication and Navigation Roadmap (Ref 4). The use of software defined radios (SDR) has been around for a long time. A typical implementation of the SDR is to use a processor and write software to implement all the functions of filtering, carrier recovery, error correction, framing etc. Even with modern high speed and low power digital signal processors, high speed memories, and efficient coding, the compute intensive nature of digital filters, error correcting and other algorithms is too much for modern processors to get efficient use of the available bandwidth to the ground. By using FPGAs, these compute intensive tasks can be done in parallel, pipelined fashion and more efficiently use every clock cycle to significantly increase throughput while maintaining low power. These methods will implement digital radios with significant data rates in the X and Ka bands. Using these state-of-the-art technologies, unprecedented uplink and downlink capabilities can be achieved in a 1/2 U sized telemetry system. Additionally, modern FPGAs have embedded processing systems, such as ARM cores, integrated inside the FPGA allowing mundane tasks such as parameter commanding to occur easily and flexibly. Potential partners include other NASA centers, industry and the DOD. These assets are associated with small satellite demonstration flights, LEO and deep space applications. MSFC currently has an SDR transponder test-bed using Hardware-in-the-Loop techniques to evaluate and improve SDR technologies.

Varnavas, Kosta↗

The TESS Science Processing Operations Center

The Transiting Exoplanet Survey Satellite (TESS) will conduct a search for Earth's closest cousins starting in early 2018 and is expected to discover approximately 1,000 small planets with R(sub p) less than 4 (solar radius) and measure the masses of at least 50 of these small worlds. The Science Processing Operations Center (SPOC) is being developed at NASA Ames Research Center based on the Kepler science pipeline and will generate calibrated pixels and light curves on the NASA Advanced Supercomputing Division's Pleiades supercomputer. The SPOC will also search for periodic transit events and generate validation products for the transit-like features in the light curves. All TESS SPOC data products will be archived to the Mikulski Archive for Space Telescopes (MAST).

Jenkins, Jon M.↗

Software Displays Data on Active Regions of the Sun

The Solar Active Region Display System is a computer program that generates, in near real time, a graphical display of parameters indicative of the spatial and temporal variations of activity on the Sun. These parameters include histories and distributions of solar flares, active region growth, coronal mass ejections, size, and magnetic configuration. By presenting solar-activity data in graphical form, this program accelerates, facilitates, and partly automates what had previously been a time-consuming mental process of interpretation of solar-activity data presented in tabular and textual formats. Intended for original use in predicting space weather in order to minimize the exposure of astronauts to ionizing radiation, the program might also be useful on Earth for predicting solar-wind-induced ionospheric effects, electric currents, and potentials that could affect radio-communication systems, navigation systems, pipelines, and long electric-power lines. Raw data for the display are obtained automatically from the Space Environment Center (SEC) of the National Oceanic and Atmospheric Administration (NOAA). Other data must be obtained from the NOAA SEC by verbal communication and entered manually. The Solar Active Region Display System automatically accounts for the latitude dependence of the rate of rotation of the Sun, by use of a mathematical model that is corrected with NOAA SEC active-region position data once every 24 hours. The display includes the date, time, and an image of the Sun in H light overlaid with latitude and longitude coordinate lines, dots that mark locations of active regions identified by NOAA, identifying numbers assigned by NOAA to such regions, and solar-region visual summary (SRVS) indicators associated with some of the active regions. Each SRVS indicator is a small pie chart containing five equal sectors, each of which is color-coded to provide a semiquantitative indication of the degree of hazard posed by one aspect of the activity at the indicated location. The five aspects in question are the history of solar flares, the history of coronal mass ejections, the growth or decay of activity, the overall size, and the magnetic configuration. Mouse-clicking on an active-region-marking dot, SRVS indicator, or NOAA region number causes the program to generate a solar-region summary table (SRT) for the active region in question. The SRT contains additional quantitative and qualitative data, beyond those contained in the SRVS: These data include the solar coordinates of the region, the area of the region and its change in area during the past 24 hours, the change in the number of sunspots in the region during the past 24 hours, the magnetic configuration, and the types, dates, and times of the most recent flare and coronal mass ejection.

Golightly, Mike↗

SatCORPS Global Cloud Composite (GCC): the Design and Delivery of A High Quality, High Resolution, Global Cloud Product Available in Near-Real Time

The NASA Satellite ClOud and Radiation Property retrieval System (SatCORPS) supports the development of an analysis ready and cloud-optimized data transformation pipeline and geospatial service enablement of a global cloud composite (GCC) product derived from global geostationary satellite imagery. This geospatial service will be available at high temporal and spatial resolution via the SatCORPS web mapping application for visualization and analysis as well as direct ingestion to common geospatial software and custom programming. The resulting global cloud composite products from the processing pipeline can then be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. Near real time global observations are created through the composition of five geostationary satellites that provides modelling and forecasting communities with the capability to provide high quality and timely information to start the projection process. The Global Cloud Composite product combines information from geostationary satellites, GOES-16, GOES-17, Himawari-8, Meteosat-11 and Meteosat-9 to create a single global composite netcdf file and images using the different products within the netcdf file. The SatCORPS team, though our Global Cloud Composite (GCC) product and web-based visualization tools including Geographical Information System (GIS) services provide near real time global cloud product information to both automated processes and traditional web users that is timely and high quality derived from geostationary satellites. The Global Cloud Composite product takes advantage of the scalable processing resources provided by the AWS batch service to provide new composites every thirty minutes. Because information from each of the low earth orbiting satellites is available on schedules tuned to the specific satellite, the processing algorithm temporally composites the final dataset as each satellite’s information becomes available. The SatCORPS team has leveraged our experience using Amazon Web Services (AWS) to build a low latency high availability tool that allows end users both human and automated to acquire high quality and high-resolution Geostationary Earth Orbiting (GEO) information at zero cost to the end user. This presentation will describe how we architected and implemented the service as well as lessons learned based on our experiences both developing and operating the system. The lessons learned include how we integrated multiple services including Amazon Batch, Amazon S3 and Amazon Lambda service to create a low cost but high-performance processing system that is capable of identifying and processing the most appropriate satellite overpass information into global cloud composites. We will also describe our web-based tools including our Geographic Information System that can be used for visualization and analysis. The products from the processing can be geospatially-service enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services for visualization and analysis via a web mapping application and common geospatial software. The SatCORPS Global Composite Cloud product provides sophisticated global composited cloud research products with very low latency that we see that as filling a rapidly growing need in the research and modelling community with no up-front nor ongoing costs associated with downloading or using the information.

AWS AMCE SMCE GCC SATCORPS GLOBAL CLOUD COMPOSITE ↗

The Caltech-NRAO Stripe 82 Survey (CNSS) Paper. I. The Pilot Radio Transient Survey in 50 Deg.(exp. 2)

We have commenced a multiyear program, the Caltech-NRAO Stripe 82 Survey (CNSS), to search for radio transients with the Jansky VLA in the Sloan Digital Sky Survey Stripe 82 region. The CNSS will deliver five epochs over the entire approx. 270 deg.(exp. 2) of Stripe 82, an eventual deep combined map with an rms noise of approx. 40 proper motion epoch y and catalogs at a frequency of 3 GHz, and having a spatial resolution of 3 inches. This first paper presents the results from an initial pilot survey of a 50 deg.(exp. 2) region of Stripe 82, involving four epochs spanning logarithmic timescales between 1 week and 1.5 yr, with the combined map having a median rms noise of 35 proper motion epoch y. This pilot survey enabled the development of the hardware and software for rapid data processing, as well as transient detection and follow-up, necessary for the full 270 deg.(exp. 2) survey. Data editing, calibration, imaging, source extraction, cataloging, and transient identification were completed in a semi-automated fashion within 6 hr of completion of each epoch of observations, using dedicated computational hardware at the NRAO in Socorro and custom-developed data reduction and transient detection pipelines. Classification of variable and transient sources relied heavily on the wealth of multiwavelength legacy survey data in the Stripe 82 region, supplemented by repeated mapping of the region by the Palomar Transient Factory. A total of 3.9(+0.5%/-0.9%) of the few thousand detected point sources werefound to vary by greater than 30%, consistent with similar studies at 1.4 and 5 GHz. Multiwavelength photometric data and light curves suggest that the variability is mostly due to shock-induced flaring in the jets of active galactic nuclei (AGNs). Although this was only a pilot survey, we detected two bona fide transients, associated with an RS CVn binary and a dKe star. Comparison with existing legacy survey data (FIRST, VLA-Stripe 82) revealed additional highly variable and transient sources on timescales between 5 and 20 yr, largely associated with renewed AGN activity. The rates of such AGNs possibly imply episodes of enhanced accretion and jet activity occurring once every approx. 40,000 yr in these galaxies. We compile the revised radio transient rates and make recommendations for future transient surveys and joint radio-optical experiments.

galaxies: active – radio continuum: galaxies –↗

RadLab and the Environmental Data Application Dashboard: Graphical and Programming Interfaces for Interrogation of Space Telemetry Data

Sensors on the International Space Station (ISS) and multiple spacecraft elsewhere in Earth orbit and in deep space continuously monitor and collect environmental data, transmitting this information back to Earth. These data include ionizing radiation and, on the ISS, CO2, relative humidity levels, and temperature, and are of great importance to space biology research. Ionizing radiation in particular has been established in ground-based experiments as being correlated with increased risk of carcinogenesis and cardiovascular and neurological effects. Looking ahead to future long duration crewed missions beyond low Earth orbit, the ability to study how factors including CO2 levels, light cycle, temperature modulate the response to ionizing radiation and microgravity is essential. To date, access to these data has been fragmented across space agencies, spacecraft, and databases. To address this issue, NASA’s Open Science Data Repository (osdr.nasa.gov) has developed two Web applications: the Environmental Data Application (EDA) and a radiation-specific RadLab. Each consists of an API (application programming interface) and an associated GUI (graphical user interface) that provide single points of access to the data. To date, OSDR has focused on the sensors from payloads and radiation detectors located on the ISS. The Web applications process telemetry information and associated data, such as spacecraft location and orientation, from multiple international databases. The applications’ request syntax enables users to interrogate these data by craft, sensor type, time range, radiation type (galactic cosmic rays, solar particle events, the contribution of the South Atlantic Anomaly), facilitating arbitrary comparisons of original source data at varying time resolutions. The applications provide programmatic access for use in computational pipelines and GUIs for data visualization and exploration, making these data FAIR (Findable, Accessible, Interoperable, and Reusable), complementing the biological data contained in OSDR, and providing the space science community with a valuable resource for scientific analyses.

radiation↗

Acoustooptic linear algebra processors - Architectures, algorithms, and applications

Architectures, algorithms, and applications for systolic processors are described with attention to the realization of parallel algorithms on various optical systolic array processors. Systolic processors for matrices with special structure and matrices of general structure, and the realization of matrix-vector, matrix-matrix, and triple-matrix products and such architectures are described. Parallel algorithms for direct and indirect solutions to systems of linear algebraic equations and their implementation on optical systolic processors are detailed with attention to the pipelining and flow of data and operations. Parallel algorithms and their optical realization for LU and QR matrix decomposition are specifically detailed. These represent the fundamental operations necessary in the implementation of least squares, eigenvalue, and SVD solutions. Specific applications (e.g., the solution of partial differential equations, adaptive noise cancellation, and optimal control) are described to typify the use of matrix processors in modern advanced signal processing.

Casasent, D.↗

A Chromaticity Analysis and PSF Subtraction Techniques for SCExAO/CHARIS Data

We present an analysis of instrument performance using new observations taken with the Coronagraphic High Angular Resolution Imaging Spectrograph (CHARIS) instrument and the Subaru Coronagraphic Extreme Adaptive Optics (SCExAO) system. In a correlation analysis of our data sets (which use the broadband mode covering the J band through the K band in a single spectrum), we find that chromaticity in the SCExAO/CHARIS system is generally worse than temporal stability. We also develop a point-spread function (PSF) subtraction pipeline optimized for the CHARIS broadband mode, including a forward modeling-based exoplanet algorithmic throughput correction scheme. We then present contrast curves using this newly developed pipeline. An analogous subtraction of the same data sets using only the H-band slices yields the same final contrasts as the full JHK sequences; this result is consistent with our chromaticity analysis, illustrating that PSF subtraction using spectral differential imaging (SDI) in this broadband mode is generally not more effective than SDI in the individual J, H, or K bands. In the future, the data processing framework and analysis developed in this paper will be important to consider for additional SCExAO/CHARIS broadband observations and other ExAO instruments which plan to implement a similar integral field spectrograph broadband mode.

Benjamin L. Gerard↗

PATOKA: Simulating Electromagnetic Observables of Black Hole Accretion

The Event Horizon Telescope (EHT) has released analyses of reconstructed images of horizon-scale millimeter emission near the supermassive black hole at the center of the M87 galaxy. Parts of the analyses made use of a large library of synthetic black hole images and spectra, which were produced using numerical general relativistic magnetohydrodynamics fluid simulations and polarized ray tracing. In this article, we describe the PATOKA pipeline, which was used to generate the Illinois contribution to the EHT simulation library. We begin by describing the relevant accretion systems and radiative processes. We then describe the details of the three numerical codes we use, iharm, ipole, and igrmonty, paying particular attention to differences between the current generation of the codes and the originally published versions. Finally, we provide a brief overview of simulated data as produced by PATOKA and conclude with a discussion of limitations and future directions.

supermassive black holes↗

Real-time Unimpeded Taxi Out Machine Learning Service

This paper describes a study on the estimation of the unimpeded taxi out time using Machine Learning (ML) tools and proposes an implementation that can be used to make real-time predictions at any airport in the National Airspace System. Kedro, an open-source pipeline framework, is used to develop the model definition and training. Models are stored in scikit-learn containers on a MLFlow server where they can be retrieved and served to make predictions in the live system. These open source frameworks provide common structures between ML services, allow for easier maintenance and updates, and overall deliver an easier CI/CD (Continuous Integration/Continuous Deployment) process. The current models were trained on data acquired at KCLT and KDFW from June 1st to December 31st, 2019 and compute taxi time in the ramp, airport movement area (AMA) and total (from gates to runways). The current versions of the models achieve relatively low uncertainties of about 10 to 15% for the total and AMA taxi times and about 20% for the ramp taxi time at both KCLT and KDFW. Initial tests on offline data from 2020 and 2021 show a small degradation (10 to 15%) in accuracy performance indicating the model’s resilience to operational changes over time.

machine learning↗

Real-time Unimpeded Taxi Out Machine Learning Service

This presentation describes a study on the estimation of the unimpeded taxi out time using Machine Learning (ML) tools and proposes an implementation that can be used to make real-time predictions at any airport in the National Airspace System. Kedro, an open-source pipeline framework, is used to develop the model definition and training. Models are stored in scikit-learn containers on a MLFlow server where they can be retrieved and served to make predictions in the live system. These open source frameworks provide common structures between ML services, allow for easier maintenance and updates, and overall deliver an easier CI/CD (Continuous Integration/Continuous Deployment) process. The current models were trained on data acquired at KCLT and KDFW from June 1st to December 31st, 2019 and compute taxi time in the ramp, airport movement area (AMA) and total (from gates to runways). The current versions of the models achieve relatively low uncertainties of about 10 to 15% for the total and AMA taxi times and about 20% for the ramp taxi time at both KCLT and KDFW. Initial tests on offline data from 2020 and 2021 show a small degradation (10 to 15%) in accuracy performance indicating the model’s resilience to operational changes over time.

Machine Learning↗

2022 Spring Internship Exit Presentation

As efforts of the National Aeronautics and Space Administration (NASA) and the Federal Aviation Administration (FAA) continue to digitize the air traffic management (ATM) domain, there is countless times of need for downstream natural language processing (NLP) tasks such as named entity recognition, text summarization, classification, and more. Although there are a plethora of open-sourced pre-trained transformer models in the NLP field such as BERT, RoBERTa, XLNet, and GPT-3, these models are trained on general corpora and perform poorly on domain-specific terminology and phraseology seen in ATM documents such as Notice to Airmen (NOTAMs) and Letters of Agreement (LoA). Our proposed research objective will be to first gather a large corpus of air traffic management related documents, orders, notices, books, technical papers, conference papers, articles, and other miscellaneous sources of text data from the FAA, NASA, and accredited conference and publication societies. After gathering this data, many steps will have to be taken to collate and preprocess the data into a format understandable by our test transformer models. Thirdly, we will set up training pipelines to train the RoBERTa model on its unsupervised training task masked language modelling (MLM) using resources provided by the NASA Advanced Supercomputing (NAS) facilities. Finally, these fine-tuned transformer models will be evaluated on their performance on down-stream NLP tasks as mentioned above, to show whether they will be effective when working with ATM related data or not. Once complete, this model could be made open-sourced on the HuggingFace website, where the rest of the ATM community can access and utilize this tool.

NLP↗

Reconfigurable Hardware for Compressing Hyperspectral Image Data

High-speed, low-power, reconfigurable electronic hardware has been developed to implement ICER-3D, an algorithm for compressing hyperspectral-image data. The algorithm and parts thereof have been the topics of several NASA Tech Briefs articles, including Context Modeler for Wavelet Compression of Hyperspectral Images (NPO-43239) and ICER-3D Hyperspectral Image Compression Software (NPO-43238), which appear elsewhere in this issue of NASA Tech Briefs. As described in more detail in those articles, the algorithm includes three main subalgorithms: one for computing wavelet transforms, one for context modeling, and one for entropy encoding. For the purpose of designing the hardware, these subalgorithms are treated as modules to be implemented efficiently in field-programmable gate arrays (FPGAs). The design takes advantage of industry- standard, commercially available FPGAs. The implementation targets the Xilinx Virtex II pro architecture, which has embedded PowerPC processor cores with flexible on-chip bus architecture. It incorporates an efficient parallel and pipelined architecture to compress the three-dimensional image data. The design provides for internal buffering to minimize intensive input/output operations while making efficient use of offchip memory. The design is scalable in that the subalgorithms are implemented as independent hardware modules that can be combined in parallel to increase throughput. The on-chip processor manages the overall operation of the compression system, including execution of the top-level control functions as well as scheduling, initiating, and monitoring processes. The design prototype has been demonstrated to be capable of compressing hyperspectral data at a rate of 4.5 megasamples per second at a conservative clock frequency of 50 MHz, with a potential for substantially greater throughput at a higher clock frequency. The power consumption of the prototype is less than 6.5 W. The reconfigurability (by means of reprogramming) of the FPGAs makes it possible to effectively alter the design to some extent to satisfy different requirements without adding hardware. The implementation could be easily propagated to future FPGA generations and/or to custom application-specific integrated circuits.

Aranki, Nazeeh↗

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

Scalable Multiprocessor for High-Speed Computing in Space

A report discusses the continuing development of a scalable multiprocessor computing system for hard real-time applications aboard a spacecraft. "Hard realtime applications" signifies applications, like real-time radar signal processing, in which the data to be processed are generated at "hundreds" of pulses per second, each pulse "requiring" millions of arithmetic operations. In these applications, the digital processors must be tightly integrated with analog instrumentation (e.g., radar equipment), and data input/output must be synchronized with analog instrumentation, controlled to within fractions of a microsecond. The scalable multiprocessor is a cluster of identical commercial-off-the-shelf generic DSP (digital-signal-processing) computers plus generic interface circuits, including analog-to-digital converters, all controlled by software. The processors are computers interconnected by high-speed serial links. Performance can be increased by adding hardware modules and correspondingly modifying the software. Work is distributed among the processors in a parallel or pipeline fashion by means of a flexible master/slave control and timing scheme. Each processor operates under its own local clock; synchronization is achieved by broadcasting master time signals to all the processors, which compute offsets between the master clock and their local clocks.

Lux, James↗