Using Computational Storage Devices: OpenMP/MPI and Charliecloud [Slides]
Originally used Spark and HadoopFS. Collected interesting results, but this method had its issues: limited application, too much overhead to gauge CSDs’ raw performance. Solution? Rewrite our benchmarks without Spark using Serial Python and Serial & Parallel C++ (Combinations of OpenMP & OpenMPI). Compared to Spark and Python, C++ implementation is a lot faster. Caveat: an expert with Spark or Python would likely be able to improve the performance of those implementations. Computational power of our CSDs seem to be much lower than the host machine. Using all 4 cores of a single CSD, the job takes ~6.8x longer than using just one core on the host machine. Host also seems to scale better with increasing file size. Resulting Question: When, if ever, would it make sense to use CSDs for compute rather than a much-faster host?