November 2017 – Garren's [Big] Data Blog

Big data [Spark] and its small files problem

Posted by Garren on 2017/11/04

Often we log data in JSON, CSV or other text format to Amazon’s S3 as compressed files. This pattern is a) accessible and b) infinitely scalable by nature of being in S3 as common text files. However, there are some subtle but critical caveats that come with this pattern that can cause quite a bit… Continue reading→

Apache Spark Best Practices, s3, Small Files, spark 6 Comments