One paper that sparked my interest in the big data domain was about remote shuffle service ⚡
I always considered Apache Spark as a black box to process massive amounts of data, but when I stumbled upon this paper I realized how complicated and engineering-heavy the entire domain is demanding you to re-imagine the approach at every single step.
If you are interested in the internals of big data processing (and not how to use it), I would highly recommend reading “Magnet: A scalable and performant shuffle architecture for Apache Spark” by LinkedIn.
I was constantly in awe while reading this paper and it blew my mind at every single line. You can find this and other papers I recommend all engineers to read on my papershelf.
arpitbhayani.me/papershelf
⚡ I keep writing and sharing these engineering nuggets, so if you are keen on learning them, follow along.
duggup.com/arpit