Member of Technical Staff, Distributed Systems at SkyPilot
United States
<h2><span style="background-color: transparent; color: rgb(0, 0, 0);">About SkyPilot</span></h2><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."</span></div><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot (</span><a href="https://github.com/skypilot-org/skypilot" rel="noopener noreferrer" target="_blank" style="background-color: transparent; color: rgb(17, 85, 204);">10k+ GitHub stars</a><span style="background-color: transparent; color: rgb(0, 0, 0);">, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.</span></div><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">The role</span></h2><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot's core is a distributed control plane that runs AI workloads across 20+ clouds, Kubernetes, Slurm, and on-prem — deciding where compute should live, keeping jobs and clusters consistent across unreliable infrastructure, and recovering automatically when instances disappear. We're looking for an engineer to own this core end to end, set its technical direction and build new features to make it more robust, performant, and user-friendly. The core you own is what lets SkyPilot run the largest AI workloads in the world.</span></div><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What you'll do</span></h2><div><br></div><ul><li class=""><strong style="background-color: transparent;">Design and build the future of AI systems: </strong><span style="background-color: transparent;">you will solve some of the hardest problems in distributed AI systems to make SkyPilot the standard solution running frontier AI workloads.</span></li><li class=""><strong style="background-color: transparent;">Shape the technical direction of SkyPilot</strong><span style="background-color: transparent;">: the modules, interfaces, and invariants the rest of engineering builds on — and raise the bar for how we design and operate the system.</span></li><li class=""><strong style="background-color: transparent;">Build enhancements and new components</strong><span style="background-color: transparent;"> to evolve SkyPilot with better support of a wide range of AI and batch workloads.</span></li><li class=""><strong style="background-color: transparent;">Engage with users:</strong><span style="background-color: transparent;"> Opportunity to work closely with our users and customers to make their use cases successful; to grow our open-source community; to gain visibility for your work via public tutorials, blog posts, and/or talks.</span></li></ul><h2><br></h2><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What we're looking for</span></h2><div><br></div><ul><li class=""><span style="background-color: transparent;">You've designed, built, and operated distributed systems at production scale.</span></li><li class=""><span style="background-color: transparent;">Deep systems fundamentals: scheduling, state machines, coordination, and fault tolerance.</span></li><li class=""><span style="background-color: transparent;">Python/Go fluency and strong reasoning skills about concurrent, distributed code.</span></li><li class=""><span style="background-color: transparent;">You reach for the simplest design that survives contact with real workloads — and can explain why.</span></li><li class=""><span style="background-color: transparent;">Experience with cloud provider control planes (AWS / GCP / Azure), Kubernetes internals, or operating across multiple clouds, and open-source or AI infrastructure contributions (e.g. KAI, KubeRay, Kueue, KServe).</span></li></ul><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What we offer</span></h2><div><br></div><ul><li class=""><span style="background-color: transparent;">Competitive compensation and equity</span></li><li class=""><span style="background-color: transparent;">Comprehensive medical, dental, vision coverage for you and your dependents</span></li><li class=""><span style="background-color: transparent;">The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.</span></li><li class=""><span style="background-color: transparent;">A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).</span></li><li class=""><span style="background-color: transparent; color: rgb(29, 28, 29);">Gourmet lunch & dinner for the team to do their best work</span></li></ul><div><br></div><div><strong style="background-color: transparent; color: rgb(0, 0, 0);">Location:</strong><span style="background-color: transparent; color: rgb(0, 0, 0);"> San Mateo, CA. Remote will be considered for exceptional candidates.</span></div>
Apply Now