Member of Technical Staff, GPU / ML Systems at SkyPilot
United States
<h2><span style="background-color: transparent; color: rgb(0, 0, 0);">About SkyPilot</span></h2><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."</span></div><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot (</span><a href="https://github.com/skypilot-org/skypilot" rel="noopener noreferrer" target="_blank" style="background-color: transparent; color: rgb(17, 85, 204);">10k+ GitHub stars</a><span style="background-color: transparent; color: rgb(0, 0, 0);">, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.</span></div><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">The role</span></h2><div><br></div><div><span style="background-color: transparent; color: rgb(0, 0, 0);">SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.</span></div><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What you'll do</span></h2><div><br></div><ul><li class=""><strong style="background-color: transparent;">Own GPU scheduling, utilization and health</strong><span style="background-color: transparent;">: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.</span></li><li class=""><strong style="background-color: transparent;">Build optimizations for training and serving</strong><span style="background-color: transparent;">: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.</span></li><li class=""><strong style="background-color: transparent;">Make the AI stack run great out of the box</strong><span style="background-color: transparent;">: deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.</span></li></ul><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What we're looking for</span></h2><div><br></div><ul><li class=""><span style="background-color: transparent;">Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.</span></li><li class=""><strong style="background-color: transparent;">Strongly preferred: </strong><span style="background-color: transparent;">Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).</span></li><li class=""><span style="background-color: transparent;">You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.</span></li><li class=""><span style="background-color: transparent;">Strong Python, and comfort reaching into systems-level and GPU-adjacent details.</span></li><li class=""><span style="background-color: transparent;">You care about squeezing most from the compute available to you</span></li><li class=""><span style="background-color: transparent;">Experience operating large-scale training or high-throughput inference in production</span></li></ul><div><br></div><h2><span style="background-color: transparent; color: rgb(0, 0, 0);">What we offer</span></h2><div><br></div><ul><li class=""><span style="background-color: transparent;">Competitive compensation and equity</span></li><li class=""><span style="background-color: transparent;">Comprehensive medical, dental, vision coverage for you and your dependents</span></li><li class=""><span style="background-color: transparent;">The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.</span></li><li class=""><span style="background-color: transparent;">A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).</span></li><li class=""><span style="background-color: transparent; color: rgb(29, 28, 29);">Gourmet lunch & dinner for the team to do their best work</span></li></ul><div><br></div><div><strong style="background-color: transparent; color: rgb(0, 0, 0);">Location:</strong><span style="background-color: transparent; color: rgb(0, 0, 0);"> San Mateo, CA. Remote will be considered for exceptional candidates.</span></div>
Apply Now