Senior Inference Optimization Engineer - Dragonfly Portfolio at Dragonfly
United States
<div><strong style="color: rgb(0, 0, 0); background-color: transparent;">Dragonfly </strong><span style="color: rgb(0, 0, 0); background-color: transparent;">is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. </span>Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.</div><div><br></div><div>This is an application to join our talent network. <strong>This is not a listing for an internal role at Dragonfly.</strong></div><div><br></div><div>We're actively sourcing for a <strong>Senior Inference Optimization Engineer</strong> for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.</div><div><br></div><div><strong>Location:</strong> Remote, USA (open to excellent candidates outside the USA)</div><div><br></div><div><strong>What We’re Looking For:</strong></div><ul><li class="">5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge</li><li class="">Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume</li><li class="">Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile</li><li class="">Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments</li><li class="">GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler</li><li class="">Proficiency in Python, Rust, or Go. C++/CUDA a strong plus</li><li class="">Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions</li></ul><div><br></div><div><strong style="background-color: transparent; color: rgb(0, 0, 0);">About the role:</strong></div><ul><li class="">Stand up and optimize GPU infrastructure including B300 nodes in owned data centers</li><li class="">Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads</li><li class="">Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU</li><li class="">Optimize multivariate inference load-balancing algorithms within the inference routing system</li><li class="">Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements</li><li class="">Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack</li></ul><div><br></div><div>Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.</div><div><br></div><div><strong style="background-color: transparent; color: rgb(0, 0, 0);">Process: </strong></div><ul><li class="">We'll review your application and assess fit for this role.</li><li class="">If there's a match, we'll facilitate a warm introduction to the team.</li><li class="">If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.</li></ul><div><br></div><div><strong style="background-color: transparent; color: rgb(0, 0, 0);">Submit your information below, and we’ll reach out if there’s a potential fit.</strong></div>
Apply Now