deployed_

Senior AI Inference Engineer - Model Optimization & Deployment

Zoox - Foster City, CA

Still listed by Zoox · checked 2 days ago · posted 11 Apr 2026 (6 months ago)Open 6+ months

OnsiteSeniorFull-timeEngineeringML / AIUSD 225K - 305K / year-salary

Skills: Cuda, Llms, Vlms, Foundation models, Model compression, Model optimization, Inference, Edge deployment, Concurrent programming, Machine learning

The Perception team is pioneering the development of a multi-modality foundation model to drive the next generation of autonomous system intelligence.

As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.

Related live jobs