Talent.com
Senior ML Storage Infrastructure Engineer

Senior ML Storage Infrastructure Engineer

ZooxFoster City, CA, United States
1 day ago
Job type
  • Full-time
Job description

Zoox is looking for a software engineer to work on our custom High-Performance Computing infrastructure and its supporting ecosystem of tools and services. This infrastructure is central to machine learning workflows across all Zoox software divisions, from data engineering to computer vision perception to simulation and more. You will take on a breadth of end-to-end responsibilities including distributed system design, algorithmic job scheduling, and adaptive cloud scaling in support of all of Zoox's computational needs.

In this role, you will :

  • Design, build, and optimize a petabyte-scale, in-house HPC storage infrastructure, ensuring high performance and reliability for our machine learning workloads across both cloud and on-premise data centers.
  • Drive GPU efficiency by strategically collocating storage and compute, architecting a storage layer that keeps tens of thousands of GPUs fully utilized and prevents bottlenecks.
  • Drive key initiatives in training and storage optimization by partnering with ML practitioners, applying your deep understanding of frameworks such as PyTorch and TensorFlow to meet their evolving demands.
  • Investigate and adopt new distributed system paradigms and cutting-edge technologies to ensure our infrastructure can scale to meet ever-growing computational and storage demands.
  • Create production-grade web service APIs, SDKs, and other essential tools to deliver a world-class developer experience for all software teams at Zoox.

Qualifications :

  • Experience designing and building high-performance, distributed storage systems (object / file) for large-scale, GPU-bound workloads.
  • Proficiency in Python, Java, or similar languages for developing data-intensive, high-performance applications.
  • Hands-on experience with cloud platforms (AWS, GCP, Azure), using their storage, GPU, and observability services to provide usage showback for ML practitioners.
  • Bachelor's degree in Computer Science or a related field with a strong foundation in data structures and systems design.
  • Bonus Qualification :

  • Experience with parallel filesystems (e.g., Lustre, FSx) and their integration with container orchestrators via Kubernetes CSI drivers.
  • Deep knowledge of ML frameworks like PyTorch and TensorFlow, and workload schedulers such as SLURM or Kubernetes.
  • Familiarity with emerging AI paradigms, including agentic systems, and observability tools like OpenTelemetry.
  • $192,000 - $300,000 a year

    Base Salary Range

    There are three major components to compensation for this position : salary, Amazon Restricted Stock Units (RSUs), and Zoox Stock Appreciation Rights. A sign-on bonus may be offered as part of the compensation package. The listed range applies only to the base salary. Compensation will vary based on geographic location and level. Leveling, as well as positioning within a level, is determined by a range of factors, including, but not limited to, a candidate's relevant years of experience, domain knowledge, and interview performance. The salary range listed in this posting is representative of the range of levels Zoox is considering for this position.

    Zoox also offers a comprehensive package of benefits, including paid time off (e.g. sick leave, vacation, bereavement), unpaid time off, Zoox Stock Appreciation Rights, Amazon RSUs, health insurance, long-term care insurance, long-term and short-term disability insurance, and life insurance.

    About Zoox

    Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We're looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.

    Follow us on LinkedIn

    Accommodations

    If you need an accommodation to participate in the application or interview process please reach out to [email protected] or your assigned recruiter.

    A Final Note :

    You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.

    We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

    Create a job alert for this search

    Senior Engineer Infrastructure • Foster City, CA, United States

    Related jobs
    • Promoted
    AI / ML Infrastructure Engineer

    AI / ML Infrastructure Engineer

    RIT Solutions, Inc.Concord, CA, United States
    Full-time
    Title : AI / ML Infrastructure Engineer, 3 days onsite, locals only.Grant St Concord California 94520 United States.Lead and design the platform and infrastructure architecture for AIML and NLP in mod...Show moreLast updated: 30+ days ago
    • Promoted
    ML Infrastructure Engineer in Oakland

    ML Infrastructure Engineer in Oakland

    Energy Jobline ZROakland, CA, United States
    Full-time
    Energy Jobline is the largest and fastest growing global Energy Job Board and Energy Hub.We have an audience reach of over 7 million energy professionals, 400,000+ monthly advertised global energy ...Show moreLast updated: 1 day ago
    • Promoted
    Senior Infrastructure Engineer

    Senior Infrastructure Engineer

    CARETSan Francisco, CA, United States
    Full-time
    Nationwide Remote - Remote, CA.Support and document IT infrastructure systems, ensuring stability and performance.Comprehend and manage system architecture involving servers, databases, APIs, load ...Show moreLast updated: 1 day ago
    • Promoted
    AI / ML Infrastructure Engineer

    AI / ML Infrastructure Engineer

    Syntricate TechnologiesConcord, CA, United States
    Full-time
    Grant St Concord California 94520 (3 days onsite in week).Lead and design the platform and infrastructure architecture for AIML and NLP in modern hybrid cloud computing. Participate in day-to-day st...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Software Engineer - ML Infrastructure

    Senior Software Engineer - ML Infrastructure

    PlaidSan Francisco, CA, United States
    Full-time
    Plaid is evolving into an AI-first company, where data and machine learning are the key enablers of smarter, more secure insight products built on top of Plaid's vast financial data network.The Mac...Show moreLast updated: 30+ days ago
    • Promoted
    ML Infrastructure Engineer in Menlo Park

    ML Infrastructure Engineer in Menlo Park

    Energy Jobline ZRMenlo Park, CA, United States
    Full-time +1
    Energy Jobline is the largest and fastest growing global Energy Job Board and Energy Hub.We have an audience reach of over 7 million energy professionals, 400,000+ monthly advertised global energy ...Show moreLast updated: 1 day ago
    • Promoted
    Senior Infrastructure Engineer

    Senior Infrastructure Engineer

    PumpSan Francisco, CA, United States
    Full-time
    Cloud spend is a whopping $500 billion / yr, the biggest growing expense category for any tech company - tackling these costs requires continuous effort and time from DevOps teams.Pump is a building ...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Software Engineer - Infrastructure Storage

    Senior Software Engineer - Infrastructure Storage

    LambdaSan Francisco, CA, United States
    Full-time
    Lambda, The Superintelligence Cloud, builds Gigawatt-scale AI Factories for Training and Inference.Lambda's mission is to make compute as ubiquitous as electricity and give every person access to a...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Software Engineer, ML Infrastructure

    Senior Software Engineer, ML Infrastructure

    LMArenaSan Francisco, CA, United States
    Full-time
    Senior Software Engineer, ML Infrastructure.Senior Software Engineer (Infrastructure).In this role, you'll architect systems that capture and process large volumes of serving requests in real time,...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Infrastructure Engineer

    Senior Infrastructure Engineer

    AngelListSan Francisco, CA, United States
    Full-time
    We exist to accelerate innovation.We do this by giving more people the opportunity to participate in the venture economy by building the financial infrastructure that makes it possible for more peo...Show moreLast updated: 1 day ago
    • Promoted
    AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - ML Compute

    AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - ML Compute

    AppleSan Francisco, CA, United States
    Full-time
    Apple is where individual imaginations gather together, committing to the values that lead to great work.Every new product we build, service we create, or Apple Store experience we deliver is the r...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Cloud Infrastructure Engineer

    Senior Cloud Infrastructure Engineer

    Omni Analytics, Inc.San Francisco, CA, United States
    Full-time
    Omni gives businesses one place to easily analyze all their data.Built by the teams behind Looker and Stitch, Omni combines data models, a point-and-click UI, spreadsheet formulas, and powerful vis...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Software Engineer - ML Infrastructure in San Francisco

    Senior Software Engineer - ML Infrastructure in San Francisco

    Energy Jobline ZRSan Francisco, CA, United States
    Full-time
    Energy Jobline is the largest and fastest growing global Energy Job Board and Energy Hub.We have an audience reach of over 7 million energy professionals, 400,000+ monthly advertised global energy ...Show moreLast updated: 1 day ago
    • Promoted
    • New!
    Senior Kubernetes & Infrastructure Engineer

    Senior Kubernetes & Infrastructure Engineer

    Third Wave AutomationUnion City, CA, United States
    Full-time
    Third Wave Automation is a rapidly growing startup that has demonstrated its core technology components, proven its market fit, and just closed its Series C funding. If you are excited about cutting...Show moreLast updated: 3 hours ago
    • Promoted
    AIML - Core Infrastructure Engineering, Core Infrastructure

    AIML - Core Infrastructure Engineering, Core Infrastructure

    AppleSan Francisco, CA, United States
    Full-time
    Do you want to make Apple products smarter for our users? The AIML Core Infra team is looking for an experienced software engineer to work on core infrastructure for information intelligence at App...Show moreLast updated: 1 day ago
    • Promoted
    MTS, Infrastructure Engineer

    MTS, Infrastructure Engineer

    DelphinaSan Francisco, CA, United States
    Full-time
    Today's Data Scientists are in pain - spending their time manually wrangling data, building models through slow trial and error, taking on painstaking rewrites for deployment, and dealing with coun...Show moreLast updated: 1 day ago
    • Promoted
    ML Infrastructure Engineer

    ML Infrastructure Engineer

    PhizenixMenlo Park, CA, United States
    Full-time +1
    Menlo Park, CA | On-Site | Full-Time / Direct Hire.Looking for ML Infra experts (Bay Area preferred) with deep experience in CUDA, GPU optimization, VLLMs, and LLM inference-pure language focus, no v...Show moreLast updated: 30+ days ago
    • Promoted
    Senior ML infrastructure engineer

    Senior ML infrastructure engineer

    KuzcoSan Francisco, CA, United States
    Full-time
    Kuzco is seeking a Senior ML Infrastructure Engineer to join our team.This role involves developing large-scale, fault-tolerant systems that handle millions of large language model inference reques...Show moreLast updated: 1 day ago