Talent.com
Databricks
Software Engineer - GenAI inferenceDatabricks • San Francisco, California
No longer accepting applications
Software Engineer - GenAI inference

Software Engineer - GenAI inference

Databricks • San Francisco, California
30+ days ago
Salary
$204,600.00 yearly
Job type
  • Full-time
Job description

P-1284

About This Role

As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.

What You Will Do

  • Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
  • Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
  • Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
  • Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
  • Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
  • Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
  • Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
  • Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams
  • Document and share learnings, contributing to internal best practices and open-source efforts when possible

What We Look For

  • BS/MS/PhD in Computer Science, or a related field
  • Strong software engineering background (3+ years or equivalent) in performance-critical systems
  • Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.
  • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
  • Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
  • Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
  • Experience building instrumentation, tracing, and profiling tools for ML models
  • Ability to work closely with ML researchers, translate novel model ideas into production systems
  • Ownership mindset and eagerness to dive deep into complex system challenges
  • Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving

Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page .

Local Pay Range$142,200—$204,600 USD
Create a job alert for this search

Software Engineer - GenAI inference • San Francisco, California

Similar jobs

Senior AI Software Engineer - Remote, High-Impact

Turquoise HealthSan Francisco, CA, United States
Remote
Full-time

A healthcare technology startup is seeking a Senior Software Engineer to join their remote AI team.This role involves designing and operating customer-facing features, building internal services, a... Show more

 • Promoted

Software Engineer, Discovery (Feed)

WhatnotSan Francisco, CA, United States
Full-time

Join the Future of Commerce with Whatnot!Whatnot is the largest livestream shopping platform in North America and Europe to buy, sell, and discover the things you love.Whether it's trading cards, f... Show more

 • Promoted

Staff Software Engineer, Android

PoshmarkRedwood City, CA, United States
Full-time

Staff Engineer, AndroidPoshmark is the leading fashion marketplace where style comes alive through discovery, self-expression, and human connection.Powered by a vibrant community of 165 million mem... Show more

 • Promoted

Software Engineer - Model APIs

BasetenSan Francisco, CA, United States
Full-time

Baseten Model Performance EngineerBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting ... Show more

 • Promoted

Software Engineer - Platform / Applied AI (Full-Stack)

Career LaunchSan Francisco, CA, USA
Remote
Full-time
Quick Apply

Career Launch is hiring for this role and related opportunities.We review candidate profiles across active roles and use resumes/profiles to understand which backgrounds may be a fit.As a Software ... Show more

Software Engineering Inference Engineer

Virtue AISan Francisco, CA, United States
Full-time

Inference EngineerVirtue AI sets the standard for advanced AI security platforms.Built on decades of foundational and award-winning research in AI security, its AI-native architecture unifies autom... Show more

 • Promoted

Software Engineer, AI Infrastructure

Fireworks AISan Mateo, CA, United States
Full-time

Software Engineer, Ai InfrastructureAt Fireworks, we're building the future of generative AI infrastructure.Our platform delivers the highest-quality models with the fastest and most scalable infer... Show more

 • Promoted

Software Engineer III/Senior, AI Gateway

ngrokSan Francisco, CA, United States
Full-time +1

Software Engineer III/Senior, AI Gatewayngrok is an all-in-one cloud networking platform that secures, transforms, and routes traffic to services running anywhere.Instead of cobbling together nginx... Show more

 • Promoted

Senior Software Engineer, AI

CentariSan Francisco, CA, United States
Full-time

Senior Software Engineer, AiWe are hiring a Senior Software Engineer, AI to join our product team.You'll be a key player in designing Centari's core AI technology, working alongside engineers who s... Show more

 • Promoted

Software Engineer, Growth

AnthropicSan Francisco, CA, United States
Full-time

Software Engineer, GrowthAnthropic's mission is to create reliable, interpretable, and steerable AI systems.We want AI to be safe and beneficial for our users and for society as a whole.Our team is... Show more

 • Promoted

Senior Software Engineer, Applied AI

TestBoxSan Francisco, CA, United States
Full-time

Zcaron; Onsite ” San Francisco ’ $180,000 USD Equity & BenefitsAbout TestBoxTestBox was founded with a bold mission:to fundamentally transform how ... Show more

 • Promoted

Software Engineer, API Engineer

OpenAISan Francisco, CA, United States
Full-time

Software EngineerOur team brings OpenAI's most capable technology to the world through our developer platform:the OpenAI API.As the leading AI development platform, our API is used by millions of d... Show more

 • Promoted

Senior Software Engineer - AI Core Engineering

Disney FranceSan Francisco, CA, United States
Full-time

Senior Software Engineer - AI Core EngineeringWe're hiring a Senior AI Engineer to build the AI core capabilities and tooling that accelerate teams across Ad Technology.You will create shared agent... Show more

 • Promoted

Software Engineer, AI & Developer Acceleration

Cartesia, Inc.San Francisco, CA, United States
Full-time

Cartesia AI-Native Software Engineer, AI & Developer AccelerationCartesia is hiring an AI-native Software Engineer, AI & Developer Acceleration to optimize developer experience and maximize... Show more

 • Promoted

Software Engineer, AI Platform

Perplexity AISan Francisco, CA, United States
Full-time

Software EngineerPerplexity is seeking an experienced Software Engineer focusing on building the next-gen AI Foundation & Platform to help revolutionize the way people search and interact onlin... Show more

 • Promoted

Senior Software Engineer, AI

UnifySan Francisco, CA, United States
Full-time

AI EngineerUnify is building the first AI-powered system of action for revenue teams, helping companies transform outbound into a top-performing growth engine by making go-to-market execution obser... Show more

 • Promoted

Software Engineer, Identity

TwilioSan Francisco, CA, US
Full-time

At Twilio, we're shaping the future of communications, all from the comfort of our homes.We deliver innovative solutions tohundreds of thousands of businessesand empower millions of developers worl... Show more

 • Promoted

Sofware Engineer

TradeJobsWorkForce94706 Albany, CA, US
Full-time

Analyze, design and develop tests and test-automation suites.Design, create and develop a processing platform using various configuration management technologies.Test software development methodolo... Show more

 • Promoted

Quantum Error Correction Software Engineer

Atom ComputingBerkeley, CA, United States
Full-time

Software EngineerAt Atom Computing, we build quantum computers using arrays of optically trapped neutral atoms that will empower customers to achieve unprecedented computational breakthroughs.Join ... Show more

 • Promoted

Senior Software Engineer - AI Platform (Remote, Equity)

LiberateSan Francisco, CA, United States
Remote
Full-time

A forward-thinking AI company is seeking a Senior Software Engineer to join their dynamic team in San Francisco.You will design and develop high-performance applications, collaborate with talented ... Show more