Talent.com
Hippocratic-Ai
LLM Inference EngineerHippocratic-Ai • Palo Alto, CA, United States
LLM Inference Engineer

LLM Inference Engineer

Hippocratic-Ai • Palo Alto, CA, United States
10 days ago
Job type
  • Full-time
Job description

About Us Hippocratic AI is the leading generative AI company in healthcare. We have the only system that can have safe, autonomous, clinical conversations with patients. We have trained our own LLMs as part of our Polaris constellation, resulting in a system with over 99.9% accuracy. Why Join Our Team Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation. Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St.Louis, Stanford, Google, Meta, Microsoft, and NVIDIA. Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others. Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative. Location Requirement We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Palo Alto office five days a week, unless otherwise specified. About the Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving infrastructure. The ideal candidate has: Extensive hands‑on experience with state‑of‑the‑art inference optimization techniques A track record of deploying efficient, scalable LLM systems in production environments What You'll Do Design and implement multi-node serving architectures for distributed LLM inference Optimize multi-LoRA serving systems Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality Implement speculative decoding and other latency optimization strategies Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases Continuously benchmark and improve system performance across various deployment scenarios and GPU types What You Bring Must-Have: Experience optimizing LLM inference systems at scale Proven expertise with distributed serving architectures for large language models Hands‑on experience implementing quantization techniques for transformer models Strong understanding of modern inference optimization methods, including: Speculative decoding techniques with draft models Eagle speculative decoding approaches Proficiency in Python and C++ Experience with CUDA programming and GPU optimization Nice-to-Have: Contributions to open‑source inference frameworks such as vLLM, SGLang, or TensorRT‑LLM Experience with custom CUDA kernels Track record of deploying inference systems in production environments Deep understanding of performance optimization systems Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters - we want to hear your story. Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value. Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale! #J-18808-Ljbffr

Create a job alert for this search

LLM Inference Engineer • Palo Alto, CA, United States

Similar jobs

Remote Lead ML & 3D Vision R&D Engineer

CoStarSunnyvale, CA, United States
Full-time

A leading real estate technology company is looking for a Lead Machine Learning R&D Engineer to innovate in spatial computing and advance their platform.This role, based in California, focuses on d... Show more

 • Promoted

Foundational ML Researcher / Engineer -- Remote, Equity

Pathway Genomics CorporationPalo Alto, CA, United States
Remote
Full-time

A cutting-edge AI startup is seeking R&D Engineers for groundbreaking work in attention-based machine learning models.Candidates should have a strong research background and experience in model... Show more

 • Promoted

ML Engineer for AI Product Tools -- Hybrid/Remote

AndiamoPalo Alto, CA, United States
Remote
Full-time

A leading technology staffing firm seeks a Machine Learning Engineer to work on real products rather than isolated research.This role emphasizes exploring AI techniques to develop intelligent featu... Show more

 • Promoted

Senior Search Engineer — ML-Driven Retrieval & Relevance

WorkatoPalo Alto, CA, United States
Full-time

A leading technology company is seeking a Senior Software Engineer specializing in Search/Retrieval.The candidate will lead the design and optimization of intelligent search systems leveraging mach... Show more

 • Promoted

MLOps and ML Engineer

APTASKSan Ramon, CA, United States
Full-time

Job TitleThe client is a leading digital transformation consultancy and engineering services company that specializes in driving digital transformation for Fortune 1000 enterprises.The company prov... Show more

 • Promoted

Director of Generative AI & LLM Research

Advanced Micro Devices, Inc.San Jose, CA, United States
Full-time

Director of AI (Generative AI) to lead their Modeling team focused on Large Language Models (LLMs).The ideal candidate will have substantial leadership experience in machine learning, particularly ... Show more

 • Promoted

Nuclear Engineer

US NavyBen Lomond, CA, US
Full-time

Nuclear Engineer (Naval Reactors Engineer).Design, regulate, and oversee the Navy’s nuclear propulsion program, including reactor design, fleet operations, and eventual defueling and decommissionin... Show more

 • Promoted

Secure AI Backend Engineer (LLM & Microservices)

Fortinet, Inc.Sunnyvale, CA, United States
Full-time

A leading cybersecurity firm is seeking a candidate to enhance LLM security by architecting monitoring and filtering systems.This role requires expertise in deploying AI systems, managing prompts, ... Show more

 • Promoted

Security Account Manager - Biopharmaceutical

Allied UniversalAptos, CA, United States
Full-time

Company Overview: Allied Universal®, North America's leading security and facility services company, offers rewarding careers that provide you a sense of purpose.While working in a dynamic, welcom... Show more

 • Promoted

Senior RDMA Networking Engineer for AI Inference Pods

EtchedSan Jose, CA, United States
Full-time

A leading AI infrastructure company based in San Jose is seeking a highly skilled Supercomputing Engineer specialized in networking.This role involves developing high-performance networking solutio... Show more

 • Promoted

Security Engineer - Data Loss Prevention (DLP)

CADENCE INCSan Jose, CA, United States
Full-time

Cadence Security EngineerAt Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.Summary:A highly skilled and experienced Security Engineer with... Show more

 • Promoted

Sr Manager - Teamcenter PLM Architect

Lam ResearchFremont, CA, United States
Full-time

Join Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure.As a crucial member of... Show more

 • Promoted

ML Engineer: LLM Research & Production

Apple Inc.Cupertino, CA, United States
Full-time

A leading technology company in Cupertino seeks a Research Scientist in AI to conduct groundbreaking research in deep learning and large language models.The ideal candidate will develop and fine-tu... Show more

 • Promoted

Senior ML Engineer — Scalable Personalization Pipelines

Adobe Inc.San Jose, CA, United States
Full-time

An innovative firm is seeking a Senior Machine Learning Engineer to lead the development of cutting-edge machine learning models that enhance personalized customer experiences.This role offers the ... Show more

 • Promoted

Principal AI Agent & ML Systems Engineer

OracleSanta Clara, CA, United States
Full-time

Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure.The... Show more

 • Promoted

Director of Engineering

Ensemble Investments, LLCSanta Cruz, CA, United States
Full-time

Nestled along the Pacific Coast, La Bahia Hotel & Spa celebrates its dramatic setting where the tip of Monterey Bay touches Sana Cruz's coveted Main Beach.Steeped in the romantic beauty of Spanish-... Show more

 • Promoted

Senior ML Infra Engineer - Training Efficiency

WaymoMountain View, CA, United States
Full-time

A leading autonomous driving technology company in Mountain View is seeking an experienced professional to enhance ML infrastructure for training workloads.Responsibilities include designing distri... Show more

 • Promoted

Member of Technical Staff, LLM Inference - MAI Superintelligence Team

Microsoft CorporationMountain View, CA, United States
Full-time

OverviewOur Inference team is responsible for building and maintaining the tools and systems that enable Microsoft AI researchers to run models easily and efficiently.Our work empowers researchers ... Show more

 • Promoted

Flexible remote AI work. Your schedule. Paid weekly, straight to your bank account.

Meridian.aiScotts Valley, CA, US
Remote
Full-time

Review and label digital content including text, images, and documents.Every task you complete helps improve how technology interprets information and performs in practical settings.Detail-oriented... Show more

 • Promoted

Project Engineer

ACCO Engineered SystemsSanta Cruz, CA, United States
Full-time

This position will collaborate closely with a dedicated Project Manager, reporting to the Sales Manager, and be responsible for the technical coordination of assigned projects.This includes fosteri... Show more