Talent.com
Hippocratic-Ai
LLM Inference EngineerHippocratic-Ai • Palo Alto, CA, United States
LLM Inference Engineer

LLM Inference Engineer

Hippocratic-Ai • Palo Alto, CA, United States
Hace 10 días
Tipo de contrato
  • A tiempo completo
Descripción del trabajo

About Us Hippocratic AI is the leading generative AI company in healthcare. We have the only system that can have safe, autonomous, clinical conversations with patients. We have trained our own LLMs as part of our Polaris constellation, resulting in a system with over 99.9% accuracy. Why Join Our Team Reinvent healthcare with AI that puts safety first. We’re building the world’s first healthcare‑only, safety‑focused LLM — a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation. Work with the people shaping the future. Hippocratic AI was co‑founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St.Louis, Stanford, Google, Meta, Microsoft, and NVIDIA. Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children’s, WellSpan Health, John Doerr, Rick Klausner, and others. Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies — ensuring our platform is powerful, trusted, and truly transformative. Location Requirement We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Palo Alto office five days a week, unless otherwise specified. About the Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving infrastructure. The ideal candidate has: Extensive hands‑on experience with state‑of‑the‑art inference optimization techniques A track record of deploying efficient, scalable LLM systems in production environments What You'll Do Design and implement multi-node serving architectures for distributed LLM inference Optimize multi-LoRA serving systems Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality Implement speculative decoding and other latency optimization strategies Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases Continuously benchmark and improve system performance across various deployment scenarios and GPU types What You Bring Must-Have: Experience optimizing LLM inference systems at scale Proven expertise with distributed serving architectures for large language models Hands‑on experience implementing quantization techniques for transformer models Strong understanding of modern inference optimization methods, including: Speculative decoding techniques with draft models Eagle speculative decoding approaches Proficiency in Python and C++ Experience with CUDA programming and GPU optimization Nice-to-Have: Contributions to open‑source inference frameworks such as vLLM, SGLang, or TensorRT‑LLM Experience with custom CUDA kernels Track record of deploying inference systems in production environments Deep understanding of performance optimization systems Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters - we want to hear your story. Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value. Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale! #J-18808-Ljbffr

Crear una alerta de empleo para esta búsqueda

LLM Inference Engineer • Palo Alto, CA, United States

Ofertas similares

Remote Lead ML & 3D Vision R&D Engineer

CoStarSunnyvale, CA, United States
A tiempo completo

A leading real estate technology company is looking for a Lead Machine Learning R&D Engineer to innovate in spatial computing and advance their platform.This role, based in California, focuses on d... Mostrar más

 • Oferta promocionada

ML Engineer for AI Product Tools -- Hybrid/Remote

AndiamoPalo Alto, CA, United States
Teletrabajo
A tiempo completo

A leading technology staffing firm seeks a Machine Learning Engineer to work on real products rather than isolated research.This role emphasizes exploring AI techniques to develop intelligent featu... Mostrar más

 • Oferta promocionada

Senior Search Engineer — ML-Driven Retrieval & Relevance

WorkatoPalo Alto, CA, United States
A tiempo completo

A leading technology company is seeking a Senior Software Engineer specializing in Search/Retrieval.The candidate will lead the design and optimization of intelligent search systems leveraging mach... Mostrar más

 • Oferta promocionada

MLOps and ML Engineer

APTASKSan Ramon, CA, United States
A tiempo completo

Job TitleThe client is a leading digital transformation consultancy and engineering services company that specializes in driving digital transformation for Fortune 1000 enterprises.The company prov... Mostrar más

 • Oferta promocionada

Senior Insurance Loss Control Consultant

Alexander & Schmidt - Getting the Job Done as a Trusted and Reliable PartnerSanta Cruz, CA, United States
A tiempo completo

Senior Insurance Loss Control Consultant.At Alexander & Schmidt, a Senior Insurance Loss Control Consultant performs inspections and prepares in-depth reports for insurance underwriting purposes.In... Mostrar más

 • Oferta promocionada

Director of Generative AI & LLM Research

Advanced Micro Devices, Inc.San Jose, CA, United States
A tiempo completo

Director of AI (Generative AI) to lead their Modeling team focused on Large Language Models (LLMs).The ideal candidate will have substantial leadership experience in machine learning, particularly ... Mostrar más

 • Oferta promocionada

Nuclear Engineer

US NavyBen Lomond, CA, US
A tiempo completo

Nuclear Engineer (Naval Reactors Engineer).Design, regulate, and oversee the Navy’s nuclear propulsion program, including reactor design, fleet operations, and eventual defueling and decommissionin... Mostrar más

 • Oferta promocionada

Secure AI Backend Engineer (LLM & Microservices)

Fortinet, Inc.Sunnyvale, CA, United States
A tiempo completo

A leading cybersecurity firm is seeking a candidate to enhance LLM security by architecting monitoring and filtering systems.This role requires expertise in deploying AI systems, managing prompts, ... Mostrar más

 • Oferta promocionada

Security Account Manager - Biopharmaceutical

Allied UniversalAptos, CA, United States
A tiempo completo

Company Overview: Allied Universal®, North America's leading security and facility services company, offers rewarding careers that provide you a sense of purpose.While working in a dynamic, welcom... Mostrar más

 • Oferta promocionada

Senior RDMA Networking Engineer for AI Inference Pods

EtchedSan Jose, CA, United States
A tiempo completo

A leading AI infrastructure company based in San Jose is seeking a highly skilled Supercomputing Engineer specialized in networking.This role involves developing high-performance networking solutio... Mostrar más

 • Oferta promocionada

Security Engineer - Data Loss Prevention (DLP)

CADENCE INCSan Jose, CA, United States
A tiempo completo

Cadence Security EngineerAt Cadence, we hire and develop leaders and innovators who want to make an impact on the world of technology.Summary:A highly skilled and experienced Security Engineer with... Mostrar más

 • Oferta promocionada

Sr Manager - Teamcenter PLM Architect

Lam ResearchFremont, CA, United States
A tiempo completo

Join Lam as an IT Engineer, where you'll be at the forefront of designing, analyzing, and implementing applications and systems that form the foundation of our infrastructure.As a crucial member of... Mostrar más

 • Oferta promocionada

ML Engineer: LLM Research & Production

Apple Inc.Cupertino, CA, United States
A tiempo completo

A leading technology company in Cupertino seeks a Research Scientist in AI to conduct groundbreaking research in deep learning and large language models.The ideal candidate will develop and fine-tu... Mostrar más

 • Oferta promocionada

Senior ML Engineer — Scalable Personalization Pipelines

Adobe Inc.San Jose, CA, United States
A tiempo completo

An innovative firm is seeking a Senior Machine Learning Engineer to lead the development of cutting-edge machine learning models that enhance personalized customer experiences.This role offers the ... Mostrar más

 • Oferta promocionada

Principal AI Agent & ML Systems Engineer

OracleSanta Clara, CA, United States
A tiempo completo

Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure.The... Mostrar más

 • Oferta promocionada

Director of Engineering

Ensemble Investments, LLCSanta Cruz, CA, United States
A tiempo completo

Nestled along the Pacific Coast, La Bahia Hotel & Spa celebrates its dramatic setting where the tip of Monterey Bay touches Sana Cruz's coveted Main Beach.Steeped in the romantic beauty of Spanish-... Mostrar más

 • Oferta promocionada

Senior ML Infra Engineer - Training Efficiency

WaymoMountain View, CA, United States
A tiempo completo

A leading autonomous driving technology company in Mountain View is seeking an experienced professional to enhance ML infrastructure for training workloads.Responsibilities include designing distri... Mostrar más

 • Oferta promocionada

Member of Technical Staff, LLM Inference - MAI Superintelligence Team

Microsoft CorporationMountain View, CA, United States
A tiempo completo

OverviewOur Inference team is responsible for building and maintaining the tools and systems that enable Microsoft AI researchers to run models easily and efficiently.Our work empowers researchers ... Mostrar más

 • Oferta promocionada

Flexible remote AI work. Your schedule. Paid weekly, straight to your bank account.

Meridian.aiScotts Valley, CA, US
Teletrabajo
A tiempo completo

Review and label digital content including text, images, and documents.Every task you complete helps improve how technology interprets information and performs in practical settings.Detail-oriented... Mostrar más

 • Oferta promocionada

Project Engineer

ACCO Engineered SystemsSanta Cruz, CA, United States
A tiempo completo

This position will collaborate closely with a dedicated Project Manager, reporting to the Sales Manager, and be responsible for the technical coordination of assigned projects.This includes fosteri... Mostrar más