| Real-Time Inference Engineering Lead |
| The Real-Time Inference Engineering Lead is responsible for designing, building, and operationalizing scalable, low-latency model serving platforms that power real-time AI and predictive decisioning use cases. This role leads the architecture, deployment, performance optimization, resiliency, and operational governance of online inference services across cloud and on-premises environments. |
| * Design and implement highly available, low-latency model serving architectures for real-time inference workloads. * Develop and standardize deployment patterns for scalable AI/ML services across Kubernetes-based platforms. * Lead API-based inference service design, integration, and lifecycle management. * Optimize model serving performance through latency tuning, caching strategies, autoscaling, and resource management. * Establish monitoring, observability, SLOs/SLAs, alerting, and operational runbooks for production services. * Drive load testing, capacity planning, resiliency engineering, and disaster recovery readiness. * Integrate inference platforms with CI/CD pipelines to enable automated deployments and controlled releases. * Partner with Data Science, Platform Engineering, MLOps, and Infrastructure teams to ensure reliable production model operations. * Govern operational best practices, security, reliability, and performance standards for enterprise AI deployments. |
| * Online inference and model-serving architectures * Real-time APIs and distributed systems design * Kubernetes, container orchestration, and service mesh technologies * Autoscaling, capacity management, and workload optimization * Performance engineering, load testing, and latency optimization * Monitoring, observability, logging, tracing, and SLO management * CI/CD, DevOps, and Infrastructure-as-Code practices * Reliability engineering, fault tolerance, and resiliency patterns * Cloud and on-premises platform operations * Python, Java, Go, or similar backend development experience |
| * Experience with MLOps platforms and enterprise AI deployment frameworks. * Hands-on experience with real-time recommendation, fraud, risk, personalization, or predictive analytics platforms. * Familiarity with GPU-based inference, model optimization, and multi-cloud deployments. |
| A successful candidate combines AI platform engineering, distributed systems expertise, and operational excellence to deliver resilient, high-performance inference platforms that enable enterprise-scale real-time AI solutions. |