- Full-time
Driving innovation in model serving and inference architectures, the full-time AI Research Engineer will optimize deployment strategies for advanced AI systems in a fully remote environment, focusing on enhancing performance across diverse applications. Key responsibilities Design and deploy efficient model serving architectures that optimize throughput and memory usage across various environments Build and monitor inference tests in production settings, analyzing key performance indicators to validate model performance Collaborate with cross-functional teams to integrate optimized inference frameworks into production pipelines for edge and on-device applications Required qualifications A degree in Computer Science or a related field, ideally a PhD in NLP, Machine Learning, or a related discipline Proven experience in low-level kernel optimizations and inference optimization on mobile devices Strong expertise in writing GPU kernels for mobile devices and modern model serving architectures Demonstrated ability to apply empirical research to enhance model serving performance Deep understanding of techniques such as Pruning, Quantization, and Diffusion Models