Machine Learning Engineer

Sarvam
Bengaluru Full-time 30-50 LPA
Batch: 2020/2021/2022/2023/2024/2025. What We're Looking For: - Strong Python and PyTorch — comfortable reading model internals, profiling inference, and debugging production failures - Hands-on experience integrating and optimising speech models (ASR or TTS) in production environments - Experience with real-time/streaming systems — WebSocket pipelines, chunked audio processing, or latency-sensitive async architectures - Solid understanding of modern speech system architectures — sequence-to-sequence models, attention mechanisms, flow-matching or diffusion-based TTS, streaming ASR - Familiarity with model serving infrastructure — Triton, TorchServe, ONNX Runtime, or equivalent - Experience with audio signal processing fundamentals: sample rates, PCM formats, spectrograms, vocoding, time-stretching - Strong async Python skills — asyncio, concurrent pipelines, managing backpressure in streaming systems - Comfort with ambiguity — the roadmap is not fully pre-specified - Undergraduate degree in a technical discipline (CS, EE, statistics, physics, or equivalent) Bonus Points: - Experience with multilingual or Indic speech systems — handling code-mixing, transliteration, tonal variation across Indian languages - Voice cloning or speaker adaptation techniques (zero-shot or few-shot) in production - Experience with real-time media protocols — WebRTC, RTMP, SRT, HLS, or real-time audio agent frameworks - Multi-participant audio systems — speaker diarization, concurrent pipeline management, per-user audio routing - Vocal source separation or speech enhancement techniques - Familiarity with FFmpeg, GStreamer, or media muxing/demuxing - Contributions to open-source speech/audio projects or a solid GitHub portfolio
Apply NowOpen in CarrerliftTry more jobs free
Posted 2026-07-27