Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization
Team: Gen AI
Location: Bengaluru
Commitment: Full-time
Workplace Type: onsite
What You’ll Do
- Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations.
- Develop and manage CI/CD pipelines for model promotion (staging → production).
- Configure autoscaling on Kubernetes (GPU pools) based on active calls or streaming sessions.
- Build health and observability dashboards: latency, token delay, WER drift, SNR/packet loss monitors.
- Integrate LM bias APIs, failover logic, and model switchers for fallback to larger/cloud models.
- Implement on-device or edge inference paths for low-latency scenarios.
- Collaborate with AI team to expose APIs for context biasing, rescoring, and diagnostics.
- Optimize GPU/CPU utilization, cost optimization, and memory footprint for concurrent ASR/TTS/Speech LLM workloads.
- Maintain data and model versioning pipelines with MLflow, DVC, or internal registries.
Desired Skills
- Experience with Triton, TensorRT, Docker, Kubernetes, and GPU scheduling.
- Familiarity with speech inference (streaming ASR, TTS pipelines).
- Proficient in Python, Bash, and cloud services (AWS/GCP/Azure).
- Understanding of observability stacks (Prometheus, Grafana, ELK).
- Knowledge of DevSecOps, access policies, and PHI-safe environments.
- Interest in inference optimization, mixed precision, and quantization.
Qualifications
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
