Back to All Openings
ML Ops Engineer
Bangalore, India
3 years exp.
Full-time
Job Overview
We are hiring a ML Ops Engineer for our GCC client — Europe’s top retail brands. SRKay Consulting Group is a consulting firm that helps Fortune 500 companies set up and scale Global Capability Centers (GCCs) in India. This opportunity gives you exposure to enterprise-scale initiatives in retail and supply chain, working alongside a peer group of talented engineers, architects, and domain specialists across geographies in a collaborative, innovation-driven environment. It’s a role that not only sharpens your technical expertise but also provides long-term visibility and growth within a global organization.
Key Responsibilities
- Pipeline Orchestration: Design, develop, and maintain complex ML workflows using Apache Airflow (Cloud Composer) to automate data ingestion, preprocessing, and model training.
- Lifecycle Management: Administer and scale MLflow for experiment tracking, model packaging, and maintaining a centralized Model Registry across the organization.
- Cloud & Hybrid Ops: Create and optimize training environments for custom ML/LLM models.
- Model Serving & Scaling: Architect high-performance inference endpoints and serve models via FastAPI/Flask with API Gateway.
- Infrastructure Management: Manage auto-scaling CUDA clusters on Google Kubernetes Engine (GKE).
- CI/CD: Manage end-to-end delivery with Continuous Integration & Continuous Delivery (CI/CD).
- Observability & Monitoring: Build dashboards to track model health, latency, and data drift.
Requirements & Skills
- Workflow Management: Experience in managing Apache Airflow and Composer to support the Data Engineering components of grounded AI solutions.
- MLflow: Deep knowledge of MLflow Tracking, Projects, and Registry. Experience migrating MLflow backends between cloud providers.
- Workflow Tools: Familiarity with Vertex AI Pipelines and Azure DevOps for automation.
- GCP AI Services: Practical experience with Vertex AI (Workbench, Model Garden, Feature Store) and BigQuery ML.
- Containerization: Expert-level Docker and Kubernetes (GKE/AKS) skills. Must understand K8s operators and resource management for ML workloads.
- Infrastructure as Code (IaC): Proficiency in Terraform to manage reproducible cloud environments.
- Programming: Advanced Python skills with a focus on software engineering best practices (unit testing, modular design).
- Data Engineering: Experience with Change Data Capture (CDC), Spark/PySpark, and optimizing data flow from BigQuery to training nodes.
- Access Control: Knowledge of IAM roles, VPC Service Controls, and securing ML endpoints.
- Experience with LLMOps (managing large-scale foundation models, prompt versioning, and vector database scaling).
- Proven track record of building enterprise-scale ML infrastructure from scratch in Azure/GCP.
- Related MLOps Engineering certifications are beneficial.