Uses vibe coding (AI-assisted development) with tools like GitHub Copilot to speed development and debugging.
About the Role
AI Operations Support Engineer responsible for maintaining reliability of enterprise AI platforms (Azure and AWS) through automation and L2 incident response. The role focuses on developing Python-based self-healing utilities, supporting Kubernetes workloads and observability stacks, and using AI-assisted 'vibe coding' to accelerate development and debugging.
Job Description
Role
Consultant-level AI Operations Support Engineer focused on maintaining and improving the reliability of enterprise AI platforms on Azure and AWS using automation and operational best practices.
Key Responsibilities
- Perform L2 troubleshooting for incidents that are not resolved by existing automation.
- Develop automation and “self-healing” utilities using Python to reduce repetitive operational tasks.
- Support Kubernetes-based workloads and ML orchestration tools (e.g., Kubeflow) and manage GitOps workflows (ArgoCD).
- Operate and maintain observability stacks, including Grafana and Prometheus.
- Use AI-assisted development tools (“vibe coding”, e.g., GitHub Copilot) to accelerate development and debugging.
- Collaborate with L1 operators, contribute to root-cause analysis, and manage incidents and tickets using ServiceNow.
Requirements
- Degree in Computer Science or a related quantitative field.
- 2–3 years of Python experience.
- Demonstrable Kubernetes skills and experience supporting Kubernetes workloads.
- Experience with Infrastructure as Code (IaC).
- Strong communication skills, resilience, and ability to work collaboratively.
- Mindset that prioritizes proactive automation over reactive firefighting.
Desirable
- Experience with agentic AI bot development (e.g., LangChain).
- Familiarity with CI/CD pipelines.
- Familiarity with ITSM processes and cloud environments (Azure/AWS).
Location & Details
- Primary location: Chennai, Tamil Nadu, India.
- Travel: No.
- Required experience: 2 years.