Vibhav AM
Vibhav A MBangalore, Karnataka

Site Reliability / DevOps Engineer

Kubernetes, AWS, and AI-Ops systems for reliable platforms.

I build cloud-native infrastructure, GitOps delivery, observability systems, and agentic automation that reduce production risk, accelerate RCA, and make engineering workflows easier to operate.

Targeting SRE / Platform EngineeringBangalore & HyderabadKubernetes + AWSAI-driven automation
60%less deployment effort with GitOps
80%reduction in incident recurrence
30%infrastructure cost reduction
75%fewer memory-pressure pod restarts

Platform Focus

Infrastructure work shaped around measurable reliability.

operateEKS

Cloud-Native Reliability

Production Kubernetes, Karpenter, Helm, cluster security, rollout safety, and performance tuning for high-availability workloads.

automateMCP

AI-Ops Automation

Agentic workflows that connect infrastructure tools, CI systems, and developer operations with traceable LLM observability.

observeSLO

Observability & RCA

New Relic, Dynatrace, Prometheus, Grafana, Langfuse, incident review loops, and production dashboards built for fast diagnosis.

Work Experience

Halodoc, Bangalore

Feb 2023 - Present / Intern to Site Reliability Engineer I to Site Reliability Engineer II

  • Managed GitOps-driven production EKS clusters through ArgoCD and operationalized canary rollouts from 20% to 100% with automated health gates.
  • Executed zero-downtime blue-green Amazon MSK migration from ZooKeeper to KRaft with 100% data consistency across mission-critical Kafka streams.
  • Built Agentic AI orchestration using Model Context Protocol to connect Jenkins, Kubernetes, and GitLab for assisted debugging and reduced tool switching.
  • Implemented Langfuse for production-grade AI observability, tracing LLM latency, cost, and prompt performance across agentic workflows.
  • Extended Checkov with server-side hooks, enforced 10+ custom policies, migrated EC2 fleets to AL2023, and reduced deployment-time risk by 40%.
  • Led deep-dive postmortems and RCA work for production outages, improving SLO/SLI monitoring and reducing recurrence by 80%.

Selected Work

Production initiatives with clear operational outcomes.

01

Amazon MSK ZooKeeper to KRaft Migration

Led a zero-downtime migration strategy with phased traffic validation, rollback safety, and full data consistency across critical messaging streams.

AWS MSK / Kafka / Blue-green / Canary
02

EKS Graviton3 to Graviton4 Migration

Directed workload compatibility checks, performance validation, and cluster right-sizing to improve efficiency and reduce infrastructure spend by 30%.

EKS / Graviton4 / Cost optimization
03

Agentic AI Platform for DevOps

Built a hackathon-winning orchestration layer that helps debug CI/CD and infrastructure workflows through connected platform tools.

MCP / Jenkins / Kubernetes / GitLab

Technical Skills

Cloud, delivery, observability, and AI-Ops toolbox.

Cloud & Kubernetes

AWS EKSEC2S3CloudFrontRoute53MSKRDSLambdaKubernetesDockerHelmTerraformKarpenter

CI/CD & Automation

JenkinsGitLabGroovyPythonShellGitOpsArgoCDn8nFastlane Match

Observability & Reliability

DynatraceNew RelicPrometheusGrafanaLangfuseIncident ManagementRCASLO/SLI

AI-Ops & Platform Engineering

Agentic AIModel Context ProtocolPrompt EngineeringLLM CI/CD ReviewAI Observability

Technical Writing

Three published Halodoc engineering notes from production work.

KRaft / 4 months ago / 8 min read

Migrating AWS MSK from ZooKeeper to KRaft: A Canary Approach

A production migration case study covering blue-green architecture, canary rollout phases, rollback strategy, and observability validation.

Read article
AWS EKS / 7 months ago / 7 min read

Reducing Amazon EKS Compute Costs by 35%: Migrating Production Workloads from Graviton3 to Graviton4

A technical guide to zero-downtime EKS migration, compatibility validation, performance improvements, and cloud cost optimization.

Read article
iOS / 2 years ago / 7 min read

A Step-by-Step Guide to Automated iOS Certificate Renewal

A CI/CD automation writeup for mobile certificate management that removes manual bottlenecks and reduces expired credential risk.

Read article
View all posts on Halodoc Blog

Education & Recognition

Certified Kubernetes practitioner with AI-Ops recognition.

  1. Certified Kubernetes Administrator, The Linux Foundation, Apr 2026
  2. Winner, Halodoc AI Hackathon, Sep 2025
  3. B.E. Information Science and Engineering, National Institute of Engineering, Mysuru

Open to SRE / Platform Engineering roles

Let us build reliable platforms with automation at the center.

Bangalore, Hyderabad, or cloud-native teams building serious production systems.

Email: amvibhav@gmail.com