Principal Engineer with 12+ years of experience building cloud-native platforms, distributed systems, and machine learning infrastructure. Experienced operating production ML workloads at scale using Kubernetes, KServe, Istio, Knative, Terraform, and GCP. Proven track record leading platform transformations, enabling data science teams, and delivering reliable ML systems in production.
Main Achievements
Operated Kubernetes-based production ML workloads serving 15M+ requests/day across ~20 ML services, with automated scaling and end-to-end observability.
Redesigned ML platform with KServe, Istio, Knative; halved infra usage and improved performance.
Integrated ClickHouse for real-time model inputs, improving pricing and revenue responsiveness.
Built pipelines on GCP (1+ TB/day, 15000+ events/sec) to build the company's data platform backbone.
Led the architecture and implementation of Cabify’s next-generation ML platform using Kubernetes, KServe, Knative and Istio, supporting ~20 production ML services and reducing infrastructure usage by 50%.
Own the platform’s production architecture on GKE, including autoscaling, service networking, traffic management and deployment automation, supporting workloads serving 15M+ requests per day.
Built the platform’s observability stack with Prometheus and Grafana, providing visibility into model-serving latency, resource utilization, autoscaling behaviour and service reliability.
Designed event-driven infrastructure using Kafka and Knative Eventing to capture high-volume ML inference audit data for downstream storage and analytics, with a focus on reliability, replayability and cost.
Developed real-time data connectivity between production ML services and ClickHouse for latency-sensitive pricing and demand use cases.
Drove platform technical direction through architectural evaluations, production prototypes and close collaboration with a 12-person Data Science team.
Designed and maintained high-throughput pipelines on GCP (GCS + BigQuery) using Scala and Apache Beam, processing 1+ TB/day and 15k+ events/sec.
Built and optimised Go-based ETLs to integrate third-party data into the Data Warehouse, improving ingestion performance and reliability.
Overhauled the Python-based machine learning platform to streamline model deployment, and operated nearly 20 live models with full MLOps responsibility.
Managed cloud infrastructure with Terraform, maintained Kubernetes clusters, and implemented automated deployments via ArgoCD.
Contained cloud spend growth despite a 4× increase in business volume, through infrastructure efficiency and proactive monitoring.
Led a 4-person Android team as Tech Lead, defining the technical roadmap and driving platform improvements for the Driver app.
Managed a cross-functional team of 7 (backend, frontend, and mobile), aligning product and engineering efforts for Driver-facing systems.
Played a key engineering leadership role during the integration of a major competitor, which doubled the number of active drivers (~50k) in key markets.
Collaborated closely with Product to ensure technical plans supported evolving business priorities and delivery timelines.
Actively supported career growth for team members through mentoring, role development, and performance coaching.
Led the rewrite of the Driver app from Java to Kotlin, designing a scalable architecture that supported a 50k-driver user base with bi-weekly releases and a 99.9% crash-free rate.
Contributed to the Rider app’s transition from Cordova to native Kotlin, helping define its RxJava-based architecture and developing native plugins. The app reached over 2 million downloads during this period.