Site Reliability & Platform Engineer with 14+ years designing, building and operating reliable systems on Kubernetes. Deep expertise in observability with the Grafana stack and others. Open-source-minded, comfortable in owning services end-to-end — from design docs and clean, maintainable code to CI/CD, IaC, on-call troubleshooting, performance engineering & tuning, incident response and mentoring. Strong in Go, Python and Java and experienced with relational stores (PostgreSQL, MySQL, SQLite).
Orchestration & IaC: Kubernetes, Docker, Terraform, Ansible, Helm, ArgoCD, Jenkins, GitOps
Platform: On-prem (OpenStack), AWS, Azure
Observability: Grafana, Prometheus, Thanos, Mimir, Loki, Tempo, OpenTelemetry, Vector, Fluentd/Bit, InfluxDB, Victoria, Instana, Splunk, ELK
Languages: Go, Python, Java, Scala, C#, Bash
Backend & APIs: Microservices, REST, gRPC, distributed systems, performance & profiling, design docs, code review
Stores: PostgreSQL, MySQL, SQLite, Redis, Kafka, S3
Performance: Gatling, Locust, k6, JMeter, LoadRunner, Neoload