BARUN MANDAL
Site Reliability & Platform Engineer

brnmndl@gmail.com

+46 709 201 126

Stockholm, Sweden

Linkedin | Github | Portfolio

Profile

Site Reliability & Platform Engineer with 14+ years designing, building and operating reliable systems on Kubernetes. Deep expertise in observability with the Grafana stack and others. Open-source-minded, comfortable in owning services end-to-end — from design docs and clean, maintainable code to CI/CD, IaC, on-call troubleshooting, performance engineering & tuning, incident response and mentoring. Strong in Go, Python and Java and experienced with relational stores (PostgreSQL, MySQL, SQLite).

Experience

Site Reliability Engineer — FDJ (La Française des Jeux)
Mar 2025 – Present · Stockholm, Sweden
  • Operate large-scale production Kubernetes — on-call, incident response and blameless postmortems.
  • Unified metrics, logs and traces; reduced MTTR via runbooks and SLO-based alerting.
  • Build and maintain backend tooling in Go and Python for observability automation, deployments and operational workflows.
  • Own reliability KPIs (SLIs/SLOs) and error-budget policy across critical services.
  • Built AI agents and MCP server integrations for intelligent metrics discovery and anomaly detection, reduced alert noise and accelerated triage.
Site Reliability Engineer — Kindred Group plc
Jun 2021 – Mar 2025 · Stockholm, Sweden
  • Modernised observability for 100+ microservices: migrated to Grafana + Prometheus + Thanos/Mimir + Loki + Tempo, with full metrics-logs-traces correlation.
  • Implemented GitOps on multiple Kubernetes cluster with ArgoCD, Terraform, Ansible, Jenkins — eliminated config drift and shortened release cycles.
  • Tuned backend services (Go, Java) for latency and throughput, working with relational stores (PostgreSQL, MySQL, Oracle) and Kafka.
  • Owned on-call rotations, incident response and blameless postmortems; mentored junior SREs and reviewed design documents across squads.
Performance Engineer — Derivco Sports
Oct 2018 – Jun 2021 · Stockholm, Sweden
  • Designed load and resilience tests (k6, Gatling, JMeter, Locust) for high-throughput betting backends; identified DB and service bottlenecks at scale.
  • Profiled and tuned JVM/Go services, reducing tail latency and infrastructure cost.
Senior Consultant — PwC India
Jan 2015 – Sep 2018 · Kolkata, India
  • Led reliability, performance and capacity engagements for enterprise clients; delivered design and remediation roadmaps end-to-end.
Performance Engineer — Cognizant
Aug 2012 – Jan 2015 · Kolkata, India
  • Performance engineering for Java backends and relational databases; automated test pipelines and reporting.

Core Skills

SREObservabilitySLI/SLO Error BudgetsDevOpsPlatform Engineering Incident ResponseAutomation Performance & CapacityResilience & DR

Tech Stack

Orchestration & IaC: Kubernetes, Docker, Terraform, Ansible, Helm, ArgoCD, Jenkins, GitOps

Platform: On-prem (OpenStack), AWS, Azure

Observability: Grafana, Prometheus, Thanos, Mimir, Loki, Tempo, OpenTelemetry, Vector, Fluentd/Bit, InfluxDB, Victoria, Instana, Splunk, ELK

Languages: Go, Python, Java, Scala, C#, Bash

Backend & APIs: Microservices, REST, gRPC, distributed systems, performance & profiling, design docs, code review

Stores: PostgreSQL, MySQL, SQLite, Redis, Kafka, S3

Performance: Gatling, Locust, k6, JMeter, LoadRunner, Neoload

Education

B.E., Engineering
Indian Institute of Engineering Science and Technology, Shibpur · 2008–2012
High School, Science
Purulia Zilla School, India · 1998–2008

Certifications