Jobiglo

No results.

Site Reliability Engineer - Telemetry

Kraken

Mid 🇬🇧 English
Prometheus VictoriaMetrics Grafana Vector Splunk Loki Grafana Alloy Tempo OpenTelemetry Pyroscope Terraform Terragrunt

Job description

About the role

Join the Telemetry team to own and evolve the shared platform that provides metrics, logs, traces, alerts, dashboards, and profiling for a global crypto trading ecosystem. You will work closely with software engineers, platform teams, security engineers, and other SREs to keep the observability stack reliable, scalable, and easy to use.

Key responsibilities

  • Operate and improve the telemetry platform covering metrics, logs, traces, alerting, dashboards, and profiling.
  • Maintain Prometheus‑compatible monitoring stacks (Prometheus, VictoriaMetrics, Grafana) and modern alerting tools.
  • Manage log pipelines using Vector, Splunk, and Loki, ensuring reliability and throughput.
  • Run distributed tracing and profiling services with Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
  • Deploy and manage telemetry services via Terraform, Terragrunt, and container orchestration across multiple environments.
  • Troubleshoot missing data, slow queries, broken alerts, pipeline back‑pressure, and capacity constraints.
  • Build reusable configuration and automation for dashboards, alerts, and telemetry integrations.
  • Participate in incident response, on‑call rotation, write runbooks, and drive platform improvements based on post‑mortems.

Required profile

  • 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production role.
  • Proven ability to manage large‑scale production systems that collect, process, store, and serve telemetry data (metrics, logs, traces, profiles).
  • Comfortable working in fast‑moving, high‑traffic financial or crypto environments.

Required skills

  • Prometheus or Prometheus‑compatible monitoring stack (including VictoriaMetrics, Grafana).
  • Log processing tools: Vector, Splunk, Loki.
  • Distributed tracing and profiling: Grafana Alloy, Tempo, OpenTelemetry, Pyroscope.
  • Infrastructure as code: Terraform, Terragrunt.
  • Container orchestration platforms (e.g., Kubernetes).

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Kraken.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 1 week from now

35 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Kraken