Jobiglo

Sin resultados.

Senior Site Reliability Engineer

Teladoc Health

Senior 🇬🇧 English
Microsoft Azure Azure Monitor Log Analytics Elastic/ELK Grafana OpenTelemetry Datadog Dynatrace New Relic

Descripcion del puesto

About the role

We are looking for a Senior Site Reliability Engineer (Sr. SRE) with deep Microsoft Azure expertise to own the reliability, scalability, and observability of our mission‑critical healthcare cloud services. You will lead an SRE team, partner with engineering, product, security and operations, and act as the technical authority for reliability across the platform.

Key responsibilities

  • Define, implement and evolve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets for critical applications and platform services.
  • Partner with application teams to improve fault tolerance, scalability, and operational readiness, and eliminate recurring reliability issues through root‑cause analysis and automation.
  • Design resilient systems that can survive Azure region, zone, network, dependency and deployment failures.
  • Build and enhance observability across applications, infrastructure and cloud services, creating dashboards, alerts, logs, traces and metrics using Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM tools.
  • Analyze performance, capacity and resilience, improve backup, disaster recovery, failover and business continuity practices.
  • Support and improve production workloads on Azure, enforce operational standards, and collaborate on secure, cost‑aware architecture.
  • Lead incident management, conduct blameless post‑mortems, and improve runbooks and response processes.
  • Work with security engineering to ensure compliance with cloud security standards and operational controls.

Required profile

  • Senior‑level professional with extensive experience in Microsoft Azure and cloud operations.
  • Proven track record of implementing SLI/SLO frameworks and driving reliability improvements.
  • Strong analytical skills to diagnose performance bottlenecks and capacity risks.
  • Experience leading or mentoring an SRE team and collaborating across engineering, product and security.

Required skills

  • Microsoft Azure (including Azure Monitor, Log Analytics, tagging, backup, recovery, identity and security).
  • Observability and APM tools: Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic.
  • Reliability engineering practices: SLIs, SLOs, error budgets, incident management, post‑mortem analysis.
  • Automation and scripting for remediation and capacity management.
  • Knowledge of resiliency patterns such as retries, circuit breakers, bulkheads and queue‑based decoupling.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Teladoc Health.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Por que reporta esta oferta?

Gracias por su reporte. Revisaremos esta oferta.

Postula en 30 segundos

Ingresa tu email para postular. Se creara una cuenta automaticamente.

Al continuar, aceptas nuestras condiciones de uso.

Ya tienes cuenta? Iniciar sesion

💬 Escríbenos en Telegram Chatear por WhatsApp

Publicado hace 1 mes

Expira en 1 semana

22 vistas · 0 interested

Aumenta tus posibilidades

Sube tu CV: te propondremos las ofertas que coinciden con tu perfil.

Analizando tu CV...

Teladoc Health