Senior Site Reliability Engineer
Teladoc Health
Descripcion del puesto
About the role
We are looking for a Senior Site Reliability Engineer (Sr. SRE) with deep Microsoft Azure expertise to own the reliability, scalability, and observability of our mission‑critical healthcare cloud services. You will lead an SRE team, partner with engineering, product, security and operations, and act as the technical authority for reliability across the platform.
Key responsibilities
- Define, implement and evolve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets for critical applications and platform services.
- Partner with application teams to improve fault tolerance, scalability, and operational readiness, and eliminate recurring reliability issues through root‑cause analysis and automation.
- Design resilient systems that can survive Azure region, zone, network, dependency and deployment failures.
- Build and enhance observability across applications, infrastructure and cloud services, creating dashboards, alerts, logs, traces and metrics using Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM tools.
- Analyze performance, capacity and resilience, improve backup, disaster recovery, failover and business continuity practices.
- Support and improve production workloads on Azure, enforce operational standards, and collaborate on secure, cost‑aware architecture.
- Lead incident management, conduct blameless post‑mortems, and improve runbooks and response processes.
- Work with security engineering to ensure compliance with cloud security standards and operational controls.
Required profile
- Senior‑level professional with extensive experience in Microsoft Azure and cloud operations.
- Proven track record of implementing SLI/SLO frameworks and driving reliability improvements.
- Strong analytical skills to diagnose performance bottlenecks and capacity risks.
- Experience leading or mentoring an SRE team and collaborating across engineering, product and security.
Required skills
- Microsoft Azure (including Azure Monitor, Log Analytics, tagging, backup, recovery, identity and security).
- Observability and APM tools: Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic.
- Reliability engineering practices: SLIs, SLOs, error budgets, incident management, post‑mortem analysis.
- Automation and scripting for remediation and capacity management.
- Knowledge of resiliency patterns such as retries, circuit breakers, bulkheads and queue‑based decoupling.
Questions fréquentes
Por que reporta esta oferta?
Explorar más
Salarios, guías y búsquedas en Argentina.
Salarios por profesión
- Sales Executive IT // Esquema Hibrido - Microcentro / CABA 25
- Líder DevOps AWS // Multinacional de Software Financiero - Híbrido/ Microcentro / CABA 16
- Engineering Manager // Software Financiero - Remoto para residentes en Argentina 14
- Project Manager 13
- Personal Shopper - Caba 12
- Community manager 11
- Ayudante de cocina 11
- Cocinero/a 11
Postula en 30 segundos
Ingresa tu email para postular. Se creara una cuenta automaticamente.
Al continuar, aceptas nuestras condiciones de uso.
Ya tienes cuenta? Iniciar sesion
Publicado hace 1 mes
Expira en 1 semana
22 vistas · 0 interested
Aumenta tus posibilidades
Sube tu CV: te propondremos las ofertas que coinciden con tu perfil.
Analizando tu CV...
Teladoc Health
Ofertas relacionadas
-
Associate Service Desk Analyst
Motivus Argentine -
Software Engineer (Rust) – Consumer Applications
Kraken Argentine -
Consultor/a Oracle Cloud - OPA (Remoto/Latam)
Quid Solutions Argentine -
Project Manager, Web Development
Power Digital Remote - Argentina; Remote - Ecuador; Remote - Nicaragua -
Consultor Funcional Odoo – Modalidad híbrida en Núñez (CABA)
ADN - Recursos Humanos Capital Federal