Senior Site Reliability Engineer
Teladoc Health
Job description
About the role
We are looking for a Senior Site Reliability Engineer (Sr. SRE) with deep Microsoft Azure expertise to own the reliability, scalability, and observability of our mission‑critical healthcare cloud services. You will lead an SRE team, partner with engineering, product, security and operations, and act as the technical authority for reliability across the platform.
Key responsibilities
- Define, implement and evolve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets for critical applications and platform services.
- Partner with application teams to improve fault tolerance, scalability, and operational readiness, and eliminate recurring reliability issues through root‑cause analysis and automation.
- Design resilient systems that can survive Azure region, zone, network, dependency and deployment failures.
- Build and enhance observability across applications, infrastructure and cloud services, creating dashboards, alerts, logs, traces and metrics using Azure Monitor, Log Analytics, Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic, or similar APM tools.
- Analyze performance, capacity and resilience, improve backup, disaster recovery, failover and business continuity practices.
- Support and improve production workloads on Azure, enforce operational standards, and collaborate on secure, cost‑aware architecture.
- Lead incident management, conduct blameless post‑mortems, and improve runbooks and response processes.
- Work with security engineering to ensure compliance with cloud security standards and operational controls.
Required profile
- Senior‑level professional with extensive experience in Microsoft Azure and cloud operations.
- Proven track record of implementing SLI/SLO frameworks and driving reliability improvements.
- Strong analytical skills to diagnose performance bottlenecks and capacity risks.
- Experience leading or mentoring an SRE team and collaborating across engineering, product and security.
Required skills
- Microsoft Azure (including Azure Monitor, Log Analytics, tagging, backup, recovery, identity and security).
- Observability and APM tools: Elastic/ELK, Grafana, OpenTelemetry, Datadog, Dynatrace, New Relic.
- Reliability engineering practices: SLIs, SLOs, error budgets, incident management, post‑mortem analysis.
- Automation and scripting for remediation and capacity management.
- Knowledge of resiliency patterns such as retries, circuit breakers, bulkheads and queue‑based decoupling.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Argentina.
Salaries by job title
- Sales Executive IT // Esquema Hibrido - Microcentro / CABA 25
- Líder DevOps AWS // Multinacional de Software Financiero - Híbrido/ Microcentro / CABA 16
- Engineering Manager // Software Financiero - Remoto para residentes en Argentina 14
- Project Manager 13
- Personal Shopper - Caba 12
- Cocinero/a 11
- Técnico de Mantenimiento 11
- Community manager 11
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 6 days from now
25 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Teladoc Health