All jobs
nameDevOps
Systems Reliability Engineer
Remote (United States)$100,000–$150,000 AnnuallyPosted today
The role is for an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production, combining software engineering principles with infrastructure and operations.
Location: Remote (United States)
Salary: $100,000–$150,000 Annually
Responsibilities
- Define, instrument, and refine service-level objectives (SLOs), SLIs, and error budgets for critical services.
- Lead incident response and resolution for production issues, including post-incident reviews.
- Design and implement monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog.
- Build and maintain on-call processes, runbooks, and escalation paths.
- Automate operational toil with production-grade tooling in Python, Go, Bash, or similar.
- Architect and operate large-scale Kubernetes clusters and container workloads.
- Design CI/CD pipelines with automated testing, canary deployments, feature flags, and progressive rollout strategies.
- Lead capacity planning and performance engineering activities.
- Partner with application development teams to embed reliability practices early in design.
- Strengthen platform resiliency through chaos engineering, fault injection, retries, timeouts, circuit breakers, and failover paths.
- Drive security improvements in collaboration with security teams.
- Contribute to the technical roadmap for reliability tooling and observability platforms.
- Mentor engineers on SRE practices and foster a culture of operational excellence.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field.
- 5+ years of SRE, DevOps, or production engineering experience supporting large-scale distributed systems.
- Strong programming skills in Python, Go, or Java.
- Deep experience operating Linux at scale, including networking and troubleshooting.
- Production experience with Kubernetes and container workloads.
- Knowledge of observability tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK.
- Experience designing and operating CI/CD pipelines.
- Understanding of distributed system design, consistency models, partitioning, and failure semantics.
- Experience leading incident response and post-incident reviews.
- Excellent communication and documentation skills.
Additional Information
- Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates encouraged to apply. No sponsorship for new H-1B visas.