workfromanywhereworkfromanywhere
All jobs
Bright Vision TechnologiesEngineering

Systems Reliability Engineer

Remote (United States)$100,000–$150,000 AnnuallyPosted today

Seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. The role involves applying software engineering principles to infrastructure and operations, automating workflows, and improving system reliability.

Location: Remote (United States)

Salary: $100,000–$150,000 Annually

Responsibilities

  • Define, instrument, and refine service-level objectives (SLOs), SLIs, and error budgets for critical services.
  • Lead incident response and resolution for production issues, including post-incident reviews.
  • Design and implement monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog.
  • Build and maintain on-call processes, runbooks, and escalation paths.
  • Automate operational workflows using programming languages such as Python, Go, Bash.
  • Architect and operate large-scale Kubernetes clusters and container workloads.
  • Design CI/CD pipelines supporting safe, frequent releases with automated testing and deployment strategies.
  • Lead capacity planning and performance engineering activities.
  • Partner with application teams to embed reliability practices early in design.
  • Strengthen platform resiliency through chaos engineering, fault injection, and failover paths.
  • Improve security posture in collaboration with security teams.
  • Contribute to the technical roadmap for reliability tooling and observability platforms.
  • Mentor engineers on SRE practices and foster a culture of operational excellence.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or related field.
  • 5+ years of SRE, DevOps, or production engineering experience supporting large-scale systems.
  • Strong programming skills in Python, Go, or Java.
  • Deep experience operating Linux at scale, including networking and troubleshooting.
  • Production experience with Kubernetes and container workloads.
  • Knowledge of observability tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK.
  • Experience designing and operating CI/CD pipelines.
  • Understanding of distributed system design, consistency, partitioning, and failure semantics.
  • Experience leading incident response and post-incident reviews.
  • Excellent communication and documentation skills.

Location

Remote (United States)

Salary

$100,000–$150,000 Annually

Category

Engineering

Source

himalayas

Posted

today

Share this job

XLinkedIn

Similar remote jobs

LOTHIAN BUSESEngineering

Job Summary

Livingston, West Lothian, United Kingdom£22.96 to £22.96 Per Hour
2d ago

Technical Specialist

Bhopal
2d ago

GPU Systems Engineer

Remote (US)$100,000–$150,000 Annually
today
TeyaNewEngineering

Senior Backend Engineer

Portugal
today