All jobs
Bright Vision TechnologiesEngineering
Systems Reliability Engineer
Remote (United States)$100,000–$150,000 AnnuallyPosted today
Seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production. The role involves applying software engineering principles to infrastructure and operations, automating workflows, and improving system reliability.
Location: Remote (United States)
Salary: $100,000–$150,000 Annually
Responsibilities
- Define, instrument, and refine service-level objectives (SLOs), SLIs, and error budgets for critical services.
- Lead incident response and resolution for production issues, including post-incident reviews.
- Design and implement monitoring, logging, and tracing strategies using tools like Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog.
- Build and maintain on-call processes, runbooks, and escalation paths.
- Automate operational workflows using programming languages such as Python, Go, Bash.
- Architect and operate large-scale Kubernetes clusters and container workloads.
- Design CI/CD pipelines supporting safe, frequent releases with automated testing and deployment strategies.
- Lead capacity planning and performance engineering activities.
- Partner with application teams to embed reliability practices early in design.
- Strengthen platform resiliency through chaos engineering, fault injection, and failover paths.
- Improve security posture in collaboration with security teams.
- Contribute to the technical roadmap for reliability tooling and observability platforms.
- Mentor engineers on SRE practices and foster a culture of operational excellence.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or related field.
- 5+ years of SRE, DevOps, or production engineering experience supporting large-scale systems.
- Strong programming skills in Python, Go, or Java.
- Deep experience operating Linux at scale, including networking and troubleshooting.
- Production experience with Kubernetes and container workloads.
- Knowledge of observability tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK.
- Experience designing and operating CI/CD pipelines.
- Understanding of distributed system design, consistency, partitioning, and failure semantics.
- Experience leading incident response and post-incident reviews.
- Excellent communication and documentation skills.
Location
Remote (United States)
Salary
$100,000–$150,000 Annually
Category
EngineeringCompany
Bright Vision TechnologiesSource
himalayas
Posted
today
Similar remote jobs
LOTHIAN BUSESEngineering
Job Summary
Livingston, West Lothian, United Kingdom£22.96 to £22.96 Per Hour
2d ago