All jobs
Intermedia Intelligent CommunicationsDevOps
Principal DevOps Engineer
UKPosted today
Intermedia is a leading provider of cloud communications and collaboration technology, fostering a fast-paced, growth-oriented environment where employees are valued and promoted from within. They are seeking a Principal DevOps Engineer to lead infrastructure, automation, and deployment initiatives.
Location: UK
Responsibilities
- Act as technical lead for DevOps/Platform/Release engineering: set direction, standards, and best practices
- Architect and govern end-to-end delivery: infrastructure provisioning, configuration management, CI/CD, release processes, and operations
- Design and support Windows-based high availability solutions, with deep ownership of Windows clustering (failover/HA patterns, maintenance, upgrades, troubleshooting)
- Lead Linux automation and platform standardization (configuration, patching, hardening, performance tuning)
- Own Infrastructure as Code strategy with Terraform (modules, environments, state, governance)
- Own automation strategy with Ansible (reusable roles, inventories, secure secrets handling, idempotency)
- Build and standardize deployments using Octopus Deploy, GitHub, and Ansible (templates, shared steps, release promotion, rollback)
- Design and mature CI/CD pipelines (artifact versioning, approvals, promotion strategy, policy-as-code where applicable)
- Establish observability standards using VictoriaMetrics/Prometheus (metrics strategy, alerting, SLO/SLA monitoring, dashboards)
- Provide production leadership: incident response, RCA/postmortems, reliability improvements, capacity planning
- Mentor engineers, review designs/code, and raise overall engineering quality across teams
- Produce and maintain architecture docs, runbooks, and platform roadmaps
Requirements
- Bachelors degree in Computer Science or related field
- 7+ years (or equivalent) in DevOps / SRE / Infrastructure Engineering, including leadership in complex environments
- Expert-level experience designing and operating Windows Server HA and clustering (Failover Clustering and related components)
- Strong Linux administration and automation experience (systemd, networking, storage, performance)
- Advanced skills with Terraform and Ansible (architecture, reusable components, secure operations)
- Strong deployment/release engineering experience with Octopus Deploy and GitHub (release governance, environment promotion, rollback)
- Monitoring/observability expertise with VictoriaMetrics and/or Prometheus (alerting strategy, metrics design, operational readiness)
- Production experience running Redis, RabbitMQ, Nginx (HA, tuning, troubleshooting)
- Strong understanding of networking and security fundamentals (TLS, DNS, load balancing, firewalling, least privilege)
- Proven ability to lead cross-team initiatives, make architectural decisions, and communicate clearly
- Kubernetes and container ecosystems (Docker, Helm)
Additional Information
- Candidates must be located in the UK