Site Reliability & Production Support Engineer
Overview
We are looking for a Site Reliability & Production Support Engineer to maintain and improve the availability, performance, and resilience of our production systems. You will own production incidents from investigation through to resolution, implementing fixes yourself where possible and coordinating with other teams or external providers where needed.
Requirements
- A Bachelor's degree or diploma in Computer Science, Information Technology, Software Engineering, Engineering, or a related technical field.
- 5+ years' relevant experience in SRE, production engineering, DevOps, technical application support, or a similar production-focused technical role.
- Experience investigating and resolving production incidents.
- Ability to troubleshoot applications, databases, operating systems, networks, and cloud infrastructure.
- Experience supporting cloud-hosted and on-premises workloads.
- Practical SQL skills, including diagnosing connectivity, query performance, locking, and connection pooling issues.
- Experience with monitoring, logging, alerting, and distributed tracing tools.
- Proficiency in at least one scripting or programming language, such as Python, Bash, PowerShell, C#, or Go.
- Ability to read application code, investigate defects, and work with developers on fixes.
- Familiarity with CI/CD, version control, infrastructure as code, and safe production deployments.
- Clear written and verbal communication, especially during incidents.
- Ability to take ownership, solve problems systematically, and prioritise under pressure.
Responsibilities
- Production support and incident resolution.
- Support production applications, APIs, databases, infrastructure, and integrations.
- Own incidents from detection and triage through investigation, service restoration, and closure.
- Prioritise incidents based on severity, business impact, and affected services.
- Investigate issues using logs, metrics, traces, database queries, application code, and infrastructure diagnostics.
- Implement configuration, script, infrastructure, and application code changes within your area of responsibility.
- Carry out controlled rollbacks, failovers, restarts, and message reprocessing with appropriate safeguards.
- Coordinate engineering teams and external providers where their expertise or access is needed, and drive issues through to resolution.
- Keep incident records and communicate impact, progress, and recovery status to stakeholders.
- Participate in an agreed on-call rotation, including support for critical incidents outside business hours.
- Conduct root cause investigations and blameless incident reviews.
- Identify causes and contributing factors across applications, infrastructure, processes, and dependencies.
- Document incident timelines, recovery actions, findings, and preventative measures.
- Assign owners to corrective actions, track completion, and verify that fixes address the underlying problem.
- Identify recurring incidents and support requests, and implement changes to prevent them.
- Define and monitor service level indicators and objectives with engineering and business stakeholders.
- Build and maintain dashboards, alerts, centralised logging, and distributed tracing.
- Monitor system health, performance, dependencies, and customer impact.
- Reduce alert noise and improve failure detection.
- Address reliability risks, performance bottlenecks, capacity constraints, and single points of failure.
- Define and support resilience tests, disaster recovery exercises, and backup and restore testing.
- Automate repetitive support tasks, diagnostics, health checks, and recovery procedures.
- Maintain operational tools, scripts, runbooks, and troubleshooting guides.
- Contribute to infrastructure as code, CI/CD pipelines, and deployment safeguards.
- Ensure services have the monitoring, health checks, rollback plans, and documentation needed for production support.
- Work with developers to improve timeout handling, retries, idempotency, and graceful failure.
- Follow production access controls, change management, and audit requirements.
How to apply
About this listing: SpanSam is an opportunity discovery service and is not the hiring employer unless explicitly stated. Application decisions and source-listing changes are controlled by the employer or institution.