AWS Site Reliability Engineer (SRE)
Prescient
Overview
We’re looking for an AWS Site Reliability Engineer (SRE) to help us build and operate highly reliable, secure, and scalable cloud platforms. This role is ideal for someone who thrives at the intersection of software engineering, cloud infrastructure, and operations, and enjoys automating everything.
Requirements
- Strong experience (> 5 years) with AWS services (EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, CloudFront, VPC, Route 53, IAM)
- Expertise in Infrastructure as Code (Terraform, AWS CDK, CloudFormation)
- Proficiency in monitoring & observability tools (CloudWatch, Grafana, ELK/OpenSearch)
- Experience with CI/CD pipelines (GitHub Actions, GitLab CI, AWS Code Pipeline)
- Knowledge of containerization & orchestration (Docker, Kubernetes, ECS, EKS)
- Strong scripting/coding skills (Python, Bash, Go, etc.)
- Experience with incident management & on-call operations
- AWS Professional certifications
- Experience running Kubernetes/EKS in production
- Knowledge of compliance frameworks (ISO27001, SOC2, PCI-DSS, POPIA)
Responsibilities
- Design, implement, and maintain highly available and resilient AWS cloud infrastructure
- Monitor system health and performance, ensuring services meet SLAs
- Respond to and resolve production incidents, performing root cause analysis and implementing long-term fixes
- Build automation for deployment, monitoring, scaling, and recovery using Infrastructure as Code
- Automate repetitive operational tasks to reduce toil and improve system reliability
- Implement CI/CD pipelines to ensure smooth and reliable delivery of applications
- Configure and manage observability solutions
- Define and track Service Level Indicators (SLIs) and Objectives (SLOs)
- Develop proactive alerting and anomaly detection mechanisms
- Apply AWS security best practices and work closely with InfoSec teams
- Perform regular audits of cloud resources
- Continuously optimize cloud infrastructure for performance, efficiency, and cost-effectiveness
- Drive post-incident reviews and develop self-healing systems
- Partner with development teams to embed reliability, scalability, and observability into applications
- Mentor engineers on AWS, DevOps, and reliability engineering practices
How to apply
Verify before you apply
SpanSam summarises opportunities for easier discovery. Always confirm the closing date, eligibility requirements and submission instructions on the original source before sending personal information or documents.
About this listing: SpanSam is an opportunity discovery service and is not the hiring employer unless explicitly stated. Application decisions and source-listing changes are controlled by the employer or institution.