Lucidya
Site Reliability Engineer
Riyadh Full-time Engineering
Job description
Lucidya is hiring a Site Reliability Engineer (SRE) in Riyadh to own the reliability, performance, and resilience of our cloud infrastructure.
What You'll Do:
- Design and maintain highly available, fault-tolerant, and scalable infrastructure.
- Manage and optimize cloud workloads across AWS, GCP, or Azure using Infrastructure as Code (Terraform).
- Operate and scale Kubernetes clusters (EKS, GKE, etc.) in production.
- Implement and refine monitoring and alerting systems using tools like Prometheus, Grafana, Datadog, or ELK.
- Write scripts and build tooling to automate repetitive operational work.
- Collaborate with DevOps and engineering teams to solve performance bottlenecks and improve CI/CD pipelines.
Requirements:
- 3+ years of experience in SRE, DevOps, or infrastructure engineering.
- Hands-on experience working with Kubernetes in production.
- Strong expertise in cloud environments (AWS, GCP, or Azure) and distributed systems.
- Proven experience using Terraform for managing infrastructure.
- Ability to write automation scripts in Python, Bash, or similar languages.
- Solid understanding of networking, load balancing, and high-availability design.
What You'll Do:
- Design and maintain highly available, fault-tolerant, and scalable infrastructure.
- Manage and optimize cloud workloads across AWS, GCP, or Azure using Infrastructure as Code (Terraform).
- Operate and scale Kubernetes clusters (EKS, GKE, etc.) in production.
- Implement and refine monitoring and alerting systems using tools like Prometheus, Grafana, Datadog, or ELK.
- Write scripts and build tooling to automate repetitive operational work.
- Collaborate with DevOps and engineering teams to solve performance bottlenecks and improve CI/CD pipelines.
Requirements:
- 3+ years of experience in SRE, DevOps, or infrastructure engineering.
- Hands-on experience working with Kubernetes in production.
- Strong expertise in cloud environments (AWS, GCP, or Azure) and distributed systems.
- Proven experience using Terraform for managing infrastructure.
- Ability to write automation scripts in Python, Bash, or similar languages.
- Solid understanding of networking, load balancing, and high-availability design.
🎤 AI interview preparation
Tailored to this job. Pick your experience level and we'll prepare the most likely questions with answer tips.
Role: Site Reliability Engineer
✨ AI cover letter
Tailored to this job. Add a few lines about your experience and we'll draft it.
Source: Manual entry