Role Overview
We are seeking an experienced Site Reliability Engineer (SRE) to support a large-scale enterprise SaaS platform operating in cloud and high-availability environments. This role is responsible for maintaining infrastructure reliability, availability, performance, and operational excellence across Microsoft Azure, Microsoft SQL Server, PostgreSQL, Windows, and Linux. This is a hands-on engineering role requiring independent investigation of complex technical issues and collaboration with engineering teams.
Responsibilities
- Maintain highly available, reliable, and scalable cloud infrastructure in Microsoft Azure.
- Monitor platform health, review technical logs, and proactively address performance and availability issues.
- Plan and support operating system patching and maintenance activities across Windows and Linux servers.
- Manage and support Microsoft SQL Server databases and prepare for PostgreSQL operational support.
- Troubleshoot and tune queries, SQL jobs, indexing, and resource utilization.
- Implement and maintain backup, restore, high availability, and disaster recovery processes.
- Perform incident response, including P1/P2 triage and Root Cause Analysis (RCA).
- Drive operational efficiency through PowerShell, Python, and Infrastructure-as-Code practices.
Requirements
- 4+ years of relevant experience in Site Reliability Engineering or Platform Operations.
- Strong production support background with Microsoft SQL Server.
- Experience or awareness of PostgreSQL operational readiness.
- Proficiency in managing Windows and Linux environments.
- Ability to work in a 24x7 production support environment, including US time zone/night shifts.
Skills
- Microsoft Azure
- Microsoft SQL Server
- PostgreSQL
- Python
- Linux