Role Overview
Site Reliability Engineering (SRE) at Equifax is a discipline that combines software and systems engineering for building and running large-scale, distributed, fault-tolerant systems. SRE ensures that internal and external services meet or exceed reliability and performance expectations while adhering to Equifax engineering principles. Our SREs are responsible for overall system operation and we use a breadth of tools and approaches to solve a broad set of problems.
Responsibilities
- Maintain and execute Infrastructure as Code (IaC) using Terraform across public cloud environments (GCP/AWS).
- Deploy, support, and troubleshoot containerized microservices operating on Kubernetes (GKE/EKS).
- GitHub action workflows and groovy scripts for automated incident remediation, task automation, and pipeline integration.
- Set up and maintain observability dashboards, alerts, and metrics using tools such as Datadog.
- Participate in a 24/7 follow-the-sun operational rotation to manage, triage, and resolve production incidents.
- Collaborate with development teams to analyze system outages and execute preventative action items via postmortems.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related technical discipline.
- 4–10 years of overall technical experience across DevOps, SRE, Systems Administration, or Software Engineering.
- 1+ years of hands-on experience running and maintaining applications in a public cloud environment (GCP preferred, AWS acceptable).
- Working experience with Infrastructure as Code using Terraform and container orchestration using Kubernetes/Docker.
- Practical scripting skills in Terraform for infrastructure automation.
- Experience with CI/CD pipeline automation (e.g., Jenkins, GitLab CI).
Skills
- Terraform
- Kubernetes
- GCP
- AWS
- Docker