Role Overview
As a Site Reliability Engineer III at JPMorgan Chase within Consumer & Community Banking, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform.
Responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Applies enterprise-authorized AI capabilities to identify patterns in operational signals that indicate reliability risk or recurring toil
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, and scalability solutions
- Implements infrastructure as code, configuration, and network as code for applications and platforms
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues
- Supports the adoption of site reliability engineering best practices within your team
Requirements
- Formal training or certification on software engineering concepts and 3+ years applied experience
- 10+ overall IT experience with minimum 5 years in SRE or DevOps
- Proficient in site reliability culture and principles
- Proficient in at least one programming language such as Python or Java/Spring Boot
- Proficient knowledge of software applications and technical processes (e.g., Cloud, artificial intelligence)
- Experience in observability using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk
- Experience with CI/CD tools like Jenkins, GitLab, or Terraform
- Familiarity with container orchestration such as ECS, Kubernetes, and Docker
- Familiarity with troubleshooting common networking technologies and issues
- Working knowledge of using enterprise-authorized AI capabilities within SRE workflows
Skills
- Python
- Kubernetes
- Terraform
- AWS ECS
- Prometheus