Role Overview
As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform.
Responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers.
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines.
- Uses enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis.
- Collaborates with teams to design, develop, test, and implement availability, reliability, and scalability solutions.
- Applies AI capabilities to identify patterns in operational signals that indicate reliability risk or recurring toil.
- Implements infrastructure, configuration, and network as code for applications and platforms.
- Collaborates with technical experts and stakeholders to resolve complex problems.
- Understands service level indicators and utilizes service level objectives to proactively resolve issues.
- Supports the adoption of site reliability engineering best practices within the team.
Requirements
- Formal training or certification on software engineering concepts and 3+ years applied experience.
- Proficient in site reliability culture and principles.
- Proficient in at least one programming language such as Python, Java/Spring Boot, or .Net.
- Good experience of building or designing applications on Cloud (AWS preferred) or Kubernetes.
- Experience with Terraform for infrastructure management.
- Experience building observability for large scale applications (e.g., Datadog, Dynatrace, AppDynamics, Grafana, Prometheus, OpenTelemetry).
- Working knowledge of using enterprise-authorized AI capabilities within SRE workflows.
- Familiarity with Linux OS and kernel concepts.
- Ability to contribute to large and collaborative teams with limited supervision.
Skills
- Python
- AWS
- Kubernetes
- Terraform
- Prometheus