Role Overview
Join a dynamic team at the forefront of technology, where your skills drive innovation and modernize mission-critical systems. As a Site Reliability Engineer II at JPMorgan Chase within the Enterprise Technology, Engineering Services and Platform team, you will solve complex business problems with straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to iteratively improve existing solutions.
Responsibilities
- Guide and assist others in building appropriate level designs and gaining consensus from peers
- Collaborate with software engineers and teams to design and implement deployment approaches using automated CI/CD pipelines
- Design, develop, test, and implement availability, reliability, and scalability solutions in applications
- Implement infrastructure, configuration, and network as code for applications and platforms
- Collaborate with technical experts, stakeholders, and team members to resolve complex problems
- Understand service level indicators and utilize service level objectives to proactively resolve issues
- Support the adoption of site reliability engineering best practices within your team
- Accelerate delivery and improve operational rigor through AI-assisted engineering adoption
- Provide 24/7 production support for business-critical applications
- Use enterprise-authorized AI capabilities to speed up incident triage, troubleshooting, and post-incident analysis
- Apply AI capabilities to identify recurring toil and reliability risks from operational signals
Requirements
- Formal training or certification on software engineering concepts and 2+ years applied experience
- Proficient in site reliability culture and principles
- Proficient in at least one programming language such as Python, Java/Spring Boot, or .Net
- Experience in observability including monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Elastic, or Splunk
- Experience with CI/CD tools like Jenkins, GitLab, or Terraform
- Experience with event streaming platforms like Kafka
- Experience with AI-assisted tools like GitHub Copilot or Claude
- Ability to assess AI-assisted operational recommendations for correctness and risk
Preferred Qualifications
- Deep understanding of TCP/IP, DNS, load balancing, firewalls, and VPN technologies
- Experience tuning Linux performance and troubleshooting system-level issues
- Certifications: AWS Certified SysOps Administrator or Professional, Certified Kubernetes Administrator (CKA), or equivalent
- Familiarity with container and orchestration technologies such as ECS, Kubernetes, and Docker
Skills
- Python
- Kubernetes
- Terraform
- AWS
- Jenkins