Role Overview
We are looking for a Platform Engineer – TechOps to join our Platform SRE team and help maintain the reliability, availability, performance, and scalability of large-scale connected-device platforms. The ideal candidate will have hands-on experience in Linux, AWS/Cloud, monitoring, incident management, networking, and production support, along with an automation-first mindset.
Responsibilities
- Monitor and maintain the reliability, availability, and performance of production platforms.
- Provide 24/7 operational support and participate in on-call rotations, including off-hours and weekends when required.
- Handle P1/P2 incidents, troubleshoot production issues, perform RCA, and contribute to post-incident reviews.
- Monitor system health, application performance, infrastructure, and connected devices using monitoring and observability tools.
- Perform Linux system troubleshooting, log analysis, service validation, and issue resolution.
- Work closely with Development, Operations, Support, and Engineering teams for incident resolution and platform improvements.
- Support deployment, commissioning, monitoring, maintenance, and decommissioning of connected devices.
- Assist with firmware changes and releases where required.
- Identify opportunities for automation and reduce repetitive manual operational tasks.
- Support capacity planning, performance tuning, resource optimization, and scalability initiatives.
- Follow security, compliance, documentation, and operational best practices.
- Troubleshoot network-related issues and connectivity problems across distributed environments.
Requirements
- 1+ year of experience in SRE, DevOps, Platform Engineering, Production Support, or Technical Support.
- Strong knowledge of Linux/Unix administration.
- Hands-on experience with AWS or other cloud platforms.
- Experience with monitoring/observability tools such as: Grafana, Prometheus, Dynatrace, Zabbix, CloudWatch.
- Good understanding of incident management, troubleshooting, RCA, and on-call operations.
- Knowledge of networking fundamentals including TCP/IP, DNS, HTTP/HTTPS, routing, firewalls, VLANs, ACLs, and subnetting.
- Scripting/automation experience using Bash, Shell, Python, or similar technologies.
- Good communication and stakeholder-management skills.
- Ability to work effectively in a fast-paced production environment.
Skills
- Linux
- AWS
- Python
- Prometheus
- Grafana