Deskripsi Pekerjaan
Informasi lengkap tentang posisi dan persyaratan
Ringkasan Yukerja
Lowongan Site Reliability Engineer (SRE) di HTC Global Services kami kurasi dari JobStreet (kategori Teknologi & IT). Perhatikan lokasi kerja (South Jakarta, Jakarta) sebelum melamar. Yukerja.com bukan pemberi kerja — lamaran diproses di situs sumber resmi.
Job Desc:
Maintain the reliability, availability, scalability and performance of Tencent Cloud's production services and infrastructure.
Engineer Tencent Cloud services to achieve 99.99%+ availability targets, depending on the service SLA.
Operate and improve highly distributed production systems supporting large numbers of customers.
Define and monitor SLIs, SLOs, error budgets and service health indicators.
Build monitoring, logging, tracing and automated alerting capabilities.
Detect and resolve live production incidents across compute, network, storage and database environments.
Lead incident response, root-cause analysis and post-mortem activities.
Develop self-healing and automated remediation mechanisms.
Automate repetitive operational activities and reduce engineering toil.
Perform capacity planning, performance optimization and resilience testing.
Design failover and disaster-recovery mechanisms.
Participate in production/on-call rotations for critical cloud services.
Work directly with Tencent Cloud product engineering teams to improve service reliability
Requirements:
Minimum 5 years experience as Site Reliability Engineer (SRE)
Core Skills: Linux, Kubernetes, containers, Python/Go/C++, distributed systems, observability, automation, CI/CD, cloud infrastructure, incident management and performance engineering.
Ready to join ASAP is preferably
Willing work in shifting (8 hours x 5 days/week)