特斯拉
Sr. Site Reliability Engineer
岗位职责
As a Senior Site Reliability Engineer at Tesla China, you will own reliability, observability, automation, and operational efficiency for critical business systems and infrastructure. You will keep services stable, secure, and scalable under production load and peak traffic; manage monitoring, change, incidents, and performance in an engineering-driven way; drive SLO/SLI adoption; partner with development, platform, and business teams on end-to-end reliability; and continuously reduce incident impact and manual toil through automation and data. Responsibilities Own stability operations for in-scope systems and platforms; build and continuously improve monitoring, alerting, logging, and tracing to speed detection and localization Define and track SLO/SLI/SLA and error budgets; use data to surface risk and drive improvements in availability, latency, capacity, and change quality Contribute to high-availability architecture and disaster recovery; land backup/restore, degradation and rate limiting, fault isolation, and emergency drills Manage configuration, deployment, scaling, inspection, and day-to-day operations through code and automation to reduce repetitive manual work Establish and improve incident response; cover on-call and major incidents; complete root cause analysis (RCA) and closed-loop remediation Partner with development, platform, network, security, and business teams to identify launch and architecture risks early and improve releasability and rollback safety (Splunk focus) Own Splunk platform architecture, deployment, data onboarding, parsing, and knowledge object management; build logging and observability capabilities for troubleshooting, audit, security, and business analytics Produce and maintain runbooks, emergency playbooks, and technical documentation; codify standards, tooling, and platform capabilities
任职要求
Bachelor’s degree or above; Computer Science, Software Engineering, Network Engineering, Information Security, or related majors preferred Five or more years of experience in SRE, DevOps/operations engineering, platform engineering, infrastructure engineering, or related reliability work Solid Linux and networking fundamentals; strong familiarity with common distributed architectures, troubleshooting methods, and performance analysis Hands-on experience building and operating observability stacks (monitoring, alerting, logging, tracing) Automation skills; ability to improve operations and delivery with Python, Go, Shell, or similar Practical experience with at least one of CI/CD, configuration management, containers, or cloud infrastructure, with strong change-risk control Strong problem decomposition, prioritization, and cross-team influence; comfortable with on-call and a fast-paced environment Strong Chinese and English communication and documentation skills; able to work with English technical materials and cross-regional teams Preferred Experience Core SRE: Reliability for large-scale distributed systems, SLO framework design, incident review mechanisms, or proven reduction of MTTR / change failure rate Splunk: Splunk architecture and cluster operations, data onboarding and parsing (props/transforms), dashboard/alert/SPL development, access control, and data governance High availability and capacity planning for Kubernetes, cloud/virtualization, or middleware (message queues, cache, databases) Hands-on use of Prometheus, Grafana, ELK/OpenTelemetry, or similar observability toolchains IaC, GitOps, platform engineering, or security/audit logging collaboration experience SRE practice in manufacturing, automotive, high-traffic internet, or large enterprise environments This job application may involve an interview with an interviewer outside of Tesla China. If you complete your application, you agree Tesla provides your application information to overseas interviewers in Tesla, Inc. for recruitment purposes. More details and contact information please see here. (here hyperlink: https://app.mokahr.com/social-recruitment/tesla/46129#/)
信息来自企业官方招聘渠道
OfferSeek 对公开岗位信息进行聚合、去重和结构化整理,最终申请以企业官方页面为准。
