特斯拉
IT Incident Response Engineer
岗位职责
This role is a senior support position within Tesla IT Infrastructure Engineering & Operations. The Incident Response team provides incident response and management support to global cross-functional engineering teams, helping maintain high availability for Tesla Manufacturing, Business Operations, Customer Service & Experience. We reduce incident occurrence through effective IT operations monitoring, risk analysis, and change management. The Tesla APAC Incident Response Center (IRC) is a growing team of professionals from diverse backgrounds, with strong development opportunities. This role is based at Giga Factory Shanghai, China, and provides global support as Tesla's business and mission scale. Senior engineer team positioning: Acts as the regional incident management lead—coordinates teams through investigation and resolution, owns incident management practices (ticket management, root cause investigation, data analysis, and management reporting), and continuously improves processes and tooling. RESPONSIBILITIES • Independently lead end-to-end, 24×7 closed-loop incident management to minimize impact and optimize response time; organize emergency response plans, post-incident reviews, and drills as needed. • Lead or drive IT service management initiatives; establish or optimize SOPs to reduce cross-team communication barriers, promote technical and skills sharing, and raise the team's incident response capability. • Oversee IT Infrastructure & Operations monitoring and operational processes; maintain day-to-day stability and provide periodic reporting on data centers, servers, networks, applications, and related systems; identify and mitigate risks early. • Proactively support team operational improvement without day-to-day supervision—including tool iteration, process optimization, and adoption of industry best practices—to accelerate operational efficiency. • Participate in Infrastructure & Operations daily operations and change management; control change risk, improve change workflows, and support execution of change events. • Use company-approved AI tools for continuous learning and innovation to empower the organization.
任职要求
Must Qualifications • Minimum 5 years of relevant experience; bachelor's degree or above in Information Technology, Software Engineering, Computer Science, or equivalent. • Fluent English; strong communication, sense of responsibility, problem-solving, and teamwork. • Solid IT infrastructure knowledge (networking, servers, virtualization, storage, Kubernetes/application services); hands-on experience preferred. • Major incident / on-call experience in an enterprise environment. • Strong operations observability: Splunk, Prometheus/Alertmanager, synthetic monitoring, Grafana (or equivalents); familiarity with ITSM ticketing practices. • Ability to lead and coordinate cross-team incident resolution while managing information recording and synchronization in parallel. • Participate in 7×24 IRC on-call duty: track response timeliness, coordinate escalation, and facilitate Bridge handling. Preferred Qualifications • Experience collaborating with infrastructure administrators and platform/application support teams. • ITIL Foundation or equivalent ITSM practice; experience in process management, change control, and project management. • Experience leading Bridge or regional / multi-team incident response coordination. • AI-assisted learning and development of operations scripts, code, or automation.
信息来自企业官方招聘渠道
OfferSeek 对公开岗位信息进行聚合、去重和结构化整理,最终申请以企业官方页面为准。
