特斯拉
Site Reliability Engineer, Traffic&Splunk
岗位职责
We are Tesla Infrastructure SRE team, responsible for ensuring the stability of Tesla's globally distributed business operations in China. Our team spans two technical pillars - Traffic and Splunk - supporting critical workloads across manufacturing, energy, connected vehicles, and global content delivery. Traffic owns Tesla core traffic infrastructure - including F5 load balancing, CDN/Edge, Varnish/Envoy, L7 proxies, and cross-site high-availability architectures - ensuring reliable delivery of Tesla official website, vehicle firmware, streaming media, and APIs. Splunk owns the build, operation, and enablement of Tesla Splunk observability platform - covering data ingestion, indexing, search, alerting, dashboards, and knowledge base curation - empowering business teams to drive incident response and service improvement with data. We are building AI-native operations capabilities - integrating Large Language Models (LLMs), MCP toolchains, and Splunk / Confluence knowledge bases to enable automated alert investigation, intelligent runbook retrieval, and troubleshooting assistance, so engineers can focus on judgment and decision-making rather than repetitive context-switching and manual queries. This is a hybrid role with a primary focus on Traffic : under the guidance of senior engineers, you will spend most of your time on traffic infrastructure operations and delivery, while also growing into Splunk platform work as a secondary pillar. Responsibilities Shared Deploy, monitor, troubleshoot, and maintain production systems to ensure platform availability. Develop automation scripts and operational tooling (Shell / Python / Go) to reduce toil and improve troubleshooting efficiency. Participate in on-call (24/7) rotations, respond to alerts per runbooks, and assist with root cause analysis and issue closure. Collaborate with global Traffic and Splunk teams and China business stakeholders to drive standardization and best practices. Traffic Assist with configuration and change execution for traffic components including F5 load balancing, Varnish, and Envoy. Participate in HTTP/TCP/SSL/TLS troubleshooting and assist with traffic anomaly and availability analysis. Configure and manage TCP/L7 proxy technologies to ensure efficient traffic management and content delivery. Participate in testing and validation of cross-site traffic architecture changes and cross-border traffic analysis. Troubleshoot and resolve critical issues across multiple layers, including CDN, Loadbalancers, storage, OS, network, K8S, virtualization, and application/DB stack. Splunk Assist with deployment, configuration, and routine health checks of Splunk cluster components (UF, Indexer, Search Head, etc.). Support business teams with machine data onboarding, field extraction, index planning, and search optimization. Help create and maintain alerts, dashboards, and SPL queries; troubleshoot common search performance issues. Build up user community to grow end-user capabilities in using Splunk as a powerful tool.
任职要求
Must Have 2-4 years of Linux systems operations or SRE experience. Solid understanding of TCP/IP, HTTP, DNS, and other core networking protocols. Deep understanding of CDN and TCP/L7 proxy technologies. Hands-on experience with load balancers such as F5 OR proxy technology. Understanding of CDN/Edge networking , including topology, protocols, hardware, and architecture. Foundational knowledge of containerization technologies such as Kubernetes. Experience with configuration and management using tools such as Ansible or Puppet. Proficiency in Shell; ability to write automation scripts in Python or Go. Strong analytical thinking, problem-solving, and communication skills; ability to read and write technical documentation in Chinese and English. High sense of ownership; comfortable with on-call rotations and a fast-paced production environment. Able to independently handle routine operational tasks with guidance shortly after onboarding. Familiarity with AI tools : experience using AI to assist with troubleshooting and automation in daily workch. Fluent spoken English for collaboration with global teams. Nice to Have AI / automation : experience using AI-assisted tools (e.g., Cursor) to improve daily troubleshooting, scripting, and documentation. Development or testing of MCP Servers, automation agents, or internal ops assistants. Traffic / networking experience : F5, Nginx, HAProxy, Varnish, Envoy, CDN, DNS, etc. Splunk experience: SPL queries, alerts/dashboards, UF deployment, or distributed architecture fundamentals. Observability platform experience: ELK, Mimir, Prometheus, Grafana, etc. Configuration management and IaC: Ansible, Terraform, Git, etc. Splunk certification (Admin or above).
信息来自企业官方招聘渠道
OfferSeek 对公开岗位信息进行聚合、去重和结构化整理,最终申请以企业官方页面为准。
