特斯拉
Site Reliability Engineer, Splunk
岗位职责
At Tesla, logs from software and hardware systems—including applications and microservices, operating systems and middleware, production-line equipment, vehicle-related systems, and network and security components—are ingested into Splunk as the standard platform for enterprise log analytics and troubleshooting. Splunk supports cross-system root-cause analysis, reduces MTTD/MTTR, helps sustain high availability of production and business systems, and provides a unified data entry point for security audit, change impact assessment, capacity planning, and compliance analysis. Platform stability, search performance, and data quality directly affect business continuity and risk control. This role is a Site Reliability Engineer (SRE) responsible for the architecture, build, and operations of the Splunk logging platform, data ingestion pipelines, and observability capabilities. The focus is on automation and platform-ready delivery, in close collaboration with application, manufacturing, and infrastructure teams, to ensure high-quality ingestion and efficient retrieval of large-scale logs that support troubleshooting and business decisions. Responsibilities
Splunk Platform & Reliability Design, build, and maintain Splunk infrastructure, including Multi-Site deployments and high-availability (HA) architectures, ensuring reliability, scalability, security, and high performance. Operate data-ingestion platforms (pipelines, routing, filtering/enrichment, and integration with Splunk) to improve ingestion quality, cost efficiency, and operability. Drive machine-data onboarding, field extraction, search and data-model optimization; create alerts, troubleshoot search performance issues, and define data storage and lifecycle policies. Develop automation, monitoring, and diagnostic tools; participate in on-call, respond quickly to bridge calls, minimize incident impact; mentor junior engineers.
Observability (Log / Metric / Trace) Design and implement enterprise observability solutions that support service health assessment, dependency analysis, and fault localization. Own architecture, ingestion, storage, query, and capacity management for Prometheus / Mimir; establish dashboard and alerting standards with Grafana. Deploy, upgrade, scale, and troubleshoot observability components and collection pipelines in Kubernetes environments. Drive correlation and unified views across Log / Metric / Trace; co-define naming, labeling, collection, and SLO/alerting standards.
Splunk AI OPS Design and deliver AI tools and services for troubleshooting, analysis, reporting, and knowledge Q&A (e.g., intelligent search assistance, alert interpretation, root-cause suggestions, runbooks, and ops assistants). Build platform AI operations capabilities—alert noise reduction and correlation, intelligent recommendations, and closed-loop feedback—integrating securely with Splunk search, alerts, dashboards, permissions, and audit; measure outcomes via adoption, accuracy, MTTR, and noise rate. Document reusable tools and best practices to improve how users leverage Splunk and observability capabilities.
任职要求
Core Skills 7+ years of systems administration experience, with strong Linux background (RHEL / CentOS / Ubuntu, etc.). 5+ years designing, maintaining, and troubleshooting mid-to-large-scale Splunk infrastructure; deep understanding of distributed Splunk architecture. Strong SPL skills for complex queries; solid grasp of best practices for reports, alerts, and dashboards. Hands-on experience operating data-ingestion platforms (collection, processing, routing, and integration with downstream systems). Experience with large-scale distributed systems and high-availability architectures; strong analytical and problem-solving skills. Preferred Qualifications Bachelor’s degree in Engineering, Mathematics, Computer Science, Information Technology, or equivalent experience. Experience extending the Splunk ecosystem (Custom Commands, Modular Inputs, App development, REST API, external service integration). AIOps experience (alert noise reduction, event correlation, anomaly detection, auto-enrichment, integration with on-call / ticketing / IM). Demonstrated ability to use AI effectively (assisted coding, problem analysis, solution design) to continuously deliver maintainable tools, scripts, or automation. Configuration management or Infrastructure-as-Code experience (e.g., Ansible, Puppet). Experience with OpenTelemetry, APM, or full distributed tracing implementations. Experience with long-term metrics storage or federated query platforms (e.g., Mimir / Thanos / Cortex / VictoriaMetrics). Experience building or evolving observability platforms on Kubernetes (Operator, Helm, GitOps, etc.). Production experience applying LLM / RAG / Agents to operations or data-analysis scenarios. Splunk Administrator or Architect level certification. Fluent English reading, writing, and speaking. Ability to define and drive platform success metrics (onboarding coverage, alert noise rate, MTTR, AI tool adoption, etc.).
信息来自企业官方招聘渠道
OfferSeek 对公开岗位信息进行聚合、去重和结构化整理,最终申请以企业官方页面为准。
