宁德时代
AI Infrastructure Engineer
岗位职责
Track 1: Data Engineering 方向一:数据工程 - Design and build scalable data pipelines for AI model development, training, evaluation, and production applications. 设计并建设可扩展的数据 Pipeline,支持 AI 模型开发、训练、评估和生产应用。 - Build ingestion, transformation, validation, versioning, and lineage capabilities for structured and unstructured data. 建设结构化与非结构化数据的采集、转换、校验、版本管理和数据血缘能力。 - Develop reliable workflows for processing text, images, documents, scientific data, time-series data, and other domain-specific datasets. 开发可靠的数据处理流程,支持文本、图像、文档、科学数据、时序数据及其他领域数据。 - Build data infrastructure for LLM and RAG applications, including vector stores, knowledge bases, chunking and embedding pipelines, and retrieval quality evaluation. 建设面向大模型与 RAG 应用的数据基础设施,包括向量存储、知识库、切分与向量化 Pipeline,以及检索质量评估。 - Improve data quality, observability, reproducibility, and accessibility across AI development workflows. 提升 AI 开发流程中的数据质量、可观测性、可复现性和可访问性。 - Work with algorithm and domain teams to translate model development
任职要求
into reusable data infrastructure. 与算法团队和业务领域团队合作,将模型开发需求转化为可复用的数据基础设施。 Track 2: AI Platform, MLOps/LLMOps and Model Infrastructure 方向二:AI 平台、MLOps/LLMOps 与模型基础设施 - Design and build AI development platforms that support the end-to-end model lifecycle, including data preparation, experimentation, training, evaluation, deployment, inference, monitoring, and iteration. 设计并建设支持模型全生命周期的 AI 开发平台,包括数据准备、实验、训练、评估、部署、推理、监控和持续迭代。 - Build standardized MLOps and LLMOps pipelines, tools, and workflows for model development, lifecycle management, and production delivery. 建设标准化的 MLOps 和 LLMOps Pipeline、工具及工作流,支持模型开发、全生命周期管理和生产交付。 - Provide reusable platform capabilities for experiment tracking, dataset and model versioning, model evaluation, model registry, release management, and reproducibility. 提供可复用的平台能力,支持实验追踪、数据集与模型版本管理、模型评估、模型注册、发布管理和结果复现。 - Build scalable training, fine-tuning, evaluation, and inference workflows across local, cloud, Kubernetes, GPU cluster, and hybrid environments. 建设适用于本地、云端、Kubernetes、GPU 集群及混合环境的可扩展训练、微调、评估和推理工作流。 - Develop standardized model deployment and serving capabilities, including model packaging, rollout, request routing, load balancing, autoscaling, monitoring, and failure recovery. 开发标准化的模型部署与服务能力,包括模型封装、发布、请求路由、负载均衡、弹性扩缩、监控和故障恢复。 - Integrate and operate mainstream model training and serving frameworks such as PyTorch, DeepSpeed, FSDP, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or equivalent technologies. 集成并运行主流模型训练和服务框架,例如 PyTorch、DeepSpeed、FSDP、vLLM、SGLang、TensorRT-LLM、Triton Inference Server 或同类技术。 - Optimize model training and inference systems for performance, scalability, GPU utilization, memory efficiency, reliability, and cost. 围绕性能、可扩展性、GPU 利用率、显存效率、可靠性和成本,优化模型训练与推理系统。 - Improve AI developer experience by providing standardized development environments, SDKs, APIs, templates, CI/CD pipelines, and self-service platform capabilities. 通过标准化开发环境、SDK、API、模板、CI/CD Pipeline 和自助式平台能力,提升 AI 开发者体验。 - Build observability, resource scheduling, capacity management, cost monitoring, security, and governance capabilities for AI workloads. 建设 AI 工作负载的可观测性、资源调度、容量管理、成本监控、安全及治理能力。 - Collaborate with research, data, and application teams to translate model development requirements into reusable platform and infrastructure capabilities. 与研发、数据及应用团队协作,将模型开发需求转化为可复用的平台与基础设施能力。 Track 3: Agent Infrastructure 方向三:智能体基础设施 - Design and build infrastructure for developing, deploying, operating, and evaluating AI agents. 设计并建设支持 AI 智能体开发、部署、运行和评估的基础设施。 - Build reusable agent runtime capabilities, including model access, tool invocation, workflow orchestration, state management, memory, and execution control. 建设可复用的 Agent Runtime 能力,包括模型接入、工具调用、工作流编排、状态管理、记忆和执行控制。 - Develop infrastructure for agent identity, permissions, sandboxing, secrets management, observability, tracing, and auditability. 建设智能体身份、权限、沙箱、密钥管理、可观测性、链路追踪和审计能力。 - Build agent evaluation and testing systems covering task completion, tool-use correctness, reliability, safety, latency, and cost. 建设智能体评估与测试系统,覆盖任务完成度、工具调用正确性、可靠性、安全性、延迟和成本。 - Provide reusable SDKs, APIs, templates, skills, connectors, and deployment workflows for agent development teams. 为智能体开发团队提供可复用的 SDK、API、模板、Skills、连接器和部署工作流。 - Support long-running, event-driven, scheduled, asynchronous and human-in-the-loop agent workflows. 支持长时间运行、事件驱动、定时执行、异步执行和以及人在回路中的智能体工作流。 - Integrate agents with enterprise systems, data platforms, model services, internal tools, and external APIs, including standardized protocols such as MCP. 将智能体与企业系统、数据平台、模型服务、内部工具及外部 API 进行集成,包括 MCP 等标准化协议。 任职要求: - Master’s degree or above, majoring in Computer Science, Software Engineering, Cloud Computing, Big Data, System Engineering, Math,Artificial Intelligence or related fields. 硕士及以上学历,计算机科学、软件工程、云计算、大数据、系统工程、数学、人工智能等相关专业 - Strong programming and software engineering skills in Python, Java, Go, Scala, C++, Rust, or another relevant programming language. 具备扎实的编程和软件工程能力,熟悉 Python、Java、Go、Scala、C++、Rust 或其他相关编程语言。 - Familiarity with relevant technologies in at least one professional track, such as data engineering, workflow orchestration systems, cloud platforms, Docker, Kubernetes, GPU clusters, distributed training frameworks, model-serving systems, or agent frameworks. 熟悉至少一个专业方向的相关技术,例如数据工程、工作流编排系统、云平台、Docker、Kubernetes、GPU 集群、分布式训练框架、模型服务系统或 Agent 框架。 - Ability to translate research, algorithm, data, or business requirements into reusable engineering platforms, infrastructure, tools, and workflows. 能够将研究、算法、数据或业务需求转化为可复用的工程平台、基础设施、工具和工作流。 - Strong system design, problem-solving, debugging, and performance-analysis capabilities. 具备较强的系统设计、问题解决、故障排查和性能分析能力。 - Strong communication and cross-functional collaboration skills, with the ability to work effectively with research, algorithm, data, product, and application teams. 具备良好的沟通和跨团队协作能力,能够与研究、算法、数据、产品及应用团队高效合作。 加分项: - 3+years of experience designing, building, deploying, or operating reliable production systems. 具备3年以上设计、建设、部署或运行高可用生产系统的经验 - Experience with AI agent frameworks, workflow engines, tool-use systems, memory systems, sandbox environments, or agent evaluation methods. 具备 AI Agent 框架、工作流引擎、工具调用系统、记忆系统、沙箱环境或智能体评估经验 - Experience supporting AI workloads in cloud, hybrid-cloud, on-premises, or multi-cluster environments. 具备在云端、混合云、本地或多集群环境中支持 AI 工作负载的经验。 - Experience optimizing data systems for performance, scalability and cost 具备数据系统的性能、可扩展性与成本优化经验 - Experience building AI infrastructure for scientific computing, materials science, energy, manufacturing, or other domain-specific applications. 具备为科学计算、材料科学、能源、制造或其他领域应用建设 AI 基础设施的经验。 - Experience adapting large models to specialized domains(scientific computing materials, energy, industrial applications) 具备将大模型适配到专业领域(科学计算、材料、能源、工业应用)的经验 - Open-source contributions, track record of academic publications, granted patents or participation in industry standards. 具备开源贡献、学术发表记录,已授权专利或行业标准参与经历。"
信息来自企业官方招聘渠道
OfferSeek 对公开岗位信息进行聚合、去重和结构化整理,最终申请以企业官方页面为准。
