Back to roles
InfrastructureFull-timeLocationShenzhen

LLM Infra Engineer

Design and build an efficient AI infrastructure platform supporting foundation-model training/inference/deployment and MLOps; optimize hardware utilization and edge deployment.

ShenzhenTeamInfrastructure
Apply now

Responsibilities

  • AI infrastructure planning and construction: design and build an efficient AI infrastructure platform; build a unified platform supporting foundation-model training/inference/deployment; design and implement an efficient, scalable MLOps system to support fast model iteration.
  • Deeply optimize platform performance and hardware-resource efficiency; optimize AI model storage and compute resource utilization (GPU/TPU, memory, bandwidth, storage) to improve reliability, performance, and scalability.
  • Deploy AI models (e.g., CNN, Transformer, Diffusion) to various edge devices (phones, embedded devices, IoT devices) and perform model quantization and acceleration.
  • Build data collection, cleaning, labeling, and management platforms for the health/fitness domain; design data-privacy protection mechanisms; establish data-quality monitoring systems.
  • Team building and collaboration: work closely with algorithm, application, and product teams to land AI capabilities; establish technical standards and best practices to improve overall engineering efficiency.

Requirements

  • Master's degree or above in CS, AI, or related fields; 3+ years of AI Infra architecture design or distributed-system development experience.
  • Proficient in Kubernetes, Docker, Hadoop, Spark, and other distributed-system technologies; experience deploying and operating large-scale compute clusters; experience with resource management and deployment on cloud platforms (AWS, Alibaba Cloud, etc.); MLOps platform experience (MLflow, Kubeflow, etc.).
  • Experience building and optimizing GPU/TPU acceleration clusters; familiar with NVIDIA CUDA, TensorRT, vLLM, and other DL inference optimization tools; strong performance-tuning skills; able to analyze and resolve performance bottlenecks in distributed environments; familiar with GPTCache, KVCache, etc.
  • Proficient in at least one of Python/Go/C++; deep understanding of Transformer architecture and foundation-model training/inference; proficient in mainstream DL frameworks (PyTorch, DeepSpeed, Megatron, etc.); experience fine-tuning and serving Llama/Qwen and other large models.
  • Proficient with inference engines (e.g., ONNX Runtime, MNN) for model conversion, quantization, pruning, and optimization; familiar with edge performance tuning, including operator fusion, memory optimization, and heterogeneous scheduling (CPU/GPU/NPU) to reduce latency and power.
  • Excellent communication and cross-team collaboration skills; experience supporting multi-team AI projects.
Send your résumé to career@speediance.com
LLM Infra Engineer | Speediance