@ManhTran.
All jobs

Infrastructure & DevOps

GPU Infrastructure Engineer

Hà NộiMiddle/SeniorFull-timeOpenings: 2Deadline: September 17, 2026

About the company

VinSmart Future (VSF) is Vingroup's technology company, formed by merging the Group's entire technology ecosystem. As a core driver of Vingroup's future growth, VSF is AI-first - with artificial intelligence as the foundation of everything we build. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.

Responsibilities

  • Design, deploy, operate, and optimize on-premise GPU infrastructure supporting AI, Machine Learning, and high-performance data processing platforms.

  • Administer Linux server systems, including operating system configuration, kernel, CPU/RAM/GPU resources, filesystem, and system services.

  • Install, configure, and troubleshoot issues related to NVIDIA GPUs, drivers, CUDA, NVIDIA Container Toolkit, NVLink, MIG, and related components.

  • Manage physical server infrastructure, including bare-metal servers, BIOS, firmware, RAID, IPMI, iDRAC, or iLO.

  • Deploy and operate GPU clusters, container systems, and workload orchestration platforms in on-premise environments.

  • Collaborate with AI, Data, and Development teams to optimize application performance, GPU resource utilization, and system stability.

  • Monitor, analyze, and resolve issues related to performance, resources, networking, storage, operating system, and hardware.

  • Build tools and scripts to automate the installation, configuration, monitoring, and operation of systems.

  • Develop technical documentation, operating procedures, configuration standards, and incident response plans.

  • Research, evaluate, and propose GPU infrastructure upgrade solutions aligned with product development needs.

Requirements

Requirements

  • Bachelor's degree in Information Technology, Computer Science, Electronics & Telecommunications, or a related field.

  • 4–5 years of experience in DevOps, System, Infrastructure, Platform Engineering, or equivalent roles.

  • Hands-on experience with on-premise infrastructure, physical servers, and data center environments.

  • Strong knowledge and skills in Linux; able to analyze and troubleshoot issues at the operating system level.

  • Solid understanding of process, memory, filesystem, networking, system services, package management, and Linux performance tuning.

  • Hands-on experience with Docker, container runtime (containerd + NVIDIA Container Toolkit), and Kubernetes in on-premise environments.

  • Experience deploying and operating GPU workloads on Kubernetes: configuring the NVIDIA device plugin, scheduling for GPU nodes (taint/toleration, node affinity), and troubleshooting pods that fail to access GPUs.

  • Ability to use Bash, Python, or an equivalent programming language to build automation tools.

  • Knowledge of networking and storage such as TCP/IP, VLAN, bonding, NFS, NAS, SAN, or Ceph.

  • Systems thinking, with strong root-cause analysis skills and the ability to resolve complex issues.

  • Ability to read and understand technical documentation in English.

Preferred

  • Experience working with NVIDIA GPU, CUDA, GPU clusters, or AI infrastructure.

  • Experience with NVIDIA GPU Operator, Node Feature Discovery (NFD), and GPU sharing mechanisms such as MIG (Multi-Instance GPU), time-slicing, or MPS.

  • Experience with batch/gang scheduling for AI/ML on Kubernetes (Volcano, Kueue) or Slurm–Kubernetes integration.

  • Experience with AI platforms on Kubernetes such as Kubeflow, KubeRay, or Training Operator for distributed training.

  • Experience with Slurm, NVIDIA DCGM, or other GPU workload management platforms.

  • Experience with InfiniBand, RoCE, NVLink, or high-speed networking (via SR-IOV / NVIDIA Network Operator).

  • Experience with monitoring using Prometheus, Grafana, Zabbix, ELK; DCGM Exporter for GPU metrics is a plus.

  • Familiarity with on-premise Kubernetes distributions: RKE2/Rancher, OpenShift, or kubeadm.

  • Experience with Ansible or other infrastructure configuration automation tools.

  • Software Development background with a transition into System, Platform, or Infrastructure Engineering.

  • Experience supporting AI/ML systems, distributed training, or high-performance computing (HPC).

Benefits

  • Income competitive with the market.
  • Lunch allowance.
  • Preferential rates across the Group's ecosystem: tuition discounts (Vinschool), healthcare (Vinmec), resorts (Vinpearl), vehicle purchase (VinFast), and home rental or purchase (Vinhomes) … under the Group's policies.
  • Full insurance coverage as required by the Labor Law (Social, Health, UI), plus Company-provided personal health insurance based on position level, and periodic health check-ups at reputable hospitals and health centers nationwide.
  • Access to strategic, large-scale key technology projects.
  • The opportunity to work in a professional technology environment that brings together scientists, experts and engineers from leading technology companies in Vietnam and worldwide.
  • Free learning resources on Udemy, Coursera and O'Reilly; internal workshops; certification sponsorship; and special mentorship programs from the Group's and Company's leadership.
  • The chance to join the Group's technology clubs and internal tech events to learn and turn personal projects and ideas into reality.
  • Training programs to become an "Internal Trainer" and share expertise, with special benefits.
  • 12 annual leave days, plus public holidays and Tết as regulated by law.

Working Hours

  • 05 official working days at the office (Monday – Friday).
  • 02 remote working days per month on Saturdays on a rotating schedule.
  • Flexible working hours with check-in window from 08:30 – 09:30.
  • Proactively manage time to complete 08 working hours/day.

Work Location

  • Technopark Tower, Ocean Park, Hà Nội

Check application status

Enter the email you used when submitting your CV — we'll send a verification code to that inbox to protect your information.

We email you a verification code so only you can view your application status.