Infrastructure & DevOps
GPU Infrastructure Engineer
About the company
VinSmart Future (VSF) is Vingroup's technology company, formed by merging the Group's entire technology ecosystem. As a core driver of Vingroup's future growth, VSF is AI-first - with artificial intelligence as the foundation of everything we build. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
Responsibilities
Design, deploy, operate, and optimize on-premise GPU infrastructure supporting AI, Machine Learning, and high-performance data processing platforms.
Administer Linux server systems, including operating system configuration, kernel, CPU/RAM/GPU resources, filesystem, and system services.
Install, configure, and troubleshoot issues related to NVIDIA GPUs, drivers, CUDA, NVIDIA Container Toolkit, NVLink, MIG, and related components.
Manage physical server infrastructure, including bare-metal servers, BIOS, firmware, RAID, IPMI, iDRAC, or iLO.
Deploy and operate GPU clusters, container systems, and workload orchestration platforms in on-premise environments.
Collaborate with AI, Data, and Development teams to optimize application performance, GPU resource utilization, and system stability.
Monitor, analyze, and resolve issues related to performance, resources, networking, storage, operating system, and hardware.
Build tools and scripts to automate the installation, configuration, monitoring, and operation of systems.
Develop technical documentation, operating procedures, configuration standards, and incident response plans.
Research, evaluate, and propose GPU infrastructure upgrade solutions aligned with product development needs.
Requirements
Requirements
Bachelor's degree in Information Technology, Computer Science, Electronics & Telecommunications, or a related field.
4–5 years of experience in DevOps, System, Infrastructure, Platform Engineering, or equivalent roles.
Hands-on experience with on-premise infrastructure, physical servers, and data center environments.
Strong knowledge and skills in Linux; able to analyze and troubleshoot issues at the operating system level.
Solid understanding of process, memory, filesystem, networking, system services, package management, and Linux performance tuning.
Hands-on experience with Docker, container runtime (containerd + NVIDIA Container Toolkit), and Kubernetes in on-premise environments.
Experience deploying and operating GPU workloads on Kubernetes: configuring the NVIDIA device plugin, scheduling for GPU nodes (taint/toleration, node affinity), and troubleshooting pods that fail to access GPUs.
Ability to use Bash, Python, or an equivalent programming language to build automation tools.
Knowledge of networking and storage such as TCP/IP, VLAN, bonding, NFS, NAS, SAN, or Ceph.
Systems thinking, with strong root-cause analysis skills and the ability to resolve complex issues.
Ability to read and understand technical documentation in English.
Preferred
Experience working with NVIDIA GPU, CUDA, GPU clusters, or AI infrastructure.
Experience with NVIDIA GPU Operator, Node Feature Discovery (NFD), and GPU sharing mechanisms such as MIG (Multi-Instance GPU), time-slicing, or MPS.
Experience with batch/gang scheduling for AI/ML on Kubernetes (Volcano, Kueue) or Slurm–Kubernetes integration.
Experience with AI platforms on Kubernetes such as Kubeflow, KubeRay, or Training Operator for distributed training.
Experience with Slurm, NVIDIA DCGM, or other GPU workload management platforms.
Experience with InfiniBand, RoCE, NVLink, or high-speed networking (via SR-IOV / NVIDIA Network Operator).
Experience with monitoring using Prometheus, Grafana, Zabbix, ELK; DCGM Exporter for GPU metrics is a plus.
Familiarity with on-premise Kubernetes distributions: RKE2/Rancher, OpenShift, or kubeadm.
Experience with Ansible or other infrastructure configuration automation tools.
Software Development background with a transition into System, Platform, or Infrastructure Engineering.
Experience supporting AI/ML systems, distributed training, or high-performance computing (HPC).
Benefits
- Income competitive with the market.
- Lunch allowance.
- Preferential rates across the Group's ecosystem: tuition discounts (Vinschool), healthcare (Vinmec), resorts (Vinpearl), vehicle purchase (VinFast), and home rental or purchase (Vinhomes) … under the Group's policies.
- Full insurance coverage as required by the Labor Law (Social, Health, UI), plus Company-provided personal health insurance based on position level, and periodic health check-ups at reputable hospitals and health centers nationwide.
- Access to strategic, large-scale key technology projects.
- The opportunity to work in a professional technology environment that brings together scientists, experts and engineers from leading technology companies in Vietnam and worldwide.
- Free learning resources on Udemy, Coursera and O'Reilly; internal workshops; certification sponsorship; and special mentorship programs from the Group's and Company's leadership.
- The chance to join the Group's technology clubs and internal tech events to learn and turn personal projects and ideas into reality.
- Training programs to become an "Internal Trainer" and share expertise, with special benefits.
- 12 annual leave days, plus public holidays and Tết as regulated by law.
Working Hours
- 05 official working days at the office (Monday – Friday).
- 02 remote working days per month on Saturdays on a rotating schedule.
- Flexible working hours with check-in window from 08:30 – 09:30.
- Proactively manage time to complete 08 working hours/day.
Work Location
- Technopark Tower, Ocean Park, Hà Nội
Check application status
Enter the email you used when submitting your CV — we'll send a verification code to that inbox to protect your information.
We email you a verification code so only you can view your application status.