Infrastructure & DevOps
GPU Infrastructure Engineer
About the company
VinSmart Future (VSF) is Vingroup's technology company, formed by merging the Group's entire technology ecosystem. As a core driver of Vingroup's future growth, VSF is AI-first - with artificial intelligence as the foundation of everything we build. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
Responsibilities
Design, deploy, operate, and optimize on-premise GPU infrastructure for AI, Machine Learning, and high-performance data processing workloads.
Administer Linux-based systems, including operating systems, kernels, CPU, memory, GPU resources, filesystems, and system services.
Install, configure, maintain, and troubleshoot NVIDIA GPUs, drivers, CUDA, NVIDIA Container Toolkit, NVLink, MIG, and related components.
Manage physical server infrastructure, including bare-metal servers, BIOS, firmware, RAID, IPMI, iDRAC, or iLO.
Deploy and operate GPU clusters, container platforms, and workload orchestration systems in on-premise environments.
Collaborate with AI, Data, and Development teams to optimize application performance, GPU utilization, and system reliability.
Monitor, analyze, and troubleshoot issues related to performance, operating systems, hardware, networking, storage, and resource utilization.
Develop scripts and automation tools for system installation, configuration, monitoring, and daily operations.
Create technical documentation, operational procedures, configuration standards, and incident response guidelines.
Research and evaluate new technologies and propose appropriate GPU infrastructure improvements.
Requirements
Bachelor’s degree in Information Technology, Computer Science, Telecommunications, or a related field.
At least 4–5 years of experience in System Engineering, Infrastructure Engineering, Platform Engineering, or a similar position.
Hands-on experience with on-premise infrastructure, physical servers, and data center environments.
Strong Linux administration and troubleshooting skills.
Solid understanding of Linux processes, memory management, filesystems, networking, system services, package management, and performance tuning.
Experience with Docker, container runtimes, and Kubernetes in on-premise environments.
Ability to use Bash, Python, or a similar programming language for automation and system tooling.
Knowledge of networking and storage technologies such as TCP/IP, VLAN, bonding, NFS, NAS, SAN, or Ceph.
Strong system-thinking, root-cause analysis, and complex troubleshooting capabilities.
Ability to read and understand technical documentation in English.
Preferred Qualifications
Hands-on experience with NVIDIA GPUs, CUDA, GPU clusters, or AI infrastructure.
Experience with Slurm, Kubernetes GPU Operator, NVIDIA DCGM, or similar GPU workload management platforms.
Knowledge of InfiniBand, RoCE, NVLink, or high-speed networking technologies.
Experience with Prometheus, Grafana, Zabbix, ELK, or similar monitoring platforms.
Experience with Ansible or other infrastructure automation tools.
A software development background followed by a transition into System, Platform, or Infrastructure Engineering.
Experience supporting AI/ML systems, distributed training, or high-performance computing workloads is a strong advantage.
Benefits
- Income competitive with the market.
- Lunch allowance.
- Preferential rates across the Group's ecosystem: tuition discounts (Vinschool), healthcare (Vinmec), resorts (Vinpearl), vehicle purchase (VinFast), and home rental or purchase (Vinhomes) … under the Group's policies.
- Full insurance coverage as required by the Labor Law (Social, Health, UI), plus Company-provided personal health insurance based on position level, and periodic health check-ups at reputable hospitals and health centers nationwide.
- Access to strategic, large-scale key technology projects.
- The opportunity to work in a professional technology environment that brings together scientists, experts and engineers from leading technology companies in Vietnam and worldwide.
- Free learning resources on Udemy, Coursera and O'Reilly; internal workshops; certification sponsorship; and special mentorship programs from the Group's and Company's leadership.
- The chance to join the Group's technology clubs and internal tech events to learn and turn personal projects and ideas into reality.
- Training programs to become an "Internal Trainer" and share expertise, with special benefits.
- 12 annual leave days, plus public holidays and Tết as regulated by law.
Working Hours
- 05 official working days at the office (Monday – Friday).
- 02 remote working days per month on Saturdays on a rotating schedule.
- Flexible working hours with check-in window from 08:30 – 09:30.
- Proactively manage time to complete 08 working hours/day.
Work Location
- Technopark Tower, Ocean Park, Hà Nội
Check application status
Enter the email you used when submitting your CV — we'll send a verification code to that inbox to protect your information.
We email you a verification code so only you can view your application status.