Data
Data Engineer
About the company
VinSmart Future (VSF) is Vingroup's technology company, formed by merging the Group's entire technology ecosystem. As a core driver of Vingroup's future growth, VSF is AI-first - with artificial intelligence as the foundation of everything we build. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.
We are looking for a Data Engineer to build and operate the core data systems for the Mapping & Mobility platform, including: Search Engine (places/addresses/autocomplete), Routing/Navigation, Real-time Traffic, Map Ads, Custom Map and Custom Geofences. The Data Engineer on the team is responsible for designing and operating data pipelines, the data warehouse/lakehouse and the data serving layer — ensuring geospatial, user-behavior and operational data are always available, accurate and exploitable at scale, serving both backend systems and analytical/reporting use cases.
Responsibilities
1) Data Platform & Pipeline for Search & Geocoding
Build and operate pipelines to ingest/transform place, address and POI data from multiple sources (third-party, internal, user-contributed).
Design data-processing flows for address/place enrichment, deduplication and normalization before indexing into search (OpenSearch/Elasticsearch).
Build quality gates: anomaly detection, schema validation, and control of data freshness and coverage by market.
Design CDC (Change Data Capture) and reindex pipelines to sync data zero-downtime with the search platform.
2) Traffic & Routing Data Pipeline
Build stream-processing pipelines for real-time traffic: ingest probe data/GPS signals, aggregate per road segment, compute speed profiles, confidence scoring and freshness.
Design batch pipelines computing historical traffic patterns, ETA baselines and speed bands per segment/time-of-day.
Build a data serving layer (time-series/KV) optimized for low-latency queries by the routing engine.
Manage road-network data: graph updates, schema versioning and delta updates for the routing engine.
3) Analytics & Reporting Platform
Build a centralized data warehouse/lakehouse across all domains: search queries, routing requests, geofence events, map-ads impressions/clicks.
Design data models (dimensional modeling / star schema) suitable for BI reporting and ad-hoc analytics.
Build and maintain data marts for Map Ads reporting: impressions, clicks, frequency, budget burn and attribution pipelines.
Automate data-quality checks and SLA monitoring on pipelines, with alerting on anomalies.
4) Geospatial Data Processing
Build geospatial data-processing pipelines: tile processing, polygon simplification, geofence ingestion and S2/H3/Geohash indexing.
Optimize spatial joins and point-in-polygon queries at scale: batching, spatial-index partitioning, caching.
Support pipelines to create and update custom map styles, layer data and feature flags per tenant/market.
5) Feature Store & ML Data Serving
Build a feature store for the team's ML models: ranking features for search, ETA correction, traffic anomaly detection.
Design offline (batch) and online (low-latency) feature pipelines, ensuring training/serving consistency.
Manage data lineage, feature versioning and audit trails.
Requirements
3–5 years of Data Engineering experience; experience with large-scale, production-grade data systems preferred.
Bachelor's degree (good grade or above) in IT, Electronics & Telecommunications, Applied Mathematics or equivalent (top universities preferred: HUST, University of Engineering and Technology, PTIT, University of Science).
Proficient in Python and/or Scala/Java for data engineering; strong understanding of distributed data processing.
Experience with distributed data-processing frameworks: Apache Spark, Flink or equivalent; solid grasp of batch and stream processing.
Experience with Kafka/Kinesis or message-streaming platforms; building real-time ingestion pipelines.
Experience designing data warehouse/lakehouse: data modeling, partitioning strategy, query optimization on Redshift/BigQuery/Snowflake or equivalent.
Experience with workflow orchestration: Apache Airflow or equivalent; managing DAGs, retry logic and alerting.
Advanced SQL experience: window functions, CTEs, query-plan optimization on large datasets.
Data-quality ownership mindset: proactively building validation, monitoring and alerting for pipelines rather than waiting for downstream to catch errors.
Preferred
Experience with geospatial data: PostGIS, GeoPandas, S2/H3/Geohash, spatial indexing.
Mapping/mobility domain experience: road-network data, GPS-trace processing, map tiles, traffic data. Experience with dbt or equivalent for data transformation, lineage and documentation.
Feature-store experience (Feast, Tecton or in-house): offline/online serving, versioning. Experience building data lakehouses: Delta Lake, Apache Iceberg, Hudi — ACID transactions, time travel, schema evolution. Experience with Kubernetes and cloud (EKS/GKE), containerizing data workloads, autoscaling.
Understanding of ad-tech data: impression/click pipelines, attribution, budget reporting.
Experience with BI tools (Superset, Metabase, Looker) or building reporting APIs. Experience with A/B testing data infrastructure and experimentation platforms.
Benefits
- Income competitive with the market.
- Lunch allowance.
- Preferential rates across the Group's ecosystem: tuition discounts (Vinschool), healthcare (Vinmec), resorts (Vinpearl), vehicle purchase (VinFast), and home rental or purchase (Vinhomes) … under the Group's policies.
- Full insurance coverage as required by the Labor Law (Social, Health, UI), plus Company-provided personal health insurance based on position level, and periodic health check-ups at reputable hospitals and health centers nationwide.
- Access to strategic, large-scale key technology projects.
- The opportunity to work in a professional technology environment that brings together scientists, experts and engineers from leading technology companies in Vietnam and worldwide.
- Free learning resources on Udemy, Coursera and O'Reilly; internal workshops; certification sponsorship; and special mentorship programs from the Group's and Company's leadership.
- The chance to join the Group's technology clubs and internal tech events to learn and turn personal projects and ideas into reality.
- Training programs to become an "Internal Trainer" and share expertise, with special benefits.
- 12 annual leave days, plus public holidays and Tết as regulated by law.
Working Hours
- 05 official working days at the office (Monday – Friday).
- 02 remote working days per month on Saturdays on a rotating schedule.
- Flexible working hours with check-in window from 08:30 – 09:30.
- Proactively manage time to complete 08 working hours/day.
Work Location
- Hanoi
- Ho Chi Minh City
Check application status
Enter the email you used when submitting your CV — we'll send a verification code to that inbox to protect your information.
We email you a verification code so only you can view your application status.