Responsibilities:
- Own and deliver end-to-end Data Migration and Data Engineering projects, independently managing all phases from Source Assessment and Data Profiling to Production deployment and post-production support.
- Assess and analyze source systems, including Source Assessment, Data Profiling, and Data Mapping, to identify data structures, dependencies, and potential migration risks.
- Design Data Architecture and Data Models aligned with business requirements, scalability, performance, and cost considerations.
- Develop and maintain Batch and Real-time/Streaming Data Pipelines using appropriate technologies and cloud platforms.
- Execute data migration across On-premise, AWS, and GCP environments, ensuring data integrity, security, and minimal disruption.
- Perform Data Quality, Validation, and Reconciliation to ensure data accuracy, completeness, and traceability.
- Plan and execute Cutover and Rollback strategies to minimize downtime and migration risks.
- Deploy, monitor, troubleshoot, and provide post-production support for data pipelines and platforms.
- Take full ownership of assigned projects with minimal dependency on other teams.
- Design solution architecture, data models, and migration strategies covering Full Load, Incremental/CDC, batch, and real-time/streaming workloads.
- Own the full delivery process single-handedly: source profiling, data mapping, schema conversion, pipeline development, orchestration, testing, data quality/reconciliation, performance tuning, deployment, monitoring, cutover, rollback, and hypercare.
- Lead and execute large-scale data migration projects (e.g., on-premise to GCP/AWS, database replatforming to AlloyDB/BigQuery) with data validation and zero/minimal-downtime cutover.
- Build and operate data pipelines and platforms on GCP and AWS using AlloyDB/PostgreSQL, BigQuery, Cloud Storage/S3, Airflow/Cloud Composer, Spark, Flink, Kafka, Pub/Sub, and Kubernetes/GKE.
- Design lakehouse/data warehouse solutions using Parquet and Apache Iceberg, including partitioning, clustering, compaction, schema evolution, and data lifecycle management.
- Establish standards for CI/CD, Infrastructure as Code, observability, logging/alerting, data lineage, metadata, IAM, encryption, privacy, and cost optimization; analyze production incidents and root causes.
- Mentor and guide data engineers on best practices; lead technical design and code reviews (Team Lead level).
- Collaborate with data analysts, data scientists, application teams, and business stakeholders to deliver reliable data products.
Technical Leadership & Solution Design:
- Translate business requirements into scalable and practical technical solutions independently.
- Evaluate and select appropriate technologies and tools based on cost, timeline, performance, scalability, and business requirements.
- Serve as a Technical Lead, establishing best practices, development standards, CI/CD processes, and Data Governance frameworks.
- Provide technical guidance and consultation to Data Engineers and continuously improve the team's technical capabilities.
- Conduct Code Reviews, Technical Coaching, knowledge sharing, and technical documentation to ensure consistent engineering standards.
Key Objectives:
- Successfully deliver large-scale Data Migration projects with high data accuracy, completeness, traceability, and minimal downtime.
- Design and build scalable Data Platforms / Lakehouse architectures on GCP and AWS to support growing data volumes, processing requirements, and user demand.
- Build highly reliable, observable, and cost-efficient Data Pipelines, with effective monitoring, alerting, and cost optimization.
- Establish engineering standards and continuously improve the team's capabilities through Code Reviews, Technical Coaching, documentation, and best practices.
- Collaborate closely with Data Analysts, Data Scientists, Application Teams, and Business Users to deliver trusted, high-quality, and business-ready data.
Qualifications:
- Strong proficiency in SQL (window functions, CTE, query optimization, execution plans) and Python (Scala/Java is a plus).
- Hands-on experience with relational/distributed databases, data warehouses, data lake/lakehouse, and messaging/stream processing systems.
- Deep understanding of data migration patterns: CDC, idempotency, exactly-once/at-least-once semantics, data quality, reconciliation, and lineage.
- Able to design and deliver production systems end-to-end independently, including architecture documents, data mapping, test evidence, runbooks, and operational handover.
- Solid understanding of data consistency, fault tolerance, disaster recovery, security, privacy, and performance/cost optimization at scale.
- Able to act as Technical Lead / Team Lead: break down work, estimate effort, review architecture/design/code, coach the team, and coordinate across teams.
- Strong leadership, ownership, and communication skills with business, application, infrastructure, security teams, and vendors; able to make technical decisions under time, quality, risk, and cost constraints.
- Experience with Debezium, Dataflow/Apache Beam, Dataproc, Datastream, dbt, Terraform, Docker, Delta Lake/Hudi, Snowflake/Redshift, or legacy/on-premise systems is a plus.
Technical Skills:
- Cloud Platform: Google Cloud Platform (หลัก), AWS (รอง)
- Data Warehouse / Database: BigQuery, AlloyDB, Cloud SQL, PostgreSQL, Oracle, SQL Server, Redshift
- Data Processing: Apache Spark (PySpark), Apache Beam / Dataflow, Apache Flink, Dataproc
- Orchestration: Apache Airflow / Cloud Composer (Dagster หรือ Prefect เป็น plus)
- Streaming & Messaging: Apache Kafka, Google Pub/Sub, Kinesis / MSK
- Storage & Table Format: Cloud Storage, Amazon S3, Parquet, Avro, ORC, Apache Iceberg, Delta Lake, Apache Hudi
- Container & Infrastructure: Docker, Kubernetes (GKE / EKS), Terraform, Helm
- Programming: Python, SQL (Scala / Java พิจารณาเป็นพิเศษ)
- DevOps & Monitoring: Git, CI/CD Pipeline, Cloud Logging / Monitoring, Prometheus, Grafana
Soft Skills & Working Style:
- Demonstrates strong end-to-end ownership, proactively managing tasks from initiation to completion and effectively handling ad-hoc issues.
- Strong systematic problem-solving and analytical skills, with the ability to identify root causes of complex data-related issues.
- Communicates effectively with both technical teams and business users, translating complex technical concepts into clear, easy-to-understand language.
- Maintains strong documentation discipline, including Technical Design, Data Dictionary, Runbook, and Handover Documents.
- Focuses on code and data quality through Code Reviews, Unit Testing, and Data Validation.
- Able to manage multiple projects simultaneously, prioritize tasks effectively, and accurately estimate timelines.
- At the Team Lead level, capable of planning team workloads, delegating tasks, coaching team members, and providing technical guidance.
- Highly adaptable and eager to learn new technologies through self-directed learning.
Data Security & Compliance:
- Good understanding of Data Security principles, including Data Masking, Encryption at Rest/In Transit, Identity and Access Management (IAM), and Row-Level/Column-Level Security.
- Good understanding of PDPA requirements and the ability to design data pipelines in compliance with data protection regulations.
General Qualifications:
- Strong ability to read and understand technical documentation in English.
- Able to communicate effectively in English when working with technical teams and international vendors.