Senior Data Engineer (L3)
Location: India
Experience: 7+ Years
Employment Type: Full-Time
About the Role
We are seeking a highly skilled Senior Data Engineer to join our growing Data Science and Analytics organization. In this role, you will work closely with Data Scientists, Machine Learning Engineers, and Analytics teams to build and maintain scalable data infrastructure that powers analytics, experimentation, machine learning models, and business decision-making.
You will be responsible for designing robust data models, developing scalable data pipelines, and ensuring the delivery of high-quality, reliable datasets across the organization. This role requires strong technical expertise in modern data platforms, distributed processing systems, and workflow orchestration.
Key Responsibilities
- Design, develop, and maintain DBT models that deliver trusted datasets, business metrics, and machine learning features.
- Build, optimize, and support scalable data pipelines in Databricks using PySpark and modern lakehouse architectures.
- Transform raw operational data into high-quality, analysis-ready datasets for reporting, experimentation, and machine learning use cases.
- Develop a deep understanding of business processes to create meaningful and scalable data solutions.
- Orchestrate and manage end-to-end workflows using Airflow and Prefect while ensuring reliability and SLA adherence.
- Collaborate closely with Data Scientists and ML Engineers to productionize feature pipelines and support model development.
- Participate in code reviews and promote engineering best practices around testing, documentation, observability, and data quality.
- Optimize DBT, Spark, and orchestration workloads for performance, scalability, and cost efficiency.
- Troubleshoot and resolve complex data platform and pipeline issues.
- Continuously evaluate and adopt modern data engineering practices and emerging technologies.
Required Qualifications
- 5+ years of experience designing, building, and maintaining production-grade data engineering systems.
- Strong expertise in SQL and data modeling.
- Hands-on experience with Databricks and PySpark in production environments.
- Strong experience with DBT for transformation and data modeling.
- Experience building and managing workflow orchestration using Airflow.
- Strong understanding of distributed data processing concepts, including scalability, fault tolerance, consistency, throughput, and latency.
- Experience with modern data lakehouse technologies such as Delta Lake and/or Apache Iceberg.
- Knowledge of Infrastructure as Code (IaC) tools such as Terraform, AWS CDK, or Pulumi.
- Strong software engineering fundamentals, including version control, testing, and CI/CD practices.
- Excellent written and verbal communication skills in English.
- Ability to work independently in a fast-paced and collaborative environment.
Preferred Qualifications
- Experience with Prefect for workflow orchestration.
- Familiarity with streaming technologies such as Kinesis, Kafka, or similar platforms.
- Exposure to MLOps workflows and supporting machine learning teams.
- Experience working with globally distributed engineering teams across multiple time zones.
- Understanding of logistics, supply chain, ecommerce, or operational analytics domains.
- Experience integrating AI or LLM-based capabilities into data platforms, workflows, or internal tools.
- Familiarity with AI-assisted development tools such as Cursor, GitHub Copilot, Claude Code, or similar solutions.
Technical Stack
Core Technologies
- Databricks
- PySpark
- DBT
- Airflow
- SQL
- Delta Lake / Apache Iceberg
Nice to Have
- Prefect
- AWS
- Kinesis
- EMR
- Pulumi
- Sigma Computing
- Terraform
What We're Looking For
- Strong ownership mindset with the ability to drive initiatives independently.
- Passion for solving complex data challenges at scale.
- Ability to learn new technologies quickly and adapt in a fast-changing environment.
- Excellent collaboration skills and willingness to work closely with cross-functional teams.
- Focus on building reliable, scalable, and maintainable data solutions.