Search Jobs

Search by job, company or skills

Lead Enterprise Lakehouse Architect Data Products & Agentic AI- Contract

Lead Enterprise Lakehouse Architect Data Products & Agentic AI- Contract

ntt singapore pte. ltd.
10-13 Years
SGD 10,000 - 12,000 per month
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

Lead Enterprise Lakehouse Architect - Open Table Formats, Data Products & Agentic AI

Contract Duration: 09 months


Seniority: L4 - More than 10 years of relevant experience


Working Arrangement: Onsite ( 5 days from office )


Headcount: 1

Role Overview

We are seeking an experienced Enterprise Data Lakehouse Architect to own the end-to-end architecture of a large-scale Lakehouse platform supporting governed data products, Data-as-a-Service, real-time analytics, knowledge layers and agentic AI workloads.

This is a senior hands-on architecture position requiring demonstrable production implementation experience. Applicants whose experience is limited to traditional data warehouses, BI reporting, general cloud architecture or data-engineering delivery without end-to-end Lakehouse ownership will not meet the requirements.

Responsibilities

  • Define the technical vision, target architecture and implementation roadmap for an enterprise-scale Lakehouse platform.
  • Architect reusable, scalable and secure platform components across on-premises, hybrid and cloud environments.
  • Design and implement Bronze, Silver and Gold medallion layers using Delta Lake, Apache Iceberg or Apache Hudi.
  • Design object-storage architecture covering lifecycle management and hot, warm and cold data-tiering strategies.
  • Architect MPP and distributed-compute workloads using Spark, Databricks, BigQuery, Dataproc, EMR, Synapse or equivalent platforms.
  • Establish foundation and business data products with formal data contracts, SLAs, ownership, lineage and data-quality rules.
  • Serve governed data products to downstream applications through REST APIs, Kafka/Pub-Sub, real-time streams, dashboards and data-marketplace capabilities.
  • Design reusable patterns for structured and unstructured content ingestion, lambda processing and retrieval-augmented data workloads.
  • Enable RAG and agentic AI workloads using embeddings, vector databases, graph databases, prompt engineering and context-management strategies.
  • Design secure hybrid-cloud connectivity using private dedicated connectivity, workload-placement strategies and data-egress cost controls.
  • Implement Infrastructure-as-Code and automated platform provisioning.
  • Lead platform performance engineering, query optimisation, capacity planning, reliability improvements and FinOps initiatives.
  • Evaluate Lakehouse, federation, query-engine, vector-database and graph-database technologies through RFPs and proofs of concept.
  • Define functional, non-functional, security and solution-design specifications.
  • Review technical designs and delivery outputs for compliance with architecture, engineering, security and quality standards.
  • Integrate the Lakehouse platform with enterprise CI/CD, testing, source-control, monitoring, scheduling and incident-management tools.
  • Lead continuous service-improvement and process-improvement initiatives.

Mandatory Requirements

Applicants must meet all the following requirements:

  • Between 10 and 15 years of relevant experience in enterprise data architecture, big-data platforms and distributed data processing.
  • At least five years of hands-on architecture ownership for enterprise-scale data platforms.
  • Personally architected and implemented at least one production-scale Lakehouse in banking or financial services.
  • Hands-on implementation experience with at least one approved platform:ClouderaHuawei CloudGoogle BigQuery, BigLake, Dataplex or DataprocAWS EMR or OutpostsAzure Synapse or Azure Databricks
  • Production implementation of Bronze, Silver and Gold medallion architecture.
  • Deep hands-on experience with at least one open-table format: Delta Lake, Apache Iceberg or Apache Hudi.
  • Ability to explain ACID transactions, schema evolution, partition evolution, time travel/snapshots, compaction and small-file management.
  • Experience designing distributed Spark/PySpark workloads and performing query, storage and compute optimisation.
  • Production experience implementing both batch and real-time/streaming pipelines.
  • Hands-on Data-as-a-Service implementation using REST APIs and Kafka/Pub-Sub.
  • Experience building reusable foundation and business data products supported by data contracts, SLAs and automated data-quality controls.
  • Experience publishing governed data products through a catalogue, exchange or data marketplace.
  • Experience with enterprise object storage and hot, warm and cold lifecycle strategies.
  • Experience implementing metadata management, data lineage, RBAC, audit logging and fine-grained access controls.
  • Production experience enabling RAG workloads using embeddings and a vector database.
  • Practical knowledge of graph databases, prompt engineering, context management and LLM governance.
  • Experience designing hybrid-cloud platforms, private connectivity, workload placement and egress-cost optimisation.
  • Hands-on Infrastructure-as-Code experience using Terraform, CloudFormation or ARM/Bicep.
  • Strong CI/CD implementation experience using Jenkins, Azure DevOps, Cloud Build, GitHub Actions or equivalent.
  • Experience with platform monitoring, incident management, performance engineering and continuous service improvement.
  • Ability to work onsite at IH2, Malaysia throughout the 12-month assignment.

Mandatory Certifications

Applicants must possess at least two current professional certifications, including:

  • One professional-level cloud architecture or data-engineering certification from Google Cloud, AWS or Microsoft Azure and
  • One Databricks Data Engineer Professional, Databricks Data Architect, CDMP or equivalent data-platform certification.

Associate-level training badges or course-completion certificates alone will not satisfy this requirement.

Preferred Experience

  • Trino, Denodo or Dremio data federation.
  • Hive, Impala or Apache Kudu query engines.
  • Migration from Teradata, Greenplum or Netezza into a modern Lakehouse.
  • Databricks Vector Search, Azure AI Search, Pinecone, Weaviate, ChromaDB or Snowflake Cortex.
  • Neo4j, JanusGraph, TigerGraph, Amazon Neptune or Stardog.
  • LangGraph, OpenAI Agents SDK, Microsoft Agent Framework, LlamaIndex Workflows or Google ADK.
  • Kubernetes or OpenShift deployment using Helm or Kustomize.
  • Banking regulatory requirements and controls covering MAS, BCBS 239, AML, data residency and auditability.

Interested candidates are kindly requested to email their CV with their experience to [Confidential Information]

We look forward to your application!

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Apache Iceberg

ARM Bicep

Cloud Build

CI CD

GitHub Actions

Delta Lake

Apache Hudi

Infrastructure-as-Code