Lead Enterprise Lakehouse Architect Data Products & Agentic AI- Contract
ntt singapore pte. ltd.- Posted 4 hours ago
- Be among the first 10 applicants
Job Description
Lead Enterprise Lakehouse Architect - Open Table Formats, Data Products & Agentic AI
Contract Duration: 09 months
Seniority: L4 - More than 10 years of relevant experience
Working Arrangement: Onsite ( 5 days from office )
Headcount: 1
Role Overview
We are seeking an experienced Enterprise Data Lakehouse Architect to own the end-to-end architecture of a large-scale Lakehouse platform supporting governed data products, Data-as-a-Service, real-time analytics, knowledge layers and agentic AI workloads.
This is a senior hands-on architecture position requiring demonstrable production implementation experience. Applicants whose experience is limited to traditional data warehouses, BI reporting, general cloud architecture or data-engineering delivery without end-to-end Lakehouse ownership will not meet the requirements.
Responsibilities
- Define the technical vision, target architecture and implementation roadmap for an enterprise-scale Lakehouse platform.
- Architect reusable, scalable and secure platform components across on-premises, hybrid and cloud environments.
- Design and implement Bronze, Silver and Gold medallion layers using Delta Lake, Apache Iceberg or Apache Hudi.
- Design object-storage architecture covering lifecycle management and hot, warm and cold data-tiering strategies.
- Architect MPP and distributed-compute workloads using Spark, Databricks, BigQuery, Dataproc, EMR, Synapse or equivalent platforms.
- Establish foundation and business data products with formal data contracts, SLAs, ownership, lineage and data-quality rules.
- Serve governed data products to downstream applications through REST APIs, Kafka/Pub-Sub, real-time streams, dashboards and data-marketplace capabilities.
- Design reusable patterns for structured and unstructured content ingestion, lambda processing and retrieval-augmented data workloads.
- Enable RAG and agentic AI workloads using embeddings, vector databases, graph databases, prompt engineering and context-management strategies.
- Design secure hybrid-cloud connectivity using private dedicated connectivity, workload-placement strategies and data-egress cost controls.
- Implement Infrastructure-as-Code and automated platform provisioning.
- Lead platform performance engineering, query optimisation, capacity planning, reliability improvements and FinOps initiatives.
- Evaluate Lakehouse, federation, query-engine, vector-database and graph-database technologies through RFPs and proofs of concept.
- Define functional, non-functional, security and solution-design specifications.
- Review technical designs and delivery outputs for compliance with architecture, engineering, security and quality standards.
- Integrate the Lakehouse platform with enterprise CI/CD, testing, source-control, monitoring, scheduling and incident-management tools.
- Lead continuous service-improvement and process-improvement initiatives.
Mandatory Requirements
Applicants must meet all the following requirements:
- Between 10 and 15 years of relevant experience in enterprise data architecture, big-data platforms and distributed data processing.
- At least five years of hands-on architecture ownership for enterprise-scale data platforms.
- Personally architected and implemented at least one production-scale Lakehouse in banking or financial services.
- Hands-on implementation experience with at least one approved platform:ClouderaHuawei CloudGoogle BigQuery, BigLake, Dataplex or DataprocAWS EMR or OutpostsAzure Synapse or Azure Databricks
- Production implementation of Bronze, Silver and Gold medallion architecture.
- Deep hands-on experience with at least one open-table format: Delta Lake, Apache Iceberg or Apache Hudi.
- Ability to explain ACID transactions, schema evolution, partition evolution, time travel/snapshots, compaction and small-file management.
- Experience designing distributed Spark/PySpark workloads and performing query, storage and compute optimisation.
- Production experience implementing both batch and real-time/streaming pipelines.
- Hands-on Data-as-a-Service implementation using REST APIs and Kafka/Pub-Sub.
- Experience building reusable foundation and business data products supported by data contracts, SLAs and automated data-quality controls.
- Experience publishing governed data products through a catalogue, exchange or data marketplace.
- Experience with enterprise object storage and hot, warm and cold lifecycle strategies.
- Experience implementing metadata management, data lineage, RBAC, audit logging and fine-grained access controls.
- Production experience enabling RAG workloads using embeddings and a vector database.
- Practical knowledge of graph databases, prompt engineering, context management and LLM governance.
- Experience designing hybrid-cloud platforms, private connectivity, workload placement and egress-cost optimisation.
- Hands-on Infrastructure-as-Code experience using Terraform, CloudFormation or ARM/Bicep.
- Strong CI/CD implementation experience using Jenkins, Azure DevOps, Cloud Build, GitHub Actions or equivalent.
- Experience with platform monitoring, incident management, performance engineering and continuous service improvement.
- Ability to work onsite at IH2, Malaysia throughout the 12-month assignment.
Mandatory Certifications
Applicants must possess at least two current professional certifications, including:
- One professional-level cloud architecture or data-engineering certification from Google Cloud, AWS or Microsoft Azure and
- One Databricks Data Engineer Professional, Databricks Data Architect, CDMP or equivalent data-platform certification.
Associate-level training badges or course-completion certificates alone will not satisfy this requirement.
Preferred Experience
- Trino, Denodo or Dremio data federation.
- Hive, Impala or Apache Kudu query engines.
- Migration from Teradata, Greenplum or Netezza into a modern Lakehouse.
- Databricks Vector Search, Azure AI Search, Pinecone, Weaviate, ChromaDB or Snowflake Cortex.
- Neo4j, JanusGraph, TigerGraph, Amazon Neptune or Stardog.
- LangGraph, OpenAI Agents SDK, Microsoft Agent Framework, LlamaIndex Workflows or Google ADK.
- Kubernetes or OpenShift deployment using Helm or Kustomize.
- Banking regulatory requirements and controls covering MAS, BCBS 239, AML, data residency and auditability.
Interested candidates are kindly requested to email their CV with their experience to [Confidential Information]
We look forward to your application!
More Info
Key Skills
Apache Iceberg
ARM Bicep
Cloud Build
CI CD
GitHub Actions
Delta Lake
Apache Hudi
Infrastructure-as-Code
