R
Data Lakehouse Architect
R
Data Lakehouse Architect
r systems (singapore) pte limited- Posted 6 hours ago
- Be among the first 10 applicants
Job Description
Responsibilities:
- You will be responsible for the end-to-end architecture of the lakehouse platform.
- This includes the design and implementation of data products, data marketplace, knowledge layer and enabling agentic workloads to run out of the lakehouse platform.
- You will also be responsible for quality assurance of the team's delivery in conformance with the Bank-defined software delivery methodology and tools.
- You will partner with other technology functions to help deliver required technology solutions.
Other responsibilities include:
- Provide technical vision and create roadmaps for the lakehouse platform.
- Create the target architecture for an application / set of applications with emphasis on platforms, reusability, scalability and security.
- Create frameworks, technical features which helps in faster operationalization of new patterns such as unstructured content extraction, lambda architecture deployment patterns, retrieval-augmented data patterns, agentic workloads etc.
- Effectively partner with business users to design data contracts, SLA, data quality rules for data products.
- Independently install, customize and integrate software packages and programs.
- Participate in selection of product/tools via RFP/POC.
- Create technical documents (functional/non-functional specification, design specification, training manual) for the solutions. Review design specifications created by development team.
- 10+ years of experience of implementing a Data Lakehouse preferably in FSI domain (using platforms such as Databricks, Snowflake, Cloudera, Huawei, Alibaba, Google Cloud, AWS, Azure)
- Experience in large scale implementations and performance optimizations in the Lakehouse using a. Open Table Formats such as Iceberg, Hudi, Delta Lake,
b. Object Storage including tiered storage (hot, warm, cold data) strategies
c. Data Federation such as Trino, Denodo, Dremio d. Multi modal Query Engines (Hive, Impala, Apache Kudu etc) - Experience in designing MPP and Distributed Compute workloads across on-premise, hybrid and cloud environments.
- Experience in serving agentic workloads using RAG, Embedding strategies, Vector DB, Graph DB, prompt engineering, context management etc.
- Experience in designing optimal hybrid and cloud workloads using private dedicated connectivity (Direct Connect, Express Route etc), workload placement strategy, egress cost optimization, Infrastructure-as-Code.
- Experience in building foundation and business data products and serving them to downstream applications via API, pub-and-sub, generative BI, real-time dashboards, etc and publishing to a data marketplace.
- Knowledge of migrating workloads out of MPP appliances such as Teradata, Greenplum, Netezza using bulk migration strategies, agentic accelerators is a plus.
- Knowledge of containerization, deploying applications to Kubernetes, Openshift using Helm package manager, Kustomize etc is a plus.
- Expertise in integrating applications with Devops too.
Requirement:
More Info
Key Skills
Quality Assurance Procedures
Data Marketing
Snowflake Cloud Data Warehouse
