E
Senior Engineer – Distributed Training Systems & Performance
E
Senior Engineer – Distributed Training Systems & Performance
Eternity Quests LLPEarly Applicant
Quick Apply
- Posted 4 days ago
- Be among the first 10 applicants
Job Description
- Location: Delhi NCR
- Exp - 4+ years
- Core Mandate: JAX/XLA/Pathways stack, sharding, FP8 numerics, checkpointing, MFU.
- Responsibilities:
- Optimize the core AI training infrastructure utilizing the JAX/XLA and Pathways stack.
- Design efficient tensor, pipeline, and data sharding strategies across massive TPU/GPU clusters.
- Implement FP8 numerics, architect reliable high-speed checkpointing systems, and continuously profile workloads to maximize Model FLOPs Utilization (MFU).
Requirements:
4+ years in distributed systems, High-Performance Computing (HPC), or ML infrastructure. Deep expertise in parallel computing, hardware accelerators, and low-level stack optimization
More Info
Key Skills
XLA
Model Parallelism
Tensor Parallelism
Pipeline Parallelism
FSDP
NCCL
GPU Optimization
Kernel Optimization
Mixed Precision
FP8
Checkpointing
MFU
ML Infrastructure
