Senior Engineer – Distributed Training Systems & Performance
Eternity Quests LLP- Posted 6 days ago
- Be among the first 10 applicants
Job Description
Senior Engineer – Distributed Training Systems & Performance
Squeeze every FLOP out of a very large cluster. Sharding, numerics, checkpointing, MFU.
Location: Delhi NCR or Bengaluru | Experience: 4 – 12 years
About the opportunity
We are retained by a well-funded AI lab building foundation models from the ground up in India — real compute, real pretraining, not fine-tuning on top of someone else's model. The team is small and the work is unusually direct: what you build ships into the model.
The client's identity is confidential at this stage and will be shared on first call.
What you will do
• Optimise core training infrastructure across the JAX/XLA stack.
• Design tensor, pipeline and data sharding strategies across large TPU/GPU clusters.
• Implement FP8 numerics and architect reliable high-speed checkpointing.
• Profile workloads continuously to maximise Model FLOPs Utilisation.
What we are looking for
• 4+ years in distributed systems, HPC, or ML infrastructure.
• Deep expertise in parallel computing and hardware accelerators.
• Low-level stack optimisation experience.
What is on offer
• Compensation is open and benchmarked to the top of the Indian market for this profile, with meaningful equity.
• Delhi NCR is the primary base, but the client is flexible on Bengaluru for the right person.
• Relocation support, including for candidates returning from outside India.
How to apply
Apply here, or write to [Confidential Information] with your CV and links to anything you have published, shipped or open-sourced. We reply to every serious application within 48 hours.
Spotlight
- Joining bonus, Performance bonus, Mobile bill reimbursements, Relocation benefits, Retirement benefits, Health insurance
Bachelor Of Technology (B.Tech/B.E), Masters in Technology (M.Tech/M.E), Master of Science (MS/M.Sc), PhD, Bachelor in Management Studies (B.M.S.)
More Info
Key Skills
XLA
Model Parallelism
Tensor Parallelism
Pipeline Parallelism
FSDP
NCCL
GPU Optimization
Kernel Optimization
Mixed Precision
FP8
Checkpointing
MFU
ML Infrastructure
AI/ML
