Before you apply
Listed location: Heidelberg
Work arrangement: remote. A remote label does not confirm worldwide eligibility or visa sponsorship.
Read the employer’s description for qualifications, compensation and work eligibility. Confirm the position is still open on the application page.
Job description supplied by Aleph Alpha; category and skill labels may be inferred. How our listings work · Report a problem
Job Description
Our Mission
Aleph Alpha is one of the few companies in Europe doing serious foundation model pre-training. Our customers - in finance, manufacturing, public administration - need models that understand German, meet European regulatory requirements, and work reliably in high-stakes settings. We're building that in Heidelberg.
We are hiring a Senior AI R&D Engineer focused on Performance to grow our pre-training efficiency team. If you are excited about making models fast, this is the role for you!
Team Culture
At Aleph Alpha, we foster a culture built on ownership, autonomy, and empowerment. Teams and individual contributors are trusted to take responsibility for their work and drive meaningful impact. We maintain a flat organizational structure with efficient, supportive management that enables quick decision‑making, open communication, and a strong sense of shared purpose.
About the role:
You will engineer the systems required to train foundation models at scale. Your objective is to maximize hardware utilization and training throughput on our large-scale GPU clusters (thousands of NVIDIA Blackwell GPUs). You will work at the intersection of deep learning frameworks, distributed systems, and GPU microarchitecture, eliminating bottlenecks from the Python layer down to the GPU kernel.
This role is for Aleph Alpha Research GmbH.
Your responsibilities:
End-to-End Optimization: Profile training loops using PyTorch Profiler, Nsight Systems and Nsight Compute to identify system- and kernel-level bottlenecks in order to maximize model throughput.
Distributed Strategy and Topology: Configure and tune composite parallelism strategies (e.g. TP, DP, HSDP/FSDP, EP), optimizing load balance, minimizing critical-path bottlenecks, and managing communication-to-computation trade-offs for large-scale LLM training.
Hardware-Aware Modeling: Partner with AI Researchers to define model architectures for hardware efficiency without compromising convergence.
Your Profile
Basic Qualifications
Are proficient in Python and the PyTorch library.
Have a strong engineering background in parallel and/or distributed systems with proven track record of excellence.
Have hands-on experience with modern machine learning techniques (especially large language models and their life cycle).
Deeply understand the CUDA programming model.
Have experience in distributed programming with APIs like NCCL or MPI.
Have experience analysing profiling traces with tools such as PyTorch Profiler and Nvidia Nsight.
Please note this role requires regular on-site collaboration in Heidelberg as a member of the Training Efficiency Team.
Preferred Qualifications
Contributions to modern distributed training frameworks (e.g., TorchTitan, Megatron-LM, DeepSpeed).
Familiarity with low-precision training formats (MXFP4, MXFP8) and their impact on numerical stability and throughput.
A deep understanding of NCCL communication primitives, NVSHMEM or CUDA IPC and their performance.
A proven track record of implementing and optimising modern transformer-based model training.
A proven track record working on the NVIDIA Blackwell architecture.
Compensation and Benefits
Become part of an AI revolution!
30 days of paid vacation
Access to a variety of fitness & wellness offerings via Wellhub
Mental health support through nilo.health
Substantially subsidized company pension plan for your future security
Subsidized Germany-wide transportation ticket
Budget for additional technical equipment
Flexible working hours for better work-life balance and hybrid working model
JobRad® Bike Lease

Skills mentioned
Categories
Frequently asked questions
Is the Senior AI R&D Engineer - Performance - Pre-training (f/m/d) position at Aleph Alpha remote?
Yes. The Senior AI R&D Engineer - Performance - Pre-training (f/m/d) role at Aleph Alpha is a remote position. Country eligibility is not specified here; check the employer listing.
What type of employment is the Senior AI R&D Engineer - Performance - Pre-training (f/m/d) role?
Aleph Alpha is hiring for a full-time Senior AI R&D Engineer - Performance - Pre-training (f/m/d) position.
Which skills are mentioned for the Senior AI R&D Engineer - Performance - Pre-training (f/m/d) job at Aleph Alpha?
Detected skill labels include Python, PyTorch, CUDA, LLM, Distributed Systems, GPU. Check the employer description to distinguish required skills from preferred experience.
How do I apply for the Senior AI R&D Engineer - Performance - Pre-training (f/m/d) position at Aleph Alpha?
You can apply for the Senior AI R&D Engineer - Performance - Pre-training (f/m/d) role directly through Aleph Alpha's official application link provided on this page.
Similar AI jobs
Strategic Account Manager
Agility Robotics · fulltime
Senior Product Marketing Manager, Risk Advisory
Fieldguide · fulltime
Software Engineer II, Mission Interface
Torc Robotics · fulltime
Research Engineer, Index Intelligence
Exa · fulltime
Research Engineer, Content Understanding
Exa · fulltime
2027 Summer Intern, MS/PhD, Machine Learning, Driver Refinement Foundations
Waymo · fulltime