Senior Applied Research Engineer - Accelerator Programming Model and Compiler
Job Description
As the world’s leading accelerated computing company, we are paving the way with innovations in self-driving cars, robotics, machine learning, super-computing, and visualization. We are building a team that will truly change the world and would love for you to join us. We are seeking a Sr Applied Research Engineer, Accelerator Programming Model and Compiler to work on the next generation programming model for NVIDIA’s Programmable Vision Accelerator (PVA) built for developers and AI coding agents. The PVA is a power-efficient, deterministic VLIW/SIMD processor used across NVIDIA DRIVE, Jetson, and IGX platforms.
This role sits at the boundary between research and production systems software. You will explore and innovate new methods for expressing PVA workloads through higher-level declarative abstractions, domain-specific languages, and compiler/runtime interfaces. These interfaces will provide well-defined targets for both developer-written and AI-generated code. The successful candidate will be comfortable moving across levels of abstraction: shaping public APIs, debugging code generation deep in an LLVM-based backend, and researching higher-level programming models. They will also integrate the compiler with agentic coding harnesses, evaluate AI-generated code for correctness and performance, and improve the compiler-feedback loops to help agents produce optimized implementations.
What you will be doing:
Work on the next-generation PVA programming model for developers and AI coding agents, making it easier to build optimized algorithms for PVA.
Enable AI agent–driven PVA development by abstracting hardware-specific details into agent accessible, declarative interfaces, supported by improvements to compile time, emulator speed, and diagnostic tooling.
Help define the architecture, and feature set of the PVA SDK, runtime APIs and programming model.
Define benchmarks and evaluations for agent-generated PVA code and use them to improve the compiler optimizations and diagnostics.
Develop and actively improve the LLVM-based VPU compiler backend targeting VLIW/SIMD architecture.
Research and define efficient integration models for PVA workloads in CUDA-based heterogeneous pipelines, including execution, memory movement and synchronization.
Drive deep technical integrations with internal and external customers to improve adoption and shape PVA runtime API for real-world workloads.
What we need to see:
BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
10+ years of experience building high-performance, low-level systems software, accelerator software, embedded software, or compiler/toolchain infrastructure.
Experience building higher-level programming abstractions, DSLs or compiler IRs that abstract hardware complexity while preserving performance.
Experience with compiler, debugger, linker, or toolchain development, particularly using LLVM.
Experience developing code with agents — Claude Code, Cursor, Open AI and inference SDKs.
Background integrating compilers and developer tools with AI coding agents or agentic development harnesses.
Experience programming SIMD/VLIW processors.
Excellent software development skills in C++, including low-level debugging and performance profiling.
Experience with Linux or QNX development environments.
Strong communication and social skills.
Ways to stand out from the crowd:
Experience with CUDA, especially integrating accelerators into a CUDA-based heterogeneous compute pipeline.
Background with OpenCL, MLIR, Halide, TVM, Triton, graph compilers, image-processing DSLs, or other declarative/compiler-based programming systems.
Familiarity with ISO 26262 and IEC 61508 or equivalent quality/safety standards.
Experience building agent harnesses and orchestrating long-running, stateful, multi-agent workflows
With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most hard-working and talented people in the world working for us. If you're creative and passionate about developing cloud services we want to hear from you!
Required Skills
Categories
Frequently asked questions
Is the Senior Applied Research Engineer - Accelerator Programming Model and Compiler position at NVIDIA remote?
The Senior Applied Research Engineer - Accelerator Programming Model and Compiler role at NVIDIA is an on-site or hybrid position.
What type of employment is the Senior Applied Research Engineer - Accelerator Programming Model and Compiler role?
NVIDIA is hiring for a full-time Senior Applied Research Engineer - Accelerator Programming Model and Compiler position.
What skills are needed for the Senior Applied Research Engineer - Accelerator Programming Model and Compiler job at NVIDIA?
Key skills for this role include CUDA, Triton.
How do I apply for the Senior Applied Research Engineer - Accelerator Programming Model and Compiler position at NVIDIA?
You can apply for the Senior Applied Research Engineer - Accelerator Programming Model and Compiler role directly through NVIDIA's official application link provided on this page.
Similar AI jobs
Software Engineer II, Vehicle Platform Integrations
Aurora · fulltime
Staff ML Engineer, Agent Training & Environments
Labelbox · fulltime
Research Engineer - Agent Memory
Mem0 · fulltime
Chief Engineer, Autonomous Flight
Anduril · fulltime
Staff Product Manager, Applied AI
Apptronik · fulltime
Copy of Senior Electrical Engineer – Power Systems
Saronic · fulltime