Senior Solutions Architect, First Time Deployment Validation - NVIS
Before you apply
Listed location: US, CA, Santa Clara | US, TX, Remote | US, VA, Remote | US, WA, Remote
Work arrangement: remote. A remote label does not confirm worldwide eligibility or visa sponsorship.
Read the employer’s description for qualifications, compensation and work eligibility. Confirm the position is still open on the application page.
Job description supplied by NVIDIA; category and skill labels may be inferred. How our listings work · Report a problem
Job Description
The First Time Deployment Team owns first-time execution of NVIDIA's latest products and systems; gathering install and bring-up evidence, operationalizing the validation process, documenting blockers and finding solutions to launch AI Factories at scale. Our results are spread across NVIDIA so we can succeed at scale.
We're looking for an ambitious Senior Solutions Architect to drive validation of NVIDIA AI factories from first rack power-on through customer handoff. You will be embedded in launches from the start, running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters using NCCL and collectives (AllReduce, AllToAll) to validate performance and scalability. When workloads or benchmarks fall short, you're the expert who digs in, partners with engineering, and drives resolution. You will operationalize observability and automation to accelerate validation, capture structured evidence across every bring-up milestone, and work directly with internal deployment teams and external customers to ensure AI factories are ready at launch. Your work directly enables the success of NVIDIA's first external product launches!
What You Will be Doing:
Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.
Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.
Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.
Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks
Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.
Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.
What We Need to See:
Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.
More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.
Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.
Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.
Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.
Proficiency with Python and Shell/Bash for scripting, automation, and tooling.
Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).
Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.
Strong communication skills and the ability to work effectively with cross-functional teams.
Ways to Stand Out From the Crowd:
Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.
Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.
Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.
Experience building automation and CI-style pipelines for running and validating benchmarks at scale.
Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.
NVIDIA is widely considered one of the technology world’s most desirable employers. Some of the world's most forward-thinking and hardworking people are working for us. If you're creative and autonomous, we want to hear from you.
#BuildTheAIFactory
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Skills mentioned
Frequently asked questions
Is the Senior Solutions Architect, First Time Deployment Validation - NVIS position at NVIDIA remote?
Yes. The Senior Solutions Architect, First Time Deployment Validation - NVIS role at NVIDIA is a remote position. Country eligibility is not specified here; check the employer listing.
What type of employment is the Senior Solutions Architect, First Time Deployment Validation - NVIS role?
NVIDIA is hiring for a full-time Senior Solutions Architect, First Time Deployment Validation - NVIS position.
Which skills are mentioned for the Senior Solutions Architect, First Time Deployment Validation - NVIS job at NVIDIA?
Detected skill labels include Python, PyTorch, TensorFlow, LLM, Distributed Systems, GPU. Check the employer description to distinguish required skills from preferred experience.
How do I apply for the Senior Solutions Architect, First Time Deployment Validation - NVIS position at NVIDIA?
You can apply for the Senior Solutions Architect, First Time Deployment Validation - NVIS role directly through NVIDIA's official application link provided on this page.
Similar AI jobs
Robotics Research Internship, Humanoid Manipulation (Summer 2027) | PhD Internship
Field AI · internship
Robotics Research Internship, Humanoid Manipulation (Spring 2027) | PhD Internship
Field AI · internship
Staff Engineer, Test Automation (R5792)
Shield AI · fulltime
Staff Technical Program Manager (R5791)
Shield AI · fulltime
Senior Engineer, Full-Stack - Forge Platform, Tools (R5790)
Shield AI · fulltime
Head of Product Design (R5776)
Shield AI · fulltime