Job Description
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About This Role:
At Crusoe, our Production Engineering team plays a pivotal role in ensuring the reliability and performance of our infrastructure. Production Engineering at Crusoe is dedicated to detecting, analyzing, and preventing issues to maintain high Service Level Agreement through Service Level Indicators (SLIs) and Service Level Objectives (SLOs). Through automation and proactive remediation, our Production Engineers not only resolve common errors automatically but also advise various engineering teams in building resilient code. We prioritize anticipating and resolving issues before they impact our customers, conducting thorough post-mortems, and driving continuous improvement. Our customer-centric approach ensures that clients always have access to the virtual machines they depend on. Join us to help build and maintain the robust systems that power Crusoe's innovative solutions.
You Will Thrive In This Role If:
5+ years of professional production engineering experience
5+ years of experience contributing to architecture and design (architecture, design patterns, reliability and scaling) of new and current systems
Bachelor's Degree in Computer Science or related field, or 8+ years relevant work experience
Solid understanding of infrastructure design, including the operational trade-offs of various designs
Experience writing high quality code with at least one programming language (Python, Go, or similar)
Experience building with modern infrastructure tools such as Docker, Kubernetes, Ansible, Cloud Formation, Terraform
Experience building with modern CI/CD practices and build systems, such as GitLab CI/CD, CircleCI, GitHub Actions
Experience with logging, monitoring and alerting systems and tools
Experience with Unix/Linux environments
Experience with TCP/IP and network programming
Experience with information security best practices
Excellent communication skills
Embody the Company values
What You’ll Be Working On:
Capacity Planning & Growth: Collaborate with Engineering and Product teams to analyze network capacity, co-develop growth plans, and contribute to network builds in data centers and Points of Presence (PoPs).
Troubleshooting & Support: Provide Tier 1 troubleshooting for switch provisioning, server build failures related to network issues, and complex network defects related to the physical layer.
Site Builds & Operations: Manage console DNS mapping, troubleshoot connectivity issues, verify automation scripts, conduct network audits, and ensure seamless handoffs to internal customers.
Hardware Management: Serve as the point of contact for network hardware failures, overseeing switch swaps and assisting with other appliance replacements. Manage RMA processes for faulty hardware.
Optical Network Expertise: Troubleshoot Layer 1 issues in the optical network using OTDR tests.
Project Collaboration: Partner with DC and Network Engineering teams on structural cabling for large-scale data center expansions and fabric switch uplifts.
Process Improvement: Identify recurring network issues and lead operational excellence projects to implement effective solutions.
Network Analysis & Optimization: Conduct network studies, analyze link and hardware stability, and recommend improvements.
Innovation & Evaluation: Participate in internal working groups to evaluate and adopt new technologies and methodologies.
What You’ll Bring to the Team:
Technical Foundation: A BS/BA in a technical field or equivalent practical experience.
Networking Expertise: Solid understanding of network fundamentals, including IP subnetting, Layer 2 and Layer 3 networking, VLANs, MAC addresses, port speeds, optics, and routing.
Protocol Proficiency: 2+ years of experience with network protocols such as BGP, MPLS, L3VPN, VPLS, Multicast, CoS, TCP, and IPv4/IPv6.
Industry Experience: 2+ years of experience as a Network Engineer for a content or network provider.
Data Center Architecture: Experience with data center network architectures, including CLOS.
Structured Cabling Skills: Experience with structured cabling in a data center environment (SMF, MMF).
Vendor Management: Experience managing vendors for logistics, infrastructure racking/cabling, and ordering, ensuring adherence to standards.
Bonus Points:
Advanced Education: BS Degree in Computer Science, Engineering, or a related technical discipline.
Scripting Proficiency: 2+ years of programming experience in Python, Perl, or other scripting languages.
Automation Expertise: Experience with manageability infrastructure tools like Ansible, REST/gRPC, and NETCONF.
Troubleshooting Acumen: Experience in network operations, with a systematic approach to troubleshooting.
Network Certifications: Network certifications (CCNP/JNCIS equivalent).
Open Source NOS Experience: Experience working with open-source Network Operating Systems.
Hardware/Software Development: A strong desire to engineer and deploy in-house network hardware and software.
Hardware Platform Experience: Experience testing and deploying new hardware platforms.
Benefits:
Crusoe offers a comprehensive benefits package designed to support well-being and financial security. This includes full social security coverage, contributions to provident, trade union, and pension funds, with options for additional pensions. Employees also have optional access to Global Life Insurance and private health insurance. Crusoe provides generous leave policies, including maternity, paternity, parental, and sick leave, ensuring you have the support you need at every stage of life.
Compensation:
Compensation will be paid as salary or hourly. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.
Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Required Skills
Categories
Frequently asked questions
Is the Staff Production Engineer position at Crusoe remote?
The Staff Production Engineer role at Crusoe is an on-site or hybrid position.
What type of employment is the Staff Production Engineer role?
Crusoe is hiring for a full-time Staff Production Engineer position.
What skills are needed for the Staff Production Engineer job at Crusoe?
Key skills for this role include Python, Kubernetes, Docker, Go, Terraform.
How do I apply for the Staff Production Engineer position at Crusoe?
You can apply for the Staff Production Engineer role directly through Crusoe's official application link provided on this page.
Similar AI jobs
Data Center - Supply Chain Manager (EMEA)
Nebius · fulltime
Revenue Accounting Manager
Baseten · fulltime
Principal Product Manager, Physical AI - Retail & 3PL Applications
Dexterity · fulltime
Senior Software Verification Engineer
NVIDIA · fulltime
Internal Auditor - Operations
NVIDIA · fulltime
Senior Software Engineer, Platforms
NVIDIA · fulltime