Job Description
About the role
Gamma's infrastructure needs to be rock-solid for millions of daily users while enabling our engineering teams to ship fast. You'll own the operational health of our full backend platform, building automation and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate. Your work directly impacts every Gamma user's experience.
This is a high-impact role where you'll balance reliability with velocity, knowing when to move fast and when to prioritize stability. You'll lead incident response, drive systemic improvements, and help shape how Gamma scales to serve its next 100 million users.
Our team has a strong in-office culture and works in person 4–5 days per week in San Francisco. We love working together to stay creative and connected, with flexibility to work from home when focus matters most.
What you'll do
Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure
Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact
Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong
Lead incident response and blameless post-mortems, then follow through on the systemic fixes that keep the same issues from coming back
Partner with engineering teams on architecture reviews, SLO and SLI design, and reliability best practices that scale with the product
Manage and optimize our compute, networking, databases, and managed services
What you'll bring
5+ years in site reliability engineering, DevOps, or systems engineering with deep, hands-on AWS expertise
Strong programming skills in Python, Go, or TypeScript/Node.js, applied to building real tools and automation
Solid experience with infrastructure-as-code (Terraform, CloudFormation) and end-to-end observability solutions
Track record of making systems meaningfully more reliable through automation, smarter monitoring, and architectural improvements
Deep understanding of networking, distributed systems, containerization (Docker, Kubernetes), and database performance at scale
Sharp incident management instincts and the debugging skills to navigate complex production failures
Experience scaling SaaS products to millions of users, or background with Kafka, chaos engineering, or service mesh technologies (Nice to have)
AWS certifications, or experience with security and compliance frameworks like SOC 2 or ISO 27001 (Nice to have)
Compensation range:
The base salary for this full-time position, which spans multiple internal levels depending on qualifications, ranges between $230K - $310K plus benefits & equity.
Final offer amounts are determined by multiple factors, including but not limited to experience and expertise in the requirements listed above.
If you're interested in this role but you don't meet every requirement, we encourage you to apply anyway! We're always excited about meeting great people.
Required Skills
Categories
Frequently asked questions
Is the Site Reliability Engineer position at Gamma remote?
The Site Reliability Engineer role at Gamma is an on-site or hybrid position.
What type of employment is the Site Reliability Engineer role?
Gamma is hiring for a full-time Site Reliability Engineer position.
What skills are needed for the Site Reliability Engineer job at Gamma?
Key skills for this role include Python, Kubernetes, Docker, AWS, Go, TypeScript, Kafka, Distributed Systems.
How do I apply for the Site Reliability Engineer position at Gamma?
You can apply for the Site Reliability Engineer role directly through Gamma's official application link provided on this page.
Similar AI jobs
Senior Site Reliability Engineer
Legora · fulltime
Member of Technical Staff - Data Platform
Reflection AI · fulltime
Account Associate- EMEA (French Speaking)
OpenAI · fulltime
Account Associate - EMEA (German Speaking)
OpenAI · fulltime
Account Associate - EMEA
OpenAI · fulltime
Incident Response Lead
Nebius · fulltime