Job Description
Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.
Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
As a Data Engineer, you'll architect and build the data infrastructure that powers everything we do—from crawling billions of pages to training our embedding models to serving real-time search. You'll have enormous autonomy in designing systems that scale to hundreds of petabytes. If you've ever wanted to build data pipelines at a scale that most companies only dream about, this is your chance.
Who You Are
Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them
Experience building and operating large-scale distributed data processing pipelines
Hands-on experience with streaming data systems (Kafka, Flink, or similar)
Familiarity with Ray, Spark, or ClickHouse at production scale
An obsessive focus on reliability and building systems that don't page you at 3am
Bonus
Experience with Lance or other vector-native storage formats
Background in GPU-accelerated data processing (RAPIDS, cuDF)
What You Could Do
Design a lakehouse architecture that handles 100+ PB of web crawl data
Build streaming pipelines that process billions of documents per day for real-time indexing
Architect the data layer for our embedding training infrastructure on Ray
Scale our ClickHouse deployment to handle analytical queries across petabytes of search logs
Logistics
Location: This is an in-person opportunity in San Francisco.
Visas: We're happy to sponsor international candidates (e.g., STEM OPT, OPT, H1B, O1, E3). While we cannot guarantee your visa, we have historically been successful in sponsoring candidates from all over the world. If you receive an offer, our team will work hard to get you a visa.
Benefits: We offer premium healthcare benefits (medical, dental, vision), fertility benefits, 16 weeks of fully paid parental leave for all new parents, and a monthly wellness stipend to all of our employees.
Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.
Required Skills
Categories
Frequently asked questions
Is the Software Engineer, Distributed Data Systems position at Exa remote?
The Software Engineer, Distributed Data Systems role at Exa is an on-site or hybrid position.
What type of employment is the Software Engineer, Distributed Data Systems role?
Exa is hiring for a full-time Software Engineer, Distributed Data Systems position.
What skills are needed for the Software Engineer, Distributed Data Systems job at Exa?
Key skills for this role include Spark, Kafka, Ray, GPU.
How do I apply for the Software Engineer, Distributed Data Systems position at Exa?
You can apply for the Software Engineer, Distributed Data Systems role directly through Exa's official application link provided on this page.
Similar AI jobs
Staff+ Software Engineer, Platform Ecosystem
Anthropic · fulltime
Product ML Engineer
Photoroom · fulltime
Senior Software Engineer, Air AI
Commure · fulltime
Marketing Operations Manager
Baseten · fulltime
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
NVIDIA · fulltime
Software Engineer, Compute Infrastructure
OpenAI · fulltime