Before you apply
Listed location: New York, NY
Work arrangement: onsite. A remote label does not confirm worldwide eligibility or visa sponsorship.
Read the employer’s description for qualifications, compensation and work eligibility. Confirm the position is still open on the application page.
Job description supplied by Fireworks AI; category and skill labels may be inferred. How our listings work · Report a problem
Job Description
About Us:
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
The Role:
As a Training Infrastructure Engineer, you'll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform. You'll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems.
Key Responsibilities:
- Architect and build scalable, resilient backend infrastructure to support distributed training, inference, and data processing pipelines
- Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems
- Design and implement core backend services with a focus on efficiency and low latency
- Drive infrastructure optimization initiatives for compute cost, storage lifecycle management, and network performance
- Collaborate with machine learning, DevOps, and product teams to translate research and product requirements into robust infrastructure solutions
- Evaluate and integrate cloud-native and open-source technologies such as Kubernetes, Ray, Kubeflow, and MLFlow to enhance platform reliability
- Own end-to-end systems from design to deployment, emphasizing reliability, fault tolerance, and operational excellence
Minimum Qualifications:
- Bachelor's degree or equivalent in Computer Science or related field plus four (4) years of experience in software engineering or related role
- 4 years of experience designing, building, and optimizing large-scale backend infrastructure and distributed data systems (e.g., PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka) in cloud environments (AWS, GCP, Azure, or equivalent), including cloud-native platforms, core infrastructure components, and optimization techniques (caching, indexing, sharding, replication, transactions, ACID)
- 4 years of experience with major server-side programming languages and frameworks (e.g., Python, C++, Go, TypeScript)
- 4 years of experience writing technical design documentation, leading cross-functional projects, and collaborating with cross-functional teams to achieve business impact
- 3 years of experience developing and maintaining data processing and API systems, including client-server communication frameworks (e.g., gRPC, Thrift)
- 3 years of experience conducting A/B testing and scientific experimentation (e.g., Statsig, Meta Deltoid, Optimizely) to measure software impact
- 3 years of experience conducting coding interviews and providing systematic feedback for engineering candidates
- 2 years of experience with cloud-native tools and infrastructure, such as Docker and Kubernetes
- 2 years of experience defining and implementing data-driven metrics to support company or team goals
How to Apply: Submit resume and apply online at http://www.fireworks.ai/careers and search for job by title.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
Total compensation for this role also includes meaningful equity in a fast-growing startup, along with a competitive salary and comprehensive benefits package. Base salary is determined by a range of factors including individual qualifications, experience, skills, interview performance, market data, and work location. The listed salary range is intended as a guideline and may be adjusted.
Why Fireworks AI?
- Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
- Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
- Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
- Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
Skills mentioned
Categories
Frequently asked questions
Is the Member of Technical Staff position at Fireworks AI remote?
The Member of Technical Staff role at Fireworks AI does not have a confirmed remote arrangement in our data. Check the employer description for its work location.
What type of employment is the Member of Technical Staff role?
Fireworks AI is hiring for a full-time Member of Technical Staff position.
Which skills are mentioned for the Member of Technical Staff job at Fireworks AI?
Detected skill labels include Python, PyTorch, Kubernetes, Docker, AWS, GCP, Azure, Go. Check the employer description to distinguish required skills from preferred experience.
How do I apply for the Member of Technical Staff position at Fireworks AI?
You can apply for the Member of Technical Staff role directly through Fireworks AI's official application link provided on this page.
Similar AI jobs
Senior Staff Product Manager, Platform Integrations & Partnerships
Twelve Labs · fulltime
Principal Engineer, AI Security
Lila Sciences · fulltime
Senior ML Systems Engineer, Inference
Runpod · fulltime
L3 Support Engineer, Data Center Infrastructure
Nebius · fulltime
L3 Support Engineer, Data Center Infrastructure
Nebius · fulltime
L3 Support Engineer, Data Center Infrastructure
Nebius · fulltime