R

Senior Systems Engineer (AI Training Infrastructure)

river ai United State
Visa Sponsorship Relocation
Apply
AI Summary

Build and optimize high-performance, distributed AI training systems at River AI, focusing on scalable GPU kernels, fault-tolerant clusters, and end-to-end performance optimization. Collaborate with researchers to accelerate model development while ensuring reliability and efficiency across thousands of nodes. Requires deep expertise in systems programming, computer architecture, and distributed systems.

Key Highlights
Architect and deploy fault-tolerant distributed systems for AI training and inference across large-scale clusters
Design high-performance GPU kernels and optimize tensor operations, memory, and networking (InfiniBand/RDMA)
Partner directly with research scientists to implement, optimize, and scale experimental model architectures
Key Responsibilities
Design and deploy fault-tolerant distributed systems for AI training and inference workloads across clusters with thousands of nodes
Develop high-performance GPU kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA
Profile and optimize systems end-to-end to resolve bottlenecks in hardware, software, data loading pipelines, and collective communication primitives
Collaborate directly with research scientists to implement, optimize, and scale experimental model architectures
Technical Skills Required
C/C++ Distributed Systems Computer Architecture
Benefits & Perks
Comprehensive health, dental, and vision insurance
Unlimited paid time off (PTO)
Relocation assistance as needed
Nice to Have
Experience with modern AI frameworks (e.g., PyTorch, JAX)
Deep familiarity with modern GPU architectures (NVIDIA/AMD) and hardware constraints (HBM, PCIe)
Proven track record of shipping and maintaining high-performance distributed systems or low-level software libraries

Job Description


At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About The Role

We are looking for exceptional systems engineers to build the high-performance engines that train our models. Your goal is to make training at River fast, reliable, and massively scalable.

You will take ownership of our core infrastructure stack; from writing custom GPU kernels to managing clusters of thousands of nodes, ensuring our researchers can focus on science rather than system bottlenecks.

What You’ll Do

  • Architect and deploy fault-tolerant distributed systems for training and inference workloads across clusters with thousands of nodes.
  • Design high-performance kernels to maximize tensor operation efficiency, memory throughput, and networking over InfiniBand/RDMA.
  • Profile systems end-to-end to resolve blockers across hardware, software, data loading pipelines, and collective communication primitives.
  • Partner directly with research scientists to rapidly implement, optimize, and scale experimental model architectures.

Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical industry experience.
  • Deep expertise in systems-level languages (C, C++, or Rust) with a track record of writing performant, maintainable code.
  • Strong foundation in computer architecture, memory management, and concurrent programming.
  • Exceptional debugging skills, especially when tackling complex, non-deterministic issues in distributed environments.
  • A highly collaborative mindset and a bias for action to push boundaries across the stack.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Hands-on experience with modern AI frameworks (e.g., PyTorch, JAX) and tooling for large-scale model training.
  • Deep familiarity with modern GPU architectures (NVIDIA/AMD) and hardware constraints (HBM bandwidth, PCIe limits).
  • A proven track record of shipping and maintaining high-performance distributed systems or low-level software libraries.

Logistics & Benefits

  • Location: Palo Alto, California.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.

Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

river ai

United State

Senior Staff Software Engineer (AI Platform) – AI/ML & Autonomous Systems

Programming
37m ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Palo Alto Networks

United State
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

intellibee inc

United State

Subscribe our newsletter

New Things Will Always Update Regularly