G

Senior Kubernetes Platform Engineer (GPU & AI Infrastructure)

gtn technical staffing Dallas-fort Worth Metroplex
Relocation
Apply
AI Summary

Design and develop Kubernetes-native software to build a next-generation GPU-accelerated compute platform for AI, ML, and HPC workloads. Focus on custom operators, controllers, CRDs, and GPU scheduling while extending Kubernetes capabilities. Requires deep expertise in Kubernetes internals, Go/Python, and GPU infrastructure.

Key Highlights
Build Kubernetes-native software (operators, controllers, CRDs) for GPU-accelerated AI/ML/HPC workloads
Extend Kubernetes to support GPU scheduling, resource allocation, and high-performance networking
Integrate NVIDIA technologies (GPU Operator, MIG, DCGM) and develop internal tools for GPU management
Key Responsibilities
Develop Kubernetes-native software (custom operators, controllers, CRDs, APIs) to orchestrate GPU infrastructure
Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads with scheduling, allocation, and resource isolation
Integrate NVIDIA technologies (GPU Operator, device plugins, MIG, DCGM) and build GPU scheduling capabilities
Automate cluster provisioning, lifecycle management, and infrastructure orchestration for GPU workloads
Develop internal tools and APIs for provisioning and managing GPU compute resources
Improve platform scalability, GPU utilization, workload performance, and reliability
Integrate Kubernetes with high-performance networking (InfiniBand, RDMA, RoCE) and bare-metal infrastructure
Build observability and automated remediation capabilities for distributed GPU environments
Technical Skills Required
Kubernetes (operators, controllers, CRDs, scheduling, RBAC) Go GPU Infrastructure (NVIDIA technologies, Slurm, Volcano, CUDA, NCCL)
Benefits & Perks
Relocation assistance available
Hybrid work arrangement (3 days onsite, 2 days remote)
Full remote may be considered for exceptional candidates
Nice to Have
Experience with Slurm, Volcano, or custom kube-scheduler extensions
Familiarity with CUDA, NCCL, PyTorch, or TensorFlow
Experience with InfiniBand, RDMA, or RoCE
Background in AI infrastructure, HPC, or large-scale distributed systems
Experience with bare-metal Kubernetes or internal developer platforms

Job Description


Senior Kubernetes Platform Developer – GPU & AI Infrastructure

Location: Dallas, TX preferred

Work Arrangement: Hybrid, 3 days onsite / 2 days remote

Remote Flexibility: Full remote may be considered for the right candidate

Relocation: Available

Employment Type: Direct Hire

Overview

We are seeking a Senior Kubernetes Platform Developer to design and build the software powering a next-generation GPU-accelerated compute platform supporting AI, machine learning, LLM, and HPC workloads.

This is a software development role focused on Kubernetes, not a traditional DevOps, SRE, or Kubernetes administration position.

The core focus is developing Kubernetes-native software including custom operators, controllers, CRDs, APIs, scheduling capabilities, and internal platform services used to orchestrate large-scale GPU infrastructure.

The ideal candidate is a strong developer who understands Kubernetes internals and has experience building software on top of Kubernetes, not simply deploying applications or maintaining clusters.

Key Responsibilities

  • Develop Kubernetes-native software using Go, Python, or similar languages.
  • Build custom operators, controllers, CRDs, APIs, and platform services.
  • Extend Kubernetes to support GPU-intensive AI/ML and HPC workloads.
  • Develop automation for cluster provisioning, lifecycle management, scheduling, and infrastructure orchestration.
  • Build GPU scheduling, allocation, workload placement, and resource-isolation capabilities.
  • Integrate NVIDIA technologies including GPU Operator, device plugins, MIG, and DCGM.
  • Develop internal tools and APIs for provisioning and managing GPU compute resources.
  • Improve platform scalability, GPU utilization, workload performance, and reliability.
  • Integrate Kubernetes with high-performance networking, storage, and bare-metal infrastructure.
  • Build observability and automated remediation capabilities for distributed compute environments.

Required Qualifications

  • Strong software development experience with Go, Python, or another modern programming language.
  • Hands-on experience building Kubernetes operators, controllers, CRDs, APIs, or other Kubernetes-native software.
  • Deep understanding of Kubernetes architecture, controllers, reconciliation, scheduling, RBAC, networking, and cluster lifecycle.
  • Experience developing platforms or distributed systems built on Kubernetes.
  • Experience with GPU infrastructure and NVIDIA technologies.
  • Experience supporting AI/ML, LLM, HPC, or other compute-intensive workloads.
  • Strong Linux and distributed systems knowledge.
  • Experience with Terraform, Helm, Kustomize, Argo CD, Flux, or similar tooling.
  • Ability to troubleshoot across Kubernetes, compute, networking, storage, GPUs, and applications.

Preferred Qualifications

  • Experience with NVIDIA GPU clusters.
  • Experience with Slurm, Volcano, kube-scheduler extensions, or custom scheduling.
  • Familiarity with CUDA, NCCL, PyTorch, or TensorFlow.
  • Experience with InfiniBand, RDMA, RoCE, or high-performance networking.
  • Experience with bare-metal Kubernetes.
  • Experience building internal developer platforms or self-service infrastructure.
  • Background in AI infrastructure, HPC, cloud infrastructure, or large-scale distributed systems.

Ideal Candidate

The ideal candidate is a platform developer who builds Kubernetes-native systems.

This person should be comfortable writing operators, controllers, APIs, schedulers, and automation that extend Kubernetes and manage complex GPU infrastructure.

Candidates whose experience is primarily DevOps, CI/CD, Terraform administration, application deployment, or Kubernetes operations without substantial software development experience are unlikely to be the right fit.

Dallas-based candidates are preferred, but full remote may be considered for candidates with exceptional Kubernetes development and GPU infrastructure experience.


Similar Jobs

Explore other opportunities that match your interests

Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

gtn technical staffing

Dallas-fort Worth Metroplex
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Mid-Senior level

SOURCE.ME

Netherlands
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

Onebrief

Namer

Subscribe our newsletter

New Things Will Always Update Regularly