AI
Engineering
AI Infrastructure Engineer
Appxpertise Inc
Full-Time
Mid-Level
Remote
Posted 2w ago
Job Description
**Role: AI Infrastructure Engineer**
**Location: Remote Role**
**Duration: 6\-12 Months**
**Interview: MS Teams**
**About the Role**
We're looking for a hybrid AI Infrastructure Engineer to help stand up and operate a new GPU compute cluster — including hands\-on work managing an 18 HGX\-enabled worker node cluster with 144 GPUs, built on NVIDIA InfiniBand and Ethernet switching. This role spans the physical/network fabric layer, broader network operations, and the platform/orchestration layer, making it a good fit for someone who wants ownership from the switch port up through the workload scheduler.
**What you'll be doing:**
*Fabric \& Networking*
· Help manage and operate the InfiniBand XDR / Quantum\-3400 rail\-aligned fabric topology supporting the 18\-node, 144\-GPU HGX cluster
· Administer UFM and NetQ for subnet management, telemetry, and firmware across the fabric
· Support 800G optics, diagnose and resolve link\-flap issues, and manage congestion isolation
· Run NCCL / HPL validation and performance regression testing to confirm cluster health and throughput after changes
· Operate Cumulus SN5600 and Arista border switching for the Ethernet side of the fabric
*Network Operations*
· Drive operational improvements in change management and daily operations
· Manage and operate large\-scale IP network technologies and infrastructure
· Work with peering and datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits
· Monitor and support network health across on\-premises and cloud infrastructure
· Document best practices and build workflow enhancements
*Platform \& Workload Enablement*
· Manage NVAIE and GPU Operator lifecycle, including driver/firmware compatibility matrices across the cluster
· Administer Run:ai for quotas, project structure, and tenant onboarding as teams ramp onto the new nodes
· Configure and support Slurm on carved\-out worker node pools for scheduled batch workloads
· Manage container images and model\-serving dependencies for production and pre\-production use
· Interface with DNA MLOps and support onboarding of first production workloads onto the new cluster
**What we need to see:**
· Hands\-on experience with NVIDIA InfiniBand (Quantum series) and high\-speed Ethernet switching in a GPU cluster environment
· Familiarity with NVIDIA HGX platform architecture and multi\-node GPU networking (rail\-aligned topologies)
· Experience with NCCL/HPL benchmarking or similar GPU cluster validation tooling
· Working knowledge of GPU Operator, driver/firmware lifecycle management, and container\-based workload deployment
· Experience with a workload orchestrator (Slurm, Run:ai, or similar) in an HPC/AI environment
· Experience with large\-scale IP networking, peering/interconnect technologies, and change management processes
· Comfort operating across network fabric, network operations, and platform/software layers. This is a hybrid role, not a pure network or pure platform position
**Nice to have:**
· Experience with UFM, NetQ, or equivalent InfiniBand fabric management tools
· Exposure to MLOps workflows and model\-serving infrastructure
· Prior work on new cluster bring\-up or greenfield GPU infrastructure deployments
· Familiarity with DWDM/wave circuit technologies in a datacenter interconnect context
Pay: From $90\.00 per hour
Expected hours: 40\.0 per week
Work Location: Remote
Get jobs like this in your inbox
Join thousands of digital nomads getting the best remote jobs delivered weekly. Free, no spam.
Similar Jobs
Senior Platform Engineer, GitLab Orbit
GitLab IncGitLab IncFull-Time · RemoteDirect
VueTypeScriptRust
$139k – $235k2w agoView details$139k – $235k2w ago
Senior Software Engineer I, Full Stack
Wpromote, LLCWpromote, LLCFull-Time · RemoteDirect
ReactDjangoPython
$135k – $155k2w agoView details$135k – $155k2w ago
Test Automation Developer
LynxLynxFull-Time · RemoteDirect
PythonDockerCI/CD
$65k – $70k2w agoView details$65k – $70k2w ago
Engineering Director
College BoardCollege BoardFull-Time · RemoteDirect
ReactNode.jsJavaScript
$140k – $175k2w agoView details$140k – $175k2w ago
Forensic Engineer
YA GroupYA GroupFull-Time · RemoteDirect
$80k – $275k2w ago
Director of Product Management, Agentic Software Delivery
GitLab IncGitLab IncFull-Time · RemoteDirect
$203k – $346k2w ago
Remote Water/Wastewater Engineer
AtwellAtwellFull-Time · RemoteDirect
$100k – $116k2w ago
Senior Forensic Engineer - Mechanical
YA GroupYA GroupFull-Time · RemoteDirect
$140k – $275k2w ago
TL
IT Security Systems Administrator
TopDog LawTopDog LawFull-Time · RemoteDirect
PythonAzure
$115k – $135k2w agoView details$115k – $135k2w ago
E
Forward Deployed Engineer
EVBEVBFull-Time · RemoteDirect
Python
$180k – $250k2w agoView details$180k – $250k2w ago