Skip to main content
Weightless
AI
Engineering

AI Infrastructure Engineer

Appxpertise Inc

Full-Time
Mid-Level
Remote
Posted 2w ago

Job Description

**Role: AI Infrastructure Engineer** **Location: Remote Role** **Duration: 6\-12 Months** **Interview: MS Teams** **About the Role** We're looking for a hybrid AI Infrastructure Engineer to help stand up and operate a new GPU compute cluster — including hands\-on work managing an 18 HGX\-enabled worker node cluster with 144 GPUs, built on NVIDIA InfiniBand and Ethernet switching. This role spans the physical/network fabric layer, broader network operations, and the platform/orchestration layer, making it a good fit for someone who wants ownership from the switch port up through the workload scheduler. **What you'll be doing:** *Fabric \& Networking* · Help manage and operate the InfiniBand XDR / Quantum\-3400 rail\-aligned fabric topology supporting the 18\-node, 144\-GPU HGX cluster · Administer UFM and NetQ for subnet management, telemetry, and firmware across the fabric · Support 800G optics, diagnose and resolve link\-flap issues, and manage congestion isolation · Run NCCL / HPL validation and performance regression testing to confirm cluster health and throughput after changes · Operate Cumulus SN5600 and Arista border switching for the Ethernet side of the fabric *Network Operations* · Drive operational improvements in change management and daily operations · Manage and operate large\-scale IP network technologies and infrastructure · Work with peering and datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits · Monitor and support network health across on\-premises and cloud infrastructure · Document best practices and build workflow enhancements *Platform \& Workload Enablement* · Manage NVAIE and GPU Operator lifecycle, including driver/firmware compatibility matrices across the cluster · Administer Run:ai for quotas, project structure, and tenant onboarding as teams ramp onto the new nodes · Configure and support Slurm on carved\-out worker node pools for scheduled batch workloads · Manage container images and model\-serving dependencies for production and pre\-production use · Interface with DNA MLOps and support onboarding of first production workloads onto the new cluster **What we need to see:** · Hands\-on experience with NVIDIA InfiniBand (Quantum series) and high\-speed Ethernet switching in a GPU cluster environment · Familiarity with NVIDIA HGX platform architecture and multi\-node GPU networking (rail\-aligned topologies) · Experience with NCCL/HPL benchmarking or similar GPU cluster validation tooling · Working knowledge of GPU Operator, driver/firmware lifecycle management, and container\-based workload deployment · Experience with a workload orchestrator (Slurm, Run:ai, or similar) in an HPC/AI environment · Experience with large\-scale IP networking, peering/interconnect technologies, and change management processes · Comfort operating across network fabric, network operations, and platform/software layers. This is a hybrid role, not a pure network or pure platform position **Nice to have:** · Experience with UFM, NetQ, or equivalent InfiniBand fabric management tools · Exposure to MLOps workflows and model\-serving infrastructure · Prior work on new cluster bring\-up or greenfield GPU infrastructure deployments · Familiarity with DWDM/wave circuit technologies in a datacenter interconnect context Pay: From $90\.00 per hour Expected hours: 40\.0 per week Work Location: Remote

Get jobs like this in your inbox

Join thousands of digital nomads getting the best remote jobs delivered weekly. Free, no spam.