HPC AI Systems Administrator
Posted: 09/02/2026
Job Id: 25322
Job Details
Houston, TX 77033
Technology/IT
On-Site
Direct Hire
$100,000-$140,000 (Dependent on Experience)
No Relocation
Openings: 1
Job Description
A fast-growing manufacturing company is seeking an HPC AI Systems Administrator to serve as the foundational architect for its growing AI infrastructure. This role will design and maintain a robust compute platform that enables the Development Team to fine-tune and deploy production-level machine learning models, while ensuring the platform complies with enterprise-level security, governance, and data privacy policies. The HPC AI Systems Administrator is responsible for building a secure, scalable, and highly optimized environment to support the company's corporate data initiatives.
Salary + Additional Benefits:
- $100,000-$140,000 (Dependent on Experience)
- Medical, Dental, Vision Insurance
- 401K
Location: Houston, TX
Type of Position: Direct Hire
Responsibilities:
- Infrastructure Architecture & Management: Lead the deployment, bare-metal configuration, maintenance, and optimization of our on-premises HPC cluster and multi-GPU architecture.
- Platform Enablement: Manage the end-to-end AI software stack, including Linux OS environments, specialized GPU drivers, runtime libraries (CUDA, NCCL), and containerization platforms.
- Developer Sandbox Orchestration: Implement and maintain workload scheduling and orchestration systems (e.g., Kubernetes, Slurm, or equivalent enterprise platforms) to manage cluster resource allocation and job prioritization for engineering teams.
- Monitoring & Performance Tuning: Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth, ensuring maximum compute efficiency.
- Security & Governance Compliance: Operationalize strict data-at-rest and data-in-transit security baselines, ensuring the compute environment aligns with corporate Zero Trust network architecture and compliance mandates.
- Vendor Relations & Support: Act as the primary technical interface for high-end hardware vendors and system integrators to manage system updates and platform maintenance.
Requirements:
- 3+ years of dedicated systems administration experience managing Linux-based High-Performance Computing (HPC) environments or enterprise-scale GPU infrastructure
- Hands-on experience configuring and maintaining modern enterprise GPU hardware (such as NVIDIA Ampere or Hopper architecture) in a data center context
- Deep expertise in Linux system engineering, container technologies (Docker, Apptainer/Singularity), and cluster resource management
- Solid baseline knowledge of high-throughput networking fabrics (e.g., InfiniBand/RoCE) and parallel or distributed enterprise storage systems
- Bachelor's degree in computer science, Computer Engineering, System Administration, or equivalent practical industry experience
- Relevant professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture solutions (preferred)
- Familiarity with the infrastructure requirements supporting modern AI frameworks, machine learning lifecycles, or Large Language Model (LLM) fine-tuning pipelines (preferred)
Due to the high volume of applications we typically receive, we regret that we are not able to personally respond to all applications. However, if you are invited to take the next step in the process, you will typically be contacted within one week of submitting your application. #LI-DNI
Apply For This Position
Apply Via
"*" indicates required fields
I want more jobs like this in my inbox.
Share This Job
Related Jobs
Meet Your Recruiter