Job Summary
Overview:
Our client is seeking a Senior HPC Systems Engineer who will design, build, and maintain advanced HPC environment. This position focuses on the reliable operation, configuration, and optimization of HPC and AI systems, including multi-node CPU and GPU clusters, high-speed InfiniBand and Ethernet networks, and large-scale parallel and object storage. The engineer implements and automates secure, efficient, and reproducible computing platforms used by faculty, researchers, and students across diverse scientific disciplines. Assignments include both ticket-based support and project-based deployments.
Specific Duties & Responsibilities:
Support and administer production systems.
Provide technical leadership/project management for system configuration, implementation, management, and user support for both new and existing systems.
Research and recommend new functionality for HPC management and administration tools by exploring system-wide impacts, working with functional users to define current and future processes.
Expertise with architecting, operating, and debugging large scale HPC network and storage infrastructure, including MPI, NCCL, RDMA, Infiniband, and parallel file systems.
Analyze results of server monitoring and implement changes to improve performance, processing, and utilization.
Propose, maintain, and enforce policies, practices and security procedures.
Provide break/fix support, setup/installation support, escalation support, and solutions support.
Collaborate closely with a variety of stakeholders, both internal and external, on all aspects of projects.
Deploy, configure, and maintain large-scale Linux-based HPC clusters comprising CPU and GPU nodes, high-speed interconnects, and parallel file systems.
Implement and optimize workload schedulers (Slurm) and job submission policies to maximize system throughput and fair-share usage.
Administer and monitor distributed storage systems (GPFS, Lustre, WekaFS, Ceph, MinIO) to ensure reliability and performance across multi-petabyte environments.
Maintain high-speed fabric and network infrastructure (Infiniband, Ethernet) to support low-latency data transfer and MPI workloads.
Support groups in deploying, testing, and optimizing scientific applications and AI/ML workflows on shared computing resources.
Develop and maintain automation and monitoring frameworks for system provisioning, metrics collection, and alerting (Prometheus, Grafana, ELK).
Participate in capacity planning, hardware lifecycle management, and evaluation of new technologies in collaboration with architects and management.
Ensure security and compliance through configuration hardening, patch management, and integration with campus identity and access control systems.
Document system designs, procedures, and troubleshooting guides to support knowledge transfer and team continuity.
Contribute to a collaborative engineering culture that emphasizes service quality, innovation, and continuous improvement in research computing operations.
Qualifications:
Eight plus years of experience in HPC systems administration or engineering, including experience with cluster management, workload scheduling (e.g., Slurm), and distributed or parallel storage.
Deep proficiency in Linux systems administration, configuration management (Ansible, Puppet, or Salt), performance monitoring, and tuning for HPC workloads.
Experience with high-speed interconnects (Infiniband, 100/400 Gb Ethernet) and parallel file systems (e.g., GPFS, Lustre, BeeGFS, or WekaFS).
Working knowledge of containerization and orchestration (Singularity, Docker, Kubernetes for HPC).
Ability to automate deployments and routine operations through scripting (Bash, Python).
Familiarity with data-center operations, GPU acceleration, and research software environments (e.g., CUDA, MPI, AI/ML frameworks).
Strong analytical and troubleshooting skills, with proven ability to support complex research workloads in multi-user, multi-tenant environments.
Experience collaborating with faculty and research groups to translate scientific requirements into practical and performant computing solutions.
Share This Job
ABOUT US
Our High Trail Verified members include elite Technical and Solutions Architects, Developers, Engineers, Administrators, Business Analysts, Project Managers and Leaders, with a proven track record of delivering solutions to complex problems.
If you are ready to Elevate Your Work, you’ve come to the right place. Contact us to schedule a consultation.
Other Jobs you might like…
Remote
$0 - $17
Data Steward / Data Analyst (Contract, Ongoing) Data Quality & Governance Role Overview Our client, a growing data analytics company…
NYC
$0 - $85
Salesforce Business Analyst Location: Hybrid – Washington, DC or New York, NY (on-site required days per week) Industry: Financial Services…
Herndon, Virginia
$170,000 - $195,000
Overview Our client is seeking a Senior Manager – Oracle Fusion Cloud ERP / Solution Architect to serve as the…
Remote
$112,000 - $146,000
Our client, a reputable financial institution is seeking an experienced AVP of Digital Marketing with a focus on Salesforce Marketing…
Remote
$70 - $80
About the Opportunity We are working with a well-respected, fast-growing social media technology company seeking a Fullstack Developer to help…
Remote
$0 - $90
About the Role Our client, a fast-growing data intelligence company serving a major retail/manufacturing industry, is looking for a Senior…
Remote
$10 - $20
Our client is seeking a highly skilled Help Desk Specialist to join a dynamic IT team supporting a global organization….
Remote
$45 - $55
Our client is seeking an experienced SAP Solution Architect to support large-scale SAP carve-out, divestiture, and business separation programs across…
Columbia, Maryland
$100,000 - $125,000
Required Skills: 4–6 years of professional experience in backend web application development using .NET Framework / .NET Core with C#…
Remote
$60 - $70
Key Responsibilities: Support and optimize Oracle Fusion Financials, focusing on Record-to-Report (R2R) processes. Collaborate with finance and business units to…