HPC Scientific Software Engineer (IT@JH Research Computing)
IT@JH Research Computing is seeking a HPC Scientific Software Engineer to support faculty, researchers, and students engaged in high-performance and AI-driven research across Johns Hopkins University . The position is responsible for deploying, optimizing, and maintaining scientific software and computational workflows on advanced HPC Systems and related infrastructure. Working primarily within Linux-based environments, the engineer manages and troubleshoots complex software stacks, containerized applications, and GPU-accelerated workloads using tools such as SLURM, Easy build, Spack, etc. The role combines ticket-based user support with long-term project work, collaborating closely with interdisciplinary research groups to enhance system performance, streamline data-intensive workflows, and integrate cutting-edge technologies. The position operates with significant independence while coordinating regularly with systems engineers and research computing leadership to ensure reliable, high-efficiency computing resources that advance the university’s scientific mission.
Specific Duties & Responsibilities
Software Deployment and Design (15%)
- Develop and refine deployment strategies for scientific software on HPC and AI systems.
- Design computational workflows, selecting optimal software configurations, and utilizing tools like Ansible for automation.
- Assist teams in implementing, tuning, and optimizing AI models and gateway applications (e.g., XDMoD, Coldfront, Open OnDemand, CryoSPARC Live, SBGrid, AI Agents).
Performance Optimization (20%)
- Analyze and optimize the performance of AI models and HPC applications, focusing on GPU-enabled computing.
- Implement parallel processing, distributed computing, and resource management techniques for efficient job execution.
Integration and Optimization (15%)
- Develop, debug, and maintain software tools, libraries, and frameworks supporting HPC and AI workloads.
- Collaborate with the system team and software vendors (e.g., NVIDIA, Intel, Matlab) to optimize systems for maximum performance.
- Utilize CUDA, DNN, TensorRT, and Intel Compilers to enhance system performance.
HPC Scientific Software Support (30%)
- Manage and support scientific software deployment across HPC, cloud-based, and colocation facilities.
- Oversee installation, configuration, and maintenance of HPC packages with tools like CMake, Make, EasyBuild, Spack, and Lua module files.
Collaboration and Mentorship (5%)
- Work closely with cross-functional teams, including researchers, data scientists, and software developers, to address complex HPC/AI challenges.
- Mentor junior engineers and foster a culture of continuous learning.
Technical Support and Training Workshops and Troubleshooting (15%)
- Resolve complex technical issues and perform root cause analysis for HPC/AI software challenges.
- Implement effective solutions to prevent recurrence and improve system reliability
- Provide training workshops for researchers and students, focusing on troubleshooting, optimizing workflows, and effectively using HPC systems.
Learning and Development (5%)
- Stay current with advances in HPC and AI technologies and methodologies.
- Incorporate new research findings into existing systems to improve performance and capabilities.
Container Orchestration (5%)
- Develop and manage container orchestration strategies to ensure scalability, reliability, and security of applications.
- Oversee the container lifecycle from creation and deployment to scaling and removal.
Documentation and Compliance (5%)
- Create comprehensive documentation for system designs, performance metrics, and project status.
- Ensure compliance with security and regulatory standards for all HPC and AI systems.
Other duties as assigned.
Minimum Qualifications
- Master’s degree in computer science or a closely related quantitative discipline.
- Five years of experience in HPC user support, software deployment, and pe