Back to dashboard

HPC Systems Engineer

Job ID 3436 | Run 20260928-184624

Changed Fields

FieldPreviousCurrent
IAPStaff Plan (target potential payout of $900, maximum of $1,800)Tier C Plan (target potential payout of 3.5%, maximum of 5%)

Job Description Diff


Previous Job Description
Current Job Description
f1Job Description:f1Job Description:
t2The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, mt2Certain terms and conditions of employment for this position, including the rate of pay, benefits, e
>aintenance, and day-to-day operations of the Institute’s HPC clusters. The HPC Systems Engineer will>tc., are currently subject to negotiation with the appropriate union The CoreHPC team at UCSF is see
>: Apply advanced systems infrastructure concepts and skills to the operations and improvement of lar>king an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day oper
>ge-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, an>ations of the Institute’s HPC clusters. The HPC Systems Engineer will: Apply advanced systems infras
>d storage systems designed to address cutting-edge research problems Apply their engineering and des>tructure concepts and skills to the operations and improvement of large-scale and highly complex res
>ign skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity >earch Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to a
>of CI systems. Select methods, techniques and evaluation criteria to develop new CI solutions to add>ddress cutting-edge research problems Apply their engineering and design skills to develop new CI so
>ress complex research problems. Be an active member of the support and maintenance efforts for the C>lutions, to develop and enhance monitoring to maintain the integrity of CI systems. Select methods, 
>oreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and m>techniques and evaluation criteria to develop new CI solutions to address complex research problems.
>aintaining systems' uptime and availability. Provides consultation, support, and guidance to researc> Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user 
>hers on how to address computational problems using standard tools, packages, and approaches. Develo>issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and 
>p enhancements of monitoring to maintain the integrity of CI systems. Participate in multiple techni>availability. Provides consultation, support, and guidance to researchers on how to address computat
>cal projects simultaneously. Applies working knowledge of security control frameworks to maintain th>ional problems using standard tools, packages, and approaches. Develop enhancements of monitoring to
>e integrity of the CI systems and the research being performed on them. Gives presentations to the a> maintain the integrity of CI systems. Participate in multiple technical projects simultaneously. Ap
>ssociated team and other technical units. Evaluates new technologies, including performing moderate >plies working knowledge of security control frameworks to maintain the integrity of the CI systems a
>to complex cost/benefit analyses. This position may lead to cross-functional technical working group>nd the research being performed on them. Gives presentations to the associated team and other techni
>s and projects in support of onboarding research customers, or making systems improvements. Departme>cal units. Evaluates new technologies, including performing moderate to complex cost/benefit analyse
>nt Overview Academic Research Systems (ARS) serves the needs of the UCSF research community by provi>s. This position may lead to cross-functional technical working groups and projects in support of on
>ding an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, >boarding research customers, or making systems improvements. Department Overview Academic Research S
>secure, professionally managed infrastructure for the storage and management of research data. ARS e>ystems (ARS) serves the needs of the UCSF research community by providing an integrated repository o
>mpowers medical scientific investigations by offering secure computing environments, data capture, m>f HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed 
>anagement and analysis tools, and support services which meet researchers’ needs. The Core HPC team >infrastructure for the storage and management of research data. ARS empowers medical scientific inve
>of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and st>stigations by offering secure computing environments, data capture, management and analysis tools, a
>orage services for UCSF researchers so they can address complex computational, AI,  and data science>nd support services which meet researchers’ needs. The Core HPC team of the Academic Research Servic
> problems.>e (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researc
 >hers so they can address complex computational, AI,  and data science problems.
33
4Qualifications:4Qualifications:
5REQUIRED QUALIFICATIONS - Bachelor's degree in a related area such as computer science or engineerin5REQUIRED QUALIFICATIONS - Bachelor's degree in a related area such as computer science or engineerin
>g, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience >g, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience 
>with large-scale or HPC systems - Expert knowledge of HPC systems infrastructure design - Strong kno>with large-scale or HPC systems - Expert knowledge of HPC systems infrastructure design - Strong kno
>wledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. - >wledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. - 
>Advanced knowledge of computer security best practices and policies including demonstrated experienc>Advanced knowledge of computer security best practices and policies including demonstrated experienc
>e securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requir>e securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requir
>ements - Demonstrated testing and test planning skills. Demonstrated ability to create automated tes>ements - Demonstrated testing and test planning skills. Demonstrated ability to create automated tes
>ting. - Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, - Demonstra>ting. - Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, - Demonstra
>ted skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband base>ted skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband base
>d clusters - Ability to elicit and communicate technical and non-technical information in a clear an>d clusters - Ability to elicit and communicate technical and non-technical information in a clear an
>d concise manner. - Self-motivated and works independently and as part of a team. Demonstrates probl>d concise manner. - Self-motivated and works independently and as part of a team. Demonstrates probl
>em-solving skills. Able to learn effectively and meet deadlines. - Understanding of system performan>em-solving skills. Able to learn effectively and meet deadlines. - Understanding of system performan
>ce monitoring and actions that can be taken to improve or correct performance. - Demonstrated advanc>ce monitoring and actions that can be taken to improve or correct performance. - Demonstrated advanc
>ed knowledge, skills and abilities associated with system problem identification and resolution. Exp>ed knowledge, skills and abilities associated with system problem identification and resolution. Exp
>erience with design, configuration, operation, repair, and tuning of technology systems. - Advanced >erience with design, configuration, operation, repair, and tuning of technology systems. - Advanced 
>experience writing and editing the most complex scripts used to perform system maintenance and admin>experience writing and editing the most complex scripts used to perform system maintenance and admin
>istration. - Ability to write technical documentation in a clear and concise manner. Ability to deve>istration. - Ability to write technical documentation in a clear and concise manner. Ability to deve
>lop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICAT>lop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICAT
>IONS - Knowledge of the design, development, and application of technology and systems to meet busin>IONS - Knowledge of the design, development, and application of technology and systems to meet busin
>ess needs. - General knowledge of other areas of IT. Thorough understanding of and experience with s>ess needs. - General knowledge of other areas of IT. Thorough understanding of and experience with s
>ystems-related issues and actions that can be taken to improve or correct performance. - Demonstrate>ystems-related issues and actions that can be taken to improve or correct performance. - Demonstrate
>d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh>d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh
>ensive understanding of how system management actions affect other systems, system users and depende>ensive understanding of how system management actions affect other systems, system users and depende
>nt/related functions.>nt/related functions.