| f | Job Description: | f | Job Description: |
| t | The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, m | t | Certain terms and conditions of employment for this position, including the rate of pay, benefits, e |
| aintenance, and day-to-day operations of the Institute’s HPC clusters. The HPC Systems Engineer will | | tc., are currently subject to negotiation with the appropriate union The CoreHPC team at UCSF is see |
| : Apply advanced systems infrastructure concepts and skills to the operations and improvement of lar | | king an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day oper |
| ge-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, an | | ations of the Institute’s HPC clusters. The HPC Systems Engineer will: Apply advanced systems infras |
| d storage systems designed to address cutting-edge research problems Apply their engineering and des | | tructure concepts and skills to the operations and improvement of large-scale and highly complex res |
| ign skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity | | earch Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to a |
| of CI systems. Select methods, techniques and evaluation criteria to develop new CI solutions to add | | ddress cutting-edge research problems Apply their engineering and design skills to develop new CI so |
| ress complex research problems. Be an active member of the support and maintenance efforts for the C | | lutions, to develop and enhance monitoring to maintain the integrity of CI systems. Select methods, |
| oreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and m | | techniques and evaluation criteria to develop new CI solutions to address complex research problems. |
| aintaining systems' uptime and availability. Provides consultation, support, and guidance to researc | | Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user |
| hers on how to address computational problems using standard tools, packages, and approaches. Develo | | issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and |
| p enhancements of monitoring to maintain the integrity of CI systems. Participate in multiple techni | | availability. Provides consultation, support, and guidance to researchers on how to address computat |
| cal projects simultaneously. Applies working knowledge of security control frameworks to maintain th | | ional problems using standard tools, packages, and approaches. Develop enhancements of monitoring to |
| e integrity of the CI systems and the research being performed on them. Gives presentations to the a | | maintain the integrity of CI systems. Participate in multiple technical projects simultaneously. Ap |
| ssociated team and other technical units. Evaluates new technologies, including performing moderate | | plies working knowledge of security control frameworks to maintain the integrity of the CI systems a |
| to complex cost/benefit analyses. This position may lead to cross-functional technical working group | | nd the research being performed on them. Gives presentations to the associated team and other techni |
| s and projects in support of onboarding research customers, or making systems improvements. Departme | | cal units. Evaluates new technologies, including performing moderate to complex cost/benefit analyse |
| nt Overview Academic Research Systems (ARS) serves the needs of the UCSF research community by provi | | s. This position may lead to cross-functional technical working groups and projects in support of on |
| ding an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, | | boarding research customers, or making systems improvements. Department Overview Academic Research S |
| secure, professionally managed infrastructure for the storage and management of research data. ARS e | | ystems (ARS) serves the needs of the UCSF research community by providing an integrated repository o |
| mpowers medical scientific investigations by offering secure computing environments, data capture, m | | f HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed |
| anagement and analysis tools, and support services which meet researchers’ needs. The Core HPC team | | infrastructure for the storage and management of research data. ARS empowers medical scientific inve |
| of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and st | | stigations by offering secure computing environments, data capture, management and analysis tools, a |
| orage services for UCSF researchers so they can address complex computational, AI, and data science | | nd support services which meet researchers’ needs. The Core HPC team of the Academic Research Servic |
| problems. | | e (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researc |
| | | hers so they can address complex computational, AI, and data science problems. |
| | | |
| Qualifications: | | Qualifications: |
| REQUIRED QUALIFICATIONS - Bachelor's degree in a related area such as computer science or engineerin | | REQUIRED QUALIFICATIONS - Bachelor's degree in a related area such as computer science or engineerin |
| g, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience | | g, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience |
| with large-scale or HPC systems - Expert knowledge of HPC systems infrastructure design - Strong kno | | with large-scale or HPC systems - Expert knowledge of HPC systems infrastructure design - Strong kno |
| wledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. - | | wledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. - |
| Advanced knowledge of computer security best practices and policies including demonstrated experienc | | Advanced knowledge of computer security best practices and policies including demonstrated experienc |
| e securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requir | | e securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requir |
| ements - Demonstrated testing and test planning skills. Demonstrated ability to create automated tes | | ements - Demonstrated testing and test planning skills. Demonstrated ability to create automated tes |
| ting. - Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, - Demonstra | | ting. - Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, - Demonstra |
| ted skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband base | | ted skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband base |
| d clusters - Ability to elicit and communicate technical and non-technical information in a clear an | | d clusters - Ability to elicit and communicate technical and non-technical information in a clear an |
| d concise manner. - Self-motivated and works independently and as part of a team. Demonstrates probl | | d concise manner. - Self-motivated and works independently and as part of a team. Demonstrates probl |
| em-solving skills. Able to learn effectively and meet deadlines. - Understanding of system performan | | em-solving skills. Able to learn effectively and meet deadlines. - Understanding of system performan |
| ce monitoring and actions that can be taken to improve or correct performance. - Demonstrated advanc | | ce monitoring and actions that can be taken to improve or correct performance. - Demonstrated advanc |
| ed knowledge, skills and abilities associated with system problem identification and resolution. Exp | | ed knowledge, skills and abilities associated with system problem identification and resolution. Exp |
| erience with design, configuration, operation, repair, and tuning of technology systems. - Advanced | | erience with design, configuration, operation, repair, and tuning of technology systems. - Advanced |
| experience writing and editing the most complex scripts used to perform system maintenance and admin | | experience writing and editing the most complex scripts used to perform system maintenance and admin |
| istration. - Ability to write technical documentation in a clear and concise manner. Ability to deve | | istration. - Ability to write technical documentation in a clear and concise manner. Ability to deve |
| lop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICAT | | lop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICAT |
| IONS - Knowledge of the design, development, and application of technology and systems to meet busin | | IONS - Knowledge of the design, development, and application of technology and systems to meet busin |
| ess needs. - General knowledge of other areas of IT. Thorough understanding of and experience with s | | ess needs. - General knowledge of other areas of IT. Thorough understanding of and experience with s |
| ystems-related issues and actions that can be taken to improve or correct performance. - Demonstrate | | ystems-related issues and actions that can be taken to improve or correct performance. - Demonstrate |
| d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh | | d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh |
| ensive understanding of how system management actions affect other systems, system users and depende | | ensive understanding of how system management actions affect other systems, system users and depende |
| nt/related functions. | | nt/related functions. |