Back to dashboard

Storage Systems Engineer

Job ID 3270 | Run 20260711-144732

Changed Fields

Only the job description changed.

Job Description Diff


Previous Job Description
Current Job Description
2This is a two-year contract of employment, inclusive of benefits. The Academic Research Services tea2This is a two-year contract of employment, inclusive of benefits. The Academic Research Services tea
>m at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to serve as a technical resource in the>m at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to serve as a technical resource in the
> design, deployment, and operation of large-scale research storage and data infrastructure. This rol> design, deployment, and operation of large-scale research storage and data infrastructure. This rol
>e will work in close partnership with the Senior Research DevOps Engineer to support UCSF’s evolving>e will work in close partnership with the Senior Research DevOps Engineer to support UCSF’s evolving
> research ecosystem, including CoreHPC, the Research Analysis Environment (RAE), and large instituti> research ecosystem, including CoreHPC, the Research Analysis Environment (RAE), and large instituti
>onal storage initiatives. This position is primarily responsible for architecture, implementation, a>onal storage initiatives. This position is primarily responsible for architecture, implementation, a
>nd lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, includin>nd lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, includin
>g support for large storage environments, NSF-funded infrastructure, and OS Nexus–aligned data platf>g support for large storage environments, NSF-funded infrastructure, and OS Nexus–aligned data platf
>orms. The role ensures seamless integration between storage systems and the CoreHPC compute cluster,>orms. The role ensures seamless integration between storage systems and the CoreHPC compute cluster,
> enabling performant, reliable, and scalable data access for AI, data science, and computational res> enabling performant, reliable, and scalable data access for AI, data science, and computational res
>earch workloads. The Storage Systems Engineer will: Work with the lead to continue supporting the de>earch workloads. The Storage Systems Engineer will: Work with the lead to continue supporting the de
>sign and evolution of storage architecture across on-prem and hybrid environments, including VAST, p>sign and evolution of storage architecture across on-prem and hybrid environments, including VAST, p
>arallel filesystems, and enterprise storage platforms Develop and maintain data movement strategies >arallel filesystems, and enterprise storage platforms Develop and maintain data movement strategies 
>and tooling (e.g., rsync, rclone, Globus, SMB workflows) to support large-scale data ingestion, migr>and tooling (e.g., rsync, rclone, Globus, SMB workflows) to support large-scale data ingestion, migr
>ation, and lifecycle management Ensure tight integration between storage and HPC compute systems, op>ation, and lifecycle management Ensure tight integration between storage and HPC compute systems, op
>timizing throughput, latency, and reliability for distributed workloads Support and scale storage sy>timizing throughput, latency, and reliability for distributed workloads Support and scale storage sy
>stems backing major institutional initiatives (FAC storage, OS Nexus integration) Collaborate closel>stems backing major institutional initiatives (FAC storage, OS Nexus integration) Collaborate closel
>y with DevOps, networking, and security teams to deliver cohesive research infrastructure solutions >y with DevOps, networking, and security teams to deliver cohesive research infrastructure solutions 
>Design and implement monitoring, performance tuning, and capacity planning strategies for storage an>Design and implement monitoring, performance tuning, and capacity planning strategies for storage an
>d data systems Troubleshoot complex issues across storage, networking, and compute boundaries Partic>d data systems Troubleshoot complex issues across storage, networking, and compute boundaries Partic
>ipate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers P>ipate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers P
>rovide guidance to researchers on data organization, transfer strategies, and performance optimizati>rovide guidance to researchers on data organization, transfer strategies, and performance optimizati
>on Evaluate and recommend emerging storage technologies and architectures This role may lead storage>on Evaluate and recommend emerging storage technologies and architectures This role may lead storage
>-focused projects and contribute to cross-functional initiatives that improve the scalability, usabi>-focused projects and contribute to cross-functional initiatives that improve the scalability, usabi
>lity, and reliability of UCSF’s research computing ecosystem. Department Overview Academic Research >lity, and reliability of UCSF’s research computing ecosystem. Department Overview Academic Research 
>Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository >Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository 
>of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed>of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed
> infrastructure for the storage and management of research data. ARS empowers medical scientific inv> infrastructure for the storage and management of research data. ARS empowers medical scientific inv
>estigations by offering secure computing environments, data capture, management and analysis tools, >estigations by offering secure computing environments, data capture, management and analysis tools, 
>and support services which meet researchers’ needs. The Research Infrastructure team of the Academic>and support services which meet researchers’ needs. The Research Infrastructure team of the Academic
> Research Service (ARS) focuses on large scale research platform support, high performance computati> Research Service (ARS) focuses on large scale research platform support, high performance computati
>onal and storage services for UCSF researchers so they can address complex computational, AI,  and d>onal and storage services for UCSF researchers so they can address complex computational, AI,  and d
>ata science problems.>ata science problems.
33
4Qualifications:4Qualifications:
t5REQUIRED QUALIFICATIONS - Bachelor's degree in a related area, such as computer science or engineerit5REQUIRED QUALIFICATIONS - Bachelor's degree in a related area, such as computer science or engineeri
>ng, and 6+ years of experience with storage infrastructure support and management, or 10+ years of r>ng, and 6+ years of experience with storage infrastructure support and management, or 10+ years of r
>elated experience with large-scale storage systems - Demonstrated skill (5 years +) deploying, manag>elated experience with large-scale storage systems - Demonstrated skill (5 years +) deploying, manag
>ing, and troubleshooting Warewulf (or similar) InfiniBand-based clusters - Strong knowledge of ZFS>ing, and troubleshooting ZFS (or similar) InfiniBand-based clusters - Strong knowledge of ZFS, high-
>high-performance parallel filesystems, and storage such as GPFS, Lustre, Vast, DDN, etc - Advanced k>performance parallel filesystems, and storage such as GPFS, Lustre, Vast, DDN, etc - Advanced knowle
>nowledge of computer security best practices and policies, including demonstrated experience securin>dge of computer security best practices and policies, including demonstrated experience securing res
>g research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPAA, or IS-3 requirements ->earch cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPAA, or IS-3 requirements - Know
> Knowledge of HPC job scheduler system design and operation, such as SLURM or PBS, - Ability to elic>ledge of HPC job scheduler system design and operation, such as SLURM or PBS, - Ability to elicit an
>it and communicate technical and non-technical information in a clear and concise manner. - Self-mot>d communicate technical and non-technical information in a clear and concise manner. - Self-motivate
>ivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to l>d and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn 
>earn effectively and meet deadlines. - Understanding of system performance monitoring and actions th>effectively and meet deadlines. - Understanding of system performance monitoring and actions that ca
>at can be taken to improve or correct performance. - Demonstrated advanced knowledge, skills, and ab>n be taken to improve or correct performance. - Demonstrated advanced knowledge, skills, and abiliti
>ilities associated with system problem identification and resolution. Experience with design, config>es associated with system problem identification and resolution. Experience with design, configurati
>uration, operation, repair, and tuning of technology systems. - Advanced experience writing and edit>on, operation, repair, and tuning of technology systems. - Advanced experience writing and editing t
>ing the most complex scripts used to perform system maintenance and administration. - Demonstrated t>he most complex scripts used to perform system maintenance and administration. - Demonstrated testin
>esting and test planning skills. Demonstrated ability to create automated testing. - Ability to writ>g and test planning skills. Demonstrated ability to create automated testing. - Ability to write tec
>e technical documentation in a clear and concise manner. Ability to develop runbooks defining comple>hnical documentation in a clear and concise manner. Ability to develop runbooks defining complex tec
>x technical processes in a clear and concise manner PREFERRED QUALIFICATIONS - Expert knowledge of V>hnical processes in a clear and concise manner PREFERRED QUALIFICATIONS - Expert knowledge of Virtua
>irtual Machines, Bare Metal Servers & HPC systems infrastructure design - Knowledge of the design, d>l Machines, Bare Metal Servers & HPC systems infrastructure design - Knowledge of the design, develo
>evelopment and application of technology and systems to meet business needs. - General knowledge of >pment and application of technology and systems to meet business needs. - General knowledge of other
>other areas of IT. E.g., Active Directory, Domain Controllers, Network Infrastructure. - Demonstrate> areas of IT. E.g., Active Directory, Domain Controllers, Network Infrastructure. - Demonstrated ski
>d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh>lls associated with adapting equipment and technology to serve user needs. Demonstrated comprehensiv
>ensive understanding of how system management actions affect other systems, system users and depende>e understanding of how system management actions affect other systems, system users and dependent/re
>nt/related functions. - Professional certification in enterprise storage technologies (e.g., NetApp,>lated functions. - Professional certification in enterprise storage technologies (e.g., NetApp, Dell
> Dell EMC PowerScale, IBM Storage Scale, VAST, Pure Storage)> EMC PowerScale, IBM Storage Scale, VAST, Pure Storage)