| n | This is a two-year contract of employment, inclusive of benefits. | n | Job Description: |
| | | This is a two-year contract of employment, inclusive of benefits. The Academic Research Services tea |
| | | m at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to serve as a technical resource in the |
| | | design, deployment, and operation of large-scale research storage and data infrastructure. This rol |
| | | e will work in close partnership with the Senior Research DevOps Engineer to support UCSF’s evolving |
| | | research ecosystem, including CoreHPC, the Research Analysis Environment (RAE), and large instituti |
| | | onal storage initiatives. This position is primarily responsible for architecture, implementation, a |
| | | nd lifecycle management for the Facility for Advanced Computing (FAC), storage and systems, includin |
| | | g support for large storage environments, NSF-funded infrastructure, and OS Nexus–aligned data platf |
| | | orms. The role ensures seamless integration between storage systems and the CoreHPC compute cluster, |
| | | enabling performant, reliable, and scalable data access for AI, data science, and computational res |
| | | earch workloads. The Storage Systems Engineer will: Work with the lead to continue supporting the de |
| | | sign and evolution of storage architecture across on-prem and hybrid environments, including VAST, p |
| | | arallel filesystems, and enterprise storage platforms Develop and maintain data movement strategies |
| | | and tooling (e.g., rsync, rclone, Globus, SMB workflows) to support large-scale data ingestion, migr |
| | | ation, and lifecycle management Ensure tight integration between storage and HPC compute systems, op |
| | | timizing throughput, latency, and reliability for distributed workloads Support and scale storage sy |
| | | stems backing major institutional initiatives (FAC storage, OS Nexus integration) Collaborate closel |
| | | y with DevOps, networking, and security teams to deliver cohesive research infrastructure solutions |
| | | Design and implement monitoring, performance tuning, and capacity planning strategies for storage an |
| | | d data systems Troubleshoot complex issues across storage, networking, and compute boundaries Partic |
| | | ipate in system upgrades, migrations, and expansion efforts with minimal disruption to researchers P |
| | | rovide guidance to researchers on data organization, transfer strategies, and performance optimizati |
| | | on Evaluate and recommend emerging storage technologies and architectures This role may lead storage |
| | | -focused projects and contribute to cross-functional initiatives that improve the scalability, usabi |
| | | lity, and reliability of UCSF’s research computing ecosystem. Department Overview Academic Research |
| | | Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository |
| | | of HIPAA compliant clinical and life sciences data and a centralized, secure, professionally managed |
| | | infrastructure for the storage and management of research data. ARS empowers medical scientific inv |
| | | estigations by offering secure computing environments, data capture, management and analysis tools, |
| | | and support services which meet researchers’ needs. The Research Infrastructure team of the Academic |
| | | Research Service (ARS) focuses on large scale research platform support, high performance computati |
| | | onal and storage services for UCSF researchers so they can address complex computational, AI, and d |
| | | ata science problems. |
| | | |
| t | The Academic Research Services team at UCSF is seeking an Storage Systems Engineer (SYS ADM 4) to s | t | Qualifications: |
| erve as a technical resource in the design, deployment, and operation of large-scale research storag | | |
| e and data infrastructure. This role will work in close partnership with the Senior Research DevOps | | |
| Engineer to support UCSF’s evolving research ecosystem, including CoreHPC, the Research Analysis Env | | |
| ironment (RAE), and large institutional storage initiatives. | | |
| | | REQUIRED QUALIFICATIONS - Bachelor's degree in a related area, such as computer science or engineeri |
| | | ng, and 6+ years of experience with storage infrastructure support and management, or 10+ years of r |
| | | elated experience with large-scale storage systems - Demonstrated skill (5 years +) deploying, manag |
| | | ing, and troubleshooting Warewulf (or similar) InfiniBand-based clusters - Strong knowledge of ZFS, |
| | | high-performance parallel filesystems, and storage such as GPFS, Lustre, Vast, DDN, etc - Advanced k |
| | | nowledge of computer security best practices and policies, including demonstrated experience securin |
| | | g research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPAA, or IS-3 requirements - |
| | | Knowledge of HPC job scheduler system design and operation, such as SLURM or PBS, - Ability to elic |
| | | it and communicate technical and non-technical information in a clear and concise manner. - Self-mot |
| | | ivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to l |
| | | earn effectively and meet deadlines. - Understanding of system performance monitoring and actions th |
| | | at can be taken to improve or correct performance. - Demonstrated advanced knowledge, skills, and ab |
| | | ilities associated with system problem identification and resolution. Experience with design, config |
| | | uration, operation, repair, and tuning of technology systems. - Advanced experience writing and edit |
| | | ing the most complex scripts used to perform system maintenance and administration. - Demonstrated t |
| | | esting and test planning skills. Demonstrated ability to create automated testing. - Ability to writ |
| | | e technical documentation in a clear and concise manner. Ability to develop runbooks defining comple |
| | | x technical processes in a clear and concise manner PREFERRED QUALIFICATIONS - Expert knowledge of V |
| | | irtual Machines, Bare Metal Servers & HPC systems infrastructure design - Knowledge of the design, d |
| | | evelopment and application of technology and systems to meet business needs. - General knowledge of |
| | | other areas of IT. E.g., Active Directory, Domain Controllers, Network Infrastructure. - Demonstrate |
| | | d skills associated with adapting equipment and technology to serve user needs. Demonstrated compreh |
| | | ensive understanding of how system management actions affect other systems, system users and depende |
| | | nt/related functions. - Professional certification in enterprise storage technologies (e.g., NetApp, |
| | | Dell EMC PowerScale, IBM Storage Scale, VAST, Pure Storage) |
| This position is primarily responsible for architecture, implementation, and lifecycle management f | | |
| or the Facility for Advanced Computing (FAC), storage and systems, including support for large stora | | |
| ge environments, NSF-funded infrastructure, and OS Nexus-aligned data platforms. The role ensures se | | |
| amless integration between storage systems and the CoreHPC compute cluster, enabling performant, rel | | |
| iable, and scalable data access for AI, data science, and computational research workloads. | | |
| | | |
| The Storage Systems Engineer will: | | |
| * Work with the lead to continue supporting the design and evolution of storage architecture across | | |
| on-prem and hybrid environments, including VAST, parallel filesystems, and enterprise storage platf | | |
| orms | | |
| * Develop and maintain data movement strategies and tooling (e.g., rsync, rclone, Globus, SMB workf | | |
| lows) to support large-scale data ingestion, migration, and lifecycle management | | |
| * Ensure tight integration between storage and HPC compute systems, optimizing throughput, latency, | | |
| and reliability for distributed workloads | | |
| * Support and scale storage systems backing major institutional initiatives (FAC storage, OS Nexus | | |
| integration) | | |
| * Collaborate closely with DevOps, networking, and security teams to deliver cohesive research infr | | |
| astructure solutions | | |
| * Design and implement monitoring, performance tuning, and capacity planning strategies for storage | | |
| and data systems | | |
| * Troubleshoot complex issues across storage, networking, and compute boundaries | | |
| * Participate in system upgrades, migrations, and expansion efforts with minimal disruption to rese | | |
| archers | | |
| * Provide guidance to researchers on data organization, transfer strategies, and performance optimi | | |
| zation | | |
| * Evaluate and recommend emerging storage technologies and architectures | | |
| | | |
| This role may lead storage-focused projects and contribute to cross-functional initiatives that imp | | |
| rove the scalability, usability, and reliability of UCSF’s research computing ecosystem. | | |
| | | |
| Department Overview | | |
| | | |
| Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an int | | |
| egrated repository of HIPAA compliant clinical and life sciences data and a centralized, secure, pro | | |
| fessionally managed infrastructure for the storage and management of research data. ARS empowers med | | |
| ical scientific investigations by offering secure computing environments, data capture, management a | | |
| nd analysis tools, and support services which meet researchers’ needs. | | |
| | | |
| The Research Infrastructure team of the Academic Research Service (ARS) focuses on large scale rese | | |
| arch platform support, high performance computational and storage services for UCSF researchers so t | | |
| hey can address complex computational, AI, and data science problems. | | |