أبلاي إيدج ابدأ البحث عن عمل

HPC Lead System Administrator

KAUST (King Abdullah University of Science and Technology) · Thuwal, Makkah, Saudi Arabia

قدّم وتابع مع أبلاي إيدج
Position SummaryServe as the Lead for the team ensuring smooth operation of the Linux cluster consisting of 300+ GPU/CPU compute nodes including parallel filesystems and high-performance network. This is partly technical and partly people leading role which involves supervision of 3-4 experienced HPC system administrators. The role involves development, implementation and supervision of standard operating procedures for the system and the team.Major ResponsibilitiesSystem operation and upgrade planning to meet laboratory and customer requirementsWorkload scheduler policy development and implementationSupport of high-performance filesystemsNetwork infrastructure management including TCP/IP and HPC networksUse of scripting languages for nodes automation and configuration managementHardware failures and spare part managementBuild effective relationships with staff, faculty and students through the Core LabsManages multiple or significant projects which may require the use of sophisticated project planning techniquesPlans, schedules, conducts, or coordinates detailed phases of the work of a major project or in a total project of moderate scopeIdentifies technical training needs for staff attached to the areaServe as a resource and as a member to respond to security and safety incidentsCreates opportunities to enhance technical methodology or content through expansion of existing, or development of, new efforts; may extend technology into new application areas; contributes or leads in major intellectual development activitiesProvides innovative problem-solving approaches to enhance organizational capabilities; uses peer network to expand technical capabilities and identify new research opportunitiesUnderstands broad strategic objectives and contributes to them; nurtures and maintains relationships with major customersMay initiate new project concepts; develops technical proposals and makes presentations to potential customersWill supervise several scientists, engineers or technicians on assigned work; provides major input to staffing of overall project teams; builds teams and staff to optimize efficiency and cost effectivenessIdentifies and evaluates candidates for open positions; mentors/trains staff in development of technical, project and business development skillsCompetenciesSLURM workload manager including GPU schedulingParallel filesystems (Weka IO, Lustre)TCP/IP and high performance networks (Infiniband)Proficient in scripting languages (i.e. Bash, Python, Ruby)Familiar with configuration management tools (Puppet)Proficient documentation skillsWill have working level contact with users and suppliersDemonstrates an analytical and systematic approach to problem solvingTakes the initiative in identifying and negotiating appropriate development opportunitiesDemonstrates effective communication skills in written and oral EnglishWorks effectively with other teams in the Supercomputing LaboratoryPlans, schedules and monitors own work (and that of others) competently within limited deadlines and according to relevant legislation and proceduresAbility to work successfully in a highly collaborative research environmentUses discretion in identifying and resolving complex problems and assignmentsPerforms a broad range of work, sometimes complex and non-routine, in a variety of environmentsMaintain expert-level knowledge in most of the laboratory systems, including high performance computing systems administration, high performance storage administration, or high performance network administrationQualifications and ExperienceBachelor of Science (or equivalent) in a relevant discipline plus 10 years’ experience, OR Master of Science (or equivalent) in a relevant discipline plus 7 years’ experience OR Doctor of Philosophy (or equivalent) in a relevant discipline plus 5 years’ experience.