Datacenter Technician
Boson AI · Barrie, Ontario, Canada
قدّم وتابع مع أبلاي إيدجBoson AI is an early-stage startup building large language tools for everyone to use. Our founders (Alex Smola, Mu Li), and a team of Deep Learning, Optimization, NLP, AutoML and Statistics scientists and engineers are working on high quality generative AI models for language and beyond.You keep the machines runningWe are looking for a Datacenter Hardware Technician to keep the physical infrastructure behind our AI research running in Barrie, ON. You will work hands-on with the latest NVIDIA GPUs, thousands of disks, terabit networking and hundreds of Supermicro servers — racking them, repairing them, keeping firmware current, and making sure every machine that should be online is online.Our researchers train models around the clock, so the difference between a good day and a bad one is often a technician who noticed a marginal cable, logged a serial number correctly, or caught a failing drive before it took a training job down with it. This is careful, methodical work, and we treat it that way.This role is for our Barrie datacenter. You must live within 50 km of the site, such that you can be onsite on short notice. Workload can be bursty, i.e. periods of smooth sailing mixed with periods of intense work during hardware failures, upgrade and maintenance cycles.A day in the lifeInstall, rack, cable and commission new Supermicro servers, storage and network equipmentDiagnose and repair hardware failures — replace DIMMs, drives, power supplies, fans, GPUs, cables and mainboards — and drive RMA cases with vendors through to resolutionInstall and update drivers, BIOS and firmware across servers, NICs, HBAs and switches, keeping fleet versions consistent and documentedTest and install network connections, including structured cabling, optics and link validation, and troubleshoot physical-layer faultsPerform preventive maintenance: inspections, cable management, airflow and filter checks, and spare-parts inventoryKeep accurate records of every asset, serial number, part swap and rack locationRespond to hardware failure alerts, and escalate to the SRE team when a fault is not purely physicalFollow runbooks and ESD and safety procedures precisely — and improve them where they are unclearIf you take pride in a tidy rack, a clean cable run, and a fleet where every machine is accounted for, we'd love to hear from you.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.Minimum Qualifications2+ years of hands-on experience with server hardware in a datacenter, colocation, IT operations or equivalent environmentPractical experience replacing and troubleshooting server components: drives, memory, power supplies, fans, GPUs and mainboardsExperience installing drivers, BIOS and firmware updates on server hardwareExperience testing and installing network connections, including structured cablingComfortable on the Linux command line — navigating the filesystem, reading logs, running diagnosticsFamiliar with remote access and operational practices: SSH, VPN, and out-of-band management (IPMI, BMC, Redfish)Exceptionally organized and meticulous. Accurate records, careful labelling and consistent procedure matter more here than raw speedAble to work onsite in Barrie, ON, living within 50 km of the siteComfortable with the physical demands of the role: lifting and racking equipment, working in cold aisles, and occasional after-hours or on-call responseClear written communication — your notes are what the next person relies onPreferred QualificationsDirect experience with Supermicro servers, chassis and IPMI toolingExperience with GPU servers (NVIDIA H100, A100 or similar) and their power and cooling requirementsExperience with high-speed networking: 100Gb+ Ethernet, InfiniBand, optics and DAC cablingFamiliarity with DCIM or asset-management tools such as NetBoxBasic scripting in Bash or Python to automate repetitive checksExperience with power distribution units, structured cabling standards, and rack and power designExperience handling RMA workflows with hardware vendorsRelevant certifications (CompTIA Server+, A+, Network+) or an equivalent hands-on track record