Reliability Manager
Gold Brand Angola ยท United Arab Emirates
Apply & track with Apply Edgeโ๏ธ We're Hiring: Reliability Manager๐ Location: United Arab Emirates (Remote) ๐ Employment Type: Full-Time ๐ผ Experience Level: Mid-Level to Senior-Level ๐ Work Arrangement: Fully RemoteAbout UsWe are an operations and asset-performance organization focused on maximizing equipment reliability, operational availability, maintenance effectiveness, safety, and lifecycle value.Our teams collaborate across Reliability Engineering, Maintenance, Engineering, Operations, Asset Management, Facilities, Utilities, Quality, HSE, Procurement, and Finance to improve asset performance and reduce unplanned downtime.The RoleWe are seeking an experienced Reliability Manager to lead reliability engineering, asset-performance improvement, failure analysis, preventive and predictive maintenance strategies, reliability programs, and continuous improvement initiatives.The successful candidate will establish reliability standards, analyze equipment failures and performance data, develop reliability-centered maintenance strategies, optimize maintenance programs, and lead initiatives that improve availability, reduce downtime, and lower total cost of ownership.Key ResponsibilitiesLead and manage the organization's reliability function.Develop and implement a comprehensive reliability strategy aligned with operational and business objectives.Establish reliability objectives, standards, processes, and performance targets.Develop asset-reliability improvement programs across facilities, plants, infrastructure, equipment, and operational assets.Collaborate with Maintenance, Engineering, Operations, Asset Management, Facilities, Utilities, HSE, Quality, Procurement, and Finance.Establish reliability-management frameworks and continuous-improvement processes.Develop reliability standards for critical equipment and systems.Identify critical assets and establish appropriate reliability strategies.Conduct asset criticality assessments.Classify assets according to safety, production, operational, financial, environmental, and customer impact.Establish criticality-based maintenance and monitoring strategies.Develop and maintain asset criticality matrices.Review asset failure history and operational consequences.Identify equipment with unacceptable failure frequency or business impact.Develop improvement plans for high-criticality assets.Analyze equipment failures and recurring breakdowns.Lead Root Cause Analysis (RCA) investigations.Use structured methodologies such as 5 Whys, Fault Tree Analysis, Fishbone Analysis, FMEA, and other appropriate techniques.Identify immediate, contributing, and systemic causes of failures.Develop corrective and preventive actions.Track RCA actions through completion.Verify effectiveness of corrective actions.Identify recurring failure modes and systemic reliability issues.Develop Failure Modes and Effects Analyses (FMEA) for critical equipment and systems.Evaluate failure consequences, probabilities, detectability, and risk.Develop mitigation strategies for significant failure modes.Maintain and periodically update FMEA documentation.Implement Reliability-Centered Maintenance (RCM) programs.Review existing preventive-maintenance programs for effectiveness.Eliminate unnecessary or ineffective maintenance tasks.Introduce condition-based and predictive-maintenance strategies where appropriate.Optimize maintenance intervals based on equipment condition, failure patterns, manufacturer requirements, and operational data.Develop maintenance strategies based on asset criticality and failure consequences.Review preventive-maintenance compliance and effectiveness.Monitor repeat failures following maintenance interventions.Analyze maintenance-related downtime and recurring work orders.Develop strategies to improve maintenance quality.Establish predictive-maintenance programs.Evaluate condition-monitoring technologies such as vibration analysis, thermography, oil analysis, ultrasound, motor-current analysis, and other applicable techniques.Develop condition-monitoring plans for critical equipment.Analyze condition-monitoring results and identify developing failures.Establish appropriate alarm thresholds and escalation criteria.Coordinate with Maintenance and Engineering teams to plan interventions before equipment failure.Evaluate opportunities for online condition monitoring and remote diagnostics.Monitor equipment health using IoT sensors and digital technologies where appropriate.Develop asset-performance dashboards and reliability analytics.Track equipment availability, reliability, downtime, failure frequency, maintenance effectiveness, and lifecycle performance.Analyze Mean Time Between Failures (MTBF).Analyze Mean Time To Repair (MTTR).Monitor equipment availability and utilization.Track Overall Equipment Effectiveness (OEE) where applicable.Analyze failure rates and downtime trends.Identify statistical patterns and reliability deterioration.Develop predictive models and risk indicators where appropriate.Establish reliability data-quality standards.Ensure failure codes, equipment hierarchies, work-order information, and maintenance records are accurately captured.Improve reliability data within CMMS, EAM, ERP, and other asset-management systems.Develop standardized equipment failure and downtime coding.Ensure maintenance history is complete and reliable.Establish appropriate asset hierarchies and functional locations.Review equipment operating data and identify performance deviations.Develop asset-health indicators.Establish early-warning indicators for critical assets.Monitor degradation trends and intervention requirements.Develop reliability risk registers.Identify assets with high operational, safety, financial, environmental, or regulatory risk.Quantify reliability-related business risks.Develop mitigation and contingency strategies.Coordinate with Risk Management and Business Continuity teams on critical asset risks.Assess single points of failure.Develop redundancy, backup, and resilience strategies where appropriate.Support business-continuity planning for critical equipment and infrastructure.Analyze spare-parts requirements for critical assets.Establish critical-spares strategies.Review spare-parts availability, lead times, obsolescence, and inventory levels.Coordinate with Procurement, Stores, and Maintenance teams to optimize spare-parts availability.Balance inventory costs against equipment criticality and failure consequences.Identify long-lead and high-risk components.Develop contingency plans for critical spare parts.Key Performance Indicators (KPIs)Performance will be measured through a combination of reliability, maintenance, operational, financial, and risk KPIs, including:Asset availabilityEquipment reliabilityMean Time Between Failures (MTBF)Mean Time To Repair (MTTR)Unplanned downtimePlanned versus unplanned maintenanceEquipment failure frequencyRepeat-failure rateOverall Equipment Effectiveness (OEE)Critical-asset availabilityPreventive-maintenance effectivenessPredictive-maintenance coverageCondition-monitoring effectivenessRCA completion rateRCA action closure rateRecurring-failure reductionMaintenance cost per assetMaintenance cost reductionDowntime cost reductionCritical-spares availabilityMaintenance backlogAsset-health improvementReliability-project ROIWarranty-recovery performanceSafety-critical equipment complianceReliability-related incidentsAsset lifecycle-cost improvementCandidate ProfileThe successful candidate should have experience in reliability engineering, maintenance management, asset management, engineering, manufacturing, utilities, facilities, energy, infrastructure, or a closely related discipline.This role is suitable for a technically strong professional who can combine engineering analysis, maintenance strategy, operational understanding, data analytics, and business judgment to improve asset reliability.What You'll BringPrevious experience in reliability engineering, maintenance, asset management, engineering, or a related discipline.Proven experience managing reliability-improvement programs.Strong understanding of reliability engineering principles and asset-performance management.Experience with Root Cause Analysis, FMEA, RCM, criticality assessment, and failure analysis.Strong understanding of preventive, predictive, and condition-based maintenance.Experience analyzing MTBF, MTTR, availability, failure rates, downtime, and maintenance performance.Strong knowledge of mechanical, electrical, instrumentation, or industrial systems as applicable.Experience with CMMS, EAM, ERP, or asset-management systems.Strong data-analysis and problem-solving capabilities.Advanced Microsoft Excel skills.Experience with reliability analytics and performance dashboards.Knowledge of predictive-maintenance technologies is advantageous.Experience with vibration analysis, thermography, oil analysis, ultrasound, or other condition-monitoring methods is advantageous.Familiarity with Lean Six Sigma, Kaizen, or continuous-improvement methodologies is advantageous.Experience with Power BI, Tableau, SQL, Python, or other analytical tools is advantageous.Strong understanding of asset lifecycle management and total cost of ownership.Experience managing critical spare parts and maintenance strategies.Strong project-management and business-case development skills.Excellent communication and stakeholder-management capabilities.Strong leadership and team-management skills.Ability to work independently and effectively within a fully remote environment.Relevant degree in Mechanical Engineering, Electrical Engineering, Reliability Engineering, Industrial Engineering, Maintenance Engineering, or a related discipline.Professional certifications such as Certified Reliability Engineer (CRE), CMRP, CRL, or equivalent qualifications are advantageous.Knowledge of UAE/GCC industrial, facilities, utilities, infrastructure, or energy sectors is advantageous.Arabic language proficiency is advantageous.