أبلاي إيدج ابدأ البحث عن عمل

NOC Engineer - Enterprise Operations

K20s - Kinetic Technologies Private Limited · Sharjah, Sharjah Emirate, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Role SummaryNOC Engineer will be responsible for continuous enterprise-wide monitoring and first-line operational support across websites, applications, network devices, Windows servers, cloud services, and end-user platforms. The role requires hands-on operation of engine-based monitoring products and Site24x7, rapid validation and escalation of alerts, accurate BMC ticket management, adherence to ITIL processes, basic network and Windows troubleshooting, vendor coordination, shift reporting, and disciplined handover. The position supports high availability, performance, reliability, and timely restoration of business services in a 24x7 enterprise environment. Key Responsibilities 2.1 24x7 Enterprise MonitoringContinuously monitor enterprise websites, applications, APIs, network devices, Windows servers, cloud services, and critical business platforms across assigned shifts.Operate engine-based monitoring products and Site24x7 dashboards, consoles, probes, agents, and alert queues on a hands-on basis.Monitor service availability, response time, transaction performance, CPU, memory, disk, interfaces, processes, services, and other defined health indicators.Acknowledge alerts promptly and follow documented procedures, runbooks, maintenance calendars, and escalation matrices.Identify monitoring tool failures, disconnected agents, stale data, disabled checks, and dashboard anomalies and escalate or correct them as authorized.Maintain awareness of planned changes, maintenance windows, known issues, and business-critical periods that may affect alert handling.2.2 Alert Validation & First-Line ResponseValidate alerts to determine whether they are genuine, duplicate, expected, transient, or related to an approved maintenance activity.Perform first-line diagnostics using approved checks for connectivity, DNS, URLs, ports, Windows services, processes, event logs, resource usage, and application availability.Attempt documented recovery actions such as service checks, approved restarts, cache or connection validation, and basic remediation within assigned authority.Capture timestamps, screenshots, logs, error messages, performance data, and troubleshooting evidence before escalation.Confirm service recovery through monitoring tools and user or technical-team validation before recommending ticket closure.Escalate immediately when service impact, alert severity, business criticality, or troubleshooting results require specialized support.2.3 Incident & Ticket ManagementCreate and manage incidents, service requests, events, and related support records in BMC Service Management / BMC Helix / BMC Remedy.Categorize and prioritize tickets accurately based on impact, urgency, affected service, configuration item, and enterprise standards.Maintain complete work notes covering alert details, checks performed, actions taken, teams contacted, updates received, and restoration evidence.Track assigned and escalated tickets until resolution, ensuring follow-up within SLA and timely reassignment when ownership changes.Provide regular updates for high-priority incidents and immediately notify the NOC Lead of potential Priority 1 or Priority 2 events.Close tickets only after resolution verification, correct categorization, meaningful closure notes, and required stakeholder confirmation.2.4 Website & Application MonitoringUse Site24x7 or equivalent platforms to monitor website and application uptime, page response, API availability, SSL certificate validity, DNS resolution, and synthetic transactions.Investigate HTTP errors, timeout conditions, slow response, failed content checks, certificate alerts, and transaction failures using approved procedures.Compare alert behaviour across locations, probes, dependencies, and time periods to support accurate diagnosis and escalation.Coordinate with application, web, database, cloud, and vendor teams when faults require second-line or specialist support.Document recurring website or application failures and provide evidence for problem management and monitoring enhancement.2.5 Infrastructure Monitoring & Basic TroubleshootingMonitor network availability, interface status, latency, packet loss, bandwidth indicators, device health, and connectivity alarms.Perform basic troubleshooting using TCP/IP, DNS, DHCP, routing and switching fundamentals and standard diagnostic utilities.Monitor Windows Server health, services, event logs, CPU, memory, disk capacity, processes, scheduled tasks, and operating system alerts.Provide basic support checks for Windows 10/11 desktops, endpoint connectivity, authentication, services, and standard enterprise applications.Engage network, server, storage, cloud, cybersecurity, database, application, and end-user support teams with clear evidence and impact details.2.6 Escalation & Vendor CoordinationEscalate incidents to internal resolver teams and third-party vendors in accordance with severity, support scope, SLA, and escalation procedures.Raise vendor cases with complete technical information, affected services, business impact, logs, screenshots, timestamps, and troubleshooting completed.Follow up with vendors and support teams until service restoration, workaround, root cause information, or next action is received.Immediately inform the NOC Lead of delayed responses, SLA risks, recurring faults, or vendor actions that may affect business services.Maintain accurate vendor case references, support contacts, status updates, and escalation records in the relevant ticket.2.7 Shift Handover & Operational DisciplineReport on time for assigned shifts and maintain full operational attention throughout the shift, including nights, weekends, and public holidays.Complete structured handovers covering active alerts, open incidents, pending vendor responses, planned changes, service risks, and required follow-ups.Follow standard operating procedures, runbooks, checklists, access controls, communication protocols, and escalation matrices.Maintain professional communication, strong responsibility, ownership, punctuality, and a solution-oriented can-do attitude.Notify the NOC Lead immediately of coverage issues, access problems, tool failures, unusual alert patterns, or any risk to continuous monitoring.2.8 Reporting & DocumentationPrepare accurate shift reports, daily status updates, incident summaries, availability observations, and pending-action lists.Contribute data for SLA, KPI, MTTA, MTTR, ticket aging, alert volume, recurring incident, and service availability reports.Update SOPs, runbooks, knowledge articles, escalation details, troubleshooting guides, and known-error documentation as assigned.Identify recurring alerts, false positives, monitoring gaps, and process issues and report them to the NOC Lead with supporting evidence.Maintain clear, concise, and audit-ready records for all operational activities.2.9 Continuous Improvement & Automation SupportSupport tuning of thresholds, alert rules, checks, dashboards, notification groups, and escalation workflows under approved change control.Participate in onboarding new services and validate that monitoring, ticketing, support ownership, escalation, and documentation are ready before go-live.Assist with testing monitoring enhancements, scripts, reports, dashboards, integrations, and automated health checks.Suggest practical improvements that reduce alert noise, improve diagnosis, shorten restoration time, and strengthen shift efficiency.Participate in training, knowledge transfer, drills, service reviews, and post-incident improvement actions. Technical Skills 3.1 Enterprise Monitoring PlatformsSite24x7 - website, application, server, network, cloud, synthetic, and availability monitoring.Engine-based monitoring products, collectors, agents, probes, dashboards, event consoles, and alert workflows.ManageEngine OpManager, AppDynamics, or equivalent enterprise monitoring platforms.Alert acknowledgement, threshold interpretation, basic dashboard administration, maintenance suppression, and monitoring health validation.3.2 ITSM & ITIL OperationsBMC Service Management / BMC Helix / BMC Remedy ticket creation, assignment, escalation, work notes, resolution, and closure.Working understanding of ITIL incident, problem, change, event, service request, and knowledge management processes.SLA awareness, ticket prioritization, escalation timelines, shift handover, and operational documentation.3.3 Network FundamentalsTCP/IP, DNS, DHCP, NAT, VPN, ports and protocols, subnetting basics, and connectivity validation.Basic routing and switching, LAN/WAN concepts, interface status, latency, packet loss, and bandwidth indicators.Troubleshooting utilities including ping, traceroute, nslookup, ipconfig, netstat, and telnet/test-netconnection.3.4 Server & Desktop PlatformsWindows Server fundamentals, services, event logs, performance counters, storage, processes, and scheduled tasks.Windows 10/11 desktop fundamentals, endpoint connectivity, authentication, services, applications, and remote support basics.Basic awareness of Active Directory, virtualization, cloud services, backup systems, storage, and endpoint security.3.5 Website & Application TechnologiesHTTP/HTTPS, URLs, DNS resolution, SSL/TLS certificates, ports, APIs, response codes, response time, and content checks.Website and application availability, synthetic transactions, user journey checks, dependency awareness, and performance indicators.Basic interpretation of monitoring, web server, operating system, and application logs.3.6 Reporting & Productivity ToolsMicrosoft Excel, Word, PowerPoint, email, collaboration tools, and standard operational reporting templates.Basic dashboard interpretation, trend comparison, ticket reporting, availability reporting, and data quality checks.Ability to prepare concise shift summaries, incident timelines, and management-ready status updates.3.7 Automation & ScriptingFamiliarity with PowerShell, Python, batch scripting, SQL, APIs, or similar tools is preferred.Ability to execute approved scripts and automated health checks and accurately interpret their outputs.Basic awareness of monitoring-to-ticketing integrations, notifications, and report automation.3.8 Project & Operational SkillsIncident troubleshooting, root cause evidence collection, change awareness, risk identification, and documentation.Ability to prioritize multiple alerts and tickets in a high-volume 24x7 environment.Clear communication, teamwork, customer focus, shift discipline, and ability to work under pressure. Educational QualificationsBachelor's degree or diploma in Computer Science, Information Technology, Computer Engineering, Electronics, or a related field. Preferred CertificationsITIL Foundation certification.CompTIA Network+ or Cisco CCNA certification.BMC Helix / BMC Remedy user or administration training.Microsoft Windows Server, Azure Fundamentals, AWS Cloud Practitioner, or equivalent certification.Site24x7, SolarWinds, PRTG, ManageEngine, AppDynamics, or equivalent monitoring training or certification. Experience Requirements2-4 years of experience in enterprise NOC, infrastructure monitoring, application support, service desk, or technical operations.Hands-on experience monitoring websites and applications in a continuous 24x7 environment.Practical familiarity with engine-based monitoring products, Site24x7, and enterprise monitoring dashboards.Experience using BMC Service Management or a comparable ITSM ticketing platform.Working experience with incident handling, alert escalation, SLA tracking, shift handover, and operational reporting.Basic troubleshooting experience across TCP/IP, DNS, DHCP, Windows Server, and Windows 10/11 environments.Experience coordinating with internal support teams or external vendors for incident resolution.Willingness and ability to work rotating shifts, including nights, weekends, and public holidays.Experience in highly available or business-critical environments is preferred. Core CompetenciesStrong alert monitoring, troubleshooting, and analytical skills.Customer-focused service delivery and ownership mindset.Clear verbal and written communication skills.Reliable, punctual, responsible, and solution-oriented behaviour.Strong attention to detail and ticket documentation quality.Team collaboration and willingness to follow escalation processes.Ability to prioritize and work calmly under pressure.Process discipline and adherence to SOPs, SLAs, and shift procedures.Continuous learning and willingness to adopt new toolSkills: management,cloud,windows,basic,troubleshooting