Apply Edge Start your job search

Observability Engineer – Telemetry Extension & ADOT Pipeline

Intellias · Cairo, Egypt

Apply & track with Apply Edge
About the RoleWe are looking for an Observability Engineer – Telemetry Extension & ADOT Pipeline to build and extend telemetry capabilities for enterprise AI and cloud-native platforms.In this role, you will focus on designing scalable observability pipelines, extending OpenTelemetry instrumentation, and developing reusable telemetry components that provide actionable insights into distributed systems.You will work closely with platform, AI, and DevOps teams to create flexible and configurable telemetry solutions that support performance monitoring, troubleshooting, distributed tracing, and operational excellence at enterprise scale.Project OverviewOur customer is a multinational corporation with more than a century of history, operating in 180+ countries and serving more than 1 billion consumers worldwide.The organization is undertaking a major transformation focused on introducing a new generation of Reduced-Risk Products (RRPs) and developing an innovative digital ecosystem around IoT, eCommerce, digital marketing, and AI.Its IT platform hosts 700+ applications, creating significant requirements for scalable, reliable, and standardized engineering capabilities.Intellias supports the engineering of a comprehensive software ecosystem for a game-changing IoT product at the intersection of innovative consumer experiences and cutting-edge technology.As an Observability Engineer, you will join the Core Architecture Team and contribute to the architecture and implementation of best practices across the Digital Engineering Enterprise Platform.The platform provides engineering teams with reusable services, technologies, and practices that accelerate software development and operations while addressing common SDLC challenges and ensuring security, compliance, and operational excellence.Key ResponsibilitiesObservability & Telemetry EngineeringDesign and implement observability solutions for AI platforms, agent-based systems, and distributed cloud-native applications.Instrument Python services using OpenTelemetry SDK and AWS Distro for OpenTelemetry (ADOT).Develop reusable telemetry components, SDK extensions, and instrumentation patterns.Implement custom metrics, spans, and configurable telemetry capabilities.Establish standardized telemetry patterns that can be adopted across multiple platform services.Build telemetry solutions that provide actionable insights into application performance, reliability, and behavior.ADOT & OpenTelemetry PipelinesDesign, configure, and maintain ADOT Collector pipelines.Configure ADOT components including:ReceiversProcessorsExportersConfigure telemetry collection, propagation, processing, and export across AWS environments.Develop and maintain OpenTelemetry instrumentation and extension patterns.Implement and maintain distributed tracing using OpenTelemetry and AWS X-Ray.Ensure reliable telemetry delivery across distributed services and agent workflows.AWS ObservabilityBuild and maintain telemetry export pipelines to AWS CloudWatch Metrics and AWS X-Ray.Implement CloudWatch metrics, dashboards, logs, and alerting strategies.Develop custom CloudWatch metrics and dimensions to provide meaningful operational insights.Support monitoring and troubleshooting across AWS-based distributed systems.Integrate observability capabilities with AWS AgentCore Observability and related platform services.AI & Agent ObservabilityEstablish monitoring standards for:AI agentsTool invocationsWorkflow executionService-to-service communicationInstrument AI services and LangChain-based applications to capture traces, metrics, and execution insights.Support observability across agent workflows and distributed AI systems.Ensure telemetry provides sufficient context for performance analysis, troubleshooting, and root-cause investigation.Collaborate with AI engineers to identify bottlenecks and optimization opportunities.Trace Context & Distributed SystemsImplement and maintain W3C Trace Context and Baggage propagation across applications, APIs, and agent workflows.Ensure trace continuity across distributed services and AI agent interactions.Develop telemetry patterns that support end-to-end distributed tracing.Troubleshoot gaps in trace propagation and telemetry collection.Support reliable correlation between logs, metrics, and traces.SDK & Library DevelopmentDesign and develop reusable Python SDK libraries for telemetry and observability.Establish extensibility patterns that allow platform teams and customers to customize telemetry.Develop configurable telemetry components and YAML-driven configuration schemas.Maintain clean, reusable, well-documented SDK interfaces.Build automated tests to validate telemetry behavior and configuration.Collaboration & Continuous ImprovementPartner with AI engineers, platform engineers, DevOps teams, and architects to define observability requirements.Identify performance, reliability, and operational challenges and propose telemetry-driven solutions.Contribute to observability standards, platform guidelines, and engineering best practices.Create technical documentation for telemetry components, configuration, pipelines, and integration patterns.Continuously improve the scalability, reliability, and usability of the enterprise observability platform.Required Qualifications & ExperienceBachelor’s degree in computer science, Software Engineering, Information Technology, or a related field.4+ years of experience in observability, platform engineering, DevOps, or a related engineering discipline.Strong hands-on experience with OpenTelemetry and/or ADOT.Experience designing and configuring ADOT / OpenTelemetry Collector pipelines, including processors and exporters.Strong understanding of OpenTelemetry SDK instrumentation and extension patterns.Experience developing Python SDKs, libraries, or reusable components.Experience implementing custom metrics, spans, and configurable telemetry.Hands-on experience with AWS CloudWatch Metrics and AWS X-Ray.Experience designing telemetry pipelines for distributed cloud-native applications.Strong understanding of distributed tracing and telemetry propagation.Experience with W3C Trace Context and Baggage propagation.Strong understanding of YAML-based configuration and schema design.Experience with structured logging, metrics, and distributed tracing.Strong troubleshooting, analytical, and problem-solving skills.Strong English communication and collaboration skills.Nice-to-Have QualificationsFamiliarity with Strands or LangGraph frameworks.Experience with AWS AgentCore Observability.Experience implementing customer-facing SDK extensibility patterns.Hands-on experience with CloudWatch custom metrics and dimensions.Experience instrumenting LangChain applications.Experience with AI agents, LLM applications, or multi-agent architectures.Knowledge of AWS Distro for OpenTelemetry deployment and operational best practices.Experience working with large-scale enterprise platforms.Familiarity with CI/CD and infrastructure-as-code practices.Key Technical SkillsOpenTelemetry | AWS Distro for OpenTelemetry (ADOT) | ADOT Collector | OpenTelemetry SDK | Python | Python SDK Development | AWS CloudWatch | AWS X-Ray | Distributed Tracing | Custom Metrics | Custom Spans | Telemetry Pipelines | W3C Trace Context | Baggage Propagation | YAML Configuration | Structured Logging | AI Observability | AWS AgentCore Observability | LangChain | LangGraph | Cloud-Native ObservabilityWhy This Position?Build enterprise-scale observability: Design telemetry capabilities supporting a global digital engineering ecosystem.Work at the intersection of AI and observability: Build monitoring and tracing capabilities for AI agents, tools, workflows, and distributed services.Cutting-edge technology: Work hands-on with OpenTelemetry, ADOT, AWS X-Ray, CloudWatch, and modern AI frameworks.Reusable platform impact: Develop SDKs, telemetry components, and standards that can be adopted across the enterprise.Global scale: Support technology platforms spanning 700+ applications across 180+ countries.Architecture influence: Join the Core Architecture Team and help establish enterprise observability standards and best practices.Solve complex distributed-system challenges: Improve visibility, troubleshooting, performance, and reliability across cloud-native and AI workloads.Intellias partnership: Join through Intellias and contribute to a major global technology transformation.EducationBachelor’s degree in computer science, Software Engineering, Information Technology, or a related technical field.