أبلاي إيدج ابدأ البحث عن عمل

AI Engineer

Intelligent Inference · Islamabad, Islāmābād, Pakistan

قدّم وتابع مع أبلاي إيدج
Company Description Intelligent Inference is building software and engineering tools that enable people to run sovereign AI inference, giving users ownership, control, and transparency over their LLM workloads. Starting with our flagship product "i2" a sovereign gateway where you can access leading open-weight AI models with agent harness and tooling, sovereignty and control. We help developers and teams run open-source models with fast, cost-effective inference and a sovereign gateway. The company offers a full-stack, self-service AI inference platform for developers and enterprises, featuring optimized open-source model APIs with regional deployment, low latency, and reduced costs. Its platform supports dedicated inference and model deployments, fine-tuning and training, as well as bring-your-own-key and bring-your-own-model workflows. Intelligent Inference provides a comprehensive control panel with deep visibility into token usage, logs, and cost attribution, plus localized billing in home currencies without foreign exchange charges. The first sovereign regional deployment is in Pakistan, with plans to expand across the AMEA region by 2030.AI Engineer, Inference PlatformIntelligent Inference (Private) Limited - 'i2'i2 currently runs a commercial inference platform inside Pakistan. We self-host open-weight frontier models on Nvidia GPUs as well as Huawei Ascend NPU infrastructure, a hybrid gateway that allows you to call model APIs on both CUDA and CANN. We expose the model APIs through an OpenAI-compatible API with all native features such as streaming etc, and bill in PKR. Our customers include developers (B2C), freelancers, startups and software-houses, regulated enterprises, banks, telcos and government, who cannot legally route workloads to foreign APIs, and developers who want frontier open-weight model access in the optimal way and without USD payments. The platform is in commercial operation. Our PK model end-points are completely local, data stays resident, and the serving stack is ours end to end. That last part is what makes this an engineering job rather than a reselling job.The roleYou will own the path a request takes from the API edge to the accelerator and back. That means the Go gateway that fronts our models, and the serving engines behind it.Concretely:Extend and operate our Go API gateway: routing, authentication, rate limiting, quota enforcement, per-model metering, streaming, request shaping, failure handling.Deploy and tune open-weight models on our Ascend 910B clusters using the CANN/MindIE stack, and on vLLM and SGLang where applicable.Own serving performance: continuous batching, KV cache configuration, tensor and expert parallel layouts, prefill and decode behaviour, context length trade-offs.Run quantisation work. Our path is BF16 to INT8 (W8A8) to INT4 (W4A16). You will run the conversions, measure the quality cost, and be honest about it.Work within a hard architectural constraint: an 8-chip HCCS coherent domain per cluster. Large MoE models have to fit that shape or be sharded around it. This is the interesting part of the job.Build and operate retrieval infrastructure for enterprise deployments, including vector stores, embedding endpoints and the ingestion path around them.Instrument everything. Per-endpoint latency, tokens per second, time to first token, utilisation, cost per million tokens. If it is not measured we do not claim it.Support dedicated enterprise deployments, including air-gapped and restricted-network environments where you cannot assume internet egress or a friendly package manager.What we need you to already haveStrong Go. You have written and operated production services in it, not just read about it.Real experience serving LLMs in production with vLLM, SGLang, TensorRT-LLM, TGI or equivalent. You should be able to explain why throughput fell when concurrency rose, and what you did about it.A working mental model of transformer inference: attention, KV cache, batching, the prefill and decode split, why MoE routing changes the memory picture.Linux, containers, and comfort at the level of drivers, kernels and hardware topology. This stack breaks below the Python line.Python for model and evaluation work.The habit of measuring before claiming. We mark unverified numbers as unverified, in customer documents and internally.What will help but is not requiredHuawei Ascend, CANN or MindIE experience. Very few people have this. If you do, say so early.Quantisation experience, particularly INT8 and INT4 on non-NVIDIA silicon.Kubernetes, and experience running stateful GPU or NPU workloads on it.Experience with regulated deployments, audit logging, or data residency requirements.Open source contributions to any inference or serving project.What we are not asking forCUDA fluency as a gate. Most of our candidates come from NVIDIA-only backgrounds and that is fine. Ascend is different silicon with a different toolchain, no FP8 or INT4 support in hardware on the 910B, and a smaller ecosystem. We will teach it. What does not transfer is patience with undocumented behaviour, so bring that.Model training or research experience. This is a serving and systems role.Prompt engineering or application layer work.Why this is worth your timeThere are perhaps a handful of teams in this country doing inference at the hardware level rather than wrapping someone else's API. You will work on a stack with genuine constraints, for customers with genuine compliance requirements, at a company where the serving layer is the product rather than a cost line. You will also learn an accelerator architecture that very few engineers anywhere have touched, at a point where that is becoming commercially relevant.Compensation and termsDepends on qualification. PKR 200,000 - 500,000 with minimum of 2 years of production level generative AI development. Equity available for the right candidate. How to applySend your CV to info@intelligentinference.ai with the subject line "AI Engineer, Inference Platform".In the body, in a few paragraphs, tell us about one inference workload you made meaningfully faster or cheaper. What the baseline was, what you changed, what you measured afterwards, and what the trade-off cost you. We read these before we read the CV.