Senior Backend Engineer - AI Platform
Scalifi Ai · Gurugram, Haryana, India
Apply & track with Apply EdgeSee what we're building: www.cognis-ai.com (product) | www.scalifiai.com (parent company)About Cognis AiCognis is an AI workspace that puts 20+ frontier models (GPT, Claude, Gemini, Grok, DeepSeek and more) into a single chat layer - with context that travels across model switches, true branching, chat-layer observability, and an adaptive UI that turns AI replies into working tools instead of walls of text. The platform is live in private beta and being rebuilt for public launch. Backend is Python/Django REST Framework. Built by Scalifi Ai, Gurugram.We are not competing with the model vendors. We bundle them and build the layers they don't ship: memory, observability, branching, and execution inside the tools work actually happens in.Work mode: In-office at our Gurugram office (Sector 65) - non-negotiable. Candidates should be in NCR or willing to relocate.Joining: Immediate to 30-day joiners strongly preferred.Why this role is differentYou'll join the founding backend team, working directly with the founder (who built the current platform). This is not a maintain-the-legacy role and not a research role - it's an ownership role. The architecture decisions on our biggest roadmap features will be yours to make, defend, and ship. You'll also help interview and mentor our next backend hire.What you'll ownBase tool layer: web search, image generation, video generation as first-class tools inside the agent loop - each with per-step observabilityFiles & media: file and image handling support across providersMemory: our persistent, user-editable memory system - pinning, releasing, compaction, and per-item token accountingRAG & knowledge bases: chunking, retrieval, reranking, and retrieval-quality evaluation, wired into the workspaceExecution environments (upcoming): code interpreters for context-efficient, parallel tool calls; sandboxed runtimes and a workspace filesystem for agent file operationsProduct analytics: deep, Cognis-specific analytics (beyond our generic platform layer)Platform growth work (shared, especially early on): adding new LLM providers and models - message serialization, streaming, tool-call handling - and new third-party integrations on our integrations layer, where OAuth, tool search, and execution are already built. You'll help turn these into runbooks our next hire executes.What you'll ship in your first 90 daysOne base tool end-to-end (e.g., web search with full step-level observability)One new LLM provider - normalized message formats, streaming, tool-call deltasThe first version of the memory/compaction architecture, co-designed with the founderMust-haves3-6 years of backend engineering with strong Python2+ years building LLM-powered features that real users used in production - not tutorials, not POCsHands-on with provider APIs (OpenAI, Anthropic, Google) - you can describe what a tool-call delta looks like in a raw streaming responseAgent orchestration in production: LangGraph/LangChain, or an equivalent system you built yourself on raw SDKsContext management: message history design, summarization/compaction trade-offs, token economics, prompt cachingRAG beyond hello-world: chunking strategy, hybrid retrieval, reranking, and how you evaluated retrieval qualitySolid REST API design; streaming via SSE/WebSocketsDjango/DRF preferred - deep FastAPI/Flask experience with willingness to work in DRF is fineStrong plusesMulti-provider or model-router systems; provider failoverTool-calling ecosystems and integration platforms (e.g., MCP)Sandboxed code execution - containers, microVMs, or code-interpreter style toolingVector databases (pgvector, Qdrant, Pinecone) and embedding pipelinesEval tooling and LLM observabilityEvent/analytics pipelinesEarly-stage startup experience or meaningful open-source workStackPython | Django REST Framework | LangGraph + LangChain (agent orchestration) | PostgreSQL | Redis | WebSockets/SSE | OpenAI/Anthropic/Google SDKs | a unified third-party integrations layer - plus whatever you make the case for.What we offerFounding-team scope with direct, daily work alongside the founder - zero bureaucracy, weekly shippingYour fingerprints on every major system of an AI product at exactly the moment the category is being definedCompetitive compensation, reviewed as the company growsIMPORTANT - How to apply (read this or your application will not be shortlisted)At the very top of your resume (or in your cover note / screening answer, if the platform has one), include 4-6 lines about ONE LLM feature you shipped to real users: what it was, the ugliest edge case you hit, and how you handled it. Add your GitHub/portfolio link if you have one. Applications without this will not be shortlisted.Keywords: LLM, GenAI, Generative AI, AI Agents, Agentic AI, LangChain, LangGraph, RAG, OpenAI, Anthropic, Claude, Gemini, Python, Django, DRF, WebSockets, SSE, PostgreSQL, Redis, Vector Database, Prompt Caching, Code Interpreter, Sandboxing, MCP