Senior Backend Engineer (AI Deployments)
Digital Place Vision Pte Ltd Β· Tangerang, Banten, Indonesia
Apply & track with Apply EdgeCompany DescriptionAt Digital Place Vision Pte Ltd, we are a team of visionaries on a mission to revolutionize the retail landscape through intelligent AI. We don't just build software, we engineer immersive, next-generation digital environments that turn complex business challenges into scalable, real-time opportunities. If you are passionate about the intersection of high-performance web applications, interactive AI, and elegant user experiences, we want to hear from you.Role DescriptionWe are seeking an autonomous, highly skilled Backend Engineer with approximately 5 years of professional experience to drive our real-time AI avatar platform's backend. You'll own the systems that turn a user's voice into a GPU-rendered, lip-synced avatar video streamed back in real time β spanning the lipsync inference pipeline, the conversational orchestration server, and the voice-cloning/TTS service, plus their deployment across GCP.You are a versatile backend engineer who thrives on asynchronous Python, real-time protocols, and GPU-backed inference services. You understand both the low-latency demands of streaming systems and the operational realities of running GPU workloads in production, allowing you to design services that are fast, resilient, and cost-aware.π οΈ What You'll DoOwn the Real-Time Streaming Backend: Develop and scale the WebSocket/Protobuf streaming layer that carries audio in and lip-synced video out with tight latency budgets.Build & Scale AI Orchestration Services: Extend the orchestration server (FastAPI, asyncio) that coordinates LLM, TTS/voice-cloning, and lipsync services per session.Manage GPU Inference Pipelines: Work with PyTorch/CUDA-based lipsync and voice-conversion models β pipeline allocation, warm pools, batching, and multi-user concurrency on shared GPUs.Develop a No-GPU Realistic Avatar Path: Design and build a CPU-only rendering path for realistic avatar output, so avatar sessions can run without dedicated GPU inference where full lipsync fidelity isn't required β reducing per-session cost and unlocking scale-to-zero / low-cost deployment tiersDesign & Integrate APIs: Build and evolve REST and WebSocket APIs, and Protocol Buffer message contracts shared across services and the frontend client.Own Deployment & Reliability: Manage Docker-based deployments to GCP Cloud Run/GCE and GPU cloud providers (RunPod, Vast.ai), including config promotion, health/readiness checks, and cold-start/scaling tradeoffs.Collaborate on System & API Design: Partner closely with the frontend engineer to co-design WebSocket/REST contracts and data flow between client and backend services.Operate Session & State Infrastructure: Work with Redis for session/state management and GCS/Vertex AI for storage and RAG-backed retrieval.β What We're Looking ForProven Track Record: 5+ years of professional software engineering experience, with a heavy emphasis on building complex, production-grade backend/distributed systems.Modern Python Mastery: Deep, production-tested expertise in Python 3.10+, asyncio, type hints, and FastAPI or an equivalent async web framework.Real-Time Systems Expertise: Practical experience implementing and debugging real-time communication β WebSockets, streaming protocols, and low-latency pipelines under concurrency.API & Message Contract Literacy: Proven competence in REST API design and Protocol Buffers/gRPC-style message contracts shared across services.Architectural Mindset: Experience designing multi-service/microservice architectures, including inter-service communication and session-state management (e.g., Redis).Cloud & Infra Fluency: Hands-on experience with GCP (Cloud Run, GCE, Cloud Storage, Artifact Registry) and Docker-based deployment pipelines; comfortable with Linux server administration and troubleshooting.GPU-Aware Engineering: Comfortable working alongside PyTorch/CUDA inference workloads β understanding GPU memory, batching, and concurrency constraints even if not training models yourself.Independence & Ownership: A self-starter mindset with the ability to take backend services from concept to deployment independently in a fast-paced environment.π Nice to HaveExperience with ML inference serving (PyTorch, ONNX Runtime, model warm-loading/batching strategies).Experience integrating LLM/genAI APIs (Vertex AI, Gemini, or equivalent) and RAG pipelines.Familiarity with audio/video streaming concepts (encoding, real-time transport, buffering/backpressure).Experience with GPU cloud providers beyond GCP (RunPod, Vast.ai) and cost/latency tradeoffs across them.Firebase or equivalent auth/identity platform experience.CI/CD pipeline design and implementation.π Why Join Us?Cutting-Edge Impact: Work directly on live, production-scale AI systems reshaping the retail industry.True Ownership: Lead front-end development with high autonomy, working within a fast-moving team that respects technical excellence.A Culture of Innovation: Join a collaborative environment where speed, creativity, and clean code are highly celebrated.π Learn more about us: digitalplace.aiπ Follow us on LinkedIn: Digital Place Vision Pte Ltdπ© Ready to apply? Send your updated CV to Career@digitalplace.ai (Indonesia citizens only) or reach out directlyβweβd love to connect!#AIJobs #Hiring #MachineLearning #RetailTech #MLOps #PyTorch #Python