أبلاي إيدج ابدأ البحث عن عمل

Distributed Inference Engineer — LLM engine, Founding Team

Cognition Commons · Belgrade, Serbia

قدّم وتابع مع أبلاي إيدج
Company Description Cognition Commons is an early-stage organization focused on building advanced large language model (LLM) infrastructure and tools that make AI systems more accessible, reliable, and efficient. The team is committed to pushing the boundaries of distributed inference, model deployment, and scalable AI operations. As part of the founding team, individuals have the opportunity to shape core technical architecture, engineering culture, and long-term product direction. Cognition Commons values collaborative problem-solving, rigorous engineering standards, and transparent decision-making. The company supports remote work and aims to attract builders who are excited about creating foundational AI technology from the ground up.Role Description As a Distributed Inference Engineer — LLM engine on the founding team, you will design, build, and optimize distributed systems that serve large language models at scale. You will develop and maintain core inference infrastructure, including model sharding, load balancing, caching, and low-latency networking components for remote production environments. You will work closely with other engineers to profile performance, reduce inference costs, improve throughput, and enhance reliability across diverse hardware configurations. Your day-to-day responsibilities will include writing production-grade code, implementing monitoring and observability, debugging complex distributed issues, and contributing to architectural decisions for the LLM engine. This is a full-time remote role that involves close collaboration via asynchronous communication, code reviews, and regular technical design discussions.Qualifications Strong software engineering skills in systems programming languages (e.g., Rust, C++, Go) and experience building high-performance, production-grade services.Experience with distributed systems concepts such as concurrency, fault tolerance, consensus, and scalable microservices architectures.Familiarity with machine learning model serving, GPU/accelerator utilization, and frameworks or runtimes used for LLM inference.Comfort with cloud infrastructure, containerization, and orchestration tools (e.g., Kubernetes, Docker, major cloud providers).Proficiency with monitoring, logging, and performance profiling tools to diagnose and resolve bottlenecks in distributed environments.Ability to write clear technical documentation, participate in rigorous code reviews, and communicate effectively in remote, distributed teams.Prior experience in early-stage startups or founding engineering teams, with a willingness to take ownership and operate in a fast-changing environment.Bachelor’s or advanced degree in Computer Science, Engineering, or a related field, or equivalent practical experience in systems and infrastructure engineering.