AI Architect
OREDATA · Istanbul, Türkiye
Apply & track with Apply EdgeOREDATA is a Digital Transformation & IT Consulting firm with 10+ years of proven expertise and hundreds of successfully implemented projects across the EMEA region. When you join the OREDATA team, you'll be working hand-in-hand with experts focused on tackling digital, operational, analytical & data science challenges with the greatest impact. We foster collaboration with proximity, an agile and autonomous approach and best practices and guiding principles.We are looking for an Data Scientist (LLM) to join our team within a leading company in the aviation industry. ✈️Read more: https://medium.com/@oredata-engineeringApply and be part of our exciting journey!ResponsibilitiesDesign, develop, test, and productionize LLM-based and Retrieval-Augmented Generation (RAG) solutions for enterprise use casesTake an active, hands-on role in the development of internal chatbot, conversational AI, knowledge assistant, and agentic AI products, from POC/MVP through production readinessOwn and contribute to the technical architecture of enterprise LLM solutions, including model selection, deployment, serving, routing, evaluation, monitoring, and integration with internal AI platforms and applicationsDeploy, operate, and optimize open-weight and commercial LLMs, with a particular focus on on-premise and private infrastructure. This includes taking a lead role in standing up and configuring on-premise platforms (such as Red Hat OpenShift AI) from scratch when necessaryEvaluate and select appropriate models based on use-case requirements, considering quality, latency, throughput, infrastructure requirements, cost, security, licensing, and operational constraintsOptimize LLM inference and infrastructure utilization through techniques such as quantization, batching, caching, model serving optimization, GPU resource management, and appropriate workload allocationAct as an advocate for AI infrastructure efficiency (AI FinOps), optimizing compute costs by balancing model performance, hardware allocation (e.g., Multi-Instance GPU), and semantic routing strategiesDesign and improve model routing and semantic routing mechanisms to ensure requests are handled by the most appropriate model based on use case, complexity, performance, and resource requirementsDesign and support agentic and tool-calling architectures, ensuring that the appropriate models, tools, and enterprise services are selected and invoked reliably and securelyContribute to the evolution of the organization's AI Gateway and shared AI platform capabilities, including model access, authorization, quotas, routing, governance, observability, and usage controlsWork with structured and unstructured data to prepare, retrieve, enrich, and optimize knowledge sources used by AI applications, including embeddings, vector search, hybrid retrieval, re-ranking, chunking, and context managementEstablish and improve LLM evaluation and monitoring practices, including benchmark datasets, offline and online evaluation, regression testing, output quality analysis, hallucination monitoring, and performance metricsCollaborate closely with platform, infrastructure, data, software engineering, architecture, product, and business teams to translate business requirements into scalable and operationally feasible AI solutionsProvide technical direction and architectural guidance on GPU capacity, AI infrastructure utilization, model serving technologies, and platform evolution, while remaining actively involved in implementation when requiredStay current with developments in Generative AI, LLMs, agentic systems, model serving, inference optimization, RAG architectures, and AI infrastructure, and evaluate their practical applicability within the enterprise environmentRequirementsMust-Haves (Minimum Qualifications)Minimum 7 years of professional experience in Artificial Intelligence, Machine Learning, Data Science, Software Engineering, Data Engineering, AI Platform Engineering, or related technical rolesStrong hands-on experience with Large Language Models (LLMs) and Generative AI solutions, including experience taking AI systems beyond experimentation and into production environmentsProven experience deploying, serving, operating, or optimizing LLMs, preferably in on-premise, private cloud, or enterprise containerized environmentsProven experience designing and scaling AI systems for high-traffic, high-concurrency environments, ensuring latency control and graceful degradation under heavy load (e.g., handling traffic spikes)Practical understanding of GPU-based LLM inference and the key factors affecting GPU memory utilization, throughput, latency, concurrency, and infrastructure efficiencyHands-on knowledge of LLM inference optimization techniques such as quantization, batching, caching, model selection, and serving optimizationExperience working with open-weight models and model ecosystems/frameworks such as Hugging Face, vLLM, NVIDIA inference technologies, TGI, Triton, or comparable technologiesExperience with containerized infrastructure and orchestration technologies such as Kubernetes and/or OpenShiftStrong practical experience designing, building, and improving RAG-based applications, including embeddings, vector databases/search, document retrieval, chunking strategies, hybrid retrieval, re-ranking, context management, and retrieval quality optimizationExperience with model routing, semantic routing, or multi-model architectures, with the ability to determine how different models should be selected and utilizedHands-on experience with agentic AI workflows, tool/function calling, orchestration patterns, and integration of LLMs with internal/external toolsStrong programming skills, preferably in Python, together with solid software engineering practices including testing, API design, version control, CI/CD, code quality, and maintainable system designStrong analytical thinking and problem-solving capability, with the ability to independently investigate technical problems, evaluate alternatives, make technical decisions, and drive solutions toward productionNice-to-Haves (Highly Preferred)Specific experience with Red Hat OpenShift AI / Red Hat AI platforms is a strong plusExperience with AI Gateway, API Gateway, model gateway, or shared enterprise AI platform architecturesGood understanding of LLMOps/MLOps and model lifecycle management, including model versioning, deployment, monitoring, observability, and production governanceExperience with chatbot, conversational AI, knowledge assistant, enterprise search, recommendation, intelligent automation, or similar AI-enabled productsUnderstanding of enterprise considerations around model licensing, open-source/open-weight usage, information security, data privacy, access control, governance, and responsible AIExperience providing technical leadership, architecture guidance, design reviews, or mentoring to other engineersGet to know usIf you want to know more about us and what we do, then visit our website: www.oredata.comWhy Oredata?Open communication, flexibility and start-up spiritLearning & Development opportunities for both personal and professional growthOpportunity to get company paid Professional Certificates (Google Cloud Platform, Confluent Kafka, etc)Access to Online Training Platforms (Udemy, Pluralsight, A Cloud Guru, Coursera, etc.)Dynamic work ecosystem where you can take initiative and responsibilityOpportunity to work on international projectsPrivate Health InsuranceBirthday Leave PolicyKişisel verileriniz işe alım sürecinin yürütülebilmesi amacıyla veri sorumlusu sıfatıyla şirketimiz Oredata Yazılım A.Ş. tarafından işlenecektir. Kişisel verilerinizin işlenmesi ve haklarınızla ilgili detaylı bilgiye https://oredata.com/personal-data-protection-policy/ bağlantısı üzerinden ulaşabilirsiniz.