أبلاي إيدج ابدأ البحث عن عمل

Principal AI Platform Architect – Enterprise GenAI

Office Beacon ASPL · Vadodara, Gujarat, India

قدّم وتابع مع أبلاي إيدج
About the RoleOffice Beacon is looking for a Principal AI Platform Architect – Enterprise GenAI to define the technical vision, architecture, and engineering strategy for an enterprise-grade AI platform.The platform will support secure and scalable deployment of Small Language Models (SLMs), Large Language Models (LLMs), Vision Language Models (VLMs), Retrieval-Augmented Generation (RAG), and AI agent workflows within enterprise environments.The Principal AI Platform Architect will own end-to-end architectural direction, from technology selection and platform design through implementation, deployment, optimization, governance, and continuous evolution.This is a hands-on architecture role requiring deep experience building and operating production AI/ML platforms, GPU infrastructure, model-serving systems, distributed systems, and cloud-native infrastructure.Key ResponsibilitiesAI Platform ArchitectureDefine the overall architecture and technical vision for the Enterprise GenAI platform.Establish architectural standards for AI applications, model serving, data flows, infrastructure, and platform services.Evaluate technologies and make architecture decisions based on scalability, reliability, security, performance, and maintainability.Define technical roadmaps for the continued evolution of the AI platform.Model Serving & GPU InfrastructureLead architecture decisions related to GPU infrastructure, model serving, and model lifecycle management.Design production-grade training and inference environments for AI workloads.Architect scalable model-serving infrastructure using technologies such as vLLM, Hugging Face TGI, TensorRT-LLM, or similar platforms.Optimize GPU utilization, inference performance, and infrastructure efficiency.GenAI, RAG & AI AgentsArchitect production-grade RAG systems involving embeddings, semantic search, retrieval, reranking, and vector databases.Design AI agent architectures and production workflows.Define architectures supporting LLMs, SLMs, VLMs, and other generative AI workloads.Establish approaches for model evaluation, deployment, monitoring, and lifecycle management.Model Adaptation & MLOpsDefine approaches for model fine-tuning and adaptation using techniques such as LoRA, QLoRA, and PEFT.Establish model development, evaluation, deployment, and monitoring workflows.Implement or guide CI/CD processes for AI models, applications, and infrastructure.Use appropriate MLOps technologies and practices to improve reproducibility and operational reliability.Cloud & InfrastructureDesign and operate cloud-native AI infrastructure on AWS, Microsoft Azure, GCP, or comparable platforms.Architect Kubernetes-based AI workloads and containerized services.Define infrastructure-as-code practices using Terraform or similar technologies.Design distributed systems capable of supporting enterprise AI workloads at scale.Security & Enterprise ArchitectureDefine security controls for AI workloads and platform infrastructure.Architect appropriate tenant isolation and data-separation strategies.Collaborate with Security and DevOps teams on infrastructure and application security.Contribute to enterprise AI governance and operational standards.Technical LeadershipProvide technical direction to AI engineers and platform engineering teams.Review architectural proposals and technical designs.Establish engineering standards and best practices across AI platform initiatives.Collaborate with DevOps, Product, QA, Security, and other engineering teams.Mentor engineers and help resolve complex architectural and technical challenges.3. Must-Have Qualifications10+ years of professional software engineering experience.5+ years of hands-on experience building and operating production AI/ML or Generative AI platforms.Demonstrated experience designing and deploying enterprise-grade AI systems.Strong production experience with multiple enterprise GenAI technologies, including LLMs, SLMs, VLMs, RAG, and AI agents.Strong proficiency in Python and PyTorch.Hands-on experience with Hugging Face Transformers.Production experience developing AI services and APIs using FastAPI or similar frameworks.Strong understanding of LoRA, QLoRA, PEFT, and model fine-tuning.Strong understanding of embeddings, semantic search, and vector databases.Hands-on experience with production model-serving or MLOps technologies such as vLLM, Hugging Face TGI, TensorRT-LLM, Ray, or MLflow.Strong experience with Kubernetes and Docker.Strong experience with Terraform or another infrastructure-as-code technology.Strong experience with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.Experience designing or operating GPU-based cloud infrastructure.Strong knowledge of software architecture and distributed systems.Proven ability to make architecture decisions for complex production systems.Preferred QualificationsExperience designing AI platforms for multi-tenant enterprise environments.Experience implementing tenant isolation and data-separation strategies.Experience with SOC 2 or ISO 27001 requirements in technology environments.Experience with document intelligence and OCR-based AI workflows.Experience with GPU optimization and inference performance tuning.Experience with model quantization techniques.Experience operating large-scale Kubernetes environments.Experience establishing AI platform governance, observability, and reliability practices.Experience leading technical architecture across multiple engineering teams.