Apply Edge Start your job search

AI Data Architect – Production AI/GenAI

Aavalar Consulting, Inc · Lincolnshire, IL

Apply & track with Apply Edge

AI Data Solutions Architect – (Open Pay Rate)Location: Lincolnshire, IL – Hybrid, 3 days onsite strongly preferred. Chicago Loop office available as an alternative.Role OverviewWe are seeking an AI Data Solutions Architect to support a rapidly evolving enterprise AI program. This is not a traditional Data Architect role. The ideal candidate will have strong enterprise data architecture fundamentals combined with hands-on, production experience architecting the data layer that supports RAG, GenAI, and agentic AI applications.The candidate must be able to operate independently from day one, translate complex AI/data requirements into scalable production architectures, and provide practical guidance on how enterprise data should be ingested, governed, secured, retrieved, evaluated, and exposed to AI systems.Key ResponsibilitiesArchitect and implement production-grade data architectures for RAG, GenAI, and agentic AI applications, supporting both structured and unstructured enterprise data.Design end-to-end AI data flows covering source ingestion, transformation, chunking, metadata enrichment, embeddings, vector indexing, retrieval, reranking, context assembly, and AI/LLM consumption.Design secure, authorization-aware retrieval architectures that enforce enterprise user entitlements at retrieval time and properly handle permission changes, document revocation, metadata synchronization, embeddings, indexes, and cached results.Integrate structured and unstructured sources including Snowflake, PostgreSQL, Salesforce, SharePoint, enterprise documents/PDFs, APIs, and cloud data platforms into AI-ready data ecosystems.Define architecture patterns for batch and real-time ingestion, data transformation, data quality, lineage, provenance, metadata, and access controls supporting AI workloads.Architect and govern vector databases, embedding pipelines, knowledge bases, semantic search, RAG pipelines, and knowledge retrieval patterns.Partner closely with AI/ML and engineering teams to design agent-to-data integration patterns, including secure access to enterprise data and tools.Design production AI observability and evaluation frameworks covering retrieval quality, source/chunk provenance, grounding, generation quality, prompt/model versions, regression testing, and production failure analysis.Evaluate and recommend AI/data technologies across AWS, Azure, GCP, Snowflake, PostgreSQL, vector databases, and related platforms, with an emphasis on architecture principles rather than dependence on a single vendor.Establish data governance standards covering data quality, lineage, metadata, PII, security, authorization, and responsible AI controls.Document architecture patterns, tradeoffs, decisions, and implementation guidance for technical and business stakeholders.Act as a hands-on architecture leader capable of identifying gaps in existing AI/data approaches and introducing production-proven patterns and best practices.Required Skills & Experience8+ years of experience in Data Architecture, Data Engineering, Cloud Architecture, or related disciplines.Demonstrated experience personally architecting and delivering production AI/GenAI data solutions, not simply exposure to AI technologies.Strong hands-on experience with RAG and/or agentic AI architectures in production.Strong understanding of the complete RAG data lifecycle, including:Document ingestionChunking strategiesMetadata enrichmentEmbedding generationVector indexingRetrievalRerankingContext assemblySource attribution/provenanceExperience designing authorization-aware RAG or secure AI data retrieval, including document-level/user-level entitlements and permission changes.Strong experience with enterprise cloud data platforms such as Snowflake, BigQuery, PostgreSQL, Databricks, or similar.Strong SQL skills and hands-on experience with Python or PySpark.Strong experience designing ETL/ELT pipelines and enterprise data models across structured and unstructured data.Experience with one or more vector technologies such as Pinecone, OpenSearch, FAISS, pgvector, Azure AI Search, or similar.Understanding of embeddings, vector search, semantic search, knowledge bases, and retrieval architectures.Experience designing AI observability, evaluation, and troubleshooting frameworks for production AI systems.Strong understanding of data lineage, data quality, metadata, governance, security, and access control in AI/data environments.Ability to explain and defend architecture decisions across different cloud and technology platforms.Strong communication skills and ability to work independently with engineering, AI/ML, data, and business stakeholders.Preferred Skills & ExperienceHands-on experience with AWS Bedrock, Azure OpenAI, Vertex AI, or similar enterprise AI platforms.Experience with LangChain, LangGraph, LlamaIndex, or comparable agentic AI frameworks.Hands-on experience with MCP (Model Context Protocol), including MCP servers, tools/resources, identity, authorization, and governance.Experience with knowledge graphs, ontologies, semantic models, and knowledge retrieval architectures, with a clear understanding of how these differ from vector databases and semantic layers.Experience with AI data lineage and source/chunk provenance.Experience managing embedding and vector index lifecycle, including updates, re-embedding, synchronization, deletion, and stale-data handling.Experience with structured and unstructured enterprise sources such as Salesforce, SharePoint, PDFs, APIs, Snowflake, and PostgreSQL.Experience with Kubernetes, Docker, CI/CD, MLflow, MLOps, or AI platform engineering.Experience in regulated industries such as financial services, insurance, healthcare, or other highly governed environments.Critical Candidate ProfileThe strongest candidate will be someone who can answer the following question in detail:“Tell us about the most sophisticated RAG or agentic AI system you personally architected and took to production. Walk us through the architecture from source data through the final model response, explain the decisions you made, how authorization was enforced, how embeddings and retrieval were managed, how the system was monitored and evaluated, what failed in production, and how you personally solved those problems.”Candidates who primarily have traditional data warehousing, ETL, Snowflake, or enterprise data architecture experience with only conceptual AI exposure are unlikely to be a strong fit for this engagement.