LLM Evaluation Engineer
HCLTech · Canada
Apply & track with Apply EdgeOverall – 8+ | AI Exp – 3+Hands-on proficiency with libraries like Ragas, DeepEval, TruLens, Promptfoo, or OpenAI Evals to measure: Context Precision & Context Recall (retrieval quality from vector stores).Faithfulness / Groundedness (checking against hallucination in bounded context).Answer Relevance and Toxicity / Guardrail Compliance.Analyzing evaluation datasets, confusion matrices, similarity score distributions, and benchmark trends.Computing classification metrics (Precision, Recall, F1, Cosine Similarity thresholds) for retrieval systems.Azure OpenAI SDK / OpenAI API and LangChain / LlamaIndex for building automated LLM-as-a-judge scoring scripts and synthetic test dataset generation.PyTest / Unittest: Writing automated unit/regression tests for prompt templates and model version upgrades.Writing complex SQL queries to extract, curate, and version golden test datasets from relational databases (e.g., PostgreSQL, Azure SQL).Querying vector stores and telemetry databases (e.g., pgvector, Splunk/log tables) to evaluate real-world candidate/user query retrieval accuracy and fallback rates.Handling structured benchmark datasets and prompt configurations using JSON / JSONL, YAML, and CSV.