Apply Edge Start your job search

Machine Learning Intern - LLM Finetuning

VIMA3YA · Bengaluru, Karnataka, India

Apply & track with Apply Edge

About the roleWe're looking for a curious, hands-on engineer to join our team as an LLM Fine-Tuning Intern. You'll help adapt open-source large language models to real product tasks: preparing training data, fine-tuning models with LoRA and QLoRA, evaluating results carefully, and helping ship fine-tuned models to production.It's a learning-heavy role. You'll work alongside experienced engineers and own real experiments from day one.Internship details- Duration: 3 months- Stipend: ₹15,000 per month- Conversion: Full-time offer based on performance- Start date: ImmediateWhat you'll do- Build and clean instruction and chat datasets: formatting, deduplication, quality checks, and train/eval splits.- Fine-tune small and medium open models (1B–8B) with supervised fine-tuning (SFT) using LoRA and QLoRA, mainly with Unsloth, Hugging Face TRL, PEFT and Transformers.- Run controlled experiments on rank, learning rate, data mix and sequence length, and record the results clearly.- Evaluate models with held-out test sets, task metrics (accuracy, exact match, format validity) and LLM-as-judge, and check that general abilities haven't regressed.- Debug training runs: loss curves, chat templates, label masking, GPU memory problems.- Export and package models (merged 16-bit, GGUF) for serving with tools like vLLM, llama.cpp or Ollama.- Keep up with post-training methods such as DPO, ORPO and GRPO, and help try them on our use cases.Must have- A degree in CS, AI/ML, Data Science or a related field (final-year students are welcome to apply).- Strong Python. Comfortable with PyTorch or similar.- A clear understanding of the basics: how transformers work at a high level, tokens, loss, overfitting, and train/validation/test splits.- Can explain LoRA and QLoRA, and why they reduce memory.- Has fine-tuned at least one model on their own: a course project, a Kaggle or Colab notebook, a hackathon or a personal project.- Careful with data, and honest about results, including negative ones.Nice to have- Hands-on experience with Unsloth, TRL/PEFT or the Hugging Face Hub.- Awareness of preference tuning (DPO) or RL-based methods (RLHF, GRPO).- Experience with experiment tracking (Weights & Biases, MLflow) and basic Linux/GPU environments.- A GitHub repo, blog or model card showing something you've built.What you'll learn- The full post-training lifecycle: data → SFT → preference tuning → evaluation → deployment.- How to run disciplined experiments and make decisions from evidence.- Serving and optimizing models in production.Hiring process1. Application review (a GitHub link or project write-up is strongly encouraged)2. Take-home or live task: fine-tune a small model on a provided dataset and compare it with the base model3. 45-minute technical interview (fundamentals, LoRA/QLoRA, data, debugging)How to applySend your resume to support@vima3ya.com, along with a GitHub link or project write-up showing a model you've fine-tuned.