LLM Engineer
Substrate AI · Valencia, Valencian Community, Spain
Apply & track with Apply EdgeAbout the roleAt Substrate AI, we build AI solutions and infrastructure for European businesses.Our Serenity platform brings AI into companies through intelligent agents that operate directly within the business, increasingly powered by our own distributed GPU infrastructure in Spain — Serenity Cloud.We are now looking for a Founding LLM Engineer to build and lead our LLM Lab.This is a unique opportunity to take ownership of the models powering Serenity’s agents: selecting, fine-tuning, evaluating, quantizing and optimizing domain-specific language models while working directly with our GPU infrastructure.You will join as a hands-on founding engineer, building the foundations yourself and progressively evolving into Head of LLM Lab as the team grows.This is not a prompt engineering or LLM API integration role.We are looking for someone who has worked directly with model weights, training pipelines, datasets and GPU infrastructure.Why this role?This is an opportunity to build an LLM capability from the ground up rather than joining an already established team.You will work across the full model lifecycle — data → training → evaluation → optimization → GPU infrastructure → production — with significant technical ownership from day one.What you will doModel development & trainingDevelop, train and fine-tune transformer-based LLMs.Adapt foundation models to verticals including enterprise, healthcare, legal, construction and hospitality.Implement supervised fine-tuning and, where relevant, preference-alignment techniques.Prepare and manage training, validation and evaluation datasets.Build reproducible evaluation pipelines and continuously monitor hallucination, overfitting, contamination and model regressions.Inference & GPU optimizationOptimize models for latency, memory, throughput and cost per token.Work with technologies such as vLLM, SGLang, TensorRT-LLM, NVIDIA NIM, Triton and TGI.Apply techniques including quantization, distillation and speculative decoding.Run single-node and distributed GPU training and inference workloads on Serenity Cloud.Deliver versioned and benchmarked models ready for production serving.Founding & leadershipDefine the model roadmap together with the CTO.Establish efficiency targets focused on maximizing useful tokens per euro.Help define the optimal training vs. inference capacity strategy across our infrastructure.Establish technical standards and best practices for the LLM Lab.Build, mentor and progressively lead the team as the Lab grows.Collaborate closely with Platform, Infrastructure / Serenity Cloud and Agents teams.What we are looking forStrong Python skills and production-quality software engineering practices.Practical experience with PyTorch and transformer architectures.Demonstrable hands-on experience training or fine-tuning LLMs, beyond consuming commercial APIs.Experience with Hugging Face Transformers, Datasets and Tokenizers.Strong understanding of attention, tokenization, embeddings, optimization, loss functions and model validation.Experience optimizing LLM inference for throughput, memory, latency and cost.Hands-on experience with NVIDIA GPUs and CUDA.Experience working with owned/on-premise GPU infrastructure is highly valued.Experience preparing and managing substantial text or multimodal datasets.Ability to independently design experiments, interpret results and communicate technical trade-offs.Ownership mindset and potential to build and lead a highly technical team.Professional English.Nice to haveExperience with some of the following will be highly valued:Models ranging from 1B to 500B parameters.Distributed training: PyTorch Distributed, DeepSpeed, FSDP or Accelerate.PEFT, LoRA and QLoRA.Quantization, distillation, pruning and speculative decoding.vLLM, SGLang, TensorRT-LLM, NVIDIA NIM, Triton or TGI.Docker, Linux and Git.On-premise or owned GPU clusters.Domain-specific models or regulated sectors.Healthcare data and data-governance environments.Open-source contributions, technical publications or demonstrable LLM experiments.RAG and agentic systems.Model efficiency metrics such as tokens/GPU-hour or €/token.ExperienceWe are primarily looking for demonstrable technical experience rather than a specific job title.As a reference, the ideal profile will typically bring:5–8+ years in Machine Learning / Deep Learning.2–4+ years of hands-on experience with transformers or LLMs.Evidence of technical leadership, ownership or experience as a founding / first engineer.A degree in Computer Science, Artificial Intelligence, Machine Learning, Mathematics, Physics or another quantitative discipline is valued.A Master’s or PhD is welcome, but practical evidence of building and optimizing models matters more than academic credentials alone.Our philosophy is simple:We prioritise demonstrable experience working directly with models, data and GPUs — and useful tokens per euro over chasing state-of-the-art for its own sake.🚀 Founding LLM Engineer / Head of LLM Lab 📍 Spain / EU · Remote / Hybrid 🌍 Working language: English 🏢 Substrate AI – Serenity LLM LabIf you have built, trained and optimized real language models and want to take ownership of an LLM Lab from its foundations, we would like to hear from you.📩 Apply through LinkedIn and, where possible, include a GitHub repository, publication, technical write-up or other evidence of relevant work.