Apply Edge Start your job search

Python / ML Developer

Sharedpro · Vadodara, Gujarat, India

Apply & track with Apply Edge
About Sharedpro - Sharedpro is an AI platform company building AI-native products on a shared foundation of computer vision, audio-ML, real-time inference, vector search, and GPU model serving. Our products include AutoPilot Interview, Emotions API, TrueFace, Emotion Quest, and Pediatric Developmental Screening. We're a small, fast-moving team where engineers take end-to-end ownership and ship real products every week.The Role - We are looking for a hands-on Python / ML Developer with 4–6 years of experience to build the ML core of our AI products: a live AI voice interviewer, a facial expression / emotion measurement API, communication-delivery scoring for interviews, and integrity / proctoring signals. You will own real-time ML pipelines end-to-end — from data preparation and model training to low-latency inference services running in production on our own GPU infrastructure. This is a hands-on individual-contributor role; you will work closely with the Technical Lead, product, and the frontend and backend teams. Prior experience in a startup or fast-moving product environment is strongly preferredWhat You'll Do - • Build and maintain production ML services in Python using FastAPI — REST APIs, WebSockets, and async pipelines designed around real-time latency budgets. • Develop real-time audio/speech pipelines: ASR with Whisper / faster-whisper, voice-activity detection (Silero VAD), prosody and speech features (praat-parselmouth, librosa), and TTS integration. • Train, fine-tune, and evaluate computer-vision models with PyTorch — facial expression / emotion recognition, face detection with YOLOv8 / OpenCV / MediaPipe. • Integrate LLMs into products: prompt engineering, structured outputs, and resilient provider handling across Groq, Mistral, and OpenAI-compatible endpoints (fallbacks, rate limits, cost and latency optimization). • Build retrieval and memory layers: embeddings, vector search with Qdrant, semantic matching across sessions. • Serve and optimize models on in-house GPU infrastructure — Triton Inference Server / vLLM, fp16, batching, and throughput/latency tuning. • Design honest, evidence-based scoring: calibration, bias checks, and explainable signals; treat privacy and PII handling as first-class requirements. • Write well-tested production code (pytest), add logging and monitoring, and take part in code reviews. • Own features end-to-end: prototype → production → measurement → iteration.Must-Have • 4–6 years of strong hands-on Python development for production systems (not just notebooks or scripts). • Solid FastAPI experience (or Flask / Django) — REST APIs, WebSockets, async programming, and concurrency. • Hands-on ML with PyTorch (training and inference) and computer vision with OpenCV. • Experience integrating LLM APIs (OpenAI-compatible endpoints) and practical prompt engineering. • Audio/speech ML exposure — ASR (Whisper), VAD, audio feature extraction — or demonstrated ability to go deep quickly. • SQL databases (PostgreSQL / MySQL) and comfort working on Linux servers. • Docker, Git, and CI/CD experience. • Testing discipline with pytest. • Strong problem-solving, debugging, and communication skills. • Startup or fast-moving product experience preferred.Good to Have • Model serving: Triton Inference Server, vLLM, ONNX, or TensorRT; GPU inference optimization, CUDA. • Vector databases (Qdrant), embeddings, and RAG pipelines. • MediaPipe, YOLO (ultralytics), and face-analysis models. • praat-parselmouth prosody analysis and speech-feature engineering. • Responsible-AI practices: calibration, bias mitigation, explainability; awareness of EU AI Act / NYC LL144. • Basic React / TypeScript for demo frontends; Streamlit or Gradio for internal tools. • Kafka / RabbitMQ / Celery / Redis / Kubernetes. • AWS / Azure / GCP.Why Sharedpro • In-house GPU infrastructure — train, serve, and optimize your own models. • Real-world AI problems: voice agents, emotion recognition, speech scoring, and proctoring — not toy projects. • Small team with high ownership; ship real products every week. • End-to-end responsibility from data to model to production service. • Modern stack: PyTorch, Whisper, LLMs, Qdrant, Triton, FastAPI. • Freedom to make meaningful technical decisions in a fast-moving startup environment.