أبلاي إيدج ابدأ البحث عن عمل

Data Scientist

VCBay · India

قدّم وتابع مع أبلاي إيدج
Role OverviewWe are looking for a versatile Data Scientist who can build robust data pipelines — from web scraping to AI/ML deployment. The ideal candidate is comfortable working across the full data and AI lifecycle: extracting data at scale, transforming and storing it reliably, and using it to train, deploy, and maintain machine learning models in production.Key ResponsibilitiesDesign, build, and maintain scalable web scraping scripts/pipelines using Python (e.g., BeautifulSoup, Scrapy, Selenium, Playwright)Handle dynamic websites, pagination, anti-bot mechanisms, proxies, and rate-limiting strategiesClean, transform, and normalize scraped data (structured/unstructured) before storageDesign MongoDB schemas and collections optimized for the type of data being handledImplement logic to identify and update unique/duplicate records efficiently (upserts, deduplication strategies)Schedule and monitor data pipelines/jobs (via cron, Airflow, or similar orchestration tools)Ensure data quality, consistency, and integrity across pipelines, including error logging, retries, and failure recoverySupport end-to-end AI/ML lifecycle: data collection, preprocessing, feature engineering, model selection, training, and validationFine-tune machine learning/deep learning models and evaluate performance against business requirementsPackage and deploy models into production environments (APIs, batch pipelines, etc.)Implement and maintain MLOps practices — versioning (models & data), CI/CD for ML, monitoring model performance/driftCollaborate with cross-functional teams to integrate AI models with existing data pipelines, including scraped/transformed data as model inputRequired SkillsCore:Strong proficiency in Python (writing clean, modular, production-grade code)Hands-on experience with MongoDB (schema design, aggregation pipelines, indexing, upsert/dedup logic)Web scraping tools/libraries: BeautifulSoup, Scrapy, Selenium, Playwright, or similarData transformation/manipulation using Pandas / NumPyAI/ML:Understanding of the end-to-end AI/ML workflow — data prep, training, evaluation, deploymentFamiliarity with ML/DL frameworks: Scikit-learn, TensorFlow, PyTorchExposure to MLOps tools: MLflow, DVC, Docker, Kubernetes (basic), CI/CD pipelinesExperience with model deployment (REST APIs via FastAPI/Flask, or cloud ML services)Good to Have:Experience with cloud platforms (AWS / GCP / Azure) for storage, compute, and ML servicesFamiliarity with LLMs / NLP (e.g., Hugging Face, LangChain) if relevant to use caseKnowledge of proxy rotation, CAPTCHA-handling, and anti-scraping evasion techniquesExperience with workflow orchestration (Airflow, Prefect)Version control (Git) and Agile development practicesSoft SkillsStrong problem-solving mindset, especially around handling unstructured/messy dataAbility to independently manage multiple workstreams across data engineering and AI/MLGood documentation habits for pipelines and modelsComfortable working in a fast-paced, evolving tech environmentQualificationsBachelor's/Master's degree in Computer Science, Data Science, Engineering, or related fieldPrior experience with production-grade web scraping and/or ML systems is a strong plusPortfolio/GitHub showcasing scraping projects and/or ML models is preferred