أبلاي إيدج ابدأ البحث عن عمل

Intern (Data Science)

Yubi · Chennai, Tamil Nadu, India

قدّم وتابع مع أبلاي إيدج
Yubi, formerly known as CredAvenue, is re-defining global debt markets by freeing the flow of finance between borrowers, lenders, and investors. We are the world's possibility platform for the discovery, investment, fulfillment, and collection of any debt solution. At Yubi, opportunities are plenty and we equip you with tools to seize it.In March 2022, we became India's fastest fintech and most impactful startup to join the unicorn club with a Series B fundraising round of $137 million.In 2020, we began our journey with a vision of transforming and deepening the global institutional debt market through technology. Our two-sided debt marketplace helps institutional and HNI investors find the widest network of corporate borrowers and debt products on one side and helps corporates to discover investors and access debt capital efficiently on the other side. Switching between platforms is easy, which means investors can lend, invest and trade bonds - all in one place. All of our platforms shake up the traditional debt ecosystem and offer new ways of digital finance.Job title : InternDuration : 6 monthStipend : Unpaid Internship (Probono)Key ResponsibilitiesJoin a high-velocity Data Science team, partnering closely with senior data scientists to design, build, and deploy reusable tools, predictive models, and intelligent automation capabilitiesTackle diverse, real-world machine learning challenges with a focus on developing personalization/recommendation systems, perfecting OCR text retrieval accuracy, automating structured document extraction, and building robust text/document classification systemsExplore, prototype, and implement Generative AI and Large Language Model (LLM) solutions to extract insights, summarize complex documents, and augment intelligent workflowAssist in fast-tracking the model development lifecycle by conducting robust data preprocessing, feature engineering, and rigorous model evaluation to ensure systems scale efficiently under heavy user trafficOptimize model performance, throughput, and latency by proactively researching, prototyping, and benchmarking the latest open-source ML libraries, state-of-the-art transformer architectures, and managed cloud service.Required Experience & ExpertiseCurrently pursuing or recently completed a degree (Bachelor's) in Computer Science, Data Science, Statistics, Mathematics, or a highly quantitative engineering fieldExceptional programming skills in Python, paired with hands-on familiarity with standard data science libraries (Pandas, NumPy, Scikit-learnStrong foundational knowledge of core Machine Learning algorithms, statistical modeling, data structures, and classification techniquesHands-on experience or strong familiarity with advanced Natural Language Processing (NLP) models, specifically fine-tuning BERT or similar encoder architectures for text classification, sentiment analysis, or semantic understandingExposure to Large Language Models (LLMs) and modern GenAI techniques, including prompt engineering, retrieval-augmented generation (RAG), or fine-tuning open-source models (e.g., Llama, Mistral) via frameworks like Hugging Face or LangChainGood conceptual or hands-on exposure to Recommendation Systems (e.g., collaborative filtering, content-based matching, matrix factorization, or vector embeddingExperience or a strong technical interest in Intelligent Document Processing (IDP), text extraction, and OCR using open-source libraries (e.g., Tesseract, EasyOCR, Keras-OCR) or cloud-native APIs (e.g., AWS Textract, GCP VisionExposure to Deep Learning architectures and frameworks (such as PyTorch, TensorFlow, or Keras) to process unstructured text and image dataFamiliarity with writing efficient SQL queries for relational databases, along with basic concepts of API development (e.g., FastAPI, Flask) for model serving