Apply Edge Start your job search

Data Scientist

Coforge · Pune District, Maharashtra, India

Apply & track with Apply Edge
Job Title: Data ScientistSkills: Artificial intelligence, Machine Learning, NLP, Gen AI, Python, Rest API, Agentic AI, LLM, RAG, Devops and AWSExperience: 4+ yearsLocation: Pune and HyderabadDuration: Full timeWe at Coforge are hiring for Data Scientist role with following skill sets:LLM & Generative AIDesign, build, and deploy LLM-powered applications using frameworks such as LangChain, LlamaIndex, or OpenAI API.Develop and optimize prompt engineering strategies (few-shot, chain-of-thought, RAG) to improve the accuracy, consistency, and reliability of LLM outputs.Implement Retrieval-Augmented Generation (RAG) pipelines using vector databases (e.g., FAISS, Pinecone, Chroma, Weaviate).Fine-tune pre-trained LLMs (e.g., GPT, LLaMA, Mistral, Falcon, Claude,Gemini) on domain-specific datasets.Validate and structure LLM outputs using Pydantic models and output parsers to ensure data integrity.Natural Language Processing (NLP)Build end-to-end NLP pipelines for real-world tasks including:Named Entity Recognition (NER)Text Classification & Sentiment AnalysisInformation & Data Extraction from DocumentsDocument Summarization & Question AnsweringSemantic Search & Document SimilarityWork with the Hugging Face Transformers ecosystem to leverage and fine-tune pre-trained models (BERT, RoBERTa, T5, etc.).Process large-scale unstructured text data from various sources such as PDFs, emails, scanned documents (OCR), and web content.Anomaly DetectionDesign and implement anomaly detection systems for various domains, including:Financial fraud detection (unusual transactions, payment anomalies).Operational anomalies (system logs, network traffic, sensor data).Text-based anomalies (unusual document patterns, suspicious NLP signals).Apply a wide range of anomaly detection techniques including:Statistical Methods: Z-score, IQR, CUSUM.ML-based Methods: Isolation Forest, One-Class SVM, Local Outlier Factor (LOF).Deep Learning Methods: Autoencoders, LSTM-based sequence anomaly detection, Variational Autoencoders (VAEs).Time-Series Methods: ARIMA, Prophet, Seasonal Decomposition.Build real-time and batch anomaly detection pipelines that can scale to large datasets.Define and tune detection thresholds and alert mechanisms in collaboration with business and operations teams.Machine Learning (ML)Design, train, evaluate, and deploy supervised and unsupervised machine learning models.Perform feature engineering, model selection, hyperparameter tuning, and cross-validation.Build and maintain end-to-end ML pipelines from data ingestion to model serving.Monitor model performance in production and implement retraining strategies to address data drift and model decay.Communicate model results, performance metrics, and business impact to technical and non-technical stakeholders.Python & Software EngineeringWrite clean, modular, production-quality, and well-documented Python code.Build and expose ML models as REST APIs using FastAPI or Flask.Collaborate with MLOps/DevOps engineers to containerize (Docker) and deploy models in cloud environments.Follow best practices in version control (Git), testing, and CI/CD pipelines.Data & AnalyticsPerform Exploratory Data Analysis (EDA) on structured and unstructured datasets to identify patterns, trends, and anomalies.Work with data from relational databases (SQL), data lakes, and cloud storage solutions.Create compelling and clear data visualizations (Matplotlib, Seaborn, Plotly) to communicate findings.