Data Engineer
TRUGlobal · Bengaluru, Karnataka, India
Apply & track with Apply EdgeData Engineer / AnalystJob DescriptionExperience Level: 5-7 YearsAbout the RoleWe are looking for a versatile Data Engineer & Analyst who can bridge the gap between robust data infrastructure and actionable business insights. The ideal candidate will design and maintain scalable data pipelines while also delivering meaningful analytics and dashboards to drive decision-making. This role requires strong hands-on expertise in modern data engineering tools combined with a working knowledge of AI/ML concepts and their practical application in data workflows.Key ResponsibilitiesDesign, build, and maintain scalable ETL/ELT pipelines using Apache Airflow for orchestration and workflow automationDevelop and optimize data models and warehousing solutions in SnowflakeWrite efficient, production-grade SQL queries for data transformation, validation, and reportingBuild data processing and automation scripts using PythonLeverage Databricks for large-scale data processing, transformation, and advanced analytics workloadsDesign and develop interactive dashboards and reports using Power BI to support business stakeholdersCollaborate with cross-functional teams (business, product, and engineering) to gather requirements and translate them into data solutionsEnsure data quality, integrity, and governance across pipelines and reporting layersIdentify opportunities to apply AI/ML techniques (e.g., predictive modeling, anomaly detection, NLP-based automation) to improve data processes or generate business insightsImplement basic AI-driven solutions or proof-of-concepts (e.g., using LLMs, forecasting models, or classification models) and integrate them into existing data workflowsMonitor pipeline performance, troubleshoot issues, and optimize for cost and efficiencyDocument technical processes, data flows, and architecture for team knowledge sharingRequired Skills & Experience5-7 years of experience in data engineering, data analytics, or a related fieldStrong hands-on experience with Apache Airflow for pipeline orchestrationProficiency in Snowflake (data modeling, performance tuning, warehouse management)Advanced SQL skills (complex queries, optimization, stored procedures)Strong programming skills in Python (data manipulation, automation, libraries like Pandas/PySpark)Experience with Databricks for big data processing and analyticsProficiency in Power BI (DAX, data modeling, dashboard design, report publishing)Practical exposure to AI/ML concepts with at least one proven implementation (e.g., a deployed model, automated AI-driven workflow, or use of LLMs/APIs in a business context)Hands-on experience with dimensional data modeling (star schema, snowflake schema, fact/dimension tables)Solid understanding of OLTP and OLAP systems and the ability to design solutions appropriate to eachSolid understanding of data warehousing concepts, ETL/ELT design patterns, and cloud data architectureFamiliarity with AWS (e.g., S3, Glue, Redshift, Lambda, or similar data-related services)Preferred QualificationsPrior experience in the semiconductor industry, working with domain-specific data such as fab/manufacturing data, yield analysis, wafer testing, supply chain, or product engineering datasetsExperience with version control (Git) and CI/CD for data pipelinesFamiliarity with data governance, security, and compliance best practicesExposure to MLOps or LLM integration frameworks (e.g., LangChain, OpenAI/Anthropic APIs)Strong problem-solving skills and ability to work independently in a fast-paced environmentExcellent communication skills to translate technical concepts for non-technical stakeholdersEducationBachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field (or equivalent practical experience)