أبلاي إيدج ابدأ البحث عن عمل

Data Scientist

Nityo Infotech · Perth, Western Australia, Australia

قدّم وتابع مع أبلاي إيدج
WHAT YOU WILL DO:AI extraction pipelinesDesign, build, and maintain multi-stage document-extraction pipelines that combine deterministic rule-based logic with LLM-powered extraction, orchestrated as durable, resumable workflows.Integrate cloud document AI services and LLMs for classification and extraction.Own prompt engineering and structured-output design, ensuring outputs are strongly typed (Pydantic), validated, and fully traceable to source documents.Structured data, mapping, and validationExtract and normalize retrieved information, classify document and pricing structures, and map results to downstream schemas.Handle document variations and amendments while preserving provenance (page/section references) for reviewer traceability.Develop and tune automated validation and business-rule checks that detect non-compliant claims.Enterprise systems integrationIntegrate with enterprise source-of-truth systems — SAP (OData), Snowflake, and Cosmos DB — to enrich, validate, and persist results.Preserve data lineage across systems so every decision is auditable.WHAT WE ARE LOOKING FOR-Strong Python: typed, well-tested, production quality.Hands-on LLM application engineering — prompt design, structured outputs, and evaluation frameworks, not just calling an API.A strong evaluation and measurement mindset: designing datasets and metrics, reasoning about precision/recall trade-offs on real data, confidence calibration, and drift detection — able to prove a system works, not just build it.Experience with cloud AI services (Azure preferred).Data-wrangling with pandas (or equivalent) and working with SQL.Git-based collaboration, code review, and CI/CD (GitLab a plus).Experience working with agentic setup/ agentic development workflows.Awareness of secure handling of sensitive enterprise data (secrets management, least-privilege auth) a plus.Experience with recommender-system algorithms and/or retrieval-augmented generation (RAG) viewed favourably — we have a recommender-oriented project starting up alongside the document-AI work.TECH ENVIRONMENTPython | Azure (Document Intelligence, OpenAI, Machine Learning, Functions) | Pydantic | MLflow | SQL — Snowflake & Cosmos DB | Streamlit & Power BI | OpenTelemetry / Grafana | GitLab CI/CD