Experience
5 - 7 yrs
Job Location
Pune, India
Vacancy
1
Designation
Senior Data Engineer
Job Type
ONSITE
Job Description
Title: Senior Specialist Data Engineering
Location: Pune
Exp: 4-7 Years
Key Responsibilities
Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) Azure AI Search (vector / RAG index).
Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals with incremental refresh, partitioning, and Z-ordering for query performance.
Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
Implement SAP cost data masking at the API / pipeline layer sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
Support historical data migration: 5 7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED Azure Data Factory (ADF) pipeline authoring Azure AI Search index management, embedding pipelines ADLS Gen2 / Delta Lake storage compute Data modelling relational + lakehouse schemas Azure SQL / SQL Server T-SQL, stored procedures Azure Key Vault secrets, connection string management Azure Synapse Analytics or Databricks Data lineage observability tools Python PySpark, pandas, data quality scripts Git / Azure DevOps for pipeline version control SAP OData / RFC / BAPI integration patterns CDC (Change Data Capture) patterns in SQL Server
GOOD TO HAVE
Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
dbt (data build tool) for transformation layer on Delta Lake.
Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
Experience with IATF 16949 or automotive quality data requirements.
Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.
", Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
Location: Pune
Exp: 4-7 Years
Key Responsibilities
Architect and implement the three-tier data storage model: Azure SQL (live transactional SNPD DB) ADLS Gen2 Delta Lake (historical analytics / ML features, nightly ADF ETL) Azure AI Search (vector / RAG index).
Build and maintain Azure Data Factory (ADF) pipelines: nightly CDC-based ETL from the SNPD SQL database, SAP (cost/PO/BOM/vendor masked at API layer), and Teamcenter PLM (BOM snapshots, part lifecycle, ECN).
Design and implement the unified SNPD data model: enforce project_id + part_number as universal pivot keys across all tables; ensure referential integrity across SNPD core domain tables and migrated portal tables (NVPC, RFQ, PPAP, CDMM, etc.).
Build the AI feature store on Delta Lake: dl_gate_cycle_times, dl_supplier_risk, dl_nvpc_benchmarks, dl_cost_variance, dl_deliverable_actuals with incremental refresh, partitioning, and Z-ordering for query performance.
Conduct data audits on SAP S/4 HANA, Teamcenter PLM, and all 8 legacy homegrown portal databases; assess data quality, identify gaps, and remediate for ML readiness.
Implement SAP cost data masking at the API / pipeline layer sensitive pricing data must be obfuscated before reaching any MCP server or AI agent.
Set up Azure AI Search vector index: embedding ingestion pipeline from the document store (SharePoint / Azure Blob), chunking strategy, metadata schema, and incremental re-indexing on document updates.
Establish data lineage, quality checks, and observability: row counts, null rates, schema drift alerts, and SLA monitoring for all ETL pipelines.
Support historical data migration: 5 7 years of legacy SNPD and portal data into the unified SNPD database; validate referential integrity and completeness post-migration.
Collaborate with the ML Engineer to serve training datasets from Delta Lake; optimize feature computation using Synapse Serverless or Databricks as compute.
Implement RBAC and data access controls at the data layer: ensure user-level and role-level scoping is enforced from Azure SQL through to Delta Lake reads and vector search results.
Maintain data catalogue and schema documentation; ensure all entities conform to the IATF 16949 audit traceability requirements.
TECHNICAL SKILLS REQUIRED Azure Data Factory (ADF) pipeline authoring Azure AI Search index management, embedding pipelines ADLS Gen2 / Delta Lake storage compute Data modelling relational + lakehouse schemas Azure SQL / SQL Server T-SQL, stored procedures Azure Key Vault secrets, connection string management Azure Synapse Analytics or Databricks Data lineage observability tools Python PySpark, pandas, data quality scripts Git / Azure DevOps for pipeline version control SAP OData / RFC / BAPI integration patterns CDC (Change Data Capture) patterns in SQL Server
GOOD TO HAVE
Hands-on SAP S/4 HANA RISE data extraction experience (ACDOCA, Material Master, BOM, MM60).
Teamcenter PLM API familiarity (REST/SOA Gateway, BOM export, ECN feeds).
dbt (data build tool) for transformation layer on Delta Lake.
Apache Kafka / Azure Event Hubs for real-time streaming from SAP change events.
Experience with IATF 16949 or automotive quality data requirements.
Familiarity with Mahindra data platform (MDP) or Azure Purview for data governance.
", Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
