Job Description
Test Engineer - Data, ETL & BI
Data Quality - Data Warehousing - ETL / BI - Cloud Data Platform on AWS
Position Title Test Engineer
Department Data Engineering
Employment Type Full-Time
Experience 5 + Years
Location Hyderabad
Role Summary
We are looking for a detail-oriented Test Engineer to own quality across our cloud-based BI and data platform on AWS - a suite of data-engineering codebases (Git repositories) covering ETL / data-pipeline processing, data extraction and job orchestration, BI / analytics, and data retention / purge. These systems extract data from source databases and document stores, load it into a cloud data warehouse via ETL tooling and Python, enforce row-level security, publish BI dashboards, and purge data per retention policies. The ideal candidate brings a strong SQL background, hands-on Python test automation, and the ability to design, execute, and automate test strategies that validate code, database, and data end-to-end across all extract, load, KPI, alerting, migration, and analytics job families. You will collaborate with cross-functional teams to ensure data integrity, system reliability, and seamless delivery of high-quality data products.
Systems & Processes in Scope
- Multiple data-engineering Git repositories - an ETL / data-warehouse code repo, a data-extraction and job-orchestration code repo, a BI / analytics code repo, and a data retention / purge code repo.
- Scheduled source-to-warehouse extract jobs - config-driven scheduling and a job-executor framework that moves data from source databases to cloud storage and into the warehouse.
- Database, custom, and dynamically-generated SQL extract jobs.
- A job restartability, dependency-check, and record-count reconciliation framework for extract pipelines.
- System health-check and data-threshold alerting jobs across multiple tenant / client configurations.
- KPI computation, import, and export jobs.
- Parallel / concurrent data-load jobs.
- BI report migration jobs (PowerShell-based).
- Predictive-analytics / ML jobs (e.g., customer-lifetime-value scoring).
- An ETL regression test suite.
- Data archival / purge and retention jobs.
- BI dashboards and datasets.
Key Responsibilities
- 5 + years of experience in QA, with a significant focus on data / ETL testing.
- Design and execute test plans, test cases, and test scripts for the data pipelines, ETL processes, and data-warehouse layers (staging, DWH, data marts) across all the codebases and job families listed above.
- Perform source-to-target reconciliation and validate data transformations, aggregations, dedup / upsert logic, and business rules across the cloud data warehouse and source databases.
- Validate ETL orchestration built with the ETL tooling and Python, including extract -> cloud-storage -> stage -> target flows.
- Validate the extract scheduling and job-executor framework - correct job sequencing, config-driven table selection, and incremental extract windows.
- Validate job restartability, dependency checks, and file / record-count reconciliation for extract jobs.
- Validate system health-check and data-alert jobs, including multi-tenant alert thresholds and configurations.
- Validate KPI computation, import, and export outputs against warehouse data.
- Validate parallel-load correctness and data integrity under concurrency.
- Validate BI report migration outputs and predictive-analytics (e.g., customer-lifetime-value) job results.
- Test row-level security policies and role / group-based data access, and validate grants and privileges.
- Validate data archival / purge and retention processes - confirm the correct data is removed or retained with no collateral impact.
- Detect, report, and track data-quality issues including duplicates, nulls, referential-integrity violations, and schema / soft-delete inconsistencies.
- Test and validate complex SQL transformations, stored procedures, views, and DDL changes (constraints, encodings, distribution / sort keys).
- Perform regression testing on BI reports, dashboards, KPIs, and aggregated datasets.
- Perform row-count checks, checksum comparisons, and field-level validation across source and target systems.
- Develop and maintain automated data-validation frameworks and SQL / pytest-based test suites; integrate them into CI pipelines for continuous data-quality monitoring.
- Participate in code / pull-request reviews; file clear, reproducible defects and verify fixes.
- Work closely with data engineers, architects, analysts, and product owners to define acceptance criteria; participate in Agile / Scrum ceremonies, and mentor junior QA members.
Technical Skills
- Advanced SQL proficiency: complex joins, window functions, CTEs, subqueries, stored procedures, and query performance tuning (Amazon Redshift / PostgreSQL).
- Hands-on experience with relational / cloud data warehouses: Amazon Redshift and Aurora PostgreSQL (SQL Server, Oracle, MySQL a plus).
- Python test automation with pytest (fixtures, mocking, assertions); ability to read and modify Python ETL, extract-scheduler, and utility code.
- Working knowledge of AWS data services: Redshift, S3, and RDS / Aurora; exposure to Amazon QuickSight for BI validation.
- ETL / ELT tools: Pentaho Data Integration (PDI / Kettle); familiarity with AWS Glue or similar is advantageous.
- Scripting in Bash and PowerShell (PowerShell is used for the BI report migration jobs) to run jobs, inspect outputs, and read logs.
- Understanding of scheduling / job-orchestration and config-driven, multi-tenant ETL frameworks.
- Familiarity with restartability, dependency-check, and record-count reconciliation concepts for extract pipelines.
- Knowledge of data modelling concepts: star schema, snowflake schema, dimensional models, and referential integrity.
- Familiarity with source / NoSQL and document stores - MongoDB (Cassandra / DynamoDB a plus).
- Exposure to validating BI reports (Amazon QuickSight, Power BI) and ML / predictive-analytics outputs (e.g., customer-lifetime-value) is a plus.
- Version control with Git (branching and pull-request workflow) and CI/CD tools (Jenkins, GitHub Actions).
- Containerization basics (Docker); security-scan gates (Trivy, AWS Inspector) and CVE triage are a bonus.
- Proficiency with test / defect management tools: JIRA (Zephyr, TestRail, or Azure DevOps a plus).
- Understanding of end-to-end data migration testing: pre-migration baseline, in-flight validation, and post-migration reconciliation.
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
