Experience
4 - 6 yrs
Job Location
India
Vacancy
1
Designation
Big Data Engineer
Job Type
Not specified
Job Description
About the Role
What Youll Do
Must-Have Requirements
Nice-to-Have
What Success Looks Like
We are looking for a Big Data Engineer to design, build, and operate large-scale data pipelines and analytical infrastructure that transform high-volume raw data into reliable, query-ready datasets for analytics, reporting, and data-driven products.
Our data platform ingests and processes data from multiple sources and serves analytics, data science, product, and downstream applications. In this role, you will own data pipelines end-to-end-from ingestion and transformation to warehousing, orchestration, data quality, and observability.
A key part of the role is owning ClickHouse as our primary analytical data store . You will be responsible for designing scalable data models, optimizing query performance, and ensuring the platform remains reliable and cost-efficient as data volumes and workloads grow.
You will work closely with data scientists, analysts, product engineers, and other engineering teams to build a modern, cloud-native data platform.
What Youll Do
- Design, build, and maintain robust batch and streaming data pipelines that ingest data from multiple sources into analytical data stores.
- Build and operate Apache Airflow DAGs , including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling.
- Develop analytics-ready datasets using dbt , following well-structured staging, intermediate, and mart layers with appropriate tests, documentation, and incremental models.
- Own ClickHouse as the primary analytical store, including:
- Schema and table design using the MergeTree family of engines
- Partitioning and sorting/primary key strategies
- Materialized views
- Distributed and replicated table architectures
- Query and memory optimization
- High-volume data ingestion and performance tuning
- Work with BigQuery where cloud data-warehouse patterns are appropriate, including data modeling and query/cost optimization.
- Design and operate NoSQL and key-value data stores , including Bigtable, DynamoDB, and Redis, based on specific access patterns and performance requirements.
- Build and maintain data-quality frameworks covering validation, testing, freshness, completeness, reconciliation, and anomaly detection.
- Implement monitoring, alerting, structured logging, and observability for data pipelines and services.
- Own pipeline SLAs, incident response, troubleshooting, and root-cause analysis.
- Manage backfills, safe re-runs, schema evolution, and data migrations while minimizing downstream impact.
- Build reproducible, containerized environments using Docker and contribute to CI/CD and Infrastructure as Code practices.
- Partner with analysts, data scientists, product managers, and product engineers to translate business and technical requirements into scalable data models and pipelines.
- Continuously improve pipeline reliability, scalability, performance, and infrastructure cost efficiency.
Must-Have Requirements
- 4-6 years of experience in data engineering or a closely related field, with strong hands-on production experience.
- Expert-level SQL and strong Python skills, with experience writing production-grade, maintainable, and well-tested code.
- Strong hands-on experience with Apache Airflow or a comparable workflow orchestration platform, including DAG design, scheduling, retries, backfills, dependency management, idempotency, and concurrency.
- Hands-on experience with dbt or a comparable transformation/ELT framework, including modular models, testing, source management, documentation, and incremental processing.
- Expert-level production experience with ClickHouse . This is a core requirement and should include:
- MergeTree engine family
- Partitioning and primary/sorting keys
- Materialized views
- Distributed and replicated tables
- Query and memory optimization
- High-volume ingestion and performance tuning
- Production experience with a cloud-based columnar/OLAP warehouse , such as BigQuery, including data modeling and performance/cost optimization.
- Hands-on experience with NoSQL and key-value databases , such as Bigtable, DynamoDB, and Redis, including data modeling, partition/key design, access patterns, and caching strategies.
- Strong understanding of ETL/ELT and dimensional/layered data modeling principles.
- Experience with GCP and/or AWS and familiarity with Docker, Git, and CI/CD .
- Strong focus on data quality, reliability, observability, and correctness .
- Strong ownership, problem-solving, and communication skills, with the ability to collaborate effectively across engineering, analytics, data science, and product teams.
Nice-to-Have
- Experience with search platforms such as Elasticsearch, OpenSearch, Apache Solr, or Vespa, including indexing pipelines, schema design, and relevance/performance tuning.
- Experience with Aerospike or other high-performance, low-latency distributed key-value/NoSQL systems.
- Experience building streaming and event-driven pipelines using Kafka, Pub/Sub, or similar technologies.
- Experience with Change Data Capture (CDC) patterns and technologies.
- Experience with Apache Spark and data-lake architectures using object storage such as GCS or S3.
- Experience with Terraform or other Infrastructure as Code tools and Kubernetes .
- Experience with data-quality and observability tools such as Great Expectations, Soda, Monte Carlo, or advanced dbt testing .
- Understanding of data platform cost optimization / FinOps practices.
- Experience handling high-volume e-commerce, product catalog, behavioral, or event data .
- Familiarity with a second backend programming language, particularly Go .
What Success Looks Like
- Data pipelines are reliable, well-tested, observable, and consistently meet freshness and completeness SLAs .
- Data models are clean, scalable, documented, and trusted by analytics, data science, and product teams.
- Data-quality issues are identified before they impact downstream consumers .
- Pipeline failures are diagnosed and resolved quickly, with clear root-cause analysis and preventive actions.
- ClickHouse and the broader analytical platform scale smoothly with growing data volumes and query workloads .
- Data infrastructure remains performant and cost-efficient as usage grows.
- Downstream teams can confidently rely on the platform for analytics, reporting, experimentation, and data-driven product experiences .
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.