Job Description
Role & responsibilities
We are looking for a LEAD GCP Data Engineer to help design, develop, test, and support cloud-based healthcare data pipelines and backend services on Google Cloud Platform. This role will work closely with senior engineers to process healthcare files from GCS, support event-driven workflows using Pub/Sub/Eventarc, implement data processing using Cloud Run/Dataflow, load data into BigQuery, relational databases, document databases, and support integration with FHIR Store and downstream applications.
The role may also support search and AI-assisted capabilities for healthcare PDFs, CCDA files, FHIR data, and clinical documents using search indexes, embeddings, NER tools, and semantic search pipelines.
Client inputs:
We need one Lead GCP Data Engineer with 10+ years experience. He should be able to handle the team and guide them. AI Experience preferred.
Required Skills
- 5+ years of hands-on database experience, data engineering experience including relational databases, Bigquery and document databases
- 5+ years of hands-on database experience, including SQL, relational database concepts, data modeling, query writing, indexing basics, and data validation.
- 5+ years of cloud experience, preferably on Google Cloud Platform.
- Strong programming experience, especially with Python, including:
- Object-oriented programming
- Modular code development
- Error handling
- Logging
- Unit testing
- Package management
- Backend or data pipeline development
- Experience with RDBMS technologies, such as PostgreSQL, MySQL, SQL Server, Oracle, or equivalent.
- Experience building or supporting data pipelines, including batch, event-driven, API-based, or streaming workflows.
- Hands-on experience with GCP services such as: GCS, Pub/sub, Cloud Run, Data flow
- Experience integrating with external APIs, including authentication, retries, error handling, and basic monitoring.
- Docker experience for containerized application development.
- Git/GitHub experience, including branching, pull requests, and code reviews.
- Experience with CI/CD pipelines, preferably using GitHub Actions, Cloud Build, or similar tools.
- Ability to debug application, pipeline, database, and cloud service issues using logs and monitoring tools.
- Ability to write clean, maintainable, testable, and well-documented code.
- Ability to work from architecture diagrams, technical specifications, and implementation tickets.
Preferred Skills
- Cloud certification or Database Certification
- Google Cloud Healthcare API experience.
- Exposure to FHIR R4, HL7, CCDA, clinical data integration, or healthcare interoperability.
- Exposure to embeddings, semantic search, ElasticSearch, OpenSearch, or vector databases.
- Experience with Gemini, Vertex AI, OpenAI API, or other LLM API integrations.
- Terraform or Infrastructure-as-Code exposure.
- HIPAA, PHI, healthcare security, or compliance awareness.
Key Responsibilities
- Design, build and support scalable healthcare data pipelines on GCP.
- Process healthcare files from GCS and route them through Pub/Sub, Eventarc, Cloud Run, Dataflow, BigQuery, FHIR Store, and downstream systems.
- Design relational and cloud database schemas for operational, analytical, and document-oriented workloads.
- Build APIs to monitor pipeline status, file processing, Pub/Sub messages, Dataflow jobs, Cloud Run services, BigQuery loads, and FHIR data validation.
- Design database models for pipeline metadata, orchestration configuration, processing history, audit logs, document metadata, and search indexes.
- Implement batch and streaming data pipelines.
- Build data validation, reconciliation, retry, and dead-letter handling processes.
- Integrate structured, semi-structured, and unstructured healthcare data.
- Support semantic search, NER, embeddings, and search result snippets for healthcare documents.
- Optimize database performance, query cost, indexing strategy, and storage design.
- Implement monitoring, logging, alerting, and operational traceability.
- Ensure secure handling of healthcare data, including PHI-aware design and least-privilege access.
- Mentor mid-level and junior engineers and define engineering best practices.
- Debug production and non-production issues using Cloud Logging, Cloud Monitoring, and application logs.
- Participate in code reviews and follow engineering best practices.
- Work with senior engineers to implement scalable, secure, and maintainable cloud data solutions.
Preferred candidate profile
Perks and benefits
No Referrers Available
There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.
