Job Description
Role Overview
We are seeking an experienced Senior Architect to lead the design and evolution of enterprise-scale distributed systems supporting 50M+ connected sensors and high-volume event processing pipelines.
This role is critical to building and operating mission-critical backend platforms that process millions of events per second across both synchronous and asynchronous architectures, with stringent requirements for scalability, reliability, security, and performance.
The ideal candidate brings a proven track record of architecting and scaling production-grade systems at extreme scale, along with the ability to drive technical strategy, governance, and cross-functional alignment in a complex enterprise environment.
Key Responsibilities
Architecture System Design
- Define and lead the architecture of large-scale distributed systems capable of ingesting and processing high-velocity data streams from 50M+ sensors
- Design resilient systems across synchronous (API-driven) and asynchronous (event-driven, streaming) paradigms
- Establish architectural standards for scalability, fault tolerance, and performance optimization
Data Platform Engineering
- Architect real-time and batch data pipelines for high-throughput ingestion, transformation, and storage
- Drive design decisions across streaming, processing, and storage layers to ensure optimal performance and cost efficiency
- Enable support for time-series, event-driven, and analytical workloads
Technology Strategy Governance
- Define and enforce enterprise architecture principles, standards, and best practices
- Evaluate and guide adoption of modern data technologies, including:
- Distributed messaging systems (e.g., Kafka, Pulsar)
- Scalable data stores (e.g., Cassandra, DynamoDB, Bigtable, ClickHouse, Elasticsearch)
- Stream and batch processing frameworks (e.g., Flink, Spark, Beam)
- Ensure alignment with security, compliance, and data governance requirements
Scalability, Reliability Observability
- Establish and operationalize SLAs, SLOs, and error budgets
- Design for high availability, multi-region resilience, and disaster recovery
- Implement enterprise-grade observability frameworks (monitoring, logging, tracing)
Leadership Collaboration
- Partner with engineering, product, security, and data teams to align architecture with organizational objectives
- Provide technical leadership, mentorship, and architectural oversight across multiple teams
- Lead design reviews and ensure adherence to architectural standards
Required Qualifications
Experience
- 10+ years of experience in distributed systems and backend architecture
- Demonstrated success in scaling systems to:
- 50M+ connected devices/sensors, or
- Comparable high-scale environments (e.g., IoT, telecom, fintech, ad-tech, infrastructure platforms)
- Proven experience with high-throughput event-driven architectures in production environments
Technical Expertise
- Deep understanding of distributed systems concepts, including:
- CAP theorem, consistency models, and trade-offs
- Partitioning, replication, and sharding strategies
- Event delivery semantics (at-least-once, exactly-once, idempotency)
- Strong experience with:
- Streaming and messaging systems (Kafka, Pulsar, or equivalent)
- Real-time and batch processing frameworks
- Scalable NoSQL and analytical data stores
System Design Engineering
- Expertise in designing:
- Low-latency, high-throughput APIs
- Event-driven and asynchronous processing systems
- Multi-region, highly available architectures
- Strong programming proficiency in one or more of: Go, Java, Scala, or Rust
Preferred Qualifications
- Experience with large-scale IoT or telemetry platforms
- Familiarity with edge-to-cloud architectures
- Experience operating in multi-cloud or hybrid environments
- Knowledge of enterprise security frameworks, data governance, and compliance (e.g., SOC2, ISO, GDPR)
- Exposure to AI/ML data pipelines and large-scale analytics platforms
Success Metrics
- Architecture supports billions of daily events with consistent performance and reliability
- Systems demonstrate horizontal scalability and fault isolation
- Clear separation and optimization of real-time vs batch workloads
- Strong adherence to enterprise architecture and governance standards