I am a Senior Data Engineer specializing in distributed computing, real-time stream processing, and Kubernetes-native data platforms. With 7+ years of experience building resilient data infrastructure, I focus on designing scalable pipelines that balance low latency, data integrity, and cost-efficiency.
- Currently building high-throughput streaming and CDC architectures.
- Deeply interested in database internals, OLAP database optimization, and cloud-native systems.
- Ask me about: Apache Spark, Apache Flink, Kafka, Kubernetes, or Infrastructure as Code.
- What it does: A lightweight, event-driven streaming pipeline designed for message routing, schema validation, and structured telemetry processing.
- Tech: Apache Kafka, Python.
- What it does: A stream processor for high-throughput clickstream event streams, performs stateful sliding-window deduplication, and writes optimized columnar data to an open-source Data Lakehouse using Apache Iceberg, MinIO, and Trino.
- Tech: Trino, Iceberg, Kafka, MinIO, Flink, Rust
- Key Achievement: Achieved near-instant replication with a transactional-to-analytical sync delay late arriving data up to 1 min.
- What it does:: A local k8 cluster that runs Apache Airflow to orchestrate and scale distributed Spark jobs to processes massive datasets
- Tech: Apache Spark, Airflow, Prometheus/Grafana, Kubernetes
- Languages: Python, Scala, Rust, Go, SQL, Bash
- Streaming & Queues: Apache Kafka, Apache Flink, Spark Streaming
- Distributed Processing: Apache Spark, YARN
- Databases & Warehouses: PostgreSQL, ClickHouse, Redis, Google Cloud Bigtable
- Storage & Lakehouse: Apache Iceberg, MinIO, Parquet
- Orchestration & Tools: Apache Airflow, dbt, Debezium, Git
- Infrastructure & DevOps: Kubernetes (Kind/Minikube), Docker, Terraform, Helm, Prometheus, Grafana


