Skip to content
View shahaba's full-sized avatar
🚀
🚀

Block or report shahaba

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shahaba/README.md

Sign of the Horns Hey, I'm Shahab Robot

About Me Rocket

I am a Senior Data Engineer specializing in distributed computing, real-time stream processing, and Kubernetes-native data platforms. With 7+ years of experience building resilient data infrastructure, I focus on designing scalable pipelines that balance low latency, data integrity, and cost-efficiency.

  • Currently building high-throughput streaming and CDC architectures.
  • Deeply interested in database internals, OLAP database optimization, and cloud-native systems.
  • Ask me about: Apache Spark, Apache Flink, Kafka, Kubernetes, or Infrastructure as Code.

Projects

Event-Driven Message Routing Engine

event-driven-kafka-pipeline

  • What it does: A lightweight, event-driven streaming pipeline designed for message routing, schema validation, and structured telemetry processing.
  • Tech: Apache Kafka, Python.

Real-Time Event Processing & Analytics Lakehouse

lakehouse-realtime-ingestion

  • What it does: A stream processor for high-throughput clickstream event streams, performs stateful sliding-window deduplication, and writes optimized columnar data to an open-source Data Lakehouse using Apache Iceberg, MinIO, and Trino.
  • Tech: Trino, Iceberg, Kafka, MinIO, Flink, Rust
  • Key Achievement: Achieved near-instant replication with a transactional-to-analytical sync delay late arriving data up to 1 min.

Kubernetes-Native Spark & Airflow (In-Progress)

k8s-spark-airflow-platform

  • What it does:: A local k8 cluster that runs Apache Airflow to orchestrate and scale distributed Spark jobs to processes massive datasets
  • Tech: Apache Spark, Airflow, Prometheus/Grafana, Kubernetes

Tech Stack

  • Languages: Python, Scala, Rust, Go, SQL, Bash
  • Streaming & Queues: Apache Kafka, Apache Flink, Spark Streaming
  • Distributed Processing: Apache Spark, YARN
  • Databases & Warehouses: PostgreSQL, ClickHouse, Redis, Google Cloud Bigtable
  • Storage & Lakehouse: Apache Iceberg, MinIO, Parquet
  • Orchestration & Tools: Apache Airflow, dbt, Debezium, Git
  • Infrastructure & DevOps: Kubernetes (Kind/Minikube), Docker, Terraform, Helm, Prometheus, Grafana

Pinned Loading

  1. event-driven-kafka-pipeline event-driven-kafka-pipeline Public

    A small but fun learning pipeline with Kafka and Python

    Python

  2. lakehouse-realtime-ingestion lakehouse-realtime-ingestion Public

    A production-grade, local real-time event ingestion, deduplication, and transformation pipeline

    HCL

  3. go-projects go-projects Public

    A small repo for go related learning and projects

    Shell

  4. ruby-projects ruby-projects Public

    A small repo for Ruby related learning and projects

  5. rust-projects rust-projects Public

    A small repo for my reading of the Rust Book

    Rust