Skip to content

Backend & infra, 2026

Distributed observability pipeline

A Datadog-style metrics pipeline in Go and Rust that keeps every data point when storage fails.

Role
Personal project
Type
Backend & infra
unit tests across Go and Rust
53
integration tests
25
containers load-tested with k6
14

How it fits together

  1. ClientsGo SDK, HTTP POST /ingest, gRPC :50051
  2. HTTP / gRPC
  3. Ingestor (Go)Gin, rate limiting, multi-tenancy, Avro validation, Redis cache
  4. metrics.raw (Avro)
  5. Kafka (KRaft)3 topics, 6 partitions, RF=2, dead-letter queue
  6. consume
  7. Processor (Rust)EWMA + Z-score anomalies, circuit breaker, batch writer
  8. commit after write
  9. StorageClickHouse (1-year TTL, hourly rollups) and Prometheus (15 days)
  10. query
  11. ObserveGrafana, AlertManager, Tempo and Jaeger via the OTel Collector

The problem

Metrics pipelines usually fail quietly. When the database slows down, consumers either drop data or stall ingestion for everyone. I wanted to build the whole path of a Datadog-style system myself and make each failure mode an explicit design decision.

How I approached it

  1. 1

    Ingest

    A Go service on Gin validates incoming metrics and publishes them to 6-partition Kafka topics, with a dead-letter queue for anything malformed. gRPC/protobuf definitions describe the same contract.

  2. 2

    Detect

    A Rust consumer on tokio runs EWMA and rolling Z-score anomaly detection on each series as it streams through.

  3. 3

    Store

    Two tiers: Prometheus keeps 15 days hot, ClickHouse MergeTree keeps a year with monthly partitions, a TTL and hourly rollups.

  4. 4

    Operate

    Per-tenant token-bucket rate limiting, W3C Trace Context propagated to Tempo and Jaeger, and multi-window SLO burn-rate alerts modeled on the Google SRE workbook.

What I built

  • Kafka offsets are committed only after a successful ClickHouse write, which gives at-least-once delivery with no silent data loss.
  • A circuit breaker opens after 5 consecutive failures and backs off exponentially to 60 seconds, so a degraded database shows up as bounded consumer lag instead of a stalled pipeline.
  • A 12-component Helm chart with HPAs and PDBs, plus a Terraform blueprint for AWS EKS, describe how the stack would run in production.

The result

The full 14-container stack runs locally and is covered by 53 unit tests (41 Go, 12 Rust), 25 integration tests and a k6 load-test script.

Built with

  • Go
  • Rust
  • Kafka
  • ClickHouse
  • Prometheus
  • Grafana
  • Kubernetes
  • Helm
  • Terraform
  • gRPC