Backend & infra, 2026
Distributed observability pipeline
A Datadog-style metrics pipeline in Go and Rust that keeps every data point when storage fails.
- Role
- Personal project
- Type
- Backend & infra
- Links
- Source on GitHub
- unit tests across Go and Rust
- 53
- integration tests
- 25
- containers load-tested with k6
- 14
How it fits together
- ClientsGo SDK, HTTP POST /ingest, gRPC :50051
- HTTP / gRPC
- Ingestor (Go)Gin, rate limiting, multi-tenancy, Avro validation, Redis cache
- metrics.raw (Avro)
- Kafka (KRaft)3 topics, 6 partitions, RF=2, dead-letter queue
- consume
- Processor (Rust)EWMA + Z-score anomalies, circuit breaker, batch writer
- commit after write
- StorageClickHouse (1-year TTL, hourly rollups) and Prometheus (15 days)
- query
- ObserveGrafana, AlertManager, Tempo and Jaeger via the OTel Collector
The problem
Metrics pipelines usually fail quietly. When the database slows down, consumers either drop data or stall ingestion for everyone. I wanted to build the whole path of a Datadog-style system myself and make each failure mode an explicit design decision.
How I approached it
- 1
Ingest
A Go service on Gin validates incoming metrics and publishes them to 6-partition Kafka topics, with a dead-letter queue for anything malformed. gRPC/protobuf definitions describe the same contract.
- 2
Detect
A Rust consumer on tokio runs EWMA and rolling Z-score anomaly detection on each series as it streams through.
- 3
Store
Two tiers: Prometheus keeps 15 days hot, ClickHouse MergeTree keeps a year with monthly partitions, a TTL and hourly rollups.
- 4
Operate
Per-tenant token-bucket rate limiting, W3C Trace Context propagated to Tempo and Jaeger, and multi-window SLO burn-rate alerts modeled on the Google SRE workbook.
What I built
- Kafka offsets are committed only after a successful ClickHouse write, which gives at-least-once delivery with no silent data loss.
- A circuit breaker opens after 5 consecutive failures and backs off exponentially to 60 seconds, so a degraded database shows up as bounded consumer lag instead of a stalled pipeline.
- A 12-component Helm chart with HPAs and PDBs, plus a Terraform blueprint for AWS EKS, describe how the stack would run in production.
The result
The full 14-container stack runs locally and is covered by 53 unit tests (41 Go, 12 Rust), 25 integration tests and a k6 load-test script.
Built with
- Go
- Rust
- Kafka
- ClickHouse
- Prometheus
- Grafana
- Kubernetes
- Helm
- Terraform
- gRPC