Concepts

Cost estimates

What 5 TB/day of logs costs to run on Plural Telemetry, on self-hosted Loki, and on Datadog. Every assumption is adjustable.

Fig. — Monthly cost of log ingestion
Plural Telemetry$2,149/mo
3 writers × 8 vCPU + 2 readers × 4 vCPU
Grafana Loki$6,860/mo
127 vCPU across 9 components, 17 ingesters3.2× Plural Telemetry
Datadog$413k/mo
234B events/month192.1× Plural Telemetry
Plural Telemetry
Compute
24 writer + 8 reader vCPU
$953
Object storage
10,313 GB incl. index
$237
Object requests
WAL, SST, compaction, cache misses
$508
Cross-AZ traffic
shard forwarding only
$450
Grafana Loki
Compute
127 vCPU (66 in RF=3 ingesters)
$3,783
Object storage
9,656 GB chunks + index
$222
Object requests
chunk flushes, query fetches
$693
Cross-AZ traffic
replication to 3 ingesters
$2,027
WAL volumes
17 × 100 GB gp3
$136
Datadog
Ingestion
$0.10 / GB ingested
$15k
Indexing
$1.70 / M events, 15-day retention
$398k
Estimates for comparison only, using AWS us-east-1 on-demand prices and Datadog list prices with an annual commitment. They exclude engineering time, Datadog Flex Logs, and committed-use discounts. The full set of assumptions is listed below.

The scenario

The default scenario is a mid-to-large platform team:

  • 5 TB/day of raw logs, about 58 MB/s on average and 87 MB/s at a 1.5× peak.
  • 15-day retention, which is Datadog's standard indexed tier.
  • 650-byte average event: 70% at 250 B, 25% at 1 KiB, 5% at 4 KiB. That's about 7.7 billion events a day.
  • 8× compression on the stored data, applied equally to Plural Telemetry and Loki.
  • AWS us-east-1 on-demand prices, three availability zones.

Where the differences come from

Datadog is priced per event indexed. At this volume, indexing is about 96% of the bill, and the price scales linearly with event count, not bytes. Flex Logs and sampling bring it down, but only by keeping less searchable data.

Loki is much cheaper than Datadog, but three things in its architecture cost real money at this scale:

  1. Replication factor 3. Every write fans out to three ingesters, so ingest CPU and memory roughly triple.
  2. Cross-AZ traffic. Two of the three replicas sit in other zones, and AWS charges $0.01/GB in each direction. At 5 TB/day that's around $2,000/month that does nothing but move bytes.
  3. Component sprawl. Frontends, schedulers, index gateways, rulers and three memcached fleets all need baseline capacity.

Plural Telemetry writes each record once, to one shard owner, and lets S3 provide durability. Its cross-AZ cost comes only from forwarding a request to the shard owner when an agent hits a different writer. An ingest buffer partitioned by shard would remove even that.

Assumptions

InputValueSource
Compute$0.0408 per vCPU-hour, memory includedm7g.2xlarge: 8 vCPU / 32 GiB at $0.3264/h
Plural Telemetry writer throughput40 MiB/s per 8-vCPU writer, sized at 85% of that rate at peakOne-third of SlateDB's 120 MiB/s sustained-ingest result, with 15% peak headroom
Plural Telemetry readers4 vCPU each; 1 per 3 writers, at least 2Moderate query load; readers do not perform ingestion or compaction
Loki distributor1 vCPU per 10 MB/s at peakGrafana sizing guidance
Loki ingester1 vCPU per 4 MB/s, ×3 replicasRF=3
Loki query path60% of Plural Telemetry's total computeGenerous to Loki
Loki other12 vCPU + 4 per TB/dayindex-gw, compactor, ruler, frontend, scheduler, memcached
S3 storage$0.023 / GB-monthS3 Standard
S3 requests$0.005 / 1k PUT, $0.0004 / 1k GETWAL at 10 PUT/s per writer, plus SST and compaction writes
Cross-AZ$0.02 / GB crossingWire compression about 3×
Datadog$0.10 / GB ingested + $1.06–$2.50 / M events indexed by retentionList price, annual commitment
These are planning numbers, not quotes

Plural Telemetry's per-writer throughput is still a provisional estimate until the product-shaped scalability benchmarks are published. Your mix of event sizes, stream cardinality, and query load will move all three columns. Run the scalability runner on your own object store and instance types before committing to a size.