Metrics

Architecture

A Prometheus-compatible time-series database on SlateDB. It takes remote write and OTLP, answers PromQL, and runs as two StatefulSets and a bucket.

Fig. 1 · Plural Metrics vs Grafana Mimir
Query load
Plural Metrics
two StatefulSets on SlateDB
AgentsPrometheus · Alloy · OTelGrafanaquerieswriter1 shard / pod×2readerreads every shard×4Kubernetes LeasesSlateDBone per shardObject storage
2
component kinds
6
pods
0
replicated disks
0
hash rings

More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.

Grafana Mimir
microservices mode
AgentsPrometheus · AlloyGrafanaqueriesdistributor×5ingesterRF=3 · TSDB head + WAL×15query-frontend×2scheduler×2querier×9store-gw×4memcached×3 pools×6memberlist rings: distributor · ingester · store-gateway · compactorcompactor×2ruler · alertmanager×3Object storage2h TSDB blocks
10
component kinds
48
pods
15 PVCs
replicated disks
4
hash rings

Each series lives in memory on three ingesters until a 2-hour block is cut and uploaded. Adding ingesters reshards series via the ring; store-gateways reshard blocks via their own ring.

Pod counts are illustrative sizing for comparison, derived from vendor capacity guides and our provisional writer envelope. They are not benchmark results.

Write path

  1. Prometheus, Alloy or an OTel Collector sends a batch to any writer, over Remote Write 1.0 or 2.0 or OTLP.
  2. Each series is hashed by its namespace and sorted labels to a storage shard, using the routing epoch in effect at each sample's timestamp. Samples for other shards are forwarded to their owner over internal gRPC.
  3. The owner buffers samples in memory and acknowledges at the configured durability.
  4. On flush, samples are appended to their series in the current hour bucket as merge operands, alongside any new series dictionary and index entries, in one atomic SlateDB write. The bucket's write generation is bumped.
  5. SlateDB compaction collapses the merge operands of each series into compressed sample records.

Read path

  1. Grafana sends PromQL to any reader. Readers open every shard read-only.
  2. The planner narrows the hour buckets by time, then resolves label matchers against each bucket's Roaring-bitmap postings, smallest first. Resolved matchers are cached per bucket.
  3. Only the matching series' sample records are loaded, through the block cache.
  4. Series are joined across shards by label fingerprint and evaluated.
  5. Range-query steps are cached and reused while the write generation of every bucket they read is unchanged, so dashboards refreshing over old data skip storage entirely.

Storage layout

Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and hour bucket:

Key layout
subsystem
0x01
version
0x03
namespace
bytes · 0x00
bucket start
u32 BE
bucket size
u8
record
u8
SlateDB segment ▲

Every key starts with this prefix, so each tenant namespace and hour bucket is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:

0x02Series dictionaryLabel fingerprint → bucket-local series ID
0x03Forward indexSeries ID → labels, type and unit
0x04Inverted indexLabel term → Roaring bitmap of series IDs
0x05Time seriesCompressed samples for one series
0x06Bucket generationWrite generation, for result-cache invalidation
0xffDiscovery catalogLabel names, values and metric metadata
Fig. 2 · Every Metrics key begins with the same prefix. Series IDs are local to one bucket of one shard.

Indexes are bucket-local. A query over a long range repeats lookups per bucket, and old buckets stay isolated from series churn in the current one. Retention drops whole buckets.

Compared with Mimir

ConcernMimirPlural Metrics
Recent dataTSDB head in ingester memory, 3× replicated, cut into 2-hour blocksWriter buffer, then SlateDB on object storage
Write fan-out3× to ingesters via the hash ring1× to the shard owner
Historical readsStore-gateways shard blocks via their own ringReaders open every shard; one block cache
CompactionSeparate compactor ring, block splittingSlateDB's built-in compactor, per shard
Query tiersFrontend, scheduler, queriers, store-gateways, cachesReaders

Compatibility

PromQL is checked against the upstream Prometheus promqltest corpus. Unsupported constructs return an error rather than a wrong answer. Supported: native histograms, out-of-order samples, staleness markers, offset and @, subqueries, every stable aggregation and function, and /federate. Not supported:

  • Exemplars are dropped, and remote read (/api/v1/read) is not served.
  • The lookback delta is fixed at 5 minutes; timeout, limit and stats are ignored.
  • There is no scraper, rule evaluator or alerting. Run Prometheus or Alloy in agent mode, and evaluate rules elsewhere.
  • Experimental upstream functions such as limitk, mad_over_time and info.