Logs

Architecture

A Loki-compatible log database on SlateDB. It speaks LogQL and Loki push, and also takes OTLP and Elasticsearch _bulk writes, from two StatefulSets and a bucket.

Fig. 1 · Plural Logs vs Grafana Loki
Query load
Plural Logs
two StatefulSets on SlateDB
AgentsAlloy · OTel · Fluent BitGrafanaquerieswriter1 shard / pod×3readerreads every shard×4Kubernetes LeasesSlateDBone per shardObject storage
2
component kinds
7
pods
0
replicated disks
0
hash rings

More writers means more shards from the next hour on. Nothing is copied or rebalanced, and readers scale on their own.

Grafana Loki
microservices mode
AgentsPromtail · Alloy · OTelGrafanaqueriesdistributor×8ingesterRF=3 · WAL on PVC×18query-frontend×2scheduler×2querier×15index-gw×3memcached×3 pools×6memberlist rings: distributor · ingester · compactor · schedulercompactor×1ruler×2Object storagechunks + TSDB index
9
component kinds
57
pods
18 PVCs
replicated disks
4
hash rings

New ingesters join the hash ring and take over token ranges; every write fans out to three of them. Scaling down means flushing and handing off in-memory chunks and WALs one ingester at a time.

Pod counts are illustrative sizing for comparison, derived from vendor capacity guides and our provisional writer envelope. They are not benchmark results.

Write path

  1. An agent sends a batch to any writer: Loki push, OTLP, or Elasticsearch _bulk.
  2. Each entry's stream is hashed to a storage shard with the routing epoch in effect at its timestamp. Entries for other shards are forwarded to their owner over internal gRPC.
  3. The owner buffers entries in memory and acknowledges at the configured durability (applied by default).
  4. Every 10 seconds, or at 64 MiB, the buffer is flushed. The flush packs every stream of a time segment into ~1 MiB multi-stream objects, updates label and term postings, and commits it all to SlateDB in one atomic write.
  5. Writer-side compaction later merges small objects of a segment, so long-lived streams end up in few objects.

Read path

  1. Grafana sends LogQL to any reader. Readers open every shard read-only and share one block cache.
  2. For each time segment in range, label matchers intersect Roaring-bitmap postings to find stream IDs.
  3. Each stream's runs carry their min and max timestamps, so out-of-range objects are rejected before any block is read.
  4. The surviving blocks are range-read through the RAM + NVMe cache and decoded, then the LogQL pipeline runs over the rows.
  5. Results from every shard are merged by timestamp and cut at the query's limit.

| match "…", a Logs extension, uses the full-text term index to nominate rows before decoding. Loki-style line filters (|=, |~) still scan decoded rows, since they must match substrings.

Storage layout

Each storage shard is one SlateDB database. Keys are scoped by tenant namespace and time segment:

Key layout
subsystem
0x03
version
0x04
namespace
bytes · 0x00
time segment
i64 BE
record
u8
SlateDB segment ▲

Every key starts with this prefix, so each tenant namespace and segment is one contiguous key range. Retention drops whole ranges, and queries scan only the ranges in their time window. The record type comes next:

0x01–03Stream dictionaryStream IDs and their full label sets
0x04Label postingsLabel term → Roaring bitmap of stream IDs
0x05RunOne stream's blocks in one object, with min/max time
0x06Object blockCompressed rows of one stream
0x08–0bSearch recordsFull-text term index for | match
0x0c–0eObject bookkeepingTombstones, rollups and object directories
Fig. 2 · Every Logs key begins with the same prefix. Segments default to one hour.

Rows of one stream are stored as independently decoded blocks of 256 rows. Replaced objects stay readable for 10 minutes after compaction, so in-flight queries never lose data underneath them.

Compared with Loki

ConcernLokiPlural Logs
Recent dataReplicated ingester memory and WALWriter buffer, then SlateDB on object storage
Write fan-out3× to ingesters via the hash ring1× to the shard owner
IndexTSDB index, index gateways, compactorPer-segment postings in the same SlateDB keyspace
Query tiersFrontend, scheduler, queriers, three cachesReaders with one block cache
Full-text searchLine filters onlyLine filters, plus indexed | match

The overview architecture guide walks through each Loki component and the failure modes on both sides.

Compatibility

Against Loki's own LogQL conformance scripts, 1373 of 1385 evaluations match. Stream selectors, line and label filters, json, logfmt, pattern, regexp, line_format, label_format, and all range and vector aggregations are supported. The main differences:

  • Empty stream labels are dropped at ingest.
  • json keeps the last value of a repeated key; Loki keeps the first.
  • approx_count_distinct is exact rather than a HyperLogLog estimate.
  • line_format and label_format also accept Jinja templates.