Works
03 · Python

Splunk Cluster Backup

A backup that knows a Splunk bucket. I wrote it for indexer clusters that stay on the premises: long retention, no object store, and disks that cannot hold the years.

  1. 01IndexerRead only. The tiers stay on the peer.
  2. 02Warm mirrorA bounded copy. The fastest way back.
  3. 03Cold stagingHolds the cold roll until the archive confirms it.
  4. 04ArchiveDisk or tape. The handoff is Veeam.
  5. 05PurgedThe local copy leaves only after that confirmation.

The same bucket can sit on several peers. One source is kept. A replica is not stored again as if it were new.

State moves forward only: synced, staged, archived, purged. A bucket that shows up again after purge is an integrity alarm, not a step backward.

Nothing on an indexer is written, deleted, or changed. The way out is find and rsync.

Restore starts as a time window. The plan names the peer and the thaweddb path, and spreads the buckets so one node is not flooded. A dry-run writes nothing. If the only copy is on tape, the plan says retrieval.

KV store and the cluster manager configs travel beside the buckets. A search-head cluster is read from its captain. A damaged archive is not restored.

Read-only terminal listing the cluster indexes, with bucket counts, size, and the newest event.
The indexes, as the cluster reports them. Read only.

Open one index and the months are the story.

Index lifecycle page for one index: bucket copies, archive coverage, and how far back a restore can reach.
The same index, as a page. Copies, what Veeam holds, and how far back a restore can reach.

Each ribbon is one bucket. Color is where it sits today.

Bucket river: one ribbon per bucket, colored by state, with the history of one bucket open beside it.
The river. Warm, staging, archive, purged.

Hot is the live edge. Warm and cold are what has to be kept.

Terminal timeline of one index: hot, warm, and cold buckets across three peers, with oldest and newest events.
Hot, warm, and cold across the same months. Three peers, one index.

Disk is the limit.

Capacity of three indexers: hot and warm on one volume, cold on another, with backup state beside the bars.
Hot and warm on one volume, cold on another. The backup state is counted beside them.

Each bucket keeps a single state. Warm, staging, archived, purged. The count above the table is the whole peer.

Read-only manifest browser: three peers, and a table of archived buckets with size and state.
The manifest. Still nothing written.

A restore starts as a window. Index, start, end. Balance across the peers comes before any copy.

Restore search screen with fields for index, start, end, and state. Nothing has been written.
The search. Dates are local. An empty index means every index.

Configs and the KV store sit in the same record.

Config snapshot and KV store archive view: sets per node, snapshot history, and how far a rollback can reach.
Configs and the KV store, beside the buckets.

A mark for each snapshot. The circles are KV store archives.

Snapshot strip of config sets, and KV store archives drawn as circles.
Height is how much changed. A hollow circle is the same digest as the one before it.
04 · Shell

Splunk AIO Backup

The single-host system this continues. One Splunk machine, not a replicated cluster.

Seven named days, Monday through Sunday, each one overwritten the next week. An unchanged bucket stays a hardlink, so the week is not seven full copies.

A failed day cannot be restored. Warm and cold buckets, the KV store, and the configs travel together. A small Splunk app watches the run.