Skip to content

Storage Setup ​

Atlas stores backups in any S3-compatible object storage. For self-hosted deployments, MinIO running in Docker is the recommended option.

Storage Backend: MinIO on Docker ​

The included docker/docker-compose.yml starts MinIO with a Docker named volume:

yaml
services:
  minio:
    image: quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z
    container_name: atlas-minio
    ports:
      - '${MINIO_API_PORT:-9000}:9000'
      - '${MINIO_CONSOLE_PORT:-9001}:9001'
    environment:
      MINIO_ROOT_USER: ${MINIO_ROOT_USER:-minioadmin}
      MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD:-minioadmin}
    volumes:
      - minio_data:/data
    command: server /data --console-address ":9001"

volumes:
  minio_data:

That is fine for development and quick testing. Real deployments need control over where the data lands.

The image comes from quay.io, not Docker Hub. Docker Hub refuses anonymous pulls of minio/minio, so a docker compose up against that registry fails with pull access denied before MinIO ever starts. quay.io/minio/minio serves the same images without authentication.

The tag is a pinned release rather than latest. Object Lock and versioning behaviour has changed between MinIO releases, and immutable backups depend on both, so an unattended docker compose pull should not be able to move the storage backend under an existing archive. Bump the tag deliberately, then re-run atlas storage-check to confirm the new release still reports the bucket as lock-capable.

Pointing MinIO to External Storage ​

To store backup data on a specific disk or mount point, replace the named volume with a bind mount. For an external drive mounted at /mnt/backup-drive:

yaml
services:
  minio:
    image: quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z
    container_name: atlas-minio
    ports:
      - '9000:9000'
      - '9001:9001'
    environment:
      MINIO_ROOT_USER: ${MINIO_ROOT_USER}
      MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD}
    volumes:
      - /mnt/backup-drive/atlas-data:/data
    command: server /data --console-address ":9001"

Make sure the directory exists and has correct permissions:

bash
sudo mkdir -p /mnt/backup-drive/atlas-data
sudo chown -R 1000:1000 /mnt/backup-drive/atlas-data

The UID 1000 matches the default MinIO container user. If your MinIO image uses a different UID, check with docker exec atlas-minio id.

Filesystem Recommendations ​

Format the storage volume with ext4 or XFS. Both are stable Linux filesystems suitable for object storage workloads:

  • ext4: the default on most Linux distributions, battle-tested, good all-around performance.
  • XFS: excels with large files and high-throughput sequential writes, which matches the pattern of backup data.

Avoid NTFS, FAT32, or network filesystems (NFS, CIFS) for the MinIO data directory. They lack the POSIX semantics MinIO relies on and will cause subtle corruption or performance issues.

Single Drive vs. RAID Storage ​

Single Drive Warning

Storing mailbox backups on a single external hard drive is acceptable only for home labs and personal testing. A single drive is a single point of failure. If it fails, all backup data is permanently lost. This setup must not be deployed in any professional or business environment.

Why Redundancy Matters ​

If your backup storage itself has no redundancy, you have not reduced risk. You have moved the single point of failure from Microsoft 365 to a USB drive on your desk.

RAID: Redundant Array of Independent Disks ​

RAID combines multiple physical drives into a single logical volume with built-in redundancy. If one drive fails, the data survives on the remaining drives and you can replace the failed drive without losing anything.

RAID LevelMinimum DrivesHow It WorksUsable CapacityRecommendation
RAID 1 (mirror)2Every write goes to both drives simultaneously. Either drive can fail without data loss.50% of totalSmall deployments (2-10 mailboxes)
RAID 5 (distributed parity)3Data is striped across all drives with parity blocks. Any single drive can fail.(N-1)/N of totalMedium deployments
RAID 6 (double parity)4Like RAID 5 but with two parity blocks. Any two drives can fail simultaneously.(N-2)/N of totalLarge deployments where downtime is unacceptable

For most Atlas deployments, RAID 1 is the practical minimum. It is simple to set up with Linux mdadm or a hardware RAID controller, and it protects fully against a single drive failure.

RAID Is Not a Backup

RAID protects against hardware failure. It does not protect against accidental deletion, ransomware, or corruption that propagates to all mirrors. For critical data, combine RAID with off-site replication (e.g. a second MinIO instance at another location, or periodic copies to cloud storage).

Setting Up RAID 1 on Linux ​

A minimal RAID 1 setup with two drives:

bash
# Create the RAID 1 array
sudo mdadm --create /dev/md0 --level=1 --raid-devices=2 /dev/sdb /dev/sdc

# Format with XFS
sudo mkfs.xfs /dev/md0

# Create mount point and mount
sudo mkdir -p /mnt/atlas-raid
sudo mount /dev/md0 /mnt/atlas-raid

# Add to /etc/fstab for persistence
echo '/dev/md0 /mnt/atlas-raid xfs defaults 0 0' | sudo tee -a /etc/fstab

# Save RAID configuration
sudo mdadm --detail --scan | sudo tee -a /etc/mdadm/mdadm.conf

Then point your MinIO bind mount to /mnt/atlas-raid/atlas-data.

MinIO Security ​

Change Default Credentials

The default MinIO credentials (minioadmin/minioadmin) are public knowledge. Anyone who can reach your MinIO port can read, modify, or delete all backup data. Change these immediately in any non-development deployment.

Set strong credentials in your docker/.env:

env
MINIO_ROOT_USER=atlas-backup-admin
MINIO_ROOT_PASSWORD=a-long-random-passphrase-at-least-32-characters

TLS for Non-Localhost Deployments ​

If MinIO is reachable from the network and not just localhost, enable TLS. Without it, S3 credentials and backup data travel in plaintext and anyone on the same network segment can intercept them.

MinIO supports TLS natively by placing certificate files in the container. See the MinIO TLS documentation for setup instructions. When TLS is enabled, update your Atlas endpoint to use https://:

env
ATLAS_S3_ENDPOINT="https://minio.internal:9000"

Bucket Provisioning and Object Lock ​

Atlas creates each tenant's bucket automatically on first backup (atlas-{tenant_id}), and since v2.1.0 creates it lock-capable: CreateBucket is issued with ObjectLockEnabledForBucket: true, which also enables versioning.

This front-loading is deliberate. On both AWS S3 and MinIO, Object Lock can only be enabled at bucket creation (AWS requires a support ticket to retrofit it; MinIO refuses outright). A lock-capable bucket without a retention policy behaves exactly like a normal versioned bucket, so nothing changes until you opt into immutability with --retention-days on a backup command.

New buckets also receive housekeeping lifecycle rules: incomplete multipart uploads are aborted after 7 days, orphaned delete markers are removed, and noncurrent object versions expire after 30 days. Versioning leaves stale versions behind whenever a delta cursor or index is overwritten, and Atlas never reads them: its file version history is stored as first-class objects, not S3 versions.

Provisioning happens on write paths only. Read-only commands (list, read, status, verify, stats, list-users) never issue CreateBucket and never bootstrap a DEK. They fail with No backups found for tenant <id> when the tenant has no stored key, so a mistyped -t cannot leave an empty billable bucket behind. See Security for the per-command-class IAM split.

Backends that reject ObjectLockEnabledForBucket get a plain bucket and a loud warning at creation time; immutability will not work there.

Run atlas storage-check to see which class a tenant's bucket is: lock-capable, versioned-only (legacy), or unversioned (legacy).

Migrating a Legacy Bucket to Object Lock ​

Buckets created before v2.1.0 are not lock-capable and cannot be upgraded in place. The migration is a tenant-by-tenant re-pointing, using Atlas's own replication machinery:

  1. Create a lock-capable target bucket manually: aws s3api create-bucket --bucket atlas-{tenant}-v2 --object-lock-enabled-for-bucket (MinIO: mc mb --with-lock).
  2. Replicate the tenant into it: atlas replicate --tenant <id> with the target configured as the replication destination. This copies the DEK first, then all snapshots, and verifies checksums. (Plain aws s3 sync also works; the DEK object _meta/dek.enc must be copied as-is.)
  3. Repoint the tenant at the new bucket (rename or alias so atlas-{tenant_id} resolves to the new bucket, e.g. delete the old bucket and rename via a final sync).
  4. Verify: atlas storage-check -t <id> should report lock-capable; run atlas verify for content integrity.
  5. Decommission the old bucket once a full backup cycle has succeeded against the new one.

Until migrated, backups to legacy buckets keep working. Only --retention-days immutability is unavailable, and the backup fails fast with ObjectLockUnsupportedError rather than pretending to be immutable.

Protecting the Tenant Encryption Key ​

Each tenant bucket stores its wrapped data-encryption key at _meta/dek.enc. Atlas writes this object exactly once, with a create-only conditional write (If-None-Match: *, supported by AWS S3 and MinIO): if two processes bootstrap the same tenant concurrently, the second write fails with 412 Precondition Failed and that process adopts the already-stored key. Overwriting a live DEK would make every object encrypted with it permanently undecryptable.

As defense in depth, a bucket policy can mandate the create-only header on the key object, so even a buggy or outdated client cannot overwrite it (AWS S3; uses the s3:if-none-match condition key):

json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "DenyDekOverwrite",
      "Effect": "Deny",
      "Principal": "*",
      "Action": "s3:PutObject",
      "Resource": "arn:aws:s3:::atlas-*/_meta/dek.enc",
      "Condition": { "StringNotEquals": { "s3:if-none-match": "*" } }
    }
  ]
}

S3-compatible stores

Some S3-compatible backends (older MinIO releases, some Ceph RGW versions, various emulators) silently ignore unknown conditional headers instead of rejecting the write. On such a backend the create-only guarantee is illusory. Verify support by writing the same key twice with If-None-Match: * and confirming the second write fails with 412. Atlas additionally reads the stored key back after bootstrap and always proceeds with what storage actually holds, which converges concurrent bootstraps even where the header is ignored.

Provider Minimum Billable Object Sizes ​

Some providers bill a small object above its stored size. A worked example: a drive with 20,000 files carries roughly 8 MB of version index data (about 400 bytes per file). Written as one object per file, which is what older Atlas versions did, those 8 MB are billed as about 1.28 GB per month on a provider with a 64 KB minimum billable size. Writing the same data as one object per backup run removes that multiplier entirely: the floor then applies to a single object instead of twenty thousand.

Atlas therefore writes version index data as one object per owner or site per backup run: onedrive/index/<owner_id>/runs/<snapshot_id>.json for OneDrive and sharepoint/index/<site_id>/runs/<snapshot_id>.json for SharePoint. A quiet run that stores nothing writes no index object at all.

Buckets created by earlier Atlas versions still hold legacy per-file index objects under .../index/<id>/files/<file_id>.json. They remain readable, and reads merge them with the per-run objects, until they are purged.

Content blobs stay one object per file, addressed by SHA-256 (onedrive/data/<owner_id>/<sha256>). Packing multiple files into one blob object is deliberately not implemented: S3 Object Lock applies to the pack rather than to the individual files inside it, and doing it safely needs ranged reads, compaction, and retention design first.

ProviderMinimum billable object sizeNote
Hetzner Object Storage64 KBAPI calls are free, so trading requests for fewer objects costs nothing (pricing)
Wasabi4 KBPlus a 90-day minimum storage duration (pricing FAQ)
AWS S3 StandardNone
AWS S3 Standard-IA / One Zone-IA / Glacier Instant Retrieval128 KBPlus a 30 or 90-day minimum duration (storage classes)
AWS S3 Glacier Flexible Retrieval / Deep ArchiveNo floor, about 40 KB metadata per objectBilled partly at the Standard rate (S3 pricing)
Cloudflare R2None documentedUsage is billed per GB-month; Infrequent Access has a 30-day minimum (pricing)
ScalewayNone documentedRecommends objects larger than 1 MB for Glacier tiers (FAQ)
Backblaze B2None documentedB2's own cost guide warns some services round up to 128 KB (cost comparison)
MinIO self-hostedNoneFilesystem block size still applies

Per-Object Size Ceiling ​

A large file is uploaded as a multipart stream to a staging key, then promoted to its content-addressed key with a server-side copy. Two S3 limits bound that:

LimitValueHow Atlas stays inside it
Single CopyObject source5 GBA source above it is promoted with ranged UploadPartCopy instead
Object size5 TBNot enforced by Atlas; a file larger than this cannot be stored in S3 at all
Parts per multipart upload10,000Copy parts are 1 GiB, which reaches 10 TiB and so never runs out

The 5 GB threshold is a constant in Atlas, not something read from the backend. MinIO does not enforce the AWS limit, so deciding from the backend's behaviour would mean the request that fails on AWS is the one never exercised locally. Both take the same branch for the same object size.

The promotion is server-side either way: no byte travels back through Atlas, and the ciphertext, its checksum, its metadata and any Object Lock retention are identical whichever request is used. Retention on a ranged copy is declared when the multipart upload is created, which is where it takes effect for a multipart destination.

The practical ceiling is therefore S3's own 5 TB per object, which is far above what Microsoft 365 will serve: OneDrive and SharePoint cap a single file at 250 GB.

Released under the Apache-2.0 License.