Storage Setup
Atlas stores backups in any S3-compatible object storage. For self-hosted deployments, MinIO running in Docker is the recommended option.
Storage Backend: MinIO on Docker
The included docker/docker-compose.yml starts MinIO with a Docker named volume:
services:
minio:
image: quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z
container_name: atlas-minio
ports:
- '${MINIO_API_PORT:-9000}:9000'
- '${MINIO_CONSOLE_PORT:-9001}:9001'
environment:
MINIO_ROOT_USER: ${MINIO_ROOT_USER:-minioadmin}
MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD:-minioadmin}
volumes:
- minio_data:/data
command: server /data --console-address ":9001"
volumes:
minio_data:That is fine for development and quick testing. Real deployments need control over where the data lands.
The image comes from quay.io, not Docker Hub. Docker Hub refuses anonymous pulls of minio/minio, so a docker compose up against that registry fails with pull access denied before MinIO ever starts. quay.io/minio/minio serves the same images without authentication.
The tag is a pinned release rather than latest. Object Lock and versioning behaviour has changed between MinIO releases, and immutable backups depend on both, so an unattended docker compose pull should not be able to move the storage backend under an existing archive. Bump the tag deliberately, then re-run atlas storage-check to confirm the new release still reports the bucket as lock-capable.
Pointing MinIO to External Storage
To store backup data on a specific disk or mount point, replace the named volume with a bind mount. For an external drive mounted at /mnt/backup-drive:
services:
minio:
image: quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z
container_name: atlas-minio
ports:
- '9000:9000'
- '9001:9001'
environment:
MINIO_ROOT_USER: ${MINIO_ROOT_USER}
MINIO_ROOT_PASSWORD: ${MINIO_ROOT_PASSWORD}
volumes:
- /mnt/backup-drive/atlas-data:/data
command: server /data --console-address ":9001"Make sure the directory exists and has correct permissions:
sudo mkdir -p /mnt/backup-drive/atlas-data
sudo chown -R 1000:1000 /mnt/backup-drive/atlas-dataThe UID 1000 matches the default MinIO container user. If your MinIO image uses a different UID, check with docker exec atlas-minio id.
Filesystem Recommendations
Format the storage volume with ext4 or XFS. Both are stable Linux filesystems suitable for object storage workloads:
- ext4: the default on most Linux distributions, battle-tested, good all-around performance.
- XFS: excels with large files and high-throughput sequential writes, which matches the pattern of backup data.
Avoid NTFS, FAT32, or network filesystems (NFS, CIFS) for the MinIO data directory. They lack the POSIX semantics MinIO relies on and will cause subtle corruption or performance issues.
Single Drive vs. RAID Storage
Single Drive Warning
Storing mailbox backups on a single external hard drive is acceptable only for home labs and personal testing. A single drive is a single point of failure. If it fails, all backup data is permanently lost. This setup must not be deployed in any professional or business environment.
Why Redundancy Matters
If your backup storage itself has no redundancy, you have not reduced risk. You have moved the single point of failure from Microsoft 365 to a USB drive on your desk.
RAID: Redundant Array of Independent Disks
RAID combines multiple physical drives into a single logical volume with built-in redundancy. If one drive fails, the data survives on the remaining drives and you can replace the failed drive without losing anything.
| RAID Level | Minimum Drives | How It Works | Usable Capacity | Recommendation |
|---|---|---|---|---|
| RAID 1 (mirror) | 2 | Every write goes to both drives simultaneously. Either drive can fail without data loss. | 50% of total | Small deployments (2-10 mailboxes) |
| RAID 5 (distributed parity) | 3 | Data is striped across all drives with parity blocks. Any single drive can fail. | (N-1)/N of total | Medium deployments |
| RAID 6 (double parity) | 4 | Like RAID 5 but with two parity blocks. Any two drives can fail simultaneously. | (N-2)/N of total | Large deployments where downtime is unacceptable |
For most Atlas deployments, RAID 1 is the practical minimum. It is simple to set up with Linux mdadm or a hardware RAID controller, and it protects fully against a single drive failure.
RAID Is Not a Backup
RAID protects against hardware failure. It does not protect against accidental deletion, ransomware, or corruption that propagates to all mirrors. For critical data, combine RAID with off-site replication (e.g. a second MinIO instance at another location, or periodic copies to cloud storage).
Setting Up RAID 1 on Linux
A minimal RAID 1 setup with two drives:
# Create the RAID 1 array
sudo mdadm --create /dev/md0 --level=1 --raid-devices=2 /dev/sdb /dev/sdc
# Format with XFS
sudo mkfs.xfs /dev/md0
# Create mount point and mount
sudo mkdir -p /mnt/atlas-raid
sudo mount /dev/md0 /mnt/atlas-raid
# Add to /etc/fstab for persistence
echo '/dev/md0 /mnt/atlas-raid xfs defaults 0 0' | sudo tee -a /etc/fstab
# Save RAID configuration
sudo mdadm --detail --scan | sudo tee -a /etc/mdadm/mdadm.confThen point your MinIO bind mount to /mnt/atlas-raid/atlas-data.
MinIO Security
Change Default Credentials
The default MinIO credentials (minioadmin/minioadmin) are public knowledge. Anyone who can reach your MinIO port can read, modify, or delete all backup data. Change these immediately in any non-development deployment.
Set strong credentials in your docker/.env:
MINIO_ROOT_USER=atlas-backup-admin
MINIO_ROOT_PASSWORD=a-long-random-passphrase-at-least-32-charactersTLS for Non-Localhost Deployments
If MinIO is reachable from the network and not just localhost, enable TLS. Without it, S3 credentials and backup data travel in plaintext and anyone on the same network segment can intercept them.
MinIO supports TLS natively by placing certificate files in the container. See the MinIO TLS documentation for setup instructions. When TLS is enabled, update your Atlas endpoint to use https://:
ATLAS_S3_ENDPOINT="https://minio.internal:9000"Bucket Provisioning and Object Lock
Atlas creates each tenant's bucket automatically on first backup (atlas-{tenant_id}), and since v2.1.0 creates it lock-capable: CreateBucket is issued with ObjectLockEnabledForBucket: true, which also enables versioning.
This front-loading is deliberate. On both AWS S3 and MinIO, Object Lock can only be enabled at bucket creation (AWS requires a support ticket to retrofit it; MinIO refuses outright). A lock-capable bucket without a retention policy behaves exactly like a normal versioned bucket, so nothing changes until you opt into immutability with --retention-days on a backup command.
New buckets also receive housekeeping lifecycle rules: incomplete multipart uploads are aborted after 7 days, orphaned delete markers are removed, and noncurrent object versions expire after 30 days. Versioning leaves stale versions behind whenever a delta cursor or index is overwritten, and Atlas never reads them: its file version history is stored as first-class objects, not S3 versions.
Provisioning happens on write paths only. Read-only commands (list, read, status, verify, stats, list-users) never issue CreateBucket and never bootstrap a DEK. They fail with No backups found for tenant <id> when the tenant has no stored key, so a mistyped -t cannot leave an empty billable bucket behind. See Security for the per-command-class IAM split.
Backends that reject ObjectLockEnabledForBucket get a plain bucket and a loud warning at creation time; immutability will not work there.
Run atlas storage-check to see which class a tenant's bucket is: lock-capable, versioned-only (legacy), or unversioned (legacy).
Migrating a Legacy Bucket to Object Lock
Buckets created before v2.1.0 are not lock-capable and cannot be upgraded in place. The migration is a tenant-by-tenant re-pointing, using Atlas's own replication machinery:
- Create a lock-capable target bucket manually:
aws s3api create-bucket --bucket atlas-{tenant}-v2 --object-lock-enabled-for-bucket(MinIO:mc mb --with-lock). - Replicate the tenant into it:
atlas replicate --tenant <id>with the target configured as the replication destination. This copies the DEK first, then all snapshots, and verifies checksums. (Plainaws s3 syncalso works; the DEK object_meta/dek.encmust be copied as-is.) - Repoint the tenant at the new bucket (rename or alias so
atlas-{tenant_id}resolves to the new bucket, e.g. delete the old bucket and rename via a final sync). - Verify:
atlas storage-check -t <id>should reportlock-capable; runatlas verifyfor content integrity. - Decommission the old bucket once a full backup cycle has succeeded against the new one.
Until migrated, backups to legacy buckets keep working. Only --retention-days immutability is unavailable, and the backup fails fast with ObjectLockUnsupportedError rather than pretending to be immutable.
Protecting the Tenant Encryption Key
Each tenant bucket stores its wrapped data-encryption key at _meta/dek.enc. Atlas writes this object exactly once, with a create-only conditional write (If-None-Match: *, supported by AWS S3 and MinIO): if two processes bootstrap the same tenant concurrently, the second write fails with 412 Precondition Failed and that process adopts the already-stored key. Overwriting a live DEK would make every object encrypted with it permanently undecryptable.
As defense in depth, a bucket policy can mandate the create-only header on the key object, so even a buggy or outdated client cannot overwrite it (AWS S3; uses the s3:if-none-match condition key):
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "DenyDekOverwrite",
"Effect": "Deny",
"Principal": "*",
"Action": "s3:PutObject",
"Resource": "arn:aws:s3:::atlas-*/_meta/dek.enc",
"Condition": { "StringNotEquals": { "s3:if-none-match": "*" } }
}
]
}S3-compatible stores
Some S3-compatible backends (older MinIO releases, some Ceph RGW versions, various emulators) silently ignore unknown conditional headers instead of rejecting the write. On such a backend the create-only guarantee is illusory. Verify support by writing the same key twice with If-None-Match: * and confirming the second write fails with 412. Atlas additionally reads the stored key back after bootstrap and always proceeds with what storage actually holds, which converges concurrent bootstraps even where the header is ignored.
Provider Minimum Billable Object Sizes
Some providers bill a small object above its stored size. A worked example: a drive with 20,000 files carries roughly 8 MB of version index data (about 400 bytes per file). Written as one object per file, which is what older Atlas versions did, those 8 MB are billed as about 1.28 GB per month on a provider with a 64 KB minimum billable size. Writing the same data as one object per backup run removes that multiplier entirely: the floor then applies to a single object instead of twenty thousand.
Atlas therefore writes version index data as one object per owner or site per backup run: onedrive/index/<owner_id>/runs/<snapshot_id>.json for OneDrive and sharepoint/index/<site_id>/runs/<snapshot_id>.json for SharePoint. A quiet run that stores nothing writes no index object at all.
Buckets created by earlier Atlas versions still hold legacy per-file index objects under .../index/<id>/files/<file_id>.json. They remain readable, and reads merge them with the per-run objects, until they are purged.
Content blobs stay one object per file, addressed by SHA-256 (onedrive/data/<owner_id>/<sha256>). Packing multiple files into one blob object is deliberately not implemented: S3 Object Lock applies to the pack rather than to the individual files inside it, and doing it safely needs ranged reads, compaction, and retention design first.
| Provider | Minimum billable object size | Note |
|---|---|---|
| Hetzner Object Storage | 64 KB | API calls are free, so trading requests for fewer objects costs nothing (pricing) |
| Wasabi | 4 KB | Plus a 90-day minimum storage duration (pricing FAQ) |
| AWS S3 Standard | None | |
| AWS S3 Standard-IA / One Zone-IA / Glacier Instant Retrieval | 128 KB | Plus a 30 or 90-day minimum duration (storage classes) |
| AWS S3 Glacier Flexible Retrieval / Deep Archive | No floor, about 40 KB metadata per object | Billed partly at the Standard rate (S3 pricing) |
| Cloudflare R2 | None documented | Usage is billed per GB-month; Infrequent Access has a 30-day minimum (pricing) |
| Scaleway | None documented | Recommends objects larger than 1 MB for Glacier tiers (FAQ) |
| Backblaze B2 | None documented | B2's own cost guide warns some services round up to 128 KB (cost comparison) |
| MinIO self-hosted | None | Filesystem block size still applies |
Per-Object Size Ceiling
A large file is uploaded as a multipart stream to a staging key, then promoted to its content-addressed key with a server-side copy. Two S3 limits bound that:
| Limit | Value | How Atlas stays inside it |
|---|---|---|
Single CopyObject source | 5 GB | A source above it is promoted with ranged UploadPartCopy instead |
| Object size | 5 TB | Not enforced by Atlas; a file larger than this cannot be stored in S3 at all |
| Parts per multipart upload | 10,000 | Copy parts are 1 GiB, which reaches 10 TiB and so never runs out |
The 5 GB threshold is a constant in Atlas, not something read from the backend. MinIO does not enforce the AWS limit, so deciding from the backend's behaviour would mean the request that fails on AWS is the one never exercised locally. Both take the same branch for the same object size.
The promotion is server-side either way: no byte travels back through Atlas, and the ciphertext, its checksum, its metadata and any Object Lock retention are identical whichever request is used. Retention on a ranged copy is declared when the multipart upload is created, which is where it takes effect for a multipart destination.
The practical ceiling is therefore S3's own 5 TB per object, which is far above what Microsoft 365 will serve: OneDrive and SharePoint cap a single file at 250 GB.