Skip to content

Performance Profiling ​

Atlas ships a CPU profiler for the backup and restore pipelines. It produces structured text reports meant for both human review and automated analysis.

Quick Start ​

bash
# Build the profiler (once, or after changes to tools/perf/)
pnpm run perf:build

# Profile a single-mailbox backup
pnpm run perf:backup -- -m user@example.com

# Profile a restore
pnpm run perf:restore -- -s <snapshot-id> -m target@example.com

# Analyze a previously captured .cpuprofile
pnpm run perf:analyze -- .perf-output/CPU.20260506.123456.12345.0.001.cpuprofile

Profiling runs the built CLI, so pnpm run build has to have produced packages/cli/dist/. The profiler resolves that entry from the CLI package's own bin field rather than a hardcoded filename, and stops with a message naming the missing file if the build is absent, instead of profiling a failed module load.

How It Works ​

The profiler uses Node.js built-in V8 CPU profiling (--cpu-prof) to sample the call stack at 500-microsecond intervals while Atlas runs. After the process exits, the captured .cpuprofile is parsed into an aggregated report.

atlas CLI process
    |
    v
node --cpu-prof --cpu-prof-dir=.perf-output packages/cli/dist/cli.mjs backup ...
    |
    v
.perf-output/CPU.*.cpuprofile   (raw V8 profile)
    |
    v
atlas-perf analyze              (parser + formatter)
    |
    v
Structured text report          (stdout)

Report Sections ​

Top Functions by Self-Time ​

The functions where CPU is actually consumed, excluding time spent in their callees. High self-time means the function is doing expensive work directly.

ColumnMeaning
Self msMilliseconds spent in this function only
Self %Proportion of total profiled time
Total msTime including all callees
FunctionFunction name
LocationFile path and line number

Domain Breakdown ​

Aggregates all functions by their Atlas package, giving a high-level view of where compute time goes:

DomainWhat it covers
@wisecom/atlas-core/cryptoKey derivation (scrypt), AES-256-GCM encrypt/decrypt
@wisecom/atlas-s3S3 PutObject/GetObject, MD5 checksum, client operations
@wisecom/atlas-m365-graphGraph client factory, rate limiting, retry logic
@wisecom/atlas-driveDownload retry, streaming restore, manifest chain folding
@wisecom/atlas-outlook/backupFolder sync, delta processing, attachment storage
@wisecom/atlas-outlook/restoreMessage reconstruction, folder creation, uploads
node:cryptoNative crypto primitives (called by core/crypto)
node:networkTLS handshakes, HTTP framing, TCP
aws-sdkAWS SDK v3 internals
ms-graph-sdkMicrosoft Graph client library

Hot Paths ​

The critical call chains from entry to the heaviest leaf. Each path follows the most expensive branch at every call site, revealing the dominant execution flow.

Observations ​

Auto-generated summary noting the proportion of time spent in crypto, S3, Graph, and network subsystems.

Flamegraph Mode ​

For interactive visual analysis, use the --flamegraph flag (requires 0x installed as a dev dependency):

bash
node tools/perf/dist/cli.js profile --flamegraph -- backup -m user@example.com

This produces the .cpuprofile text report and an interactive HTML flamegraph in .perf-output/.

The elliptic audit finding ​

0x pulls a browserify chain to render its HTML output, and pnpm audit reports one low-severity advisory from the bottom of it: tools/perf > 0x > browserify > crypto-browserify > browserify-sign > elliptic (GHSA-848j-6mx2-7j84). It is accepted rather than fixed. The advisory has no patched version to move to, and elliptic is only reachable when a developer renders a flamegraph on their own machine: no published package depends on it, and Atlas never loads it at runtime. Any other advisory pnpm audit reports is a real finding and belongs in an issue.

Limitations ​

CPU profiles only capture compute time. Network I/O (waiting for Graph API responses, waiting for S3 uploads to acknowledge) appears as idle time and is NOT reflected in the profile. The profile answers "what is burning CPU?" not "what is the process waiting on?"

For I/O-bound bottleneck analysis:

  • Use the elapsed_ms timers already present in backup/restore output
  • Compare total wall-clock time against CPU time. A large gap means I/O dominates
  • Add targeted performance.now() spans around suspected network operations

Profiling Tips ​

  • Profile with realistic data: A single-message backup won't reveal concurrency bottlenecks. Use a mailbox with 50+ messages and attachments.
  • Compare before/after: Always capture a baseline profile before optimizing, then re-profile after to validate the improvement.
  • Check sample count: If the report shows very few samples (<100), the operation completed too fast for meaningful profiling. Use a larger dataset.
  • Mind the overhead: CPU profiling adds ~5% overhead. The absolute numbers are slightly inflated, but relative proportions remain accurate.

Architecture ​

The profiling tool lives in tools/perf/ (not a published package):

tools/perf/
  src/
    cli.ts                 # Commander CLI: 'profile' and 'analyze' subcommands
    profiler.ts            # Spawns node with --cpu-prof, manages artifacts
    profile-parser.ts      # Parses .cpuprofile JSON, builds call tree
    domain-classifier.ts   # Maps V8 script URLs to Atlas domain names
    report-formatter.ts    # Renders analysis as structured text
    types.ts               # TypeScript interfaces
  package.json
  tsconfig.json

Output artifacts are written to .perf-output/ (git-ignored).

Released under the Apache-2.0 License.