Shared flowchart

Software · L5 · Distributed Training Architectures

How data parallelism, the parameter server, ring-AllReduce, and pipeline parallelism relate as distinct scaling strategies.

by @openstemUpdated Software
Model too large or dataset too largeData parallelism: replicate model, shard dataModel parallelism: shard model across devicesHow to aggregate gradients?Parameter server: workers push/pull vs. server stateRing-AllReduce: O(1) bandwidth per workerSynchronous SGD: barrier, no staleness, straggler-limitedAsynchronous SGD: no barrier, staleness τPipeline parallelism: layers as sequential stagesPipeline bubble (idle devices)GPipe fill-drain / PipeDream 1F1B schedulingcentraliseddecentralisedmitigate

We use privacy-friendly product analytics (no session recording, PII masked) to improve OpenStem. Load analytics? Privacy Policy