Status: active.
This crate is the first executable companion for CS336-style distributed parallelism in the CS336 Rust Equivalent track.
It teaches parallelism as typed partitioning:
WorldSize + RankIndex -> RankId
GlobalBatchSize / WorldSize -> LocalBatchSize
ModelWidth / WorldSize -> ShardWidth
LayerCount / WorldSize -> LayersPerRank
CollectiveTrace + ParallelTraceVisibility -> PublicParallelismReport
- lecture direction: parallelism in CS336 Rust Equivalent
- package:
rust_ml_parallelism
- active teaching crate
- typed ranks, world sizes, batch sizes, model widths, layer counts, micro-batches, shard starts, shard lengths, and communication bytes
- data-parallel, tensor-parallel, and pipeline-parallel layout summaries
- tiny all-reduce trace over rank-owned shard sums
- typed
std::opsarithmetic for exact splits, pipeline schedule length, communication addition, and communication-round multiplication - public-report review that blocks restricted or private collective traces before they reach learner-facing material
- expressive
thiserrordiagnostics throughParallelismError
src/
error.rs
lib.rs
examples/
01_data_parallel_batch.rs
02_tensor_parallel_width.rs
03_collective_all_reduce.rs
04_pipeline_schedule.rs
05_public_report.rs
01_data_parallel_batchsplits a batch axis into rank-owned shards.02_tensor_parallel_widthsplits model width across ranks.03_collective_all_reduceshows why data parallelism needs gradient communication.04_pipeline_scheduleconnects layers, stages, micro-batches, and pipeline bubbles.05_public_reportseparates public toy traces from restricted or private distributed-training evidence.
Read parallelism as a map from one global object into rank-indexed local objects:
GlobalBatchSize / WorldSize -> LocalBatchSize
TensorLine / WorldSize -> RankShard*
RankShard* -> CollectiveTrace
ReviewedCollectiveTrace -> PublicParallelismReport
LayerCount / WorldSize -> PipelineLayout
The composition rule is ownership preservation. A distributed plan is only trustworthy when every rank has a valid identity, every shard has a clear origin, every communication estimate carries units, and every learner-facing report has passed an explicit visibility review.
- Which global object is being split?
- Which rank owns each local object?
- Which reviewed traces are safe for the public learning surface?
cargo test --manifest-path code/Cargo.toml -p rust_ml_parallelism --all-targetscargo run --manifest-path code/Cargo.toml -p rust_ml_parallelism --example 01_data_parallel_batch
cargo run --manifest-path code/Cargo.toml -p rust_ml_parallelism --example 02_tensor_parallel_width
cargo run --manifest-path code/Cargo.toml -p rust_ml_parallelism --example 03_collective_all_reduce
cargo run --manifest-path code/Cargo.toml -p rust_ml_parallelism --example 04_pipeline_schedule
cargo run --manifest-path code/Cargo.toml -p rust_ml_parallelism --example 05_public_reportThis crate intentionally does not use GPUs, NCCL, PyTorch, MPI, or real network communication.
The goal is to teach the invariants first: ranks must fit the world, splits must divide evenly in the simple examples, communication has units, and each parallel strategy splits a different axis of the training problem. Public reports are built only from toy public traces so private cluster details, benchmarks, and operational evidence cannot leak into the learner-facing repo.