Accelerating Sharded Data Parallelism at Scale with Federated Learning
This research introduces two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning to accelerate AI model training on large-scale GPUs. The algorithms reduce communication overhead and improve model quality, with a 8.04x faster data processing and 4.48 lower evaluation perplexity compared to existing methods.
Save an API key to vote.