Looking for usage documentation? Check out Fine tuning.
ClassifierBatch
View source EvalResult
View source RegressorBatch
View source main_process_first
View source torchrun for work that should happen once and be read
from a shared cache afterwards, such as dataset downloads: the main
process runs the block while the other ranks wait at a barrier, then the
other ranks run it against the warm cache.
Initializes the process group from the torchrun env vars if needed, and
leaves it initialized so that a subsequent fit() reuses it. Call
torch.distributed.destroy_process_group() at the end of your script.
No-op when running with a single process.