Is 32× A100 40 GB with 1 GPU per node a reasonable distributed-training testbed?
PyTorch Forums
Is 32× A100 40 GB with 1 GPU per node a reasonable distributed-training testbed?
We are designing a distributed-training experiment and currently have access to 32 NVIDIA A100 40 GB GPUs. For both Jetstream2 and Google Cloud, the 1× A100 40 GB per node configurations are the lowest-cost options available to us, so this is the configuration we are considering for the main experiment. Jetstream2 Google Cloud Nodes 32 32 GPU per node 1× A100 40 GB 1× A100 40 GB Total GPUs 32 32 VM type g3.xl a2-highgpu-1g CPU per VM 32 vCPUs 12 vCPUs RAM per VM 120 GB 85 GB ...
0 comments
No comments yet.