the.bay.news

Scheduling every job on a GPU that can only hold one model

DEV Community
Scheduling every job on a GPU that can only hold one model
I run this homelab on two small boxes. One never turns off and just serves DNS and background services; the other has a GPU with a memory pool that can hold exactly one loaded model at a time, nothing more. This post covers the queue and gateway that constraint forced me to build, and the network shape wrapped around both boxes.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.