vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
DEV Community
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
TL;DR: vLLM 0.30.0 can keep a model's weights in a long-running daemon so that engines restart...
0 comments
No comments yet.