g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
DEV Community
g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput
Serving Gemma 4 E2B in pure JAX on AWS g5g.2xlarge and g6.2xlarge with a byte-identical payload. The older instance loses 87% of decode to dtype conversion, and nothing in the logs says so.
0 comments
No comments yet.