"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
DEV Community
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
A llama.cpp error that says "turn flash attention on" and means "your V cache is stored transposed" — plus the two-boot probe that showed my context-window formula was off by a factor of two.
0 comments
No comments yet.