the.bay.news

A Better FP4 Gradient Quantizer That Training Couldn't Notice

DEV Community
A Better FP4 Gradient Quantizer That Training Couldn't Notice
A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the published state of the art, two rented-GPU training runs that landed within noise anyway, and the measurement that explains both: the gap only shows at million-token batches, 35 to 643 times larger than anything I ran.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.