the.bay.news

The Embedding Table Was 72% of the Model

DEV Community
The Embedding Table Was 72% of the Model
Nearly three quarters of my small transformer was a lookup table, so the dial that mattered was embedding precision, not network precision. int4 with one scale per row costs 16 KB more than one scale per tensor and recovers 2.2 of the 2.5 points per-tensor int4 loses.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.