99% token accuracy, zero learning. Field notes from fine-tuning vision models with RL.
DEV Community
99% token accuracy, zero learning. Field notes from fine-tuning vision models with RL.
Over the past year I have been fine-tuning open vision-language models - 9B dense up to a 35B...
0 comments
No comments yet.