A specialized web agent scored 41.7 on WebRetriever while GPT and Claude failed the same form-filling task
DEV Community
A specialized web agent scored 41.7 on WebRetriever while GPT and Claude failed the same form-filling task
General-purpose LLMs are bad at browser automation. I spent the last few months working with various...
0 comments
No comments yet.