How to choose an LLM for your task
On the text leaderboard of Arena, a public scoreboard that ranks AI models by user votes, the top three scores sat about five points apart at the time of writing [1]. The site's own margin of error, its estimate of how far off each score could be, was about five points too. Read strictly, the top of the table was a tie, printed as a ranking all the same. That is the starting problem for anyone working out how to choose an LLM.
