Which model wins when someone picks between real answers to the same prompt. A model is only counted in rounds it actually answered and someone voted on.
Speed and time-to-first-token aren't here yet. Both are currently measured in a way that misreports models which reason before answering, and a wrong number averaged across models would be worse than no number.