Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> The same model will benchmark very differently

This is surprising to me. Does anyone have a definitive answer that accounts for this difference between providers?

Do the benchmarks that Open Router runs not account for the stochastic nature of the models? Do the providers lie about quantization or context window sizing? Does Open Router not take into account variations for a given model?

I could understand latency/cost benchmarks varying but not the actual generated token responses of the models given the exact same model parameters.

 help



As said in the article, all of the providers run different runtimes with proprietary, sometimes untested and weird optimisations

There are multiple types of quantisation: weights and KV Cache. Quantizing kv cache can drastically hurt performance




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: