This is surprising to me. Does anyone have a definitive answer that accounts for this difference between providers?
Do the benchmarks that Open Router runs not account for the stochastic nature of the models? Do the providers lie about quantization or context window sizing? Does Open Router not take into account variations for a given model?
I could understand latency/cost benchmarks varying but not the actual generated token responses of the models given the exact same model parameters.
This is surprising to me. Does anyone have a definitive answer that accounts for this difference between providers?
Do the benchmarks that Open Router runs not account for the stochastic nature of the models? Do the providers lie about quantization or context window sizing? Does Open Router not take into account variations for a given model?
I could understand latency/cost benchmarks varying but not the actual generated token responses of the models given the exact same model parameters.