Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

  if you want to control the routing, you'd lock down providers
I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" section says it works, not randomly have a 100x error rate.

If OpenRouter is a marketplace to pick a provider while avoiding huge problems, it is terrible at that job. It surfaces literally none of that info in the top-level list, the graphs below are mislabeled and useless at best, and doesn't notify you of this horrifying situation anywhere, even in passing. There's not even a way to compare providers, AFAICT -- you can only compare models.

  This is quite literally the point of OpenRouter.
Their tagline is "better prices, better uptime, no subscriptions". The first two of these directly and inherently contradict your understanding -- neither would be possible if OpenRouter was just a fancy way to change something in their GUI rather than changing the target url of your gateway.
 help



> I don't want to control routing. I want the model to work how the giant, prominent "BENCHMARKS" section says it works, not randomly have a 100x error rate.

Why do you care about the public benchmarks at all?

The way companies and effective individual developers use OpenRouter, is that you first create your evaluation framework/benchmark, for your specific tasks and use cases, and make that real easy to run and use various models and providers with it.

Then you run this to gather data. Then you use said data to figure out what works and what the quality/cost tradeoff you want to make is. Then you lock that down in production while you keep iterating on your benchmark to make it match with real-world use cases and keep adding the new models that pop up.

I don't think anyone serious is just willy-nilly making individual requests against OpenRouter and similar platforms, get a "feel for a provider" then use only that provider. Not only would it be wildly inefficient, but also you need hard numbers to compare so you can make informed choices.

For this process and workflow, OpenRouter is great, because adding/changing providers and models is essentially changing two strings, rather than having a adapter for each platform you want to try out.

If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRouter, it doesn't make much sense, but use the provider the model maker has available, as almost all of them run their own endpoints.


If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity.

  I don't think anyone serious is just willy-nilly making individual requests against OpenRouter
Despite your confidence, that is indeed the basis of this massive corporations entire business plan.

  If you just want best accuracy requests from SOTA models for your agent you run locally or whatever, then don't use OpenRoute
If OpenRouter is only for bad accuracy, they should say as much and fade into deserved obscurity.

I think you're focusing only on the general coding agent aspect of LLMs.

> If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity.

That is the way LLMs have to be used for highest reliability. While the term stochastic parrot has been co-opted by unreasonable LLM skeptics, that is indeed what LLMs are. You have to have grounded evals that check outcomes if you want to use them reliably - or a human in the loop works too.

The more general your family of tasks, the less likely you can make automated evals. So for "general" coding agents, you need a human in the loop that can verify and it's not that easy to write an eval.

But if you have specific tasks, then you can spend the time to make a eval, and then you can optimise the way you use the LLM and get extremely good success rates. It's not like it's black magic. Nor does it need large budgets.

> that is indeed the basis of this massive corporations entire business plan.

No. Individual developers using codex (for extremely underspecified general engineering) needs human in the loop, is not amenable to evals but is only a fraction of all LLM usecases.


[flagged]


> Hmm, yeah, good and condense version of what my previous comment said. I'm much impressed by your reading ability.

10month old account with 20k karma. Low value rubbish postings as a professional user. Sad.


You'd be surprised. Not enough teams still have their own evals. India[1] and the US[2] is my experience. To get many of them to understand the benefit of putting a couple of people to do data labelling and write a couple of verifiers for just a few days every few months was so difficult. Many folks have understood it all wrong and made allotments like "big model for this task" "small model for this task". Small was sometimes parameters, sometimes brand version number, or sometimes because it has "mini" in its name.

Proper LLM adoption beyond fucking around with claude code or github copilot is so low and is only going to go up as people figure it out. I also think the new cloud agents thing might accelerate adoption among these companies, since it's a bit more plug and play. But they are more likely to be stingy about it, so I'm not sure about the high margin expectations of certain model families.

[1] not witch, but BFSI. Surprisingly parts of witch companies have it down to science already, they embedded openai or anthropic or startups like devrev more than a year back.

[2] again not high fly SF companies, BFSI.


> There's not even a way to compare providers, AFAICT

That's not quite true. The only thing they don't show per-provider is benchmark data, cause I don't think they are doing continuous benchmarking of each model from each provider, as I assume they feel that's too expensive. You can see hugely detailed breakdowns for near-time metrics per provider for any model by visiting the page for that model on Openrouter. For example see the page for Qwen 3.8 27B: https://openrouter.ai/qwen/qwen3.8-27b

Some of the killer stats they show per provider:

- Pricing: Effective price accounting for cache hit rate, by provider

- Performance: Throughput in tok/s, latency, E2E latency, tool call error rate, structured output error rate, and more; all per provider.

- Uptime: You have to click on the provider to see their specific uptime, but doing so does show the last-7-days uptime, and you can click to see more.


The thing they don't show is the one we really need, especially because model providers can skimp on quality (run lower quantization, lower kv cache precision, etc) to improve their pricing and performance. I agree that it's probably too expensive to keep running the benchmark, but we need some way to hold the providers to a certain standard, otherwise every user has to discover the problems on their own.

We are doing continuous benchmarking of each endpoint, for each provider, and it is very expensive :)

No I know, I was in the GUI as I wrote that lol. As the other person said: if they vary this much in quality, not including that way above updtime and performance is absurd. What would you use a fast, always-up, broken endpoint for?



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: