Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

LLM ops is not trivial. The systems for running inference are very complex, and running across multiple GPUs and nodes adds tons more complexity. And when LLMs are run incorrectly, they still work, just not at optimal performance. Even noticing that something is wrong is not trivial, and finding the problem is far far harder.

So I guess I shouldn't be surprised at all to see these benchmarks, but still I am!

There are such huge economies of scale with batched inference that it's clear this sort of service will continue, but it has a lot of growing up to do. Even AWS Bedrock has a Claude that feels different to me, but I haven't had a chance to do actual benchmarks that would show that.

 help



Bedrock Claude is absolutely not identical to 1P Claude.

They’re close enough to not matter though.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: