D2OQZG8l5BI1S06 · 879 points · 609 comments · anteontem · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
SOSol-anteontem
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
THthefourthchimeontem
It does vey well at one shotting a PacMan clone, pretty much perfect.
https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...
2nd only to Opus 5.5, which is perfect.
https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results:
https://jonclegg.github.io/pacman-bakeoff/
ABabejoraanteontem
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
AZazuanrbontem
Unless you’re using frontier models like Astra, Sol, Fable, or Opus, I think you’re often better off using Chinese models for a fraction of the price. I’m not sure people realises just how competitive they’ve become.
GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.
I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.
The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
VIvincengomesontem
For the folks asking what is the point of Sonnet 5.5 when Opus 5.5 is better in every way, Sonnet is now the new default in the free tier and now people using Claude Web in the free tier get access to an almost close to the frontier model.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5. More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity). So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day. So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point. Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son... 2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu... Up until very recently, all models struggled with this. All results: https://jonclegg.github.io/pacman-bakeoff/
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange. Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap. [1] Section 8.5 of the Sonnet 5.5 System Card
Unless you’re using frontier models like Astra, Sol, Fable, or Opus, I think you’re often better off using Chinese models for a fraction of the price. I’m not sure people realises just how competitive they’ve become. GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case. I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up. The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
For the folks asking what is the point of Sonnet 5.5 when Opus 5.5 is better in every way, Sonnet is now the new default in the free tier and now people using Claude Web in the free tier get access to an almost close to the frontier model.