We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.
Data at https://gertlabs.com/rankings
MIminimaxir昨天
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
WHwhatifitoldyou23小时前
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
GRgradus_ad昨天
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
THthe_duke昨天
The GPT 6 release was ... not great.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
评论
5 条预览评论 · 正在加载完整讨论请先登录 h4cker 账号,然后连接 Hacker News 后发表评论。
We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work. Data at https://gertlabs.com/rankings
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
The GPT 6 release was ... not great. Sol 6 was so bad that I switched over to Opus 5.5 exclusively. Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna. Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things. I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different. (Note: this is after preferring and shilling Codex/OpenAI models for the last half year)