GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
crorella · 1K points · 917 comments · kemarin · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
GEgertlabs2 jam yang lalu
We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.
Data at https://gertlabs.com/rankings
MIminimaxirkemarin
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
WHwhatifitoldyou23 jam yang lalu
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
GRgradus_adkemarin
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
THthe_dukekemarin
The GPT 6 release was ... not great.
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work. Data at https://gertlabs.com/rankings
> Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
I must say that this AI thing is going more or less as I felt it would back about a year ago. I think there is no real moat in AI models. It's a commodity and the big labs have predictably been caught in a race to the bottom. Not sure if this is going to turn better or worse for all of us common folks. I must say I'm a bit happy though in the sense that "intelligence" is not going to be controlled and be rented out by a small minority.
Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
The GPT 6 release was ... not great. Sol 6 was so bad that I switched over to Opus 5.5 exclusively. Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna. Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things. I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different. (Note: this is after preferring and shilling Codex/OpenAI models for the last half year)