bryan0 · 826 points · 344 comments · 19 godzin temu · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
JUjug18 godzin temu
We also have Nerf Bench:
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
CHChristmasTomer2 godziny temu
Glad someone's keeping track. These threads always turn into “they made it dumber” vs “nah, you're imagining it.”
A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
SHsheepscreek14 godzin temu
> It could also mean nothing happened and people are pattern-matching on noise.
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
MSmsejas10 godzin temu
As a claude code power user, when I get the 'rate the feedback on Claude' pop up, I used to say good or fine out of habit, and immediately after sending this feedback, I felt an instant degradation and mistakes that usually don't happen.
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
CHchrisss3951 godzinę temu
This is great. The business these companies are in is actually quite simple, and managing the expense of running the hardware is an enormous lever for them when it comes to profitability. They have many developers whose sole focus is on this optimization problem.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
We also have Nerf Bench: https://www.bridgebench.ai/nerf-bench They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra. This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
Glad someone's keeping track. These threads always turn into “they made it dumber” vs “nah, you're imagining it.” A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
> It could also mean nothing happened and people are pattern-matching on noise. It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday. The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests. What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output. But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
As a claude code power user, when I get the 'rate the feedback on Claude' pop up, I used to say good or fine out of habit, and immediately after sending this feedback, I felt an instant degradation and mistakes that usually don't happen. Now I dismiss it every time and the quality is more consistent. Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
This is great. The business these companies are in is actually quite simple, and managing the expense of running the hardware is an enormous lever for them when it comes to profitability. They have many developers whose sole focus is on this optimization problem.