news.ycombinator.com

Untitled

jug · 0 points · 0 comments · 19 hours ago

We also have Nerf Bench: https://www.bridgebench.ai/nerf-bench They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra. This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

Comments

5 preview comments · loading full thread
nsarrazin10 hours ago

Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time. It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

rplnt12 hours ago

> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects. I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days. It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".

comboy18 hours ago

But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.

throwitaway2223 hours ago

Often times people think of "nerfs" as my first prompt (which was greenfield - no or little code existed) used 5% of my plan usage. And then 2 weeks later (as the agent is busy reading hundreds of .rs and .ts files it previously generated) the user complains the usage is going down 30% for a single prompt instead of 5%. Attributing this to a "NERF" makes little sense because it's the same model.

scrollop13 hours ago

There's also this one which has been around for a while https://marginlab.ai/trackers/claude-code/