news.ycombinator.com

Untitled

jug · 0 points · 0 comments · 20 時間前

We also have Nerf Bench: https://www.bridgebench.ai/nerf-bench They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra. This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

Comments

5 preview comments · loading full thread
nsarrazin12 時間前

Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time. It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

rplnt13 時間前

> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects. I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days. It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".

comboy19 時間前

But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.

scrollop14 時間前

There's also this one which has been around for a while https://marginlab.ai/trackers/claude-code/

rednb13 時間前

I'd take this kind of benchmark with a grain of salt. At this point, I have a set of comprehensive guidelines covering both backend and frontend work, and for the frontend we go as far as explaining what we a good design is in our visual system, and even how to conduct a visual review when screenshots are handed to the model. Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied. My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.