news.ycombinator.com

Untitled

jug · 0 points · 0 comments · vor 19 Stunden

We also have Nerf Bench: https://www.bridgebench.ai/nerf-bench They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra. This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

Comments

5 preview comments · loading full thread
nsarrazinvor 11 Stunden

Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time. It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.

rplntvor 13 Stunden

> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects. I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days. It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".

comboyvor 18 Stunden

But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.

scrollopvor 13 Stunden

There's also this one which has been around for a while https://marginlab.ai/trackers/claude-code/

rednbvor 12 Stunden

I'd take this kind of benchmark with a grain of salt. At this point, I have a set of comprehensive guidelines covering both backend and frontend work, and for the frontend we go as far as explaining what we a good design is in our visual system, and even how to conduct a visual review when screenshots are handed to the model. Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied. My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.