nicowaltz · 239 points · 93 comments · hier · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
SHsharihhier
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
TNTN1ckhier
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
[1] https://tn1ck.com/blog/jevdit
THthmhier
Ask Jeeves - Only took us 30 years to come full circle.
ITitzikkatzhier
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long. It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine. [1] https://tn1ck.com/blog/jevdit
Ask Jeeves - Only took us 30 years to come full circle.
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
This reminds me of "on-premise cloud".