Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
tgluck · 63 points · 15 comments · yesterday · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
RIricardobeat20 hours ago
This will only work for simple text classification tasks, which is the least interesting possible use of Jev.
UNunknownyesterday
[deleted]
GIgingersnapyesterday
Is the local model similar to model2vec?
UNunknownyesterday
[deleted]
TGtgluckyesterday
Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev's answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn't have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains.
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
This will only work for simple text classification tasks, which is the least interesting possible use of Jev.
[deleted]
Is the local model similar to model2vec?
[deleted]
Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev's answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn't have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains. Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.