Language models for text classification: From bag-of-words to Jev
Anon84 · 195 points · 10 comments · yesterday · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
MAmalshe4 hours ago
Sebastian is the author of two excellent books related to LLMs
Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/
Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/
TOTopfi9 hours ago
Solid assessment, very much what I assumed after announcement.
The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.
Hype really is the worst aspect of this industry.
AEaesthesia1 hour ago
I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
NZnzoschke17 hours ago
Great article. This matches my feelings:
> Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.
We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email
KEkevinwang4 hours ago
Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
Sebastian is the author of two excellent books related to LLMs Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/ Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/
Solid assessment, very much what I assumed after announcement. The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state. Hype really is the worst aspect of this industry.
I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
Great article. This matches my feelings: > Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task. We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email
Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.