Language models for text classification: From bag-of-words to Jev
Anon84 · 193 points · 10 comments · เมื่อวาน · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
MAmalshe3 ชั่วโมงที่ผ่านมา
Sebastian is the author of two excellent books related to LLMs
Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/
Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/
AEaesthesia43 นาทีที่ผ่านมา
I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
TOTopfi8 ชั่วโมงที่ผ่านมา
Solid assessment, very much what I assumed after announcement.
The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.
Hype really is the worst aspect of this industry.
NZnzoschke16 ชั่วโมงที่ผ่านมา
Great article. This matches my feelings:
> Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.
We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email
KEkevinwang4 ชั่วโมงที่ผ่านมา
Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
Sebastian is the author of two excellent books related to LLMs Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/ Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/
I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
Solid assessment, very much what I assumed after announcement. The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state. Hype really is the worst aspect of this industry.
Great article. This matches my feelings: > Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task. We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email
Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.