PSSA: A non-transformer language model written from scratch in Rust
sparticle62 · 84 points · 37 comments · hace 16 horas · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
JPJPLeRouzichace 11 horas
I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).
For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:
[0] https://aclanthology.org/2023.findings-emnlp.936/
https://arxiv.org/abs/2503.00392
https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...
PRprospero_hace 12 horas
Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
MLmllev15hace 15 horas
You can just tell when the idea itself was generated
JAjanalsncmhace 14 horas
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.
A1a19486hace 14 horas
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets). For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token: [0] https://aclanthology.org/2023.findings-emnlp.936/ https://arxiv.org/abs/2503.00392 https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...
Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
You can just tell when the idea itself was generated
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation. You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run. Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it? In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.