Show HN: Offlineisbetter: Efficient Models for Text on CPU
chronicallyoffl · 1 points · 1 comments · 4 uur geleden · Open original
Comments
1 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
CHchronicallyoffl4 uur geleden
hi folks,
i'm an engineer and amateur hacker who's been working on a little side project, and i want to show it off and also solicit your feedback.
i'm building efficient machine learning models for natural language processing on consumer cpus: sentiment analysis, task routing, document retrieval, etc. i know there are a lot of models for these tasks already, but too many of them are:
(1) difficult to configure, or
(2) parameter-inefficient, or
(3) too high-latency for most applications.
instead of building academically "novel" architectures, my project is making parameter-efficient and low-latency models easy to use through distillation or finetuning, computation graph optimization, and quantization. the eventual vision is a single terminal command to install the model. (i've not quite achieved this, it's like three or four commands right now).
i'm posting here because i've completed a beta version of my first model and i'd be very grateful for community feedback. the first model is called `offline-sentiment-small` and it's for binary sentiment analysis (positive/negative). you can install the runtime using `pip install offlinedemo` and download the checkpoint from the releases page of
https://github.com/offlineisbetter/offlinedemo
the readme of the above repository shows benchmarks against three other models for sentiment analysis: distilbert, roberta, and modernbert. they all do about the same on sst2 (~0.92-0.95 f1 score) because it's a relatively easy dataset. you'll notice that, though my model is 230m parameters, you need to go down to distilbert (67m) to get the same kind of latency. i want to stress-test these models on datasets with longer-context or more difficult sentiment tasks, to see how performance degrades with both my model and these models. if you know about any good datasets for this, please do let me know!
for those curious, my model is a liquid foundation model (lfm2.5) encoder (on hf, LiquidAI/LFM2.5-Encoder-230M) finetuned by low-rank adaptation on the sst2 training set, optimized with onnx, and quantized to int8. i chose lfm over bert as the base model because it uses grouped-query attention and convolution layers instead of dense attention. these changes over the "vanilla" transformer architecture make its latency subquadratic in the input context (and i've verified this empirically). lfms are often considered more parameter-efficient than transformers: for example, liquid ai's 2.6b autoregressive model competes with llama 7b.
and disclaimer, i have no affiliation to liquid, i just think their work is cool!
i want to emphasize that i'm not claiming technical novelty! rather, i'm building a way to make parameter-efficient, distilled/finetuned, and quantized models more accessible to more people. right now, if you wanted to run a distilled and quantized text model, it takes a non-negligible amount of compute and time and effort to set it up. not everybody wants to do that, and so out of convenience they'll just opt for a big llm in the cloud.
i worry that so many people these days pay for openai/anthropic tokens just to do a simple task, such as sort their emails into categories. not only is this wasteful, high-latency, bad for the environment, etc. but you're giving somebody else your data. so i've started the offlineisbetter project to help raise awareness to wasteful model use and provide the community with an alternative.
please let me know if you have issues testing the model on your machine. i'd be very grateful to hear any feedback you might have!
Comments
1 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
hi folks, i'm an engineer and amateur hacker who's been working on a little side project, and i want to show it off and also solicit your feedback. i'm building efficient machine learning models for natural language processing on consumer cpus: sentiment analysis, task routing, document retrieval, etc. i know there are a lot of models for these tasks already, but too many of them are: (1) difficult to configure, or (2) parameter-inefficient, or (3) too high-latency for most applications. instead of building academically "novel" architectures, my project is making parameter-efficient and low-latency models easy to use through distillation or finetuning, computation graph optimization, and quantization. the eventual vision is a single terminal command to install the model. (i've not quite achieved this, it's like three or four commands right now). i'm posting here because i've completed a beta version of my first model and i'd be very grateful for community feedback. the first model is called `offline-sentiment-small` and it's for binary sentiment analysis (positive/negative). you can install the runtime using `pip install offlinedemo` and download the checkpoint from the releases page of https://github.com/offlineisbetter/offlinedemo the readme of the above repository shows benchmarks against three other models for sentiment analysis: distilbert, roberta, and modernbert. they all do about the same on sst2 (~0.92-0.95 f1 score) because it's a relatively easy dataset. you'll notice that, though my model is 230m parameters, you need to go down to distilbert (67m) to get the same kind of latency. i want to stress-test these models on datasets with longer-context or more difficult sentiment tasks, to see how performance degrades with both my model and these models. if you know about any good datasets for this, please do let me know! for those curious, my model is a liquid foundation model (lfm2.5) encoder (on hf, LiquidAI/LFM2.5-Encoder-230M) finetuned by low-rank adaptation on the sst2 training set, optimized with onnx, and quantized to int8. i chose lfm over bert as the base model because it uses grouped-query attention and convolution layers instead of dense attention. these changes over the "vanilla" transformer architecture make its latency subquadratic in the input context (and i've verified this empirically). lfms are often considered more parameter-efficient than transformers: for example, liquid ai's 2.6b autoregressive model competes with llama 7b. and disclaimer, i have no affiliation to liquid, i just think their work is cool! i want to emphasize that i'm not claiming technical novelty! rather, i'm building a way to make parameter-efficient, distilled/finetuned, and quantized models more accessible to more people. right now, if you wanted to run a distilled and quantized text model, it takes a non-negligible amount of compute and time and effort to set it up. not everybody wants to do that, and so out of convenience they'll just opt for a big llm in the cloud. i worry that so many people these days pay for openai/anthropic tokens just to do a simple task, such as sort their emails into categories. not only is this wasteful, high-latency, bad for the environment, etc. but you're giving somebody else your data. so i've started the offlineisbetter project to help raise awareness to wasteful model use and provide the community with an alternative. please let me know if you have issues testing the model on your machine. i'd be very grateful to hear any feedback you might have!