AI Engineer Notebooks – free, framework-free RAG/agents/evals on Colab
calmrocks · 92 points · 9 comments · 9시간 전 · Open original
Comments
4 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
GHghm21991시간 전
One thing that evals are super important from the get go are where the harness+model inference is part of the product, e.g. if you are doing voice ai, building out a test harness to test the system is a non trivial first step.
COconsumer4516시간 전
Man, I could have used something like this last year.
Upon review, it appears that fumbling my way through the dark with Claude and Google led to the same place, in nearly all cases.
However, this is all written by Claude — it has too many em-dashes to not be, does it not? So, maybe that's why we ended up in the same places.
Does anyone know of any other resources in this vein?
KOKolibriFly3시간 전
Glad to hear they are prioritizing evaluation right from the start. Usually people just throw together a rag pipeline on the knee and then judge the metrics by eye, skimming three responses in the terminal
NYnycdatasci4시간 전
Why does applied AI intentionally exclude a framework/harness around AI? The job is to harness the power of AI, and a harness is a critical part of that.
Comments
4 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
One thing that evals are super important from the get go are where the harness+model inference is part of the product, e.g. if you are doing voice ai, building out a test harness to test the system is a non trivial first step.
Man, I could have used something like this last year. Upon review, it appears that fumbling my way through the dark with Claude and Google led to the same place, in nearly all cases. However, this is all written by Claude — it has too many em-dashes to not be, does it not? So, maybe that's why we ended up in the same places. Does anyone know of any other resources in this vein?
Glad to hear they are prioritizing evaluation right from the start. Usually people just throw together a rag pipeline on the knee and then judge the metrics by eye, skimming three responses in the terminal
Why does applied AI intentionally exclude a framework/harness around AI? The job is to harness the power of AI, and a harness is a critical part of that.