One thing that evals are super important from the get go are where the harness+model inference is part of the product, e.g. if you are doing voice ai, building out a test harness to test the system is a non trivial first step.
COconsumer4516小时前
Man, I could have used something like this last year.
Upon review, it appears that fumbling my way through the dark with Claude and Google led to the same place, in nearly all cases.
However, this is all written by Claude — it has too many em-dashes to not be, does it not? So, maybe that's why we ended up in the same places.
Does anyone know of any other resources in this vein?
KOKolibriFly3小时前
Glad to hear they are prioritizing evaluation right from the start. Usually people just throw together a rag pipeline on the knee and then judge the metrics by eye, skimming three responses in the terminal
NYnycdatasci4小时前
Why does applied AI intentionally exclude a framework/harness around AI? The job is to harness the power of AI, and a harness is a critical part of that.
评论
4 条预览评论 · 正在加载完整讨论请先登录 h4cker 账号,然后连接 Hacker News 后发表评论。
One thing that evals are super important from the get go are where the harness+model inference is part of the product, e.g. if you are doing voice ai, building out a test harness to test the system is a non trivial first step.
Man, I could have used something like this last year. Upon review, it appears that fumbling my way through the dark with Claude and Google led to the same place, in nearly all cases. However, this is all written by Claude — it has too many em-dashes to not be, does it not? So, maybe that's why we ended up in the same places. Does anyone know of any other resources in this vein?
Glad to hear they are prioritizing evaluation right from the start. Usually people just throw together a rag pipeline on the knee and then judge the metrics by eye, skimming three responses in the terminal
Why does applied AI intentionally exclude a framework/harness around AI? The job is to harness the power of AI, and a harness is a critical part of that.