news.ycombinator.com

Untitled

luka_a · 0 points · 0 comments · 4小时前

I ran a statistical analysis of my Claude Code API request data. In the linked blog post, I show that prompt caching leads to input token cost savings of around 85 percent, even though I'm not running fully autonomous sessions often. I show how cost relates to the cache miss rate, and I model the request dynamics with two timescales, one corresponding to fast processes such as agentic tool calls, and the other to slow processes such as reviewing the LLM output. These two timescales also appear in the data. The cache misses come primarily from the slow timescale. Running these measurements on your data can help you determine how cost-efficient your usage pattern is and what drives your cost. You can use this to adjust your work pattern (or configure the system you're building) by, for example, setting a non-default cache expiry time, or consciously trimming the mean time of your slow timescale.

评论

0 条预览评论 · 正在加载完整讨论