Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
matt_d · 68 points · 17 comments · 7 uur geleden · Open original
Comments
5 preview comments · loading full thread
Log in to use comments
Log in to h4cker, then connect Hacker News to publish comments.
BOboorang3 uur geleden
I got pretty good mileage out of context engineering, adding my personal coding heuristics to my AGENTS.md and referencing subdocuments on a "when doing X, consult Y" pattern. I assume others are doing similar things, but I was pretty surprised when I was able to get it to generate code that is pretty close to what I would do if I was doing it manually. I'm curious if scientists and mathematicians are doing things like that. "When I see X, I typically immediately check Y" or whatever their domain heuristics look like.
JOjohnnyApplePRNG4 uur geleden
Not surprised to see Claude significantly higher in scientific intelligence than Sol.
You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it.
That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length.
Don't get me wrong, codex is great at finding bugs and building games. It's great.
AKakshay_akula6 uur geleden
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
JEjerpint4 uur geleden
The fact that opus 5 is outperforming fable is odd to me
From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
RErespectattentio3 uur geleden
Like you were reading my mind. I was waiting for such benchmark to land. This will improve models for such scientific research workflows.
AI should have started with science from the beginning, not after 4 years.
I am building on top of it with agents to improve scientific workflows.
Comments
5 preview comments · loading full threadLog in to h4cker, then connect Hacker News to publish comments.
I got pretty good mileage out of context engineering, adding my personal coding heuristics to my AGENTS.md and referencing subdocuments on a "when doing X, consult Y" pattern. I assume others are doing similar things, but I was pretty surprised when I was able to get it to generate code that is pretty close to what I would do if I was doing it manually. I'm curious if scientists and mathematicians are doing things like that. "When I see X, I typically immediately check Y" or whatever their domain heuristics look like.
Not surprised to see Claude significantly higher in scientific intelligence than Sol. You can tell that Claude really does grasp a wide array of highly specific scientific and mathematical nuances... where's codex is just basically for coding and that's it. That's the feel I get from the both of them anyways and I've used both on the 20x plan for the past week at length. Don't get me wrong, codex is great at finding bugs and building games. It's great.
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
The fact that opus 5 is outperforming fable is odd to me From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Like you were reading my mind. I was waiting for such benchmark to land. This will improve models for such scientific research workflows. AI should have started with science from the beginning, not after 4 years. I am building on top of it with agents to improve scientific workflows.