opsen
opsen

Which model is actually better

Comparing models on your own runs, with the honesty to say when there is not enough data.

On your task, not a benchmark

A task page compares models on that task's own runs: success rate, cost per run, and a lower bound that accounts for how few runs there are.

Two out of two is not a hundred percent

It is two runs. opsen does not recommend from it. A recommendation appears when there are enough runs on more than one model for the difference to mean something, and until then the page says so.

This is the part most cost dashboards get wrong. Ranking by observed success rate over small samples will confidently recommend whichever model got lucky first.

Recording an outcome

POST /v1/tasks/{task_id}/outcome

opsen knows whether a run exited zero. Only you know whether it did the job. Reporting that is what turns cost-per-run into cost-per-success, which is the number worth optimising.