Which model is actually better
Comparing models on your own runs, with the honesty to say when there is not enough data.
On your task, not a benchmark
A task page compares models on that task's own runs: success rate, cost per run, and a lower bound that accounts for how few runs there are.
Two out of two is not a hundred percent
It is two runs. opsen does not recommend from it. A recommendation appears when there are enough runs on more than one model for the difference to mean something, and until then the page says so.
This is the part most cost dashboards get wrong.
Ranking by observed success rate over small samples will confidently
recommend whichever model got lucky first.
Recording an outcome
POST /v1/tasks/{task_id}/outcomeopsen knows whether a run exited zero. Only you know whether it did the job. Reporting that is what turns cost-per-run into cost-per-success, which is the number worth optimising.