How pass@k changes with sampling temperature: careful sampling wins at one attempt, adventurous sampling wins at many.

Two different questions, one benchmark

Code has a property prose does not: you can run it. So grading is honest — but the model does not answer the same way twice, and “did it work?” turns out to hide two separate questions.

attempts you are willing to pay for1

chance at least one attempt passes the tests

pass@1 asks “can I trust it once?”. pass@k asks “can I find a working answer at all?”. They are different jobs, they want opposite settings, and reporting one number for both hides that.