How pass@k changes with sampling temperature: careful sampling wins at one attempt, adventurous sampling wins at many.
Two different questions, one benchmark
Code has a property prose does not: you can run it. So grading is honest — but the model does not answer the same way twice, and “did it work?” turns out to hide two separate questions.
chance at least one attempt passes the tests
pass@1 asks “can I trust it once?”. pass@k asks “can I find a working answer at all?”. They are different jobs, they want opposite settings, and reporting one number for both hides that.