Measuring Reasoning Without Leaking the Test: Procedural Generation for LLM Evaluation
Bilal Ahmed, Omar Farooq, Hamid Raza
ICLR 2026
Abstract
Static reasoning benchmarks decay as their contents enter training corpora. We propose procedurally generated evaluation families with controllable difficulty and show that model rankings remain stable across generations while absolute scores reveal contamination in three widely used benchmarks.
Cite
Bilal Ahmed, Omar Farooq, Hamid Raza (2026). Measuring Reasoning Without Leaking the Test: Procedural Generation for LLM Evaluation. ICLR 2026.