Measuring Reasoning Without Leaking the Test: Procedural Generation for LLM Evaluation

Bilal Ahmed, Omar Farooq, Hamid Raza

ICLR 2026

Abstract

Static reasoning benchmarks decay as their contents enter training corpora. We propose procedurally generated evaluation families with controllable difficulty and show that model rankings remain stable across generations while absolute scores reveal contamination in three widely used benchmarks.

Read the full text

Cite

Bilal Ahmed, Omar Farooq, Hamid Raza (2026). Measuring Reasoning Without Leaking the Test: Procedural Generation for LLM Evaluation. ICLR 2026.