Does Quantisation Hurt Low-Resource Languages More? Evidence from Urdu and Sindhi
Omar Farooq, Bilal Ahmed, Hira Yousaf
arXiv preprint
Abstract
Post-training quantisation is evaluated across languages. We find degradation is 1.8–2.6× larger for Urdu and Sindhi than for English at 4-bit precision and identify tokenizer fragmentation as the main mechanism.
Cite
Omar Farooq, Bilal Ahmed, Hira Yousaf (2026). Does Quantisation Hurt Low-Resource Languages More? Evidence from Urdu and Sindhi. arXiv preprint.