UrduQA: why we built a benchmark that models cannot have memorised
Static benchmarks decay the moment they enter training data. Here is how we built one that stays honest, and what it tells us about the state of Urdu language technology.
Journal
Research findings, what admissions is planning, and honest perspectives on where the field is going.
The Letter
One considered note a month. No noise.
Static benchmarks decay the moment they enter training data. Here is how we built one that stays honest, and what it tells us about the state of Urdu language technology.
Forty places, two scholarships for women in engineering, and a new capstone partnership with the provincial health department.
What we learned deploying crop-disease detection on sub-$50 hardware across forty-one farms in Punjab — including everything that broke.
We reviewed the documentation for twelve models in use at financial institutions in the region. Most would not pass a basic procurement check.
Highlights from the institute’s annual forum, held this year with the central-bank working group and three provincial departments.
A note from the Dean on the pedagogy behind the academy — and why we grade validation plans more heavily than model accuracy.