On the use of years 2016-2024 in the Danish citizenship test #1181
Replies: 2 comments 6 replies
-
|
Previously it was from 2016 to 2023, and we added 2024 to ensure that the test set keeps changing, hopefully reducing overfitting. The newer years are used in the test split. We haven't conducted any year-by-year performance, nope! How come? |
Beta Was this translation helpful? Give feedback.
-
Aha! I see. Yeah, that was definitely a misunderstanding then 🙂
You mean sample-level information? I'm not sure what they would get out of it though, as for most of the benchmarks they would just see the samples and the model's structured output. It's only on reading comprehension and summarisation that the models get to output (almost) what they want. I can see the use of it to detect errors in the evaluation, but that's exactly what the
I agree with you here, but only on tasks where the models have free reign to output what they want, as otherwise you won't really get much useful information. And again, this is accessible to anyone using the |
Beta Was this translation helpful? Give feedback.
Uh oh!
There was an error while loading. Please reload this page.
-
We use all the years in evaluating models. I have a few questions on this:
Beta Was this translation helpful? Give feedback.
All reactions