We chose end-capability like OLMES because you'd have to measure perplexity against some fixed dataset, and in this experiment, one of the main things we're doing is changing the training dataset, so it wouldn't be an apples to apples comparison.
Hi, thanks for the reaction. Although it seems viable to create a large and varied held-out dataset that is stratified w.r.t. the training sets, do conditional generation and calculate GenPPL with a strong model. I believe this would cause a fairer comparison, it is often used in scaling papers that compare across architectures. Finally, great work anyways and I hope to see more of this, perhaps you tried this already, but based on my experience this would produce more stable measurements.
We did also use held-out loss across the various architectures too, primarily for (1) determining the compute optimal allocation point (as this metric is much more stable than the OLMES eval), and also (2) as another (correlated) measure of the compute multipliers.
At the time of doing the experiment we thought it was simply easier to use the end-eval to compare across the model recipes and data corpuses (as done in previous studies like DataDecide). But a more sophisticated construction of the hold-out dataset would be really additive too!
The metric for comparison should have been generative perplexity or an analogue, which is the common way to compare pretrains.
We chose end-capability like OLMES because you'd have to measure perplexity against some fixed dataset, and in this experiment, one of the main things we're doing is changing the training dataset, so it wouldn't be an apples to apples comparison.
Hi, thanks for the reaction. Although it seems viable to create a large and varied held-out dataset that is stratified w.r.t. the training sets, do conditional generation and calculate GenPPL with a strong model. I believe this would cause a fairer comparison, it is often used in scaling papers that compare across architectures. Finally, great work anyways and I hope to see more of this, perhaps you tried this already, but based on my experience this would produce more stable measurements.
Thanks for the suggestion!
We did also use held-out loss across the various architectures too, primarily for (1) determining the compute optimal allocation point (as this metric is much more stable than the OLMES eval), and also (2) as another (correlated) measure of the compute multipliers.
At the time of doing the experiment we thought it was simply easier to use the end-eval to compare across the model recipes and data corpuses (as done in previous studies like DataDecide). But a more sophisticated construction of the hold-out dataset would be really additive too!
It’s great to hear that you did implement and try it! And it seems like you thought it out well and did not overlook this metric, thanks for engaging.
Dwarkesh counted the groceries, declared metabolism solved, and somehow missed the digestive system. 🐾