6 Comments
User's avatar
GertLek's avatar

The metric for comparison should have been generative perplexity or an analogue, which is the common way to compare pretrains.

Dwarkesh Patel's avatar

We chose end-capability like OLMES because you'd have to measure perplexity against some fixed dataset, and in this experiment, one of the main things we're doing is changing the training dataset, so it wouldn't be an apples to apples comparison.

GertLek's avatar
1hEdited

Hi, thanks for the reaction. Although it seems viable to create a large and varied held-out dataset that is stratified w.r.t. the training sets, do conditional generation and calculate GenPPL with a strong model. I believe this would cause a fairer comparison, it is often used in scaling papers that compare across architectures. Finally, great work anyways and I hope to see more of this, perhaps you tried this already, but based on my experience this would produce more stable measurements.

Jerry Han's avatar

Thanks for the suggestion!

We did also use held-out loss across the various architectures too, primarily for (1) determining the compute optimal allocation point (as this metric is much more stable than the OLMES eval), and also (2) as another (correlated) measure of the compute multipliers.

At the time of doing the experiment we thought it was simply easier to use the end-eval to compare across the model recipes and data corpuses (as done in previous studies like DataDecide). But a more sophisticated construction of the hold-out dataset would be really additive too!

GertLek's avatar

It’s great to hear that you did implement and try it! And it seems like you thought it out well and did not overlook this metric, thanks for engaging.

cyberkittens's avatar

Dwarkesh counted the groceries, declared metabolism solved, and somehow missed the digestive system. 🐾