undefined | Better HN

0 pointsf38zf5vdt1y ago0 comments

1. Ignore the benchmarks. I've been A/Bing 11B today with Molmo 72B [1], which itself has an ELO neck-and-neck with GPT4o, and it's even. Because everyone in open source tends to train on validation benchmarks, you really can not trust them.

2. The method of tokenization/adapter is novel and uses many fewer tokens than all comparable CLIP/SigLIP-adapter models, making it _much_ faster. Attention is O(n^2) on memory/compute per sequence length.

[1] https://molmo.allenai.org/blog

0 comments

espadrine1y ago

> I've been A/Bing 11B today with Molmo 72B

How are you testing Molmo 72B? If you are interacting with https://molmo.allenai.org/, they are using Molmo-7B-D.

benreesman1y ago

It’s not just open source that trains on the validation set. The big labs have already forgotten more about gaming MMLU down to the decimal than the open source community ever knew. Every once in a while they get sloppy and Claude does a faux pas with a BIGBENCH canary string or some other embarrassing little admission of dishonesty like that.

A big lab gets exactly the score on any public eval that they want to. They have their own holdouts for actual ML work, and they’re some of the most closely guarded IP artifacts, far more valuable than a snapshot of weights.

sumedh1y ago

I tried some OCR use cases, Claude Sonnet just blows Molmo.

knicholes1y ago

When you say "blows," do you mean in a subservient sense or more like, "it blows it out of the water?"

grahamj1y ago

yeah does it suck or does it suck?

GaggiX1y ago

How about its performance compare to Qwen-2-72B tho?

f38zf5vdtOP1y ago

Refer to the blog post I linked. Molmo is ahead of Qwen2 72b.

j / k navigate · click thread line to collapse