Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's pretty much just scale, either via Dataset size or parameter size. Before GPT-4, the general SOTA model was not in fact from Open AI (Flan-PaLM from Google).

The attention from GPT-4 is a little different (probably some kind of flash attention) so that memory requirements for longer contexts are no longer quadratic. But there's nothing to suggest the intellectual gains from 4 isn't just bigger scale.

Google could have made a 4 equivalent I'm sure. It's not like there wasn't a road to take. We already knew 3 was severely undertrained even from a computer optimal perspective. And then of course, you can just train on even more tokens to get them even better.



How do you know it’s pretty much just scale? “Open”AI has been pretty tight-lipped about the details of its training and merely “claims” it was scale. It hired a ton of humans to train the model in little ways. If that’s the “scale” you’re talking about then it’s humans all the way down:

https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexa...


Also human feedback is not the only way to get a competitive model. Claude's models (Anthropic) are trained from AI feedback and they work just fine.

https://crfm.stanford.edu/helm/latest/?group=core_scenarios

Again there's nothing suggesting any big architectural changes. It's just scale.


It's just been scale up to before GPT-4. Open AI hasn't said anything specific about 4 so feel free to think there's something major going on if you want but there's basically nothing to support that.


How do we know it’s just been scale up to before GPT-4? As I said, OpenAI hasn’t told us why ChatGPT is so much better than other models (Bloom, LaMDA, LLaMa) and yet we know that they employ thousands of people to do RLHF, including for the coding models: https://www.semafor.com/article/01/27/2023/openai-has-hired-...

Doesn’t quite sound like it’s “just scale”. I asked ChatGPT about its training and corpus and it explicitly disavows having that information.


>OpenAI hasn’t told us why ChatGPT is so much better than other models

Here's the thing.... It's not. Unless you restrict the pool to models you know about.

https://crfm.stanford.edu/helm/latest/?group=core_scenarios

And Again, chatGPT was not the general SOTA LLM. https://arxiv.org/abs/2210.11416

So in the research world with detailed papers explaining what and what not, we had models that were better.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: