Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yeah, he's explaining how you would create the base model, which is actually one of the more straightforward parts given that they've published their architecture (though I'm sure they've withheld a bit of their special sauce).

In reality, putting aside the millions of $$ needed to pay for the GPUs to train the model, the complexity actually lies in the training data acquisition/cleaning and the infrastructure needed to harness the 1000s of GPUs to train it in a remotely reasonable timeframe.

That being said there are a number of companies (Google, AI21, Cohere, and probably others) who have successfully created large language models like GPT3, so it's definitely not impossible when you have the resources.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: