Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

For DPO I only had around 700 preference examples, so not much data at all. That took about 12 minutes to train on a single GPU.

Pretraining was obviously a a lot slower, the 125M model took roughly half a day.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: