Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

By SAAS I assume you mean public LLMs, the problem is the hand-wringing occurring over intellectual property leaking from the company. Companies are actually writing policies banning their use.

In regards to Private LLMs, the situation has become disappointing in the 6 months.

I can only think of Mistral as being a genuine vendor.

But given the limitations in context window size, fine tuning is still necessary, and even that requires capex that I rarely see.

But my comment comes from the fact that I heard from several sources, smart people say "we tried language models at work and it failed".

However in my discussion with them, they have no concept of the size of the datacentres used by the webscalers.



It's not clear to me that fine-tuning is even capex. If you fine tune new models regularly, that's opex. If you mean literally just the GPUs, you would presumably just rent them right? (Either from cloud providers for small runs or the likes of sfcompute for large runs) Or do you imagine 24/7 training?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: