Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I tested a consortium of qwens on the brainfuck test and it solved it, while the single models fail.

MOEs are a single model. An 'expert' is a subset of layers chosen by a router model for each token. This makes them run faster. A consortium is a type of parallel reasoning that uses multiple of the same or different models to generate parallel response and find the best one.

All models have a jagged frontier with weird skill gaps. A consortium can bridge those gaps and increase performance on the frontier.



Has anyone compared a consortium of leading edge 3B-20B models compared to the most powerful models?

I'd love to see how they performed.


Do you have a favourite benchmark? I may just have the budget for testing some 3b models




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: