Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

So if this is true - which is a big if since this looks like speculation rather than real information - could this work with even smaller models?

For example, what about 20 x 65B = 1.3T params? Or 100 x 13B = 1.3T params?

Hell, what about 5000 x 13B params? Thousands of small highly specialized models, with maybe one small "categorization" model as the first pass?



Well at the end of the day you’ll also need a model for ranking the candidates, which becomes harder as the number of candidates grows. And the mean quality of any one candidate response will drop as the model size decreases, as will the max quality.





Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: