Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

llama.cpp (or rather G.Gerganov et. al.) are trying to avoid cuBLAS entirely, using ins own kernels. not sure how jart's effort relates, and whether jart intends to upstream these into llama.cpp which seems to still be the underlying tech behind the llamafile.


Here are links to the most recent pull requests sent

    https://github.com/ggerganov/llama.cpp/pull/6414
    https://github.com/ggerganov/llama.cpp/pull/6412


This doesn't relate to GPU kernels unfortunately.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: