The Linux kernel does not break userspace. > What's wrong with using an older we...

zozbot234 · 2025-08-12T06:43:55 1754981035

> Yeah, they tried this, this was the old setup as I understand it. But every time they needed support for a new model and had to update llama.cpp, an old model would break and one of their partners would go ape on them.

Shouldn't any such regressions be regarded as bugs in llama.cpp and fixed there? Surely the Ollama folks can test and benchmark the main models that people care about before shipping the update in a stable release. That would be a lot easier than trying to reimplement major parts of llama.cpp from scratch.

tarruda · 2025-08-11T23:11:12 1754953872

> every time they needed support for a new model and had to update llama.cpp, an old model would break and one of their partners would go ape on them. They said it happened more than once, but one particular case (wish I could remember what it was) was so bad they felt they had no choice but to reimplement. It's the lowest risk strategy.

A much lower risk strategy would be using multiple versions of llama-server to keep supporting old models that would break on newer llama.cpp versions.

MarkSweep · 2025-08-11T23:30:38 1754955038

The Ollama distribution size is already pretty big (at least on Windows) due to all the GPU support libraries and whatnot. Having to multiple that by the number of llama.cpp versions supported would not be great.

jychang · 2025-08-12T01:10:12 1754961012

?

    llamacpp> ls -l \*llama\*
    -rwxr-xr-x 1 root root 2505480 Aug  7 05:06 libllama.so
    -rwxr-xr-x 1 root root 5092024 Aug  7 05:23 llama-server

That's a terrible excuse, Llama.cpp is just 7.5 megabytes. You can easily ship a couple copies of that. The current ollama for windows download is 700MB.

I don't buy it. They're not willing to make an 700MB download a few megabytes bigger to ~730MB, but they are willing to support a fork/rewrite indefinitely (and the fork is outside of their core competency, as seen by the current issues)? What kind of decisionmaking is that?

MarkSweep · 2025-08-23T17:20:04 1755969604

Sorry, I forgot to include in my comment this part:

If you include multiple version of llama, and each of those llama version depends on different GPU libraries, that could balloon the download size.

If these GPU libraries change rarely, then yes, you are correct, it might not be a problem.

jychang · 2025-08-24T11:29:29 1756034969

Well llama.cpp requires minimum CUDA 11 from 2020, or if you need CUDA C++17 support, then CUDA 12 from 2022.

It'll compile fine against latest CUDA 12.8 or 12.9 etc, but there's zero need to pack whatever the latest CUDA version is.

vlovich123 · 2025-08-12T04:54:06 1754974446

It’s 700mib because they’re likely redistributing the CUDA libraries so that users don’t need to separately run that installer. Llama.cpp is a bit more “you are expected to know what you’re doing” on that front. But yeah, you could plausibly ship multiple versions of the inference engine although from a maintenance perspective that sounds like hell for any number of reasons