Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.bugs.dist > #1231921 > unrolled thread
| Started by | Petter Reinholdtsen <pere@hungry.com> |
|---|---|
| First post | 2025-02-05 22:00 +0100 |
| Last post | 2025-02-06 09:30 +0100 |
| Articles | 6 — 2 participants |
Back to article view | Back to linux.debian.bugs.dist
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-05 22:00 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 01:40 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 02:50 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:00 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 15:20 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:30 +0100
| From | Petter Reinholdtsen <pere@hungry.com> |
|---|---|
| Date | 2025-02-05 22:00 +0100 |
| Subject | Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ |
| Message-ID | <KcUF3-eFLJ-1@gated-at.bofh.it> |
Where can I find the draft packaging for llama.cpp now? Is there a public git repo somewhere? -- Happy hacking Petter Reinholdtsen
[toc] | [next] | [standalone]
| From | Petter Reinholdtsen <pere@hungry.com> |
|---|---|
| Date | 2025-02-06 01:40 +0100 |
| Message-ID | <KcY5X-eHYo-3@gated-at.bofh.it> |
| In reply to | #1231921 |
[Christian Kastner] > Repo is here [1]. Very good. Built just fine here. I checked in a few minor fixes. I noticed llama.cpp depend on llama.cpp-backend with no concrete dependency first. This lead to unpredictable behaviour, and I suggest depending on for example 'llama.cpp-cpu | llama.cpp-backend' to make sure 'apt install llama.cpp' behave predictably. I was sad to discover the server example is missing, as it is the llama.cpp progam I use the most. Without it, I will have to continue using my own build. > I thought it best to upload now and fix the remaining issues above while > the package sits in NEW. Very good. I hope to get whisper.cpp to the same state, so it can have a fighting chance to get into testing before the freeze. -- Happy hacking Petter Reinholdtsen
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-02-06 02:50 +0100 |
| Message-ID | <KcZbI-eIzF-3@gated-at.bofh.it> |
| In reply to | #1231943 |
On Thu, 2025-02-06 at 01:33 +0100, Petter Reinholdtsen wrote:
>
> I was sad to discover the server example is missing, as it is the
> llama.cpp progam I use the most. Without it, I will have to continue
> using my own build.
I second this. llama-server is also the service endpoint for DebGPT.
I pushed a fix for ppc64el. The hwcaps works correctly for power9, given the baseline is power 8.
(chroot:unstable-ppc64el-sbuild) root@debian-project-1 /h/d/llama.cpp.pkg [1]# ldd (which llama-cli)
linux-vdso64.so.1 (0x00007fffa4810000)
libeatmydata.so => /lib/powerpc64le-linux-gnu/libeatmydata.so (0x00007fffa4600000)
libllama.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libllama.so (0x00007fffa4450000)
libggml.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml.so (0x00007fffa4420000)
libggml-base.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml-base.so (0x00007fffa4330000)
libstdc++.so.6 => /lib/powerpc64le-linux-gnu/libstdc++.so.6 (0x00007fffa3fd0000)
libm.so.6 => /lib/powerpc64le-linux-gnu/libm.so.6 (0x00007fffa3e90000)
libgcc_s.so.1 => /lib/powerpc64le-linux-gnu/libgcc_s.so.1 (0x00007fffa3e50000)
libc.so.6 => /lib/powerpc64le-linux-gnu/libc.so.6 (0x00007fffa3be0000)
/lib64/ld64.so.2 (0x00007fffa4820000)
libggml-cpu.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml-cpu.so (0x00007fffa3b20000)
libgomp.so.1 => /lib/powerpc64le-linux-gnu/libgomp.so.1 (0x00007fffa3a90000)
[toc] | [prev] | [next] | [standalone]
| From | Petter Reinholdtsen <pere@hungry.com> |
|---|---|
| Date | 2025-02-06 09:00 +0100 |
| Message-ID | <Kd4XL-eM7v-9@gated-at.bofh.it> |
| In reply to | #1231945 |
[M. Zhou] > I second this. llama-server is also the service endpoint for DebGPT. So, what exactly need to happen for llama-server to be included in the package? I found this in d/copyright: DFSG compliance --------------- The server example contains a number of minified and generated files in the frontend. These seem to be essential to the example, so the server example has been removed entirely, for now. I guess some build mechanics need to be included to build the minified and generated files, but do not know which one are the problem. According to examples/server/README.md the "Web UI" is buitl using 'npm run build', so I guess some nodejs dependencies are needed. Sadly, I do not know how to convince npm to not download random stuff from the Internet. -- Happy hacking Petter Reinholdtsen
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-02-06 15:20 +0100 |
| Message-ID | <KdaTv-eQwc-3@gated-at.bofh.it> |
| In reply to | #1231945 |
On Thu, 2025-02-06 at 09:13 +0100, Christian Kastner wrote: > > I meant to ask anyway: performance-wise, is it comparable to your local > build? I mean, I wouldn't know what in the code would alter this, but I > built and tested this on platti.d.o and performance was poor, so another > data point would be useful. For ppc64el, the llama.cpp-blas backend is way slower than the -cpu backend. I did not test on amd64. But on ppc64el the package does not feel different than local build. CPU is slow anyway. How does HIP performs? phi-4-q4.gguf | power9, cpu (8-threads) | 0.62 tokens/s phi-4-q4.gguf | amd64, 13900H | 6.7 tokens/s GPU is way faster than this. The phi-4 model does not fit in my nvidia GPU. No number for GPU this time.
[toc] | [prev] | [next] | [standalone]
| From | Petter Reinholdtsen <pere@hungry.com> |
|---|---|
| Date | 2025-02-06 09:30 +0100 |
| Message-ID | <Kd5qN-eMwX-9@gated-at.bofh.it> |
| In reply to | #1231943 |
[Christian Kastner] > Look fine, though I deliberately skipped the poetry dependency for now > as it looked more like a false positive. Aha. I just trusted lintian-brush on this one, did not investigate. > This was my intention, but I initially wasn't sure what the default > would be (-cpu or -blas). Looks like I forgot to add one before > upload. Given that every machine it can be installed on got a CPU, but not all of them got a supported GPU, I beieve -cpu is the most sensible default. > It'll be re-enabled soon. The were a few generated and minified files > in that example, so I just opted to skip those for now, and focus on > the build process. Great to hear. :) > Seeing as how closely llama.cpp and whisper.cpp are related, in the > ideal case, you should be able to just carry over some patches, and > mostly just copy d/rules, as llama.cpp and whisper.cpp share the ggml > library on a source basis. I hope so too, but I guess we will soon find out. My initial draft on <URL: https://salsa.debian.org/deeplearning-team/whisper.cpp > will need a lot of updates to bring it in line with this new approach. :) -- Happy hacking Petter Reinholdtsen
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.bugs.dist
csiph-web