Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.bugs.dist > #1231921 > unrolled thread

Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++

Started byPetter Reinholdtsen <pere@hungry.com>
First post2025-02-05 22:00 +0100
Last post2025-02-06 09:30 +0100
Articles 6 — 2 participants

Back to article view | Back to linux.debian.bugs.dist

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-05 22:00 +0100
    Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 01:40 +0100
      Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 02:50 +0100
        Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:00 +0100
        Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 15:20 +0100
      Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:30 +0100

#1231921 — Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++

FromPetter Reinholdtsen <pere@hungry.com>
Date2025-02-05 22:00 +0100
SubjectBug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++
Message-ID<KcUF3-eFLJ-1@gated-at.bofh.it>
Where can I find the draft packaging for llama.cpp now?  Is there a
public git repo somewhere?

-- 
Happy hacking
Petter Reinholdtsen

[toc] | [next] | [standalone]


#1231943

FromPetter Reinholdtsen <pere@hungry.com>
Date2025-02-06 01:40 +0100
Message-ID<KcY5X-eHYo-3@gated-at.bofh.it>
In reply to#1231921
[Christian Kastner]
> Repo is here [1].

Very good.  Built just fine here.  I checked in a few minor fixes.

I noticed llama.cpp depend on llama.cpp-backend with no concrete
dependency first.  This lead to unpredictable behaviour, and I suggest
depending on for example 'llama.cpp-cpu | llama.cpp-backend' to make
sure 'apt install llama.cpp' behave predictably.

I was sad to discover the server example is missing, as it is the
llama.cpp progam I use the most.  Without it, I will have to continue
using my own build.

> I thought it best to upload now and fix the remaining issues above while
> the package sits in NEW.

Very good.

I hope to get whisper.cpp to the same state, so it can have a fighting
chance to get into testing before the freeze.

-- 
Happy hacking
Petter Reinholdtsen

[toc] | [prev] | [next] | [standalone]


#1231945

From"M. Zhou" <lumin@debian.org>
Date2025-02-06 02:50 +0100
Message-ID<KcZbI-eIzF-3@gated-at.bofh.it>
In reply to#1231943
On Thu, 2025-02-06 at 01:33 +0100, Petter Reinholdtsen wrote:
> 
> I was sad to discover the server example is missing, as it is the
> llama.cpp progam I use the most.  Without it, I will have to continue
> using my own build.

I second this. llama-server is also the service endpoint for DebGPT.

I pushed a fix for ppc64el. The hwcaps works correctly for power9, given the baseline is power 8.

(chroot:unstable-ppc64el-sbuild) root@debian-project-1 /h/d/llama.cpp.pkg [1]# ldd (which llama-cli)                                                                                               
        linux-vdso64.so.1 (0x00007fffa4810000)                                                                                                                                                     
        libeatmydata.so => /lib/powerpc64le-linux-gnu/libeatmydata.so (0x00007fffa4600000)                                                                                                         
        libllama.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libllama.so (0x00007fffa4450000)                                                                               
        libggml.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml.so (0x00007fffa4420000)                                                                                 
        libggml-base.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml-base.so (0x00007fffa4330000)                                                                       
        libstdc++.so.6 => /lib/powerpc64le-linux-gnu/libstdc++.so.6 (0x00007fffa3fd0000)                                                                                                           
        libm.so.6 => /lib/powerpc64le-linux-gnu/libm.so.6 (0x00007fffa3e90000)
        libgcc_s.so.1 => /lib/powerpc64le-linux-gnu/libgcc_s.so.1 (0x00007fffa3e50000)
        libc.so.6 => /lib/powerpc64le-linux-gnu/libc.so.6 (0x00007fffa3be0000)
        /lib64/ld64.so.2 (0x00007fffa4820000)
        libggml-cpu.so => /usr/lib/powerpc64le-linux-gnu/llama.cpp/glibc-hwcaps/power9/libggml-cpu.so (0x00007fffa3b20000)
        libgomp.so.1 => /lib/powerpc64le-linux-gnu/libgomp.so.1 (0x00007fffa3a90000)

[toc] | [prev] | [next] | [standalone]


#1231961

FromPetter Reinholdtsen <pere@hungry.com>
Date2025-02-06 09:00 +0100
Message-ID<Kd4XL-eM7v-9@gated-at.bofh.it>
In reply to#1231945
[M. Zhou]
> I second this. llama-server is also the service endpoint for DebGPT.

So, what exactly need to happen for llama-server to be included in the
package?

I found this in d/copyright:

 DFSG compliance
 ---------------
 The server example contains a number of minified and generated files in the
 frontend. These seem to be essential to the example, so the server example
 has been removed entirely, for now.

I guess some build mechanics need to be included to build the minified
and generated files, but do not know which one are the problem.
According to examples/server/README.md the "Web UI" is buitl using 'npm
run build', so I guess some nodejs dependencies are needed.  Sadly, I do
not know how to convince npm to not download random stuff from the
Internet.

-- 
Happy hacking
Petter Reinholdtsen

[toc] | [prev] | [next] | [standalone]


#1231996

From"M. Zhou" <lumin@debian.org>
Date2025-02-06 15:20 +0100
Message-ID<KdaTv-eQwc-3@gated-at.bofh.it>
In reply to#1231945
On Thu, 2025-02-06 at 09:13 +0100, Christian Kastner wrote:
> 
> I meant to ask anyway: performance-wise, is it comparable to your local
> build? I mean, I wouldn't know what in the code would alter this, but I
> built and tested this on platti.d.o and performance was poor, so another
> data point would be useful.

For ppc64el, the llama.cpp-blas backend is way slower than the -cpu backend.
I did not test on amd64. But on ppc64el the package does not feel different
than local build.

CPU is slow anyway. How does HIP performs?

phi-4-q4.gguf | power9, cpu (8-threads) | 0.62 tokens/s
phi-4-q4.gguf | amd64, 13900H           | 6.7 tokens/s

GPU is way faster than this. The phi-4 model does not fit in my nvidia GPU.
No number for GPU this time.

[toc] | [prev] | [next] | [standalone]


#1231967

FromPetter Reinholdtsen <pere@hungry.com>
Date2025-02-06 09:30 +0100
Message-ID<Kd5qN-eMwX-9@gated-at.bofh.it>
In reply to#1231943
[Christian Kastner]
> Look fine, though I deliberately skipped the poetry dependency for now
> as it looked more like a false positive.

Aha.  I just trusted lintian-brush on this one, did not investigate.

> This was my intention, but I initially wasn't sure what the default
> would be (-cpu or -blas). Looks like I forgot to add one before
> upload.

Given that every machine it can be installed on got a CPU, but not all
of them got a supported GPU, I beieve -cpu is the most sensible default.

> It'll be re-enabled soon. The were a few generated and minified files
> in that example, so I just opted to skip those for now, and focus on
> the build process.

Great to hear. :)

> Seeing as how closely llama.cpp and whisper.cpp are related, in the
> ideal case, you should be able to just carry over some patches, and
> mostly just copy d/rules, as llama.cpp and whisper.cpp share the ggml
> library on a source basis.

I hope so too, but I guess we will soon find out.  My initial draft on
<URL: https://salsa.debian.org/deeplearning-team/whisper.cpp > will need
a lot of updates  to bring it in line with this new approach. :)

-- 
Happy hacking
Petter Reinholdtsen

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.bugs.dist


csiph-web