Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.bugs.dist > #1231996
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Newsgroups | linux.debian.bugs.dist |
| Subject | Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ |
| Date | 2025-02-06 15:20 +0100 |
| Message-ID | <KdaTv-eQwc-3@gated-at.bofh.it> (permalink) |
| References | (18 earlier) <KcZbI-eIzF-3@gated-at.bofh.it> <KdaTv-eQwc-5@gated-at.bofh.it> <KdaTv-eQwc-7@gated-at.bofh.it> <I6Vyx-9PtM-15@gated-at.bofh.it> <KdaTv-eQwc-9@gated-at.bofh.it> |
| Organization | linux.* mail to news gateway |
On Thu, 2025-02-06 at 09:13 +0100, Christian Kastner wrote: > > I meant to ask anyway: performance-wise, is it comparable to your local > build? I mean, I wouldn't know what in the code would alter this, but I > built and tested this on platti.d.o and performance was poor, so another > data point would be useful. For ppc64el, the llama.cpp-blas backend is way slower than the -cpu backend. I did not test on amd64. But on ppc64el the package does not feel different than local build. CPU is slow anyway. How does HIP performs? phi-4-q4.gguf | power9, cpu (8-threads) | 0.62 tokens/s phi-4-q4.gguf | amd64, 13900H | 6.7 tokens/s GPU is way faster than this. The phi-4 model does not fit in my nvidia GPU. No number for GPU this time.
Back to linux.debian.bugs.dist | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-05 22:00 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 01:40 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 02:50 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:00 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ "M. Zhou" <lumin@debian.org> - 2025-02-06 15:20 +0100
Bug#1063673: ITP: llama.cpp -- Inference of Meta's LLaMA model (and others) in pure C/C++ Petter Reinholdtsen <pere@hungry.com> - 2025-02-06 09:30 +0100
csiph-web