Llama.cpp : Tuxedo 17 (en Ubuntu 24) + GIGABYTE AORUS RTX 5060 Ti AI Box .

Sur le Tuxedo 17 ( https://www.tuxedocomputers.com/ ) j’ai ajouté via le Thunderbold un « GIGABYTE AORUS RTX 5060 Ti AI Box ».
Mon installation :
– Ubuntu : 24.04
– CUDA : 13.3
– NVIDIA : 610.43.03

# lsb_release -a
No LSB modules are available.
Distributor ID: Tuxedo
Description:    TUXEDO OS
Release:        24.04
Codename:       noble
# nvidia-smi 
Thu Jul 16 17:29:22 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.43.03              KMD Version: 610.43.03     CUDA UMD Version: 13.3     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 3060 ...    Off |   00000000:01:00.0 Off |                  N/A |
| N/A   46C    P8             11W /  115W |       1MiB /   6144MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+
|   1  NVIDIA GeForce RTX 5060 Ti     Off |   00000000:05:00.0 Off |                  N/A |
|  0%   34C    P0             15W /  180W |       2MiB /  16311MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

Ma version de Thunderbold est 4 :

# lspci | grep -i thunderbolt
00:07.0 PCI bridge: Intel Corporation Tiger Lake-H Thunderbolt 4 PCI Express Root Port #0 (rev 05)
00:0d.0 USB controller: Intel Corporation Tiger Lake-H Thunderbolt 4 USB Controller (rev 05)
00:0d.2 USB controller: Intel Corporation Tiger Lake-H Thunderbolt 4 NHI #0 (rev 05)

A noter que pour le build de llama.cpp j’ai ajouté une directive de compilation : CMAKE_CUDA_ARCHITECTURES qui a pour valeur par défaut 75 donc 7.5.

# nvidia-smi --query-gpu=compute_cap --format=csv
compute_cap
8.6
12.0
...
# cmake -B build -DGGML_CUDA=ON  -DCMAKE_CUDA_COMPILER=`which nvcc` -DLLAMA_CURL=ON -DCMAKE_CUDA_ARCHITECTURES=86
...

J’ai d’abord essayé le soft de Gigabyte : https://www.gigabyte.com/Support/Utility/Graphics-Card . Mais il n’est pas fonctionnel.

Puis j’ai essayé une utilisation normale, et j’avais des crash sans arrêts.Avec des erreurs du type :

... nvidia-modeset: ERROR: GPU:1: Error while waiting for GPU progress: ...

Ensuite je suis tombé sur ce site : https://github.com/mmhorda/ai-box-rtx-5060ti-egpu-guide , et j’ai compris qu’il y avait un problème de débit.

En fait il y a un bug ouvert sur le sujet : https://github.com/NVIDIA/open-gpu-kernel-modules/issues/979#issuecomment-4172636691 : « RTX 5080 via Thunderbolt 5 eGPU: Hard lock on CUDA operations (nvidia-smi works at idle) « .

J’ai finalement trouvé un projet avec un fix pour Ubuntu : https://github.com/lokmantsui/aorus-5090-egpu/tree/nvidia-610.43.02-ubuntu-5060ti

# git clone https://github.com/lokmantsui/aorus-5090-egpu.git -b nvidia-610.43.02-ubuntu-5060ti

J’ai du faire des modifications car il détectait pas la bonne carte NVIDIA. Et finalement c’est tombé en marche.

Pour mon installation la commande grup ressemble à ceci :

GRUB_CMDLINE_LINUX="iommu.passthrough=1 thunderbolt.host_reset=false pcie_aspm.policy=performance thunderbolt.clx=0 pcie_port_pm=off pci=resource_alignment=35@0000:00:05.0"

J’ai donc fait mes premiers tests avec gemma-3-1b-it-q4_k_m.gguf :

# llama-bench --list-devices 
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
Available devices:
  CUDA0: NVIDIA GeForce RTX 5060 Ti (15888 MiB, 15752 MiB free)
  CUDA1: NVIDIA GeForce RTX 3060 Laptop GPU (5803 MiB, 5685 MiB free)

# llama-bench -m /models/gemma-3-1b-it-q4_k_m.gguf -dev CUDA0,CUDA1
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------------ | --------------: | -------------------: |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 | CUDA0        |           pp512 |   19994.19 ± 1751.94 |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 | CUDA0        |           tg128 |        268.49 ± 0.38 |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 | CUDA1        |           pp512 |    11158.41 ± 527.54 |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 | CUDA1        |           tg128 |        229.10 ± 0.37 |

build: c3d47e696 (10030)
# llama-bench -m /models/gemma-3-1b-it-q4_k_m.gguf 
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 |           pp512 |   16582.32 ± 1005.41 |
| gemma3 1B Q4_K - Medium        | 762.49 MiB |   999.89 M | CUDA       |  -1 |           tg128 |        261.64 ± 1.81 |

build: c3d47e696 (10030)

Et donc il est préférable d’utiliser la carte la plus puissante et non pas les deux cartes en même temps.

J’ai refais un test avec un modèle plus gros : Qwen3.6-27B-Q3_K_M.gguf

# llama-bench -m /models/Qwen3.6-27B-Q3_K_M.gguf 
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --------------: | -------------------: |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       |  -1 |           pp512 |        707.38 ± 5.11 |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       |  -1 |           tg128 |         18.31 ± 0.01 |

build: c3d47e696 (10030)

# llama-bench -m /models/Qwen3.6-27B-Q3_K_M.gguf -dev CUDA0
ggml_cuda_init: found 2 CUDA devices (Total VRAM: 21692 MiB):
  Device 0: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
  Device 1: NVIDIA GeForce RTX 3060 Laptop GPU, compute capability 8.6, VMM: yes, VRAM: 5803 MiB
| model                          |       size |     params | backend    | ngl | dev          |            test |                  t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | ------------ | --------------: | -------------------: |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       |  -1 | CUDA0        |           pp512 |        841.06 ± 9.58 |
| qwen35 27B Q3_K - Medium       |  12.64 GiB |    26.90 B | CUDA       |  -1 | CUDA0        |           tg128 |         20.50 ± 0.02 |

build: c3d47e696 (10030)

Prochaine étape, l’ajout d’Hermes Agent IA.