Dockerfile : link llama-server sans GPU (--allow-shlib-undefined)

Le builder n'a pas de libcuda.so.1 (API driver) : le link de llama-server
échouait sur des références cu* non résolues. Comme le cuda.Dockerfile
officiel de llama.cpp, on autorise les symboles non résolus des .so au
link — le runtime NVIDIA injecte le vrai libcuda à l'exécution.
Ajout aussi des flags UI amont (LLAMA_BUILD_UI=OFF, LLAMA_USE_PREBUILT_UI=OFF).
This commit is contained in:
Loki committed 2026-08-14 22:21:45 +00:00
1 parent bb9aa9f559
commit 0846a70eb0
1 file changed
+5
+5
View File
@@ -27,6 +27,10 @@ ARG CUDA_ARCHS=75;86;89
RUN git clone --depth 1 --branch "${LLAMACPP_REF}" \ RUN git clone --depth 1 --branch "${LLAMACPP_REF}" \
https://github.com/ggml-org/llama.cpp /src/llama.cpp https://github.com/ggml-org/llama.cpp /src/llama.cpp
# --allow-shlib-undefined : pas de GPU (donc pas de libcuda.so.1) sur le
# builder — l'API driver (cuMemCreate…) reste non résolue au link et sera
# fournie à l'exécution par le runtime NVIDIA. Même parade que le
# cuda.Dockerfile officiel de llama.cpp.
RUN cmake -S /src/llama.cpp -B /src/llama.cpp/build \ RUN cmake -S /src/llama.cpp -B /src/llama.cpp/build \
-DCMAKE_BUILD_TYPE=Release \ -DCMAKE_BUILD_TYPE=Release \
-DGGML_CUDA=ON \ -DGGML_CUDA=ON \
@@ -38,6 +42,7 @@ RUN cmake -S /src/llama.cpp -B /src/llama.cpp/build \
-DLLAMA_BUILD_EXAMPLES=OFF \ -DLLAMA_BUILD_EXAMPLES=OFF \
-DLLAMA_BUILD_UI=OFF \ -DLLAMA_BUILD_UI=OFF \
-DLLAMA_USE_PREBUILT_UI=OFF \ -DLLAMA_USE_PREBUILT_UI=OFF \
-DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined \
&& cmake --build /src/llama.cpp/build --target llama-server -j"$(nproc)" \ && cmake --build /src/llama.cpp/build --target llama-server -j"$(nproc)" \
&& mkdir -p /opt/llama.cpp \ && mkdir -p /opt/llama.cpp \
&& find /src/llama.cpp/build -name 'llama-server' -type f -exec cp {} /opt/llama.cpp/ \; \ && find /src/llama.cpp/build -name 'llama-server' -type f -exec cp {} /opt/llama.cpp/ \; \