mirror of
https://github.com/R0m1k3/Loki.git
synced 2026-10-11 17:26:57 +02:00
Simplification radicale : moteur pris de l'image officielle llama.cpp
Plus aucune compilation CUDA : le runtime part de ghcr.io/ggml-org/llama.cpp:server-cuda (llama-server précompilé et maintenu par l'équipe amont, backends .so chargés dynamiquement, archs GPU courantes, repli CPU fonctionnel). Le build complet passe de ~40 min à ~4 min et tous les pièges du build CUDA sans GPU (stubs libcuda, espace disque du runner, choix des architectures) disparaissent. - Dockerfile : 2 étapes (Go + image officielle), build-arg LLAMACPP_IMAGE pour épingler une version ou passer en CPU/Vulkan - entrypoint : BIN=/app/llama-server - workflow GHCR : purge disque et CUDA_ARCHS supprimés, cache max - compose/.env/README à l'avenant Validé : build 3 min 55, UI HTTP 200, healthcheck healthy, superviseur PID OK, llama-server démarre (repli CPU) et n'échoue que sur un modèle factice.
This commit is contained in:
6 files changed
+49
-89
No files matched your search
+22
-54
@@ -1,54 +1,20 @@
|
||||
# ── Loki — image GPU autonome (fork d'AJEAN, https://github.com/nathaninline/ajean)
|
||||
# ── Loki — image GPU (fork d'AJEAN, https://github.com/nathaninline/ajean)
|
||||
#
|
||||
# Trois étapes :
|
||||
# 1. compilation de llama.cpp avec CUDA (mêmes flags que backend_build.go,
|
||||
# sauf GGML_NATIVE=OFF : l'image est bâtie sur un runner GitHub, pas sur
|
||||
# la machine qui l'exécutera — des instructions CPU « natives » du runner
|
||||
# provoqueraient un Illegal instruction ailleurs) ;
|
||||
# 2. compilation du binaire Go `loki` (UI embarquée via go:embed) ;
|
||||
# 3. runtime CUDA léger : llama-server + loki + entrypoint.
|
||||
# Deux étapes seulement :
|
||||
# 1. compilation du binaire Go `loki` (UI embarquée via go:embed) ;
|
||||
# 2. runtime = l'image serveur CUDA OFFICIELLE de llama.cpp — llama-server
|
||||
# y est précompilé et maintenu par l'équipe amont (base nvidia/cuda
|
||||
# runtime, backends .so dans /app, architectures GPU courantes).
|
||||
# Aucune compilation CUDA ici : le build complet prend quelques minutes.
|
||||
#
|
||||
# Pré-requis d'exécution : NVIDIA Container Toolkit sur l'hôte.
|
||||
# docker build -t loki --build-arg CUDA_ARCHS=86 .
|
||||
# CUDA_ARCHS : 75=Turing, 86=Ampere, 89=Ada — restreindre à sa carte divise
|
||||
# le temps de build et la taille de l'image.
|
||||
# Épingler une version : --build-arg LLAMACPP_IMAGE=ghcr.io/ggml-org/llama.cpp:server-cuda-b10423
|
||||
# Variante CPU (test sans GPU) : --build-arg LLAMACPP_IMAGE=ghcr.io/ggml-org/llama.cpp:server
|
||||
# Pré-requis d'exécution GPU : NVIDIA Container Toolkit sur l'hôte.
|
||||
|
||||
# ── Étape 1 : llama.cpp CUDA ────────────────────────────────────────────
|
||||
FROM nvidia/cuda:12.6.3-devel-ubuntu24.04 AS llamacpp
|
||||
# Déclaré avant le premier FROM : requis pour être utilisable dans un FROM.
|
||||
ARG LLAMACPP_IMAGE=ghcr.io/ggml-org/llama.cpp:server-cuda
|
||||
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
git cmake build-essential ca-certificates \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# Épingler LLAMACPP_REF sur un tag (ex. b4991) pour des builds reproductibles.
|
||||
ARG LLAMACPP_REF=master
|
||||
ARG CUDA_ARCHS=75;86;89
|
||||
|
||||
RUN git clone --depth 1 --branch "${LLAMACPP_REF}" \
|
||||
https://github.com/ggml-org/llama.cpp /src/llama.cpp
|
||||
|
||||
# --allow-shlib-undefined : pas de GPU (donc pas de libcuda.so.1) sur le
|
||||
# builder — l'API driver (cuMemCreate…) reste non résolue au link et sera
|
||||
# fournie à l'exécution par le runtime NVIDIA. Même parade que le
|
||||
# cuda.Dockerfile officiel de llama.cpp.
|
||||
RUN cmake -S /src/llama.cpp -B /src/llama.cpp/build \
|
||||
-DCMAKE_BUILD_TYPE=Release \
|
||||
-DGGML_CUDA=ON \
|
||||
-DGGML_CUDA_F16=ON \
|
||||
-DGGML_NATIVE=OFF \
|
||||
-DCMAKE_CUDA_ARCHITECTURES="${CUDA_ARCHS}" \
|
||||
-DLLAMA_CURL=OFF \
|
||||
-DLLAMA_BUILD_TESTS=OFF \
|
||||
-DLLAMA_BUILD_EXAMPLES=OFF \
|
||||
-DLLAMA_BUILD_UI=OFF \
|
||||
-DLLAMA_USE_PREBUILT_UI=OFF \
|
||||
-DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined \
|
||||
&& cmake --build /src/llama.cpp/build --target llama-server -j"$(nproc)" \
|
||||
&& mkdir -p /opt/llama.cpp \
|
||||
&& find /src/llama.cpp/build -name 'llama-server' -type f -exec cp {} /opt/llama.cpp/ \; \
|
||||
&& find /src/llama.cpp/build -name '*.so*' -exec cp -P {} /opt/llama.cpp/ \;
|
||||
|
||||
# ── Étape 2 : binaire Go loki ───────────────────────────────────────────
|
||||
# ── Étape 1 : binaire Go loki ───────────────────────────────────────────
|
||||
FROM golang:1.25 AS gobuild
|
||||
WORKDIR /src
|
||||
COPY go.mod go.sum ./
|
||||
@@ -58,8 +24,8 @@ COPY internal/ internal/
|
||||
COPY tools/ tools/
|
||||
RUN CGO_ENABLED=0 go build -trimpath -ldflags "-s -w" -o /out/loki ./cmd/loki
|
||||
|
||||
# ── Étape 3 : runtime ───────────────────────────────────────────────────
|
||||
FROM nvidia/cuda:12.6.3-runtime-ubuntu24.04 AS runtime
|
||||
# ── Étape 2 : runtime sur l'image serveur CUDA officielle ───────────────
|
||||
FROM ${LLAMACPP_IMAGE} AS runtime
|
||||
|
||||
# Marqueur de build (sha court), injecté par le workflow GitHub.
|
||||
ARG LOKI_VERSION=dev
|
||||
@@ -67,21 +33,21 @@ LABEL org.opencontainers.image.revision="${LOKI_VERSION}" \
|
||||
org.opencontainers.image.source="https://github.com/R0m1k3/Loki" \
|
||||
org.opencontainers.image.description="Loki — fork conteneurisé d'AJEAN (github.com/nathaninline/ajean, MIT)"
|
||||
|
||||
# curl pour le HEALTHCHECK ; git pour les outils de l'agent ; nodejs/npm pour
|
||||
# les serveurs MCP lancés via npx ; libgomp1 pour llama-server ; tini en PID 1.
|
||||
# git pour les outils de l'agent ; nodejs/npm pour les serveurs MCP lancés
|
||||
# via npx ; tini en PID 1. curl et libgomp1 sont déjà dans l'image amont.
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
ca-certificates curl git libgomp1 tini nodejs npm \
|
||||
git nodejs npm tini \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
COPY --from=llamacpp /opt/llama.cpp /opt/llama.cpp
|
||||
COPY --from=gobuild /out/loki /usr/local/bin/loki
|
||||
COPY docker-entrypoint.sh /usr/local/bin/docker-entrypoint.sh
|
||||
RUN chmod +x /usr/local/bin/docker-entrypoint.sh && mkdir -p /data /models
|
||||
|
||||
# llama-server et ses .so vivent dans /app (convention de l'image amont).
|
||||
ENV LOKI_CONTAINER=1 \
|
||||
LOKI_HOME=/data \
|
||||
LOKI_MODEL_DIRS=/models \
|
||||
LD_LIBRARY_PATH=/opt/llama.cpp \
|
||||
LD_LIBRARY_PATH=/app \
|
||||
NVIDIA_VISIBLE_DEVICES=all \
|
||||
NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
||||
|
||||
@@ -89,8 +55,10 @@ ENV LOKI_CONTAINER=1 \
|
||||
# conteneur — il n'est pas authentifié par défaut.
|
||||
EXPOSE 8090
|
||||
|
||||
# Remplace le HEALTHCHECK de l'image amont (qui vise le moteur sur :8080).
|
||||
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
|
||||
CMD curl -fsS "http://localhost:${LOKI_WEB_PORT:-8090}/" >/dev/null || exit 1
|
||||
|
||||
WORKDIR /data
|
||||
# Remplace l'ENTRYPOINT de l'image amont (/app/llama-server).
|
||||
ENTRYPOINT ["tini", "--", "docker-entrypoint.sh"]
|
||||
Reference in new issue
Block a user