mirror of
https://github.com/R0m1k3/Loki.git
synced 2026-10-11 17:26:57 +02:00
llama-server garde en RAM hôte une copie exacte des états de slot qu'il quitte (KV, état récurrent, points de reprise, brouillon MTP) et la recharge octet pour octet. Mais son défaut de 8 Gio ne tient pas une conversation de 30 à 65 k jetons à côté de l'état d'un vérificateur, d'un sous-agent, d'une tâche ou d'un bench : à la requête suivante il évince la conversation pour sauver l'annexe, et tout est recalculé (30 à 80 s sur un 27B). Rien de ce que voit le modèle ne change ici : seul le bruit de découpage des lots, comme le cache_prompt actuel. - CACHE_RAM (Mio ; vide/auto, nombre passé tel quel, -1, 0) → --cache-ram, seulement si l'aide du moteur le connaît. Loki se tait devant -cram d'EXTRA_ARGS et LLAMA_ARG_CACHE_RAM. L'auto ne descend JAMAIS sous le défaut : il agrandit seulement un modèle tout-GPU (VRAM NVIDIA connue, aucun poids ni KV sur CPU, NGL complet), d'après l'état mesuré dans le GGUF (2,5 × KV à CTX + état récurrent et points de reprise des hybrides), plafonné à 30 % de la RAM (limite cgroup comprise) et à la moitié de la RAM libre. - --slot-save-path LOKI_HOME/slots (0700, purgé au lancement) uniquement si le moteur est en boucle locale ou protégé par clé, et si le dossier existe. Les proxys /v1 et du relais refusent désormais toute action /slots (405). - engineSideJob efface le slot 0 APRÈS le sous-agent, la passe de vérification, la tâche planifiée et le bench (pas après la compaction : rien à y gagner). Synchrone, borné à 2 s, et seulement si : moteur local, build ≥ 8660 lu dans /props, total_slots == 1, slot 0 au repos, aucune requête de Loki en vol (compteur tenu par le chat, les résumés, le bench et les proxys). CACHE_ISOLATE=off le coupe. - usage.prompt_tokens_details.cached_tokens : si la conversation revient avec moins de la moitié en cache, conseil (une fois) de relever CACHE_RAM. - Éditeur de preset : champ « Cache de prompts » à côté d'UBATCH, avec la valeur auto calculée par le serveur (/api/preset/cacheram). - Lecteur GGUF : embedding_length, head_count et dimensions ssm.*. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
250 lines
12 KiB
Go
250 lines
12 KiB
Go
package loki
|
|
|
|
import (
|
|
"bytes"
|
|
"encoding/json"
|
|
"fmt"
|
|
"net/http"
|
|
"path/filepath"
|
|
"strconv"
|
|
"strings"
|
|
"time"
|
|
)
|
|
|
|
// benchResult captures the timings llama.cpp returns from /completion.
|
|
type benchResult struct {
|
|
PromptN int `json:"prompt_n"`
|
|
PromptMs float64 `json:"prompt_ms"`
|
|
PromptPerSecond float64 `json:"prompt_per_second"`
|
|
PredictedN int `json:"predicted_n"`
|
|
PredictedMs float64 `json:"predicted_ms"`
|
|
PredictedPerSec float64 `json:"predicted_per_second"`
|
|
Elapsed float64 `json:"elapsed_sec"`
|
|
}
|
|
|
|
// benchCorpus is a varied passage used to defeat speculative decoding
|
|
// (MTP / n-gram draft) — repetitive text inflates decode tok/s because every
|
|
// drafted token gets accepted, which is unlike real chat. We pull a chunk of
|
|
// natural-looking content and tile it to reach the target prompt size.
|
|
const benchCorpus = `In the early hours of an October morning, Camille walked along the canal, watching the cargo barges slip past the iron bridge that spanned the water. She thought about the meeting she had skipped, the unanswered messages on her phone, the way the city always seemed to forget her name after summer ended. Three streets away, a pâtisserie opened its shutters and the smell of warm butter mixed with diesel exhaust from the waiting bus.
|
|
Pendant ce temps, à Marseille, un chercheur en biologie marine prépare son matériel pour une plongée. Il étudie les herbiers de posidonie, ces prairies sous-marines vieilles de plusieurs milliers d'années qui stockent autant de carbone qu'une forêt amazonienne. Le bateau quitte le port à six heures vingt-trois.
|
|
Quantum computers, properly engineered, can solve certain classes of problems exponentially faster than classical machines. The catch is that decoherence ruins everything. Engineers use dilution refrigerators to drop superconducting qubits to fifteen millikelvin, colder than deep space. The wires connecting the chip to room-temperature electronics must dissipate almost no heat, or the qubit state collapses before any useful computation finishes.
|
|
Le boulanger lève la pâte à quatre heures. Il regarde la balance numérique en plissant les yeux : six cent vingt-trois grammes, presque le compte. Son chien dort sur le tapis de farine près du four. Dehors, deux chats se disputent un poisson abandonné par le pêcheur de nuit.
|
|
Consider a recursive descent parser written in Go. The lexer emits tokens; the parser consumes them and produces an abstract syntax tree. Error recovery is hard: after a syntax error, the parser must resynchronize at a known boundary—a semicolon, a closing brace—without losing track of subsequent diagnostics. Tree-sitter solves this with incremental parsing and a glr-like algorithm.
|
|
Le philosophe stoïcien disait : "Ce qui nous trouble, ce n'est pas ce qui nous arrive, mais l'opinion que nous nous en faisons." Vingt siècles plus tard, la phrase apparaît dans un livre de poche au rayon développement personnel d'une librairie d'aéroport, à côté d'un roman policier suédois.
|
|
Mitochondria descended from ancient bacteria engulfed by archaeal cells roughly two billion years ago. They still keep their own ring of DNA, separate from the nuclear genome. Mutations in mitochondrial DNA accumulate with age and have been implicated in everything from Parkinson's disease to ordinary muscle fatigue. Yet they remain stubbornly difficult to repair therapeutically because each cell contains hundreds.
|
|
La marée descend lentement, exposant des rochers couverts d'huîtres et d'algues vertes. Un héron immobile surveille les flaques laissées par l'eau. Plus loin, deux enfants courent avec un cerf-volant rouge qui refuse de monter à cause de l'humidité dans la voile.
|
|
Compilers translate high-level languages into machine code through several intermediate representations. LLVM IR sits in the middle: typed, mostly static-single-assignment, suitable for both aggressive optimization and direct lowering to x86 or ARM. The optimizer runs dozens of passes—dead code elimination, loop-invariant code motion, induction variable simplification—each touching the IR in carefully ordered ways.
|
|
Le cuisinier ferme les yeux pour goûter la sauce. Trop salée. Il ajoute une pomme de terre crue coupée en quartiers, sachant qu'elle absorbera l'excès en mijotant vingt minutes. Sa grand-mère lui a appris ce geste un dimanche de novembre il y a très longtemps.
|
|
`
|
|
|
|
// runBench fires a prompt of roughly `nPrompt` tokens at /completion with
|
|
// cache_prompt:false (so prefill is actually measured, not cached). The prompt
|
|
// is a varied corpus to keep speculative decoding (MTP / n-gram draft) from
|
|
// inflating decode numbers — what you measure here is close to what you'll
|
|
// see in real chat at the same context length.
|
|
func runBench(nPrompt, nPredict int) (*benchResult, error) {
|
|
// Un bench mesure le moteur LOCAL : sur un preset externe, healthCheck dit
|
|
// « prêt » sans moteur, et la mesure taperait un port arrêté ou un moteur
|
|
// resté en vie — attribuée à tort à ce preset.
|
|
if externalActive() {
|
|
return nil, fmt.Errorf("benchmark indisponible : le preset actif est une API externe")
|
|
}
|
|
port := LLMPort()
|
|
if !healthCheck() {
|
|
return nil, fmt.Errorf("serveur injoignable sur :%d", port)
|
|
}
|
|
// Le bench prend le slot de la conversation : une fois fini, on l'efface
|
|
// pour qu'elle soit rechargée depuis le cache RAM (llm_slots.go).
|
|
defer engineSideJob()()
|
|
if nPrompt <= 0 {
|
|
nPrompt = 2000
|
|
}
|
|
// 1 word ≈ 1.3 tokens roughly. Tile the varied corpus until we exceed
|
|
// nPrompt, then truncate to characters so the server tokenises a passage
|
|
// close to the requested size.
|
|
corpusWords := strings.Fields(benchCorpus)
|
|
target := nPrompt * 5 // ~5 chars/token gives a generous over-estimate
|
|
var b strings.Builder
|
|
for b.Len() < target {
|
|
for _, w := range corpusWords {
|
|
b.WriteString(w)
|
|
b.WriteByte(' ')
|
|
if b.Len() >= target {
|
|
break
|
|
}
|
|
}
|
|
}
|
|
prompt := strings.TrimSpace(b.String())
|
|
// Use the same endpoint your real chat hits, so the comparison is honest
|
|
// (chat template, reasoning, OpenAI-compat layer all included).
|
|
payload := map[string]any{
|
|
"model": "loki",
|
|
"messages": []Message{{Role: "user", Content: prompt + "\n\nContinue this passage with another 1000+ words of original varied prose, mixing French and English narrative paragraphs on different topics."}},
|
|
"max_tokens": nPredict,
|
|
"stream": false,
|
|
"temperature": 0.7,
|
|
"cache_prompt": false,
|
|
}
|
|
body, _ := json.Marshal(payload)
|
|
url := fmt.Sprintf("http://localhost:%d/v1/chat/completions", port)
|
|
t0 := time.Now()
|
|
req, _ := http.NewRequest("POST", url, bytes.NewReader(body))
|
|
req.Header.Set("Content-Type", "application/json")
|
|
authHeader(req)
|
|
client := &http.Client{Timeout: 5 * time.Minute}
|
|
defer engineRequestStart()() // avant l'effacement, différé plus haut
|
|
resp, err := client.Do(req)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
defer resp.Body.Close()
|
|
var parsed struct {
|
|
Timings struct {
|
|
PromptN int `json:"prompt_n"`
|
|
PromptMs float64 `json:"prompt_ms"`
|
|
PromptPerSecond float64 `json:"prompt_per_second"`
|
|
PredictedN int `json:"predicted_n"`
|
|
PredictedMs float64 `json:"predicted_ms"`
|
|
PredictedPerSec float64 `json:"predicted_per_second"`
|
|
} `json:"timings"`
|
|
Usage struct {
|
|
PromptTokens int `json:"prompt_tokens"`
|
|
CompletionTokens int `json:"completion_tokens"`
|
|
} `json:"usage"`
|
|
}
|
|
if err := json.NewDecoder(resp.Body).Decode(&parsed); err != nil {
|
|
return nil, err
|
|
}
|
|
// /v1/chat/completions may report timings at top level or omit them; if
|
|
// missing, fall back to wall-clock derived from `usage` so we always show
|
|
// numbers comparable to what chat displays in its label.
|
|
if parsed.Timings.PredictedN == 0 && parsed.Usage.CompletionTokens > 0 {
|
|
elapsed := time.Since(t0).Seconds()
|
|
parsed.Timings.PromptN = parsed.Usage.PromptTokens
|
|
parsed.Timings.PredictedN = parsed.Usage.CompletionTokens
|
|
// Distribute elapsed time using a rough split (prefill is usually <20% at this size).
|
|
parsed.Timings.PromptMs = elapsed * 1000 * 0.15
|
|
parsed.Timings.PredictedMs = elapsed * 1000 * 0.85
|
|
if parsed.Timings.PromptMs > 0 {
|
|
parsed.Timings.PromptPerSecond = float64(parsed.Timings.PromptN) / (parsed.Timings.PromptMs / 1000)
|
|
}
|
|
if parsed.Timings.PredictedMs > 0 {
|
|
parsed.Timings.PredictedPerSec = float64(parsed.Timings.PredictedN) / (parsed.Timings.PredictedMs / 1000)
|
|
}
|
|
}
|
|
elapsed := time.Since(t0).Seconds()
|
|
t := parsed.Timings
|
|
res := &benchResult{
|
|
PromptN: t.PromptN, PromptMs: t.PromptMs, PromptPerSecond: t.PromptPerSecond,
|
|
PredictedN: t.PredictedN, PredictedMs: t.PredictedMs, PredictedPerSec: t.PredictedPerSec,
|
|
Elapsed: elapsed,
|
|
}
|
|
saveLastBench(res)
|
|
saveBenchForActivePreset(res)
|
|
return res, nil
|
|
}
|
|
|
|
// savedBench is a benchResult plus the model it was run against and a timestamp.
|
|
type savedBench struct {
|
|
Result benchResult `json:"result"`
|
|
Model string `json:"model"`
|
|
At int64 `json:"at"`
|
|
}
|
|
|
|
// saveLastBench enregistre le dernier benchmark (best-effort) pour que l'UI
|
|
// puisse l'afficher sans le relancer.
|
|
func saveLastBench(res *benchResult) {
|
|
sb := savedBench{Result: *res, Model: filepath.Base(ReadConfig()["MODEL"]), At: time.Now().Unix()}
|
|
_ = putJSON(bkState, "last_bench", sb)
|
|
}
|
|
|
|
// loadLastBench relit le benchmark enregistré, ou nil s'il n'y en a pas.
|
|
func loadLastBench() *savedBench {
|
|
var sb savedBench
|
|
if !getJSON(bkState, "last_bench", &sb) {
|
|
return nil
|
|
}
|
|
return &sb
|
|
}
|
|
|
|
// benchMatchesPreset : le bench a-t-il été mesuré sur le modèle ACTUEL du
|
|
// preset ? Les benchs sont rangés par id de preset : sans ce contrôle, un
|
|
// preset dont le modèle a changé affichait les mesures d'un autre modèle.
|
|
// Repris d'AJEAN 0.16.3.
|
|
func benchMatchesPreset(sb savedBench, cfg map[string]string) bool {
|
|
m := strings.TrimSpace(cfg["MODEL"])
|
|
return m != "" && sb.Model == filepath.Base(m)
|
|
}
|
|
|
|
// deletePresetBench oublie le bench d'un preset supprimé.
|
|
func deletePresetBench(id string) {
|
|
m := loadBenchStore()
|
|
if _, ok := m[id]; ok {
|
|
delete(m, id)
|
|
_ = putJSON(bkState, "bench_presets", m)
|
|
}
|
|
}
|
|
|
|
// loadBenchStore renvoie les benchmarks par preset (vide s'il n'y en a pas).
|
|
func loadBenchStore() map[string]savedBench {
|
|
m := map[string]savedBench{}
|
|
getJSON(bkState, "bench_presets", &m)
|
|
if m == nil {
|
|
m = map[string]savedBench{}
|
|
}
|
|
return m
|
|
}
|
|
|
|
// saveBenchForActivePreset records res under the name of the currently active
|
|
// preset (celui qui correspond à la configuration active). Sans effet si aucun
|
|
// preset ne correspond — le benchmark reste enregistré par saveLastBench.
|
|
func saveBenchForActivePreset(res *benchResult) {
|
|
list, err := ListPresets()
|
|
if err != nil {
|
|
return
|
|
}
|
|
id := ""
|
|
for _, p := range list {
|
|
if p.Active {
|
|
id = p.ID
|
|
break
|
|
}
|
|
}
|
|
if id == "" {
|
|
return
|
|
}
|
|
m := loadBenchStore()
|
|
m[id] = savedBench{Result: *res, Model: filepath.Base(ReadConfig()["MODEL"]), At: time.Now().Unix()}
|
|
_ = putJSON(bkState, "bench_presets", m)
|
|
}
|
|
|
|
func cmdBench(args []string) error {
|
|
nPredict, nPrompt := 300, 2000
|
|
if len(args) >= 1 && args[0] != "" {
|
|
if n, err := strconv.Atoi(args[0]); err == nil {
|
|
nPredict = n
|
|
} else {
|
|
return fmt.Errorf("argument invalide: %s", args[0])
|
|
}
|
|
}
|
|
if len(args) >= 2 && args[1] != "" {
|
|
if n, err := strconv.Atoi(args[1]); err == nil {
|
|
nPrompt = n
|
|
} else {
|
|
return fmt.Errorf("argument invalide: %s", args[1])
|
|
}
|
|
}
|
|
fmt.Printf("[bench] prompt ~%d tokens, n_predict=%d…\n", nPrompt, nPredict)
|
|
r, err := runBench(nPrompt, nPredict)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
fmt.Println()
|
|
fmt.Printf(" %s %7.1f tok/s (%d tokens en %.2fs)\n", cyan("Prefill"), r.PromptPerSecond, r.PromptN, r.PromptMs/1000)
|
|
fmt.Printf(" %s %7.1f tok/s (%d tokens en %.2fs)\n", cyan("Decode "), r.PredictedPerSec, r.PredictedN, r.PredictedMs/1000)
|
|
fmt.Printf(" Total %.2fs\n", r.Elapsed)
|
|
fmt.Println()
|
|
return nil
|
|
}
|