Ajoute jean link : acces distant via le relais ajean.link

- nouvelle commande `jean link <token>` (link.go) : connexion sortante
  WebSocket + yamux vers le relais, sert l'UI web Jean ET l'endpoint
  compatible OpenAI (/v1) a travers le tunnel, avec injection des cles
  locales, keepalive et reconnexion automatique.
- web.go : extraction de newWebMux() pour reutiliser le routeur HTTP
  cote `jean link` sans repasser par un ecouteur TCP local.
- ui/index.html : base d'URL relative (API_BASE) + script marked relatif.
- main.go : cablage de la commande `link` + aide.
- README : section acces distant (ajean.link) + commande link.
- .gitignore : secrets runtime (.web_key/.link_token/.link_machine).

Inclut aussi les evolutions en cours du coeur de jean (llm.go, skills.go,
tools.go, chat.go).
This commit is contained in:
nathaninline committed 2026-06-20 20:11:32 +02:00
1 parent 6ea2e91810
commit e606206d90
12 files changed
+724 -91

No files matched your search

+3
View File
@@ -11,6 +11,9 @@
# Config locale / secrets runtime (générés à l'install, jamais versionnés)
config.env
.api_key
.web_key
.link_token
.link_machine
# OS / éditeurs
.DS_Store
+22 -1
View File
@@ -23,6 +23,7 @@ Faire tourner llama.cpp comme un vrai service, ça veut dire d'habitude : trouve
- **Intégration systemd** — `jean install` écrit l'unité, une règle sudoers `systemctl` sans mot de passe, et les dossiers de données
- **Interface web** (`jean web`) — chat, changement de modèle/preset, activation des skills & tools
- **Chat terminal** (`jean chat`) — réponses en streaming
- **Accès distant** (`jean link`) — connecte ce serveur à [ajean.link](https://ajean.link) par une connexion **sortante** (aucun port à ouvrir, marche même en CGNAT) : retrouve l'UI web et un endpoint **compatible OpenAI** depuis n'importe où
- **Presets** (`jean switch`) — garde plusieurs profils `config.env` et bascule entre eux
- **Protection par clé API** (`jean set-api-key`) — auth Bearer pour exposer le serveur publiquement ; la clé est stockée à part pour survivre aux changements de preset
- **Benchmark** (`jean bench`) — tok/s prefill/decode honnêtes avec un corpus varié
@@ -112,9 +113,13 @@ Interaction :
chat [system-prompt] chat terminal streamé
web [PORT] interface web (défaut :8090)
Accès distant (ajean.link) :
link <token> connecte ce Jean au relais (accès web + API OpenAI depuis partout)
link status | logout état du lien / oublier le token
Outils côté LLM :
skills [on|off|list] laisse le modèle lire SKILLS/<nom>/SKILL.md
tools [on|off|status] active run_shell (le modèle exécute des commandes shell)
machine [on|off|status] active l'accès machine (le modèle dispose d'un shell complet sur le serveur)
Backend (llama.cpp) :
llamacpp install clone + compile llama.cpp (détecte CUDA/ROCm/Metal/CPU), règle BIN
@@ -190,6 +195,22 @@ La clé est stockée dans `$JEAN_HOME/.web_key`, distincte de `.api_key` (pilota
> ⚠️ La clé voyage en clair en HTTP. Pour une exposition publique, place Jean derrière un reverse-proxy HTTPS (Caddy, nginx) ou un tunnel (Tailscale, Cloudflare Tunnel).
### Accès distant via ajean.link
Plutôt que d'exposer un port, `jean link` ouvre une connexion **sortante** vers le relais [ajean.link](https://ajean.link) : ton serveur reste injoignable depuis l'extérieur, mais tu y accèdes quand même depuis n'importe où — idéal derrière une box ou en CGNAT.
```bash
jean web # l'UI doit tourner localement
jean link <token> # token fourni sur ajean.link
```
Une fois lié, tu retrouves depuis le portail :
- l'**interface web** de ton serveur, à distance ;
- un **endpoint compatible OpenAI** (`https://ajean.link/oai/<machine>/v1`, clé = ton token) pour brancher n'importe quel outil (OpenCode, etc.) ;
- la gestion de **plusieurs serveurs** et d'**agents** depuis un tableau de bord.
C'est un service optionnel et payant ; tout le reste de Jean est et restera open source et gratuit.
### Variables d'environnement
| Variable | Signification | Défaut |
+5 -2
View File
@@ -72,8 +72,11 @@ func cmdChat(args []string) error {
case ev.Stats != nil:
stats = ev.Stats
case ev.ToolUsed != nil:
icon := "📖"
verb := "lecture du skill"
if ev.ToolUsed.Done || ev.ToolUsed.Typing {
break // résultat / frappe live affichés côté web ; en terminal on garde l'annonce seule
}
icon := "📚"
verb := "skill"
if ev.ToolUsed.Name == "run_shell" {
icon = "⚙️"
verb = "exécution"
+6 -1
View File
@@ -1,3 +1,8 @@
module github.com/nathaninline/jean
go 1.22
go 1.23
require (
github.com/coder/websocket v1.8.15
github.com/hashicorp/yamux v0.1.2
)
+4
View File
@@ -0,0 +1,4 @@
github.com/coder/websocket v1.8.15 h1:6B2JPeOGlpff2Uz6vOEH1Vzpi0iUz20A+lPVhPHtNUA=
github.com/coder/websocket v1.8.15/go.mod h1:NX3SzP+inril6yawo5CQXx8+fk145lPDC6pumgx0mVg=
github.com/hashicorp/yamux v0.1.2 h1:XtB8kyFOyHXYVFnwT5C3+Bdo8gArse7j2AQ0DA0Uey8=
github.com/hashicorp/yamux v0.1.2/go.mod h1:C+zze2n6e/7wshOZep2A70/aQU6QBRWJO/G6FT1wIns=
+228
View File
@@ -0,0 +1,228 @@
package main
// link.go — `jean link <token>` : connecte ce serveur Jean au relais public
// (ajean.link) par une connexion SORTANTE persistante, pour qu'un utilisateur
// y accède depuis n'importe où sans ouvrir de port (CGNAT, box, etc.).
//
// Principe : l'agent ouvre un WebSocket vers le relais, l'authentifie avec le
// token d'abonnement, puis multiplexe (yamux) ce lien unique en un stream par
// requête navigateur. Chaque stream est reverse-proxyfié vers le `jean web`
// local. Keepalive + reconnexion automatique avec backoff.
//
// Le token est fourni par la boutique à l'achat. Il est mémorisé dans
// $JEAN_HOME/.link_token pour que `jean link` (sans argument) reprenne la
// connexion.
import (
"context"
"crypto/rand"
"encoding/hex"
"fmt"
"io"
"net/http"
"net/http/httputil"
"net/url"
"os"
"path/filepath"
"strings"
"time"
"github.com/coder/websocket"
"github.com/hashicorp/yamux"
)
// defaultRelayURL est l'endpoint WebSocket du relais. Surchageable via
// $JEAN_LINK_URL (utile pour tester contre un relais local).
const defaultRelayURL = "wss://ajean.link/agent"
func linkTokenPath() string { return filepath.Join(JeanHome(), ".link_token") }
func linkMachinePath() string { return filepath.Join(JeanHome(), ".link_machine") }
// machineID returns a stable per-machine identifier, creating one on first use.
// Permet au relais de regrouper les connexions d'une même machine sous le compte.
func machineID() string {
if b, err := os.ReadFile(linkMachinePath()); err == nil {
if id := strings.TrimSpace(string(b)); id != "" {
return id
}
}
buf := make([]byte, 8)
_, _ = rand.Read(buf)
id := hex.EncodeToString(buf)
_ = os.MkdirAll(JeanHome(), 0o755)
_ = os.WriteFile(linkMachinePath(), []byte(id+"\n"), 0o600)
return id
}
// readLinkToken returns the saved subscription token, or "".
func readLinkToken() string {
b, err := os.ReadFile(linkTokenPath())
if err != nil {
return ""
}
return strings.TrimSpace(string(b))
}
func saveLinkToken(tok string) error {
if err := os.MkdirAll(JeanHome(), 0o755); err != nil {
return err
}
return os.WriteFile(linkTokenPath(), []byte(tok+"\n"), 0o600)
}
// relayURL resolves the relay WebSocket endpoint (env override → default).
func relayURL() string {
if u := strings.TrimSpace(os.Getenv("JEAN_LINK_URL")); u != "" {
return u
}
return defaultRelayURL
}
func cmdLink(args []string) error {
// Sous-commandes utilitaires.
if len(args) > 0 {
switch args[0] {
case "status":
tok := readLinkToken()
if tok == "" {
fmt.Println(yellow("[info]") + " aucun token enregistré — lance: jean link <token>")
return nil
}
fmt.Printf("%s token enregistré (%s…), relais: %s\n", green("[ok]"), tok[:min(8, len(tok))], relayURL())
return nil
case "logout":
_ = os.Remove(linkTokenPath())
fmt.Println(green("[ok]") + " token supprimé")
return nil
}
}
token := readLinkToken()
if len(args) > 0 && args[0] != "" {
token = strings.TrimSpace(args[0])
if err := saveLinkToken(token); err != nil {
return err
}
}
if token == "" {
return fmt.Errorf("aucun token. Usage: jean link <token> (token fourni à l'achat sur la boutique)")
}
fmt.Printf("%s connexion au relais %s …\n", cyan("[link]"), relayURL())
fmt.Printf(" (UI Jean + endpoint OpenAI servis dans le tunnel — pas besoin de 'jean web')\n")
// On construit le handler une seule fois ; il est servi à travers chaque tunnel.
handler := newLinkHandler()
backoff := time.Second
for {
err := runLinkSession(token, handler)
if err != nil {
fmt.Printf("%s lien perdu: %v — reconnexion dans %s\n", yellow("[link]"), err, backoff)
}
time.Sleep(backoff)
if backoff < 30*time.Second {
backoff *= 2
if backoff > 30*time.Second {
backoff = 30 * time.Second
}
}
}
}
// newLinkHandler construit le handler servi à travers le tunnel :
// - /v1/*, /health, /props, /metrics, /slots → llama-server local (endpoint
// compatible OpenAI, avec injection de la clé API locale) → permet de
// brancher OpenCode, Hermes, etc. sur ajean.link/oai/<machine>/v1
// - tout le reste → l'UI web de Jean (avec injection de la clé de pilotage)
func newLinkHandler() http.Handler {
web := withLocalAuth(newWebMux())
llama := &url.URL{Scheme: "http", Host: fmt.Sprintf("127.0.0.1:%d", LLMPort())}
lp := httputil.NewSingleHostReverseProxy(llama)
lp.FlushInterval = -1 // streaming SSE des complétions
apiKey := readAPIKey()
base := lp.Director
lp.Director = func(req *http.Request) {
base(req)
// Le client distant n'a pas la clé API de llama-server ; on l'injecte ici
// (l'auth réelle est faite par le relais via la clé de liaison du compte).
if apiKey != "" {
req.Header.Set("Authorization", "Bearer "+apiKey)
}
}
lp.ErrorHandler = func(w http.ResponseWriter, r *http.Request, e error) {
http.Error(w, "llama-server injoignable: "+e.Error(), http.StatusBadGateway)
}
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
p := r.URL.Path
if strings.HasPrefix(p, "/v1") || p == "/health" || p == "/props" || p == "/metrics" || strings.HasPrefix(p, "/slots") {
lp.ServeHTTP(w, r)
return
}
web.ServeHTTP(w, r)
})
}
// withLocalAuth injecte la clé de pilotage locale dans chaque requête arrivant
// par le tunnel. Le navigateur distant ne connaît que le token (vérifié par le
// relais) ; c'est ici, en local, qu'on satisfait l'auth de l'API web sans
// exposer la clé au client.
func withLocalAuth(next http.Handler) http.Handler {
webKey := readWebKey()
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if webKey != "" && r.Header.Get("Authorization") == "" {
r.Header.Set("Authorization", "Bearer "+webKey)
}
next.ServeHTTP(w, r)
})
}
// runLinkSession opens one WebSocket→yamux session and serves it until it dies.
// It blocks for the lifetime of the connection and returns the error that ended
// it (so the caller can reconnect).
func runLinkSession(token string, handler http.Handler) error {
ctx := context.Background()
dialCtx, cancel := context.WithTimeout(ctx, 20*time.Second)
defer cancel()
// Identifie la machine auprès du relais (id stable + hostname).
host, _ := os.Hostname()
dialURL := relayURL()
q := url.Values{"m": {machineID()}, "h": {host}}
if strings.Contains(dialURL, "?") {
dialURL += "&" + q.Encode()
} else {
dialURL += "?" + q.Encode()
}
c, _, err := websocket.Dial(dialCtx, dialURL, &websocket.DialOptions{
HTTPHeader: http.Header{"Authorization": {"Bearer " + token}},
})
if err != nil {
return fmt.Errorf("dial: %w", err)
}
// Pas de limite de taille : on fait passer des streams arbitraires (SSE).
c.SetReadLimit(-1)
conn := websocket.NetConn(ctx, c, websocket.MessageBinary)
// L'agent est le côté "serveur" yamux : c'est le relais qui ouvre un stream
// par requête navigateur, et nous on les accepte.
ycfg := yamux.DefaultConfig()
ycfg.EnableKeepAlive = true
ycfg.KeepAliveInterval = 25 * time.Second
ycfg.ConnectionWriteTimeout = 30 * time.Second
ycfg.LogOutput = io.Discard
sess, err := yamux.Server(conn, ycfg)
if err != nil {
return fmt.Errorf("yamux: %w", err)
}
defer sess.Close()
fmt.Printf("%s lien établi ✓\n", green("[link]"))
// sess implémente net.Listener : chaque Accept() = un stream = une requête.
// On sert directement l'UI Jean (le handler gère lui-même le flush SSE).
srv := &http.Server{Handler: handler}
return srv.Serve(sess)
}
+190 -35
View File
@@ -42,18 +42,22 @@ type ToolFunction struct {
// readSkillTool / runShellTool: OpenAI-shaped function definitions advertised
// to the model when the corresponding feature flag is on.
func readSkillTool() Tool {
// skillTool : un seul outil pour toute la gestion des skills (lecture +
// auto-gestion création/maj/suppression), ce qui évite d'envoyer trois schémas.
func skillTool() Tool {
return Tool{
Type: "function",
Function: ToolFunction{
Name: "read_skill",
Description: "Lit le contenu détaillé d'un skill listé dans le system prompt.",
Name: "skill",
Description: "Gère tes skills (guides Markdown réutilisables). action: read=lire, write=créer/remplacer, append=ajouter à la fin sans réécrire, delete=supprimer. content requis pour write/append (1re ligne = titre court #).",
Parameters: map[string]any{
"type": "object",
"properties": map[string]any{
"name": map[string]any{"type": "string", "description": "Nom exact du skill"},
"action": map[string]any{"type": "string", "enum": []string{"read", "write", "append", "delete"}},
"name": map[string]any{"type": "string", "description": "Nom du skill (alphanum, ._-)"},
"content": map[string]any{"type": "string", "description": "Contenu Markdown (write/append)"},
},
"required": []string{"name"},
"required": []string{"action", "name"},
},
},
}
@@ -64,12 +68,12 @@ func runShellTool() Tool {
Type: "function",
Function: ToolFunction{
Name: "run_shell",
Description: "Exécute une commande shell sur le serveur (bash). Retourne stdout, stderr et le code de sortie. Utilise-le pour inspecter le système, lancer des scripts, consulter des logs. Évite les commandes destructrices sauf si l'utilisateur le demande explicitement.",
Description: "Exécute une commande shell (bash) sur cette machine et retourne stdout, stderr et le code de sortie. Pour inspecter le système, lancer des scripts, lire des logs. Évite les commandes destructrices sauf demande explicite.",
Parameters: map[string]any{
"type": "object",
"properties": map[string]any{
"command": map[string]any{"type": "string", "description": "La commande bash à exécuter"},
"timeout": map[string]any{"type": "integer", "description": fmt.Sprintf("Timeout en secondes (défaut %d, max %d)", toolDefaultTimeout, toolMaxTimeout)},
"command": map[string]any{"type": "string", "description": "La commande bash"},
"timeout": map[string]any{"type": "integer", "description": fmt.Sprintf("Timeout s (défaut %d, max %d)", toolDefaultTimeout, toolMaxTimeout)},
},
"required": []string{"command"},
},
@@ -77,26 +81,34 @@ func runShellTool() Tool {
}
}
// InjectSkills prepends the lightweight skills directory message to msgs when
// skills are enabled. Merges with an existing system message if present.
// InjectSkills prepends context system messages to msgs: the machine briefing
// (when machine access is enabled) and the lightweight skills directory (when
// skills are enabled). Merges with an existing system message if present.
func InjectSkills(msgs []Message) []Message {
sp := skillsSystemPrompt()
if sp == "" {
var parts []string
if mp := machineSystemPrompt(); mp != "" {
parts = append(parts, mp)
}
if sp := skillsSystemPrompt(); sp != "" {
parts = append(parts, sp)
}
if len(parts) == 0 {
return msgs
}
prefix := strings.Join(parts, "\n\n")
if len(msgs) > 0 && msgs[0].Role == "system" {
existing, _ := msgs[0].Content.(string)
merged := append([]Message{{Role: "system", Content: sp + "\n\n" + existing}}, msgs[1:]...)
merged := append([]Message{{Role: "system", Content: prefix + "\n\n" + existing}}, msgs[1:]...)
return merged
}
return append([]Message{{Role: "system", Content: sp}}, msgs...)
return append([]Message{{Role: "system", Content: prefix}}, msgs...)
}
// EnabledTools returns the tools to advertise on the next inference call.
func EnabledTools() []Tool {
tools := []Tool{}
if skillsEnabled() && len(ListSkills()) > 0 {
tools = append(tools, readSkillTool())
if skillsEnabled() {
tools = append(tools, skillTool())
}
if toolsEnabled() {
tools = append(tools, runShellTool())
@@ -114,8 +126,58 @@ type StreamEvent struct {
Err error
}
type ToolUsedEvent struct {
Name string
Label string // user-visible summary (skill name or first ~80 chars of command)
Name string
Label string // user-visible summary (skill name or the command)
Result string // tool output (stdout/stderr/exit for run_shell, skill body for read_skill)
Done bool // false = call announced (command only); true = result is ready
Typing bool // true = command still being written (partial), no spinner yet
}
// previewArg pulls the (possibly incomplete) string value of key out of a
// streaming tool-call arguments JSON, so the UI can show the command being
// typed live. Best-effort: it tolerates a truncated tail and basic escapes.
func previewArg(args, key string) string {
i := strings.Index(args, "\""+key+"\"")
if i < 0 {
return ""
}
rest := args[i+len(key)+2:]
if j := strings.Index(rest, ":"); j >= 0 {
rest = rest[j+1:]
} else {
return ""
}
q := strings.Index(rest, "\"")
if q < 0 {
return ""
}
rest = rest[q+1:]
var b strings.Builder
for x := 0; x < len(rest); x++ {
c := rest[x]
if c == '\\' && x+1 < len(rest) {
switch rest[x+1] {
case 'n':
b.WriteByte('\n')
case 't':
b.WriteByte('\t')
case 'r':
case '"':
b.WriteByte('"')
case '\\':
b.WriteByte('\\')
default:
b.WriteByte(rest[x+1])
}
x++
continue
}
if c == '"' {
break
}
b.WriteByte(c)
}
return b.String()
}
// StatsEvent carries llama.cpp's per-completion timing (final chunk).
@@ -168,6 +230,10 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
// bubble works regardless of backend. The ik_llama.cpp fork already sends
// reasoning_content, in which case we leave content untouched.
reasoningOn := reasoningActive(ReadConfig()["REASONING"])
// When llama.cpp fails to parse a model-generated tool call (HTTP 500), we
// retry the same turn once with tools removed so the model answers in plain
// text from the tool results already gathered, instead of dying mid-chat.
disableTools := false
for iter := 0; iter < 8; iter++ {
payload := map[string]any{
"model": "jean",
@@ -175,8 +241,13 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
"stream": true,
"temperature": temperature,
}
if len(tools) > 0 {
if len(tools) > 0 && !disableTools {
payload["tools"] = tools
// The model sometimes emits parallel tool calls, which this llama.cpp
// build serialises as two concatenated JSON objects in one arguments
// string ("{...}{...}") and then fails to parse (HTTP 500). Forcing a
// single tool call per turn avoids that.
payload["parallel_tool_calls"] = false
}
body, _ := json.Marshal(payload)
url := fmt.Sprintf("http://localhost:%d/v1/chat/completions", LLMPort())
@@ -191,9 +262,36 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
cb(StreamEvent{Err: err})
return err
}
// A non-200 here (e.g. context window exceeded after several large tool
// outputs) is NOT valid SSE: without this check we'd scan an empty/HTML
// body, find no data lines, and return silently — the chat just stops
// with no answer. Surface the body so the cause is visible instead.
if resp.StatusCode != http.StatusOK {
b, _ := io.ReadAll(io.LimitReader(resp.Body, 2000))
resp.Body.Close()
msg := strings.TrimSpace(string(b))
if msg == "" {
msg = resp.Status
}
// Most common 500 here: llama.cpp couldn't parse a malformed tool call
// the model emitted. Retry the turn once without tools so it answers
// in plain text rather than leaving the chat dead.
if !disableTools && len(tools) > 0 {
disableTools = true
// Nudge the model to answer in plain text from what it already
// gathered, so it doesn't immediately re-emit a tool call that
// llama.cpp would again fail to parse.
messages = append(messages, Message{Role: "system", Content: "N'appelle plus d'outil. Réponds maintenant directement en français à partir des informations déjà obtenues."})
continue
}
err := fmt.Errorf("llama-server a renvoyé %d : %s", resp.StatusCode, msg)
cb(StreamEvent{Err: err})
return err
}
toolCalls := map[int]*ToolCall{}
assistantContent := strings.Builder{}
finishReason := ""
lastPreview := "" // last command preview emitted (to stream the typing)
// Per-completion reasoning-split state (see reasoningOn comment above).
sawReasoningField := false
thinkOpen := reasoningOn
@@ -246,6 +344,21 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
}
cur.Function.Arguments += tc.Function.Arguments
}
// Stream the command being typed: extract the partial value and
// emit it whenever it grows, so the UI shows it appear live.
if cur := toolCalls[0]; cur != nil {
key := "command"
if cur.Function.Name == "skill" {
key = "name"
}
if p := previewArg(cur.Function.Arguments, key); p != "" && p != lastPreview {
lastPreview = p
if !cb(StreamEvent{ToolUsed: &ToolUsedEvent{Name: cur.Function.Name, Label: p, Typing: true}}) {
aborted = true
break
}
}
}
continue
}
if ch.Delta.ReasoningContent != "" {
@@ -308,7 +421,10 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
return nil
}
if finishReason == "tool_calls" && len(toolCalls) > 0 {
// Treat any accumulated tool calls as a tool turn even if the backend set
// finish_reason to "stop" instead of "tool_calls" (some llama.cpp builds
// do this) — otherwise we'd skip execution AND skip answering.
if len(toolCalls) > 0 {
// 1. Append assistant message with tool_calls so the model sees its own decision next turn.
idxs := make([]int, 0, len(toolCalls))
for k := range toolCalls {
@@ -335,19 +451,52 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
for _, tc := range tcs {
var args map[string]any
_ = json.Unmarshal([]byte(tc.Function.Arguments), &args)
result := ""
// Derive the human label (command / skill name) up front so we can
// announce the call BEFORE running it — otherwise the UI shows
// nothing while a slow shell command runs and looks frozen.
label := ""
switch tc.Function.Name {
case "read_skill":
name, _ := args["name"].(string)
label = name
if c := SkillContent(name); c != "" {
result = c
} else {
result = fmt.Sprintf("[erreur] skill '%s' introuvable", name)
case "skill":
label, _ = args["name"].(string)
case "run_shell":
label, _ = args["command"].(string)
}
cb(StreamEvent{ToolUsed: &ToolUsedEvent{Name: tc.Function.Name, Label: label}})
result := ""
switch tc.Function.Name {
case "skill":
action, _ := args["action"].(string)
content, _ := args["content"].(string)
switch action {
case "read":
if c := SkillContent(label); c != "" {
result = c
} else {
result = fmt.Sprintf("[erreur] skill '%s' introuvable", label)
}
case "write":
if werr := SaveSkill(label, "", content); werr != nil {
result = "[erreur] " + werr.Error()
} else {
result = fmt.Sprintf("[ok] skill '%s' enregistré", label)
}
case "append":
if werr := AppendSkill(label, content); werr != nil {
result = "[erreur] " + werr.Error()
} else {
result = fmt.Sprintf("[ok] skill '%s' enrichi (append)", label)
}
case "delete":
if derr := DeleteSkill(label); derr != nil {
result = "[erreur] " + derr.Error()
} else {
result = fmt.Sprintf("[ok] skill '%s' supprimé", label)
}
default:
result = "[erreur] action skill inconnue: " + action
}
case "run_shell":
cmd, _ := args["command"].(string)
to := 0
switch v := args["timeout"].(type) {
case float64:
@@ -355,19 +504,25 @@ func runChat(messages []Message, temperature float64, cb ChatCallback) error {
case int:
to = v
}
label = cmd
if len(label) > 80 {
label = label[:80] + "…"
}
result = runShell(cmd, to)
result = runShell(label, to)
default:
result = "[erreur] outil inconnu: " + tc.Function.Name
}
cb(StreamEvent{ToolUsed: &ToolUsedEvent{Name: tc.Function.Name, Label: label}})
shown := result
if r := []rune(shown); len(r) > 4000 {
shown = string(r[:4000]) + "\n…[tronqué]"
}
cb(StreamEvent{ToolUsed: &ToolUsedEvent{Name: tc.Function.Name, Label: label, Result: shown, Done: true}})
messages = append(messages, Message{Role: "tool", ToolCallID: tc.ID, Content: result})
}
continue
}
// Normal end of turn. If the model produced no visible answer at all
// (empty content, e.g. it stopped right after a tool result), say so
// instead of leaving the user staring at a silent, finished chat.
if strings.TrimSpace(assistantContent.String()) == "" {
cb(StreamEvent{Content: "_(le modèle n'a pas produit de réponse — finish: " + finishReason + ")_"})
}
return nil
}
cb(StreamEvent{Content: "\n\n[stop: trop d'appels d'outils]"})
+8 -2
View File
@@ -38,9 +38,11 @@ func main() {
mustExit(cmdChat(args))
case "web":
mustExit(cmdWeb(args))
case "link":
mustExit(cmdLink(args))
case "skills":
mustExit(cmdSkills(args))
case "tools":
case "machine", "tools":
mustExit(cmdTools(args))
case "serve":
mustExit(cmdServe(args))
@@ -89,9 +91,13 @@ Interaction:
chat [system-prompt] chat terminal streamé
web [PORT] UI web (défaut :8090) — chat + presets + skills + tools
Accès distant (ajean.link) :
link <token> connecte ce Jean au relais public (accès web depuis partout, sans ouvrir de port)
link status | logout état du lien / oublier le token
LLM-side outils:
skills [on|off|list] active la lecture de SKILLS/<nom>/SKILL.md par l'IA
tools [on|off|status] active run_shell (l'IA exécute des commandes bash)
machine [on|off|status] active l'accès machine (l'IA dispose d'un shell complet sur le serveur)
Backend llama.cpp :
llamacpp install clone + compile llama.cpp (détecte CUDA/ROCm/Metal/CPU), pointe BIN dessus
+32 -7
View File
@@ -117,6 +117,32 @@ func SaveSkill(name, old, content string) error {
return os.WriteFile(filepath.Join(d, "SKILL.md"), []byte(content), 0o644)
}
// AppendSkill ajoute du contenu à la fin de SKILL.md (le crée si absent).
// C'est le mode "offset" : l'IA peut consigner une solution trouvée sans
// réécrire tout le skill.
func AppendSkill(name, content string) error {
d, err := safeSkillDir(name)
if err != nil {
return err
}
if err := os.MkdirAll(d, 0o755); err != nil {
return err
}
p := filepath.Join(d, "SKILL.md")
existing, _ := os.ReadFile(p)
var b strings.Builder
if len(existing) > 0 {
b.Write(existing)
if !strings.HasSuffix(string(existing), "\n") {
b.WriteString("\n")
}
b.WriteString("\n")
}
b.WriteString(strings.TrimRight(content, "\n"))
b.WriteString("\n")
return os.WriteFile(p, []byte(b.String()), 0o644)
}
func DeleteSkill(name string) error {
d, err := safeSkillDir(name)
if err != nil {
@@ -135,14 +161,13 @@ func skillsSystemPrompt() string {
return ""
}
list := ListSkills()
if len(list) == 0 {
return ""
}
var b strings.Builder
b.WriteString(`Tu as accès à des skills (guides spécialisés). Pour lire le contenu détaillé d'un skill quand c'est pertinent, appelle la fonction read_skill(name="<nom>"). N'appelle pas read_skill si la question est sans rapport avec un skill listé.
Skills disponibles :
`)
b.WriteString(`Skills = guides Markdown que tu gères toi-même via le tool skill(action, name, content) : read=lire, write=créer/remplacer, append=ajouter à la fin sans réécrire, delete=supprimer. De ta propre initiative, crée ou enrichis un skill quand tu découvres une solution réutilisable (problème enfin résolu, procédure, astuce) ; sa 1re ligne est un titre court (#) servant de description. N'utilise read que si la question concerne un skill listé.`)
if len(list) == 0 {
b.WriteString("\nAucun skill pour l'instant.")
return b.String()
}
b.WriteString("\nSkills :\n")
for _, s := range list {
fmt.Fprintf(&b, "- %s: %s\n", s.Name, s.Desc)
}
+37 -4
View File
@@ -5,6 +5,8 @@ import (
"fmt"
"os"
"os/exec"
"os/user"
"runtime"
"strings"
"time"
)
@@ -35,6 +37,37 @@ func setToolsEnabled(on bool) error {
return nil
}
// machineSystemPrompt returns a short briefing about the host the model is
// running on, so that when machine access is enabled it knows *which* machine
// run_shell acts upon (and doesn't claim it has no access to "your PC").
// Returns "" when machine access is off.
func machineSystemPrompt() string {
if !toolsEnabled() {
return ""
}
host, _ := os.Hostname()
if host == "" {
host = "inconnu"
}
who := ""
if u, err := user.Current(); err == nil {
who = u.Username
}
cwd, _ := os.Getwd()
var b strings.Builder
b.WriteString("Accès machine actif : le tool run_shell exécute des commandes sur cette machine — c'est ton accès réel, ne dis jamais que tu n'y as pas accès. Pour toute question système (OS, disque, process, fichiers, réseau), utilise-le au lieu de supposer.\n")
b.WriteString(fmt.Sprintf("Machine : hôte=%s, %s/%s", host, runtime.GOOS, runtime.GOARCH))
if who != "" {
b.WriteString(", user=" + who)
}
if cwd != "" {
b.WriteString(", cwd=" + cwd)
}
b.WriteString(".")
return b.String()
}
// runShell executes a command via the platform shell (bash -c on Unix, cmd /C
// on Windows — see newShellCmd in platform_*.go) with a clamped timeout,
// returning a single string formatted "exit: N\n\nstdout:\n...\n\nstderr:\n..."
@@ -95,21 +128,21 @@ func cmdTools(args []string) error {
if err := setToolsEnabled(true); err != nil {
return err
}
fmt.Println(green("[ok]") + " tool run_shell activé — exécution bash possible")
fmt.Println(green("[ok]") + " accès machine activé — l'IA dispose d'un shell complet sur le serveur")
case "off":
if err := setToolsEnabled(false); err != nil {
return err
}
fmt.Println(green("[ok]") + " tools désactivés")
fmt.Println(green("[ok]") + " accès machine désactivé")
case "", "status":
state := dim("off")
if toolsEnabled() {
state = green("on")
}
fmt.Printf("%s run_shell état: %s\n", cyan("Tools"), state)
fmt.Printf("%s état: %s\n", cyan("Accès machine"), state)
fmt.Printf(" timeout défaut: %ds, max: %ds\n", toolDefaultTimeout, toolMaxTimeout)
default:
return fmt.Errorf("usage: jean tools [on|off|status]")
return fmt.Errorf("usage: jean machine [on|off|status]")
}
return nil
}
+152 -17
View File
@@ -96,6 +96,14 @@ button:disabled{opacity:.5;cursor:not-allowed}
#scrollbtn{position:absolute;bottom:80px;right:20px;background:var(--panel);border:1px solid var(--border);color:var(--accent);width:36px;height:36px;border-radius:50%;cursor:pointer;display:none;font-size:16px;z-index:5}
#scrollbtn.show{display:block}
.msg.reasoning{align-self:flex-start;background:transparent;border:1px dashed var(--border);color:var(--mag);font-size:12px;max-width:90%}
.msg.tool{align-self:flex-start;background:var(--panel);border:1px solid var(--border);max-width:100%;white-space:normal;padding:8px 12px}
.msg.tool .tool-head{font-size:12px;color:var(--accent);font-weight:600;margin-bottom:6px}
.msg.tool .tool-sub{font-size:10px;color:var(--dim);text-transform:uppercase;letter-spacing:.1em;margin:8px 0 4px}
.msg.tool pre{position:relative;background:#0d1117;border:1px solid var(--border);border-radius:6px;padding:8px 10px;overflow:auto;margin:0;max-height:280px}
.msg.tool pre code{background:transparent;padding:0;font-size:.82em;line-height:1.4;white-space:pre}
.msg.tool .tool-caret{color:var(--accent);animation:toolblink 1s step-end infinite}
.msg.tool .tool-wait{font-size:12px;color:var(--dim);margin-top:8px;animation:toolblink 1.2s ease-in-out infinite}
@keyframes toolblink{0%,100%{opacity:.45}50%{opacity:1}}
.msg .label{font-size:10px;color:var(--dim);text-transform:uppercase;letter-spacing:.1em;margin-bottom:4px;display:block}
#inputbar{border-top:1px solid var(--border);padding:12px;display:flex;gap:8px}
#input{flex:1;background:var(--panel);color:var(--text);border:1px solid var(--border);border-radius:6px;padding:10px;font:inherit;resize:none;min-height:44px;max-height:200px}
@@ -126,9 +134,9 @@ button:disabled{opacity:.5;cursor:not-allowed}
<details><summary>Presets<button class="preset-add" onclick="event.preventDefault();event.stopPropagation();openPreset('')" title="ajouter">+</button></summary>
<div id="presets"></div>
</details>
<details><summary>Tools <span id="tools-badge" style="margin-left:auto;font-size:10px"></span></summary>
<label style="display:flex;align-items:center;gap:6px;font-size:12px;padding:4px 0"><input type="checkbox" id="tools-toggle" onchange="toggleTools()"> activer <code style="color:var(--warn);font-size:11px">run_shell</code></label>
<div class="muted" style="font-size:11px">⚠️ permet à l'IA d'exécuter des commandes bash sur le serveur (timeout 30s, max 300s)</div>
<details><summary>Accès machine <span id="tools-badge" style="margin-left:auto;font-size:10px"></span></summary>
<label style="display:flex;align-items:center;gap:6px;font-size:12px;padding:4px 0"><input type="checkbox" id="tools-toggle" onchange="toggleTools()"> activer l'accès machine</label>
<div class="muted" style="font-size:11px">⚠️ donne à l'IA un accès shell complet au serveur — elle peut inspecter, lancer des scripts et tout faire via le terminal (timeout 30s, max 300s)</div>
</details>
<details><summary>Skills <span id="skills-badge" style="margin-left:auto;font-size:10px"></span><button class="preset-add" onclick="event.preventDefault();event.stopPropagation();openSkill('')" title="ajouter">+</button></summary>
<label style="display:flex;align-items:center;gap:6px;font-size:12px;padding:4px 0"><input type="checkbox" id="skills-toggle" onchange="toggleSkills()"> activer</label>
@@ -155,6 +163,13 @@ button:disabled{opacity:.5;cursor:not-allowed}
<div class="main" style="position:relative">
<div id="chat"></div>
<button id="scrollbtn" onclick="jumpBottom()" title="aller en bas">↓</button>
<div id="ctxbar" style="padding:4px 12px 0;font-size:11px;display:flex;align-items:center;gap:8px">
<div style="flex:1;height:6px;background:var(--panel);border:1px solid var(--border);border-radius:4px;overflow:hidden">
<div id="ctx-fill" style="height:100%;width:0%;background:var(--ok,#3a7);transition:width .3s,background .3s"></div>
</div>
<span id="ctx-text" class="muted" style="white-space:nowrap">contexte 0 / 0</span>
<button id="ctx-compact" onclick="compactContext()" style="display:none;margin:0;padding:2px 10px;font-size:11px;background:#3d2f1f;border-color:var(--warn,#c93)">compacter</button>
</div>
<div id="inputbar">
<textarea id="input" placeholder="message…" onkeydown="onKey(event)"></textarea>
<button id="send" onclick="send()">send</button>
@@ -214,7 +229,7 @@ button:disabled{opacity:.5;cursor:not-allowed}
</div>
</div>
</div>
<script src="/marked.min.js"></script>
<script src="marked.min.js"></script>
<script>
// marked: GFM on, no auto-breaks (single newlines stay inline).
marked.setOptions({ gfm: true, breaks: false });
@@ -237,15 +252,19 @@ let msgs = [];
let busy = false;
function toggleSide(){ document.getElementById('side').classList.toggle('open'); document.getElementById('backdrop').classList.toggle('open'); document.body.classList.toggle('drawer-open'); }
function saveSys(){ localStorage.setItem('jean.sys', document.getElementById('sysprompt').value); }
document.addEventListener('DOMContentLoaded', ()=>{ document.getElementById('sysprompt').value = localStorage.getItem('jean.sys') || ''; });
document.addEventListener('DOMContentLoaded', ()=>{ document.getElementById('sysprompt').value = localStorage.getItem('jean.sys') || ''; restoreChat(); });
function toast(m){ const t=document.getElementById('toast'); t.textContent=m; t.classList.add('show'); setTimeout(()=>t.classList.remove('show'),1800); }
// Clé de pilotage : envoyée en Authorization: Bearer sur chaque appel /api/*.
// Sur un 401 on la (re)demande et on rejoue la requête. Stockée en localStorage.
let TOKEN = localStorage.getItem('jean.key') || '';
function authHeaders(h){ h = Object.assign({}, h||{}); if(TOKEN) h['Authorization']='Bearer '+TOKEN; return h; }
// Base d'URL : "" en local (page à "/"), "/u/<id>" servie via le relais ajean.link.
// Garde les chemins absolus (/api/…) corrects derrière le tunnel.
const API_BASE = location.pathname.replace(/\/(index\.html)?$/, '');
async function jfetch(u, opts){
opts = opts || {};
opts.headers = authHeaders(opts.headers);
if(u.charAt(0) === '/') u = API_BASE + u;
let r = await fetch(u, opts);
if(r.status === 401){
const k = prompt('Clé de pilotage jean requise :', '');
@@ -259,6 +278,20 @@ async function loadStatus(){
const s=await jget('/api/status');
const t=s.active?'<span class="tag ok">● active</span>':'<span class="tag err">○ '+s.state+'</span>';
document.getElementById('status-svc').innerHTML = t + ' <span class="muted">:'+s.port+'</span>';
if(s.ctx){ CTX_MAX=s.ctx; updateCtxMeter(); }
}
// Compteur de contexte : CTX_USED estimé via les stats serveur (prefill+decode
// du dernier tour ≈ taille du prochain prompt). À 90% on propose de compacter.
let CTX_MAX=0, CTX_USED=0;
function setCtxUsed(n){ CTX_USED=n||0; updateCtxMeter(); }
function updateCtxMeter(){
if(!CTX_MAX) return;
const pct=Math.min(100, Math.round(CTX_USED*100/CTX_MAX));
const fill=document.getElementById('ctx-fill');
fill.style.width=pct+'%';
fill.style.background = pct>=90 ? 'var(--err,#c44)' : pct>=70 ? 'var(--warn,#c93)' : 'var(--ok,#3a7)';
document.getElementById('ctx-text').textContent='contexte '+CTX_USED+' / '+CTX_MAX+' ('+pct+'%)';
document.getElementById('ctx-compact').style.display = pct>=90 ? 'inline-block' : 'none';
}
async function loadVram(){
const gpus=await jget('/api/vram');
@@ -349,7 +382,7 @@ async function loadTools(){
}
async function toggleTools(){
const on=document.getElementById('tools-toggle').checked;
if(on && !confirm("Activer run_shell ? L'IA pourra exécuter des commandes bash sur le serveur.")){ document.getElementById('tools-toggle').checked=false; return; }
if(on && !confirm("Activer l'accès machine ? L'IA aura un accès shell complet au serveur.")){ document.getElementById('tools-toggle').checked=false; return; }
await jpost('/api/tools/toggle',{on}); loadTools();
}
async function loadAll(){ await Promise.all([loadStatus(),loadVram(),loadCfg(),loadPresets(),loadSkills(),loadTools()]); }
@@ -645,6 +678,34 @@ function setLabel(el, text){ el.querySelector('.label').textContent = text; }
function bodyOf(el){ return el.querySelector('.body'); }
// Render markdown into a message body in place; safe because md() escapes HTML.
function renderBody(el, text){ const b=bodyOf(el); b.innerHTML = md(text); addCopyButtons(b); scrollMaybe(); }
// Render a tool call as its own conversation message: the command the model
// wrote, then the response it got back. textContent keeps it injection-safe.
function renderToolMsg(el, tu){
const isShell = tu.name==='run_shell';
setLabel(el, isShell ? 'terminal' : 'skill');
const body=bodyOf(el); body.innerHTML='';
const head=document.createElement('div'); head.className='tool-head';
head.textContent = isShell ? '⚙️ commande' : '📚 skill';
body.appendChild(head);
if(tu.label){
const pre=document.createElement('pre'); pre.className='tool-cmd';
const code=document.createElement('code'); code.textContent=tu.label;
if(tu.typing){ const car=document.createElement('span'); car.className='tool-caret'; car.textContent='▋'; code.appendChild(car); }
pre.appendChild(code); body.appendChild(pre);
}
const hasResult = tu.result!==undefined && tu.result!=='';
if(hasResult){
const sub=document.createElement('div'); sub.className='tool-sub'; sub.textContent='réponse';
body.appendChild(sub);
const pre=document.createElement('pre');
const code=document.createElement('code'); code.textContent=tu.result;
pre.appendChild(code); body.appendChild(pre);
} else if(!tu.done && !tu.typing){
const wait=document.createElement('div'); wait.className='tool-wait'; wait.textContent='⏳ exécution en cours…';
body.appendChild(wait);
}
addCopyButtons(body); scrollMaybe();
}
// Inject a "copier" button into every <pre> code block (idempotent).
function addCopyButtons(root){
root.querySelectorAll('pre').forEach(pre=>{
@@ -662,7 +723,63 @@ function addCopyButtons(root){
pre.appendChild(btn);
});
}
function resetChat(){ msgs=[]; document.getElementById('chat').innerHTML=''; toast('chat vidé'); }
function resetChat(){ msgs=[]; document.getElementById('chat').innerHTML=''; saveChat(); try{localStorage.removeItem('jean.ctxused');}catch(e){} setCtxUsed(0); toast('chat vidé'); }
// Compaction : on demande à l'IA un résumé de la conversation destiné à la
// reprendre dans une session neuve, puis on repart d'un contexte propre seedé
// avec ce résumé. Réduit drastiquement les tokens tout en gardant le fil.
async function compactContext(){
if(busy) return;
if(!msgs.length){ toast('rien à compacter'); return; }
if(!confirm('Compacter le contexte ? L\'IA va résumer la conversation puis repartir d\'une session propre basée sur ce résumé.')) return;
busy=true;
const btn=document.getElementById('ctx-compact'); btn.textContent='compaction…'; btn.disabled=true;
const instruction='Résume de façon structurée et complète TOUTE notre conversation jusqu\'ici, dans le but de poursuivre le travail dans une nouvelle session sans rien perdre d\'important : objectifs, décisions prises, état actuel, fichiers/commandes/éléments clés, et ce qu\'il reste à faire. Écris uniquement le résumé, sans préambule.';
let summary='';
try{
const sys=(document.getElementById('sysprompt').value||'').trim();
const base=[...msgs,{role:'user',content:instruction}];
const out = sys ? [{role:'system',content:sys},...base] : base;
const r=await jfetch('/api/chat',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({messages:out})});
const reader=r.body.getReader(); const dec=new TextDecoder(); let buf='';
while(true){
const {done,value}=await reader.read(); if(done) break;
buf+=dec.decode(value,{stream:true}); let i;
while((i=buf.indexOf('\n\n'))>=0){
const chunk=buf.slice(0,i); buf=buf.slice(i+2);
for(const line of chunk.split('\n')){
if(!line.startsWith('data:')) continue;
const data=line.slice(5).trim(); if(data==='[DONE]') continue;
try{ const o=JSON.parse(data); const d=(o.choices&&o.choices[0]&&o.choices[0].delta)||{}; if(d.content) summary+=d.content; }catch(e){}
}
}
}
}catch(e){ toast('échec compaction: '+e.message); busy=false; btn.textContent='compacter'; btn.disabled=false; return; }
summary=summary.trim();
if(!summary){ toast('résumé vide'); busy=false; btn.textContent='compacter'; btn.disabled=false; return; }
// Session propre seedée avec le résumé.
msgs=[{role:'user',content:'Résumé de notre session précédente (reprends le travail à partir de là) :\n\n'+summary},
{role:'assistant',content:'Compris, je reprends à partir de ce résumé.'}];
document.getElementById('chat').innerHTML='';
const note=addMsg('tool',''); note.querySelector('.label').textContent='contexte compacté';
renderBody(note, '🗜️ **Contexte compacté.** Résumé de reprise :\n\n'+summary);
saveChat();
setCtxUsed(0); try{localStorage.removeItem('jean.ctxused');}catch(e){}
busy=false; btn.textContent='compacter'; btn.disabled=false;
toast('contexte compacté');
}
// Persistance de la conversation : on garde user+assistant en localStorage pour
// survivre à un refresh (les bulles tool/reasoning sont éphémères, non stockées).
function saveChat(){ try{ localStorage.setItem('jean.chat', JSON.stringify(msgs)); }catch(e){} }
function restoreChat(){
let raw=null; try{ raw=localStorage.getItem('jean.chat'); }catch(e){}
if(!raw) return;
try{ msgs=JSON.parse(raw)||[]; }catch(e){ msgs=[]; return; }
for(const m of msgs){
if(m.role==='user') addMsg('user', m.content);
else if(m.role==='assistant'){ const el=addMsg('assistant',''); renderBody(el, m.content||''); }
}
const cu=parseInt(localStorage.getItem('jean.ctxused')||'0',10); if(cu>0) setCtxUsed(cu);
}
function onKey(e){ if(e.key==='Enter' && !e.shiftKey){ e.preventDefault(); send(); } }
let abortCtrl=null;
function stopGen(){ if(abortCtrl){ abortCtrl.abort(); toast('stop'); } }
@@ -674,8 +791,8 @@ async function send(){
document.getElementById('send').style.display='none';
document.getElementById('stop').style.display='inline-block';
addMsg('user', text);
msgs.push({role:'user',content:text});
let reasonEl=null, contentEl=null, fullContent='', fullReason='';
msgs.push({role:'user',content:text}); saveChat();
let reasonEl=null, contentEl=null, pendingToolEl=null, fullContent='', fullReason='';
let tokCount=0, reasonTokCount=0, t0=performance.now(), firstTok=0, statsTimer=null;
let serverStats=null;
const updateStats=(final)=>{
@@ -721,20 +838,33 @@ async function send(){
const d=(o.choices&&o.choices[0]&&o.choices[0].delta)||{};
if(d.stats){ serverStats = d.stats; }
if(d.tool_used){
const ne=addMsg('reasoning','');
const icon = d.tool_used.name==='run_shell' ? '⚙️' : '📖';
const verb = d.tool_used.name==='run_shell' ? 'exécution' : 'lecture du skill';
bodyOf(ne).textContent = icon+' '+verb+' : '+d.tool_used.label;
setLabel(ne, 'tool · '+d.tool_used.name);
// Close out the current assistant/reasoning bubbles so any text the
// model streams *after* the tool call starts a fresh bubble BELOW it
// (and isn't appended above, into the pre-call bubble).
contentEl=null; reasonEl=null;
const tu=d.tool_used;
// One bubble per call: the live "typing" updates, the "en cours"
// spinner, and the final result all re-render the SAME element.
if(!pendingToolEl) pendingToolEl=addMsg('tool','');
renderToolMsg(pendingToolEl, tu);
if(tu.done){
pendingToolEl=null;
// L'IA gère ses skills elle-même : rafraîchir le menu en direct
// dès qu'un skill est créé/modifié/supprimé (pas besoin de F5).
if(tu.name==='skill') loadSkills();
}
}
if(d.reasoning_content){
if(!reasonEl) reasonEl=addMsg('reasoning','');
// Start a fresh buffer per bubble — otherwise a reasoning block that
// follows a tool call re-renders ALL previously accumulated reasoning
// into the new bubble (duplicated dozens of times over a turn).
if(!reasonEl){ reasonEl=addMsg('reasoning',''); fullReason=''; }
if(!firstTok) firstTok=performance.now();
reasonTokCount++;
fullReason+=d.reasoning_content; renderBody(reasonEl, fullReason);
}
if(d.content){
if(!contentEl){ contentEl=addMsg('assistant',''); firstTok=performance.now(); }
if(!contentEl){ contentEl=addMsg('assistant',''); fullContent=''; firstTok=performance.now(); }
tokCount++;
fullContent+=d.content; renderBody(contentEl, fullContent);
}
@@ -742,7 +872,11 @@ async function send(){
}
}
}
msgs.push({role:'assistant',content:fullContent});
msgs.push({role:'assistant',content:fullContent}); saveChat();
// Le KV cache s'accumule : chaque tour ajoute les tokens NOUVELLEMENT traités
// (prompt_tokens, le préfixe caché n'est pas recompté) + ceux générés. La
// taille réelle du contexte = somme cumulée, donc monotone croissante.
if(serverStats){ setCtxUsed(CTX_USED + (serverStats.prompt_tokens||0) + (serverStats.gen_tokens||0)); try{ localStorage.setItem('jean.ctxused', CTX_USED); }catch(e){} }
}catch(e){
if(e.name==='AbortError'){
if(contentEl) setLabel(contentEl, 'assistant · '+tokCount+' tok · interrompu');
@@ -751,6 +885,7 @@ async function send(){
} else {
addMsg('assistant','[erreur] '+e.message); msgs.pop();
}
saveChat();
}
if(statsTimer) clearInterval(statsTimer);
updateStats(true);
+37 -22
View File
@@ -28,6 +28,34 @@ func cmdWeb(args []string) error {
}
port = n
}
mux := newWebMux()
addr := fmt.Sprintf("0.0.0.0:%d", port)
ln, err := net.Listen("tcp", addr)
if err != nil {
// Port occupé : on identifie le process qui le tient et on propose de
// le terminer pour relancer à sa place.
if !resolvePortConflict(port) {
return err
}
if ln, err = net.Listen("tcp", addr); err != nil {
return err
}
}
fmt.Printf("[jean web] http://%s (Ctrl-C pour arrêter)\n", addr)
if readWebKey() == "" {
fmt.Printf("%s API de pilotage NON protégée (aucune clé). Avant de l'exposer sur internet :\n", yellow("[!]"))
fmt.Printf(" %s\n", bold("jean set-web-key"))
} else {
fmt.Printf("%s API protégée par clé (Authorization: Bearer …)\n", green("[ok]"))
}
return http.Serve(ln, mux)
}
// newWebMux construit le routeur HTTP de l'UI web. Extrait de cmdWeb pour être
// réutilisé par `jean link`, qui sert ce même mux à travers le tunnel sans
// repasser par un écouteur TCP local.
func newWebMux() *http.ServeMux {
mux := http.NewServeMux()
// Pages publiques : le HTML et le JS ne contiennent aucun secret. Toute la
// donnée et toutes les actions passent par /api/* qui, lui, exige la clé.
@@ -67,27 +95,7 @@ func cmdWeb(args []string) error {
api("/api/bench", handleBench)
api("/api/bench/last", handleBenchLast)
api("/api/chat", handleChat)
addr := fmt.Sprintf("0.0.0.0:%d", port)
ln, err := net.Listen("tcp", addr)
if err != nil {
// Port occupé : on identifie le process qui le tient et on propose de
// le terminer pour relancer à sa place.
if !resolvePortConflict(port) {
return err
}
if ln, err = net.Listen("tcp", addr); err != nil {
return err
}
}
fmt.Printf("[jean web] http://%s (Ctrl-C pour arrêter)\n", addr)
if readWebKey() == "" {
fmt.Printf("%s API de pilotage NON protégée (aucune clé). Avant de l'exposer sur internet :\n", yellow("[!]"))
fmt.Printf(" %s\n", bold("jean set-web-key"))
} else {
fmt.Printf("%s API protégée par clé (Authorization: Bearer …)\n", green("[ok]"))
}
return http.Serve(ln, mux)
return mux
}
// resolvePortConflict identifies the process listening on `port`, asks the user
@@ -200,11 +208,18 @@ func handleStatus(w http.ResponseWriter, r *http.Request) {
if active {
health = healthCheck()
}
ctx := 32768
if v := ReadConfig()["CTX"]; v != "" {
if n, err := strconv.Atoi(v); err == nil && n > 0 {
ctx = n
}
}
sendJSON(w, 200, map[string]any{
"state": state,
"active": active,
"health": health,
"port": LLMPort(),
"ctx": ctx,
})
}
@@ -584,7 +599,7 @@ func handleChat(w http.ResponseWriter, r *http.Request) {
return true
}
if ev.ToolUsed != nil {
emit(map[string]any{"tool_used": map[string]any{"name": ev.ToolUsed.Name, "label": ev.ToolUsed.Label}})
emit(map[string]any{"tool_used": map[string]any{"name": ev.ToolUsed.Name, "label": ev.ToolUsed.Label, "result": ev.ToolUsed.Result, "done": ev.ToolUsed.Done, "typing": ev.ToolUsed.Typing}})
return true
}
if ev.Stats != nil {