AI practice & chat¶
Practice and Chat talk to the Soju backend (OpenAI-compatible API),
not directly to Ollama. uv run poe up publishes Vite :14321, API :14322,
and docs :14323. Prod (poe up-prod) publishes only nginx on :8080 —
UI at /, API at /api/, docs at /docs/ — see Docker.
Set PUBLIC_AI_ENABLED=true (default in Compose).
Host Ollama (desktop app) — usual setup:
uv run poe up
# or: uv run poe up-prod
Pull models on the host (defaults match backend YAML / ollama-pull):
uv run poe setup-ollama
# or: ollama pull gemma4:e4b && ollama pull nomic-embed-text
The backend reaches host Ollama at http://host.docker.internal:11434
(see docker/soju/backend.yaml). The browser calls PUBLIC_AI_BASE_URL
(default http://localhost:14322 in poe up; /api in poe up-prod).
Ollama in Docker Compose:
docker compose --profile ollama up ollama ollama-pull web
The ollama services mount your host Ollama cache (default ~/.ollama).
Override with OLLAMA_DATA_DIR if needed.
Browser env (Compose)¶
Variable |
Values |
Purpose |
|---|---|---|
|
|
Show Practice & Chat |
|
e.g. |
Soju API root (dev direct / prod via nginx) |
|
|
Default speech engine (Settings can override) |
Chat model, embed model, tutor name/prompt, and TTS voice are loaded at runtime from
GET /v1/soju/config/client (backend YAML). Legacy PUBLIC_OLLAMA_* / older
PUBLIC_AI_MODEL names remain as fallbacks when set.
Backend config (YAML)¶
Edit packaged defaults or overrides:
Packaged:
src/soju/backend/config/files/default_config.yamlUser:
~/.config/soju/backend.yamlCompose:
docker/soju/backend.yaml(prod) ordocker/soju/backend.dev.yaml(dev)
Typical keys: llm.chat_model, llm.embed_model, llm.base_url,
client.system_prompt, client.tutor_name, tts.engine (edge / piper),
tts.voice.
Run the API on the host (optional):
uv run soju backend --config docker/soju/backend.yaml
See soju backend.
Practice generation¶
Practice builds a level- and theme-grounded session (one exercise type at a time) instead of dumping the full registry into the prompt.
Content
Levels come from
data/content/levels.yaml(guidance + optionalgrammar_summaryandinclude_levelsfor parent bands).Vocabulary and grammar in the embedding cache carry optional
level. Tagged entries match the selected course band (including parents viainclude_levels). Unassigned entries (missing/nulllevel) are excluded unless the Practice UI checkbox Include supplemental content is enabled (includeUnassigned).Themes come from
data/content/practice/themes.yaml(café, directions, family, daily routine, shopping). The UI also accepts a custom free-text theme.
Embedding index (offline CLI)
Before Generate can retrieve vocabulary, build the cache once (and again after large registry/grammar changes):
uv run poe embed-index
This writes data/cache/embeddings/ (gitignored). See soju embed-index.
Python uses OLLAMA_HOST + SOJU_EMBED_MODEL (default nomic-embed-text).
Keep that model in sync with backend llm.embed_model / /v1/soju/config/client.
Retrieve (dev-only API)
Flow on Generate session:
Browser embeds the theme via Soju
/v1/embeddings.Browser
POSTs{ level, queryVector, includeUnassigned? }to/api/practice/retrieve.The SvelteKit endpoint (dev server only) filters the cache by course band (and optionally unassigned vocab and grammar), ranks by cosine similarity, and returns hangul + grammar snippets.
Browser calls chat completions through the Soju API with that RAG payload.
/api/practice/retrieve is dev-only (same gate as /api/staging): it runs under
vite dev / docker compose up, not in the static production build. Missing or
corrupt cache returns 503 with a hint to run poe embed-index.
UI controls
Level → Theme (or Custom) → Exercise type (sentences / questions / fill-in-blank /
story / vocabulary candidates) → Count → optional Include supplemental content →
Generate. Results show the active type; any vocabulary_candidates are listed as
hangul — english (display only; no add-to-vocab).