Models & Training
Kaba runs small models on your own hardware and personalizes them with LoRA adapters trained on your memories. This page covers the command-line side. For the same things in the client, see machine learning.
Two layers
Section titled “Two layers”- A base model is a whole set of weights in GGUF format. kabactl ships with two variants of Gemma:
e2b(about 2B effective parameters, the default) ande4b(about 4B, stronger for code, needs more memory). Other GGUF models with a chat template can be imported from the client. - An adapter is a small LoRA delta applied on top of a base. Adapters are what training produces. They pair with a base by architecture family.
Both live in <storage>/kaba-engine/; adapters are in kaba-engine/loras/.
Get a model
Section titled “Get a model”kabactl get --model e2bThis is the fastest path on a new machine: one serving file (about 2.8 GB for e2b, 4.4 GB for e4b). It is enough to ask questions and to serve adapters.
To train locally you need the full preparation, which also downloads the original weights and converts them:
kabactl getDownloads try a mirror first and fall back to Hugging Face. Set HF_TOKEN if your account needs it. All get options are in the CLI reference.
The client does this for you: the first time a policy needs a model that is not present, it offers to download it and holds your messages until it is ready.
# base model, no memorykabactl ask "Explain QUIC in two sentences"
# grounded on your memorieskabactl ask "What did I read about QUIC last week?" --workspace you@example.com
# on another devicekabactl ask "Review this function" --peer gpu-tower --base-model e4bTokens stream to stdout, so ask composes with other tools:
kabactl ask "Write a commit message for this diff: $(git diff --staged)" > msg.txtLocal ask runs the engine inside the command and expects kabactl server to be stopped. If the server is running, use the client or --peer instead.
Train an adapter
Section titled “Train an adapter”Training turns a window of your memories into an adapter.
kabactl train \ --workspace you@example.com \ --last 30d \ --output-name october- Memories in the window are selected. Pages and domains on your Memory Ignore List are excluded, and votes you gave in Hippocampus (Boost or Suppress in training) weight the examples.
- Memories are summarised into training examples (skip with
--skip-summaries). - The adapter is trained and written to
kaba-engine/loras/october.mpkwith a manifestoctober.json.
Training refuses to start with fewer than default_min_records memories (200 by default). Local training needs a build with GPU support.
Train on a peer
Section titled “Train on a peer”kabactl train --workspace you@example.com --node gpu-towerYour corpus is streamed to the peer, progress is reported back, and the finished adapter is pulled to this device. Date filters and hyperparameters apply on the peer. From the client, a training run on a peer can be paused, resumed or stopped from Settings → Devices.
Use an adapter
Section titled “Use an adapter”kabactl loras # what is installed herekabactl loras --peer gpu-tower # what a peer haskabactl ask "…" --lora octoberTo make an adapter the default for a node, set lora = "october" under kaba_engine_options. In the client, pick it in a policy (Settings → Policies → LoRA adapter).
On a peer, the peer’s policy decides which adapter answers unless you name one: --lora <name> selects one present on the peer, and --lora "" forces its base model.
Import an adapter
Section titled “Import an adapter”kabactl lora-import https://huggingface.co/<owner>/<repo>kabactl lora-import https://example.com/adapter.gguf --base-model e4bThe download is converted so kabactl can run it and a manifest is written next to it. The original file is kept as provenance.
Tuning the serving node
Section titled “Tuning the serving node”These environment variables are read by the node that runs the model.
| Variable | Effect |
|---|---|
KABA_BACKEND | Force the backend: auto, cuda, rocm, wgpu, cpu. |
KABA_LLAMA_GPU_LAYERS | Cap how many layers go to the GPU, whatever a request asks for. |
KABA_LLAMA_MAX_CTX | Upper bound on the context window. |
KABA_LLAMA_THREADS | CPU threads for inference. |
KABA_MODEL_IDLE_SECS | Unload the model after this long idle. |
KABA_INFER_MAX_SECS | Hard limit on one generation. |
More are listed in the environment reference.