Skip to content

Models & Training

Kaba runs small models on your own hardware and personalizes them with LoRA adapters trained on your memories. This page covers the command-line side. For the same things in the client, see machine learning.

  • A base model is a whole set of weights in GGUF format. kabactl ships with two variants of Gemma: e2b (about 2B effective parameters, the default) and e4b (about 4B, stronger for code, needs more memory). Other GGUF models with a chat template can be imported from the client.
  • An adapter is a small LoRA delta applied on top of a base. Adapters are what training produces. They pair with a base by architecture family.

Both live in <storage>/kaba-engine/; adapters are in kaba-engine/loras/.

Terminal window
kabactl get --model e2b

This is the fastest path on a new machine: one serving file (about 2.8 GB for e2b, 4.4 GB for e4b). It is enough to ask questions and to serve adapters.

To train locally you need the full preparation, which also downloads the original weights and converts them:

Terminal window
kabactl get

Downloads try a mirror first and fall back to Hugging Face. Set HF_TOKEN if your account needs it. All get options are in the CLI reference.

The client does this for you: the first time a policy needs a model that is not present, it offers to download it and holds your messages until it is ready.

Terminal window
# base model, no memory
kabactl ask "Explain QUIC in two sentences"
# grounded on your memories
kabactl ask "What did I read about QUIC last week?" --workspace you@example.com
# on another device
kabactl ask "Review this function" --peer gpu-tower --base-model e4b

Tokens stream to stdout, so ask composes with other tools:

Terminal window
kabactl ask "Write a commit message for this diff: $(git diff --staged)" > msg.txt

Local ask runs the engine inside the command and expects kabactl server to be stopped. If the server is running, use the client or --peer instead.

Training turns a window of your memories into an adapter.

Terminal window
kabactl train \
--workspace you@example.com \
--last 30d \
--output-name october
  1. Memories in the window are selected. Pages and domains on your Memory Ignore List are excluded, and votes you gave in Hippocampus (Boost or Suppress in training) weight the examples.
  2. Memories are summarised into training examples (skip with --skip-summaries).
  3. The adapter is trained and written to kaba-engine/loras/october.mpk with a manifest october.json.

Training refuses to start with fewer than default_min_records memories (200 by default). Local training needs a build with GPU support.

Terminal window
kabactl train --workspace you@example.com --node gpu-tower

Your corpus is streamed to the peer, progress is reported back, and the finished adapter is pulled to this device. Date filters and hyperparameters apply on the peer. From the client, a training run on a peer can be paused, resumed or stopped from Settings → Devices.

Terminal window
kabactl loras # what is installed here
kabactl loras --peer gpu-tower # what a peer has
kabactl ask "…" --lora october

To make an adapter the default for a node, set lora = "october" under kaba_engine_options. In the client, pick it in a policy (Settings → Policies → LoRA adapter).

On a peer, the peer’s policy decides which adapter answers unless you name one: --lora <name> selects one present on the peer, and --lora "" forces its base model.

Terminal window
kabactl lora-import https://huggingface.co/<owner>/<repo>
kabactl lora-import https://example.com/adapter.gguf --base-model e4b

The download is converted so kabactl can run it and a manifest is written next to it. The original file is kept as provenance.

These environment variables are read by the node that runs the model.

VariableEffect
KABA_BACKENDForce the backend: auto, cuda, rocm, wgpu, cpu.
KABA_LLAMA_GPU_LAYERSCap how many layers go to the GPU, whatever a request asks for.
KABA_LLAMA_MAX_CTXUpper bound on the context window.
KABA_LLAMA_THREADSCPU threads for inference.
KABA_MODEL_IDLE_SECSUnload the model after this long idle.
KABA_INFER_MAX_SECSHard limit on one generation.

More are listed in the environment reference.