Skip to content

Machine Learning

AI is everywhere in Kaba but not in the way

Kaba runs small models on your own devices, personalizes them with adapters trained on your memories, and puts a policy between every model and everything it can touch. All of it is optional.

Kaba harness

Settings → Features → Enable AI. It is off by default. Turning it on adds:

  • the AI actions in the right-click menu,
  • Projects and Toolbench,
  • the Policies and Models & Adapters pages in Settings.

The first time something needs a model that is not installed, Kaba tells you which one and how big it is, and waits for you to press the button. Nothing downloads on its own. Messages you send in the meantime are held until the model is ready.

To free memory at any time, right-click the header and choose Unload System Model.

  • /projects or /ask in the Omnibar (local ai chat)
  • super + enter in the Omnibar, for an answer in a toast
  • right click menu

Details for the first two are under omnibar; the right click menu is next.

Select text and right-click. The actions run on your device and answer in a toast; from the toast you can continue in a project.

GroupActions
AnalysisExplain, Fallacies of Logic, Synthesize, Reverse Engineer
QuestionGenerate Questions, Bull Case, Bear Case, If true, then?
RespondRebuke, Agree, Auto Debate
DebugLinux Debug, Code Debug
FormatProofread, Summarize, Rewrite
ConvertJSON, YAML, HTML, Go, Rust, Lua, JavaScript, Ruby, C
Writing HelpersFix Spelling & Grammar, Change with Prompt…

In a text field with nothing selected, Writing Helpers write into the field instead: Write What Belongs Here, Write with Prompt…, Continue Writing, Make It Shorter, Make It More Formal, Make It Friendlier, Draft a Reply to This Page, and Summarize This Page Here.

[todo: add screenshot of the AI context menu and of parts the Analysis, Question, Respond, Debug, Format, Convert and Writing Helpers groups, and an answer toast.]

A policy bundles every decision about a model in one named, reviewable object: where it runs, which model and adapter answer, how it generates, and what it is allowed to do. One policy is active for your account; a project can choose another.

Settings → Policies lists them. Create one with Create Policy, or Import policy from a .kabap bundle.

[todo: add screenshot of the Policies list and of parts Active marker, Name, Inference target, Create Policy, Import policy.]

Where it runs

SettingOptions
Inference targetThis device, or a peer in your cluster.
Sandbox target & containerWhich device and which container image all file, code and execution work uses. Only images imported on that target are listed.

What answers

SettingOptions
Base modelDefault (the device decides), small (about 2B, fast), large (about 4B, stronger for code, more memory), or an imported model.
LoRA adapterNone, or one of the adapters on the target. For a peer: the peer’s default, or a named one.
System promptOptional. Empty uses the model’s own default, or the built-in Kaba persona.

How it generates

SettingNotes
Tune forPresets: Tool work, Chat & information, Writing help, Code generation, Brainstorming & ideas, or Custom.
Context windowAuto sizes each request to fit, up to 32768 tokens. Auto-slide starts tool runs at 8192 and grows only when needed.
Max output tokensDefault 1024. Up to half the context window.
SeedOff for chat. Set it, with temperature 0, for tool work and benchmarking, so the same prompt gives the same answer.
Turbo cacheCompresses the model’s working memory on the serving node so larger contexts and more simultaneous runs fit, at a small quality cost. Off by default.
GPU offloadAuto fits layers to free graphics memory. Or CPU only, half, or the whole model.
Expert-model prefillFor mixture-of-experts models: Auto, Prefer speed or Prefer stability.

What it may do

SettingOptions
Command executionAsk (propose, you click Run), Auto-run (read-only allowlisted commands run immediately), or Off.
Sandbox networkingOff, Model may request, or Always on.
Max tool steps per runThe ceiling for one run, up to 400.
Workflow choiceAutomatic, Standard or Testing.
Steer controlHow the coding loop recovers when a step goes wrong. All on by default.
Prevent overriding policy settingsProjects under this policy run it exactly as configured. Their own controls are disabled and saved overrides are ignored.

Tool-by-tool permissions are edited in Toolbench, opened from the policy with Toolbench →. See projects & tool loop.

[todo: add screenshot of the policy editor and of parts Inference target, Base model, LoRA adapter, Context window, Generation settings, Command execution, Sandbox networking, Prevent overriding policy settings.]

A policy’s denied tools and its step limit are applied last, and nothing below can undo them. A project can narrow what a policy allows. It cannot widen it.

Export writes a .kabap file: the policy’s settings, its tool overlay and its flows. It does not include model files, and it leaves out anything that names a specific machine. Whoever imports it chooses where it runs. See file formats.

Settings → Models & Adapters is the library for one device. Use the Device selector to look at a peer’s.

Two layers personalize inference:

  • Models are whole base weights in GGUF format.
  • Adapters are LoRA deltas applied on top of a base.

Which one actually serves a request is chosen by the policy. This page is where they are stored.

Import a model by pasting a Hugging Face repository URL or a direct .gguf link. When a repository has several quantizations, the best fit is picked automatically. Any GGUF with a chat template works:

  • Gemma-family models add local training and image and audio input.
  • Other architectures are served through their own chat template, text only.

One model can be marked the node default. It serves any request that does not name a model. Clear it to fall back to the stock weights.

Adapters come from three places: training on your memories, importing by URL, or browsing Hugging Face from the page. The list shows each adapter’s name, base and format.

Adapters pair with a base by architecture family: a Gemma adapter for Gemma models, a Qwen adapter for a Qwen import. Policies filter the choices automatically, and the engine verifies the shapes before loading one.

[todo: add screenshot of Models & Adapters and of parts the Device selector, the model list with node default, Import model, the adapter list (Name, Base, Format), Import adapter, Browse HuggingFace.]

Training turns a selection of your memories into a LoRA adapter on top of the base model. You can start it from the client or from the command line.

Start it from Settings → Devices: open a device’s menu and choose to train there, on this device or a peer with a GPU.

Kaba remote compute

The training dialog asks for:

Section
Your profileYour name, a line about you, interests, preferred tone, and anything to always keep in mind. This is built into the adapter so answers read as your own recall.
Date rangeAll records, the last period (such as 7d or 24h), or a range.
Only these domains / Exclude domainsSuffix match, for example nature.com.
TopicsSemantic match against each memory, with a threshold (default 0.30).
KeywordsLiteral match in title, address or text.

All filtering happens on this device. When you train on a peer, the peer receives only the filtered set.

While a run is in progress, the device’s row shows its progress, with Pause, Resume and Stop. When it finishes, the adapter appears under Models & Adapters and can be chosen in a policy.

Enrich memories… in the same menu runs the summarisation and tagging pass on its own, which fills in categories and tags in Hippocampus.

What goes in is shaped by what you did in Hippocampus: boosted and suppressed memories, your thoughts, and the ignore list.

[todo: add screenshot of the training dialog and of parts Your profile, Date range, domain filters, Topics with threshold, Keywords; and a device row showing training progress with Pause and Stop.]

Manually Triggering Kaba’s Training Task

Section titled “Manually Triggering Kaba’s Training Task”

To manually trigger Kaba’s Language Model (LLM) training task, you can use the following command:

Terminal window
~/.config/kaba/bin/kabactl train --help

This command provides information about various options available for Kaba’s training task. Here is a brief explanation of the most relevant options:

  • --workspace: The account (your email) whose encrypted memories to train on. Without it, training opens the global store, not your saved memories.
  • --since, --until, --last: Only train on memories in a window. Accepts YYYY-MM-DD, an ISO instant, or a duration like 7d or 24h.
  • --min-records: Minimum number of memories required before training will start.
  • --node: Train on a cluster peer instead of this machine.
  • --output-name: Name for the resulting adapter.
  • --epochs: Number of passes over the data.
  • --lr: Peak learning rate.
  • --warmup-ratio: Fraction of steps used to warm the learning rate up.
  • --lora-rank: Size of the adapter.
  • --batch-size: Number of samples per training step. Directly impacts VRAM usage.
  • --seq-len: Number of tokens the model processes at once (Context length).
  • --grad-accum: Number of batches to aggregate before updating model weights.
  • --no-qlora: Train without quantization.
  • --skip-summaries: Skip the summarisation pass.

Local training expects the kabactl server to be stopped, and needs a GPU. The full list, with defaults, is in the CLI reference.

The above is the manual process for training Kaba’s personal language models. If configured to do so, Kaba will also run this nightly to have rolling weights that continue to train on your memory as it progresses autonomously.

[todo: confirm how nightly rolling training is configured in kaba 0.146 / kabactl 0.87 and document the setting.]

  • run nightly
  • updates
  • etc

[todo: write the Autonomous section: what runs on its own (memory capture, enrichment, scheduled training), how often, and how to turn each off.]

Every request goes to kabactl, on this device or on the peer the policy names. Prompts, outputs and adapters do not go to a cloud service. On a peer, that peer’s own settings decide whether it accepts the work at all. See cluster & mesh.

Everything on this page has a command-line form: kabactl get, ask, train, loras and lora-import. See models & training.