CLI Reference
kabactl [GLOBAL OPTIONS] <COMMAND>| Command | Purpose |
|---|---|
server | Run the server in the foreground. |
service | Run the server as a background daemon, or install it. |
cluster | Invite, join, list and evict peers. |
get | Download base-model weights and the voice bundle. |
ask | Run inference locally or on a peer. |
train | Fine-tune a LoRA adapter on your memories. |
loras | List adapters. |
lora-import | Import a third-party adapter by URL. |
security | Manage the local Safe Browsing database. |
doctor | Health-check the installation. |
self-update | Replace the binary with the latest release. |
Global options
Section titled “Global options”| Option | Description |
|---|---|
--trace | Crash and memory diagnostics: heap-profile stack traces (Linux), memory sampling, full panic backtraces and raised core-dump limits. Writes a bundle to <storage>/perf/<timestamp>/ and points perf/CURRENT at it. Same as KABA_TRACE=1, which the service daemon inherits. |
-c, --config-path <FILE> | Path to config.toml. |
-v / -q | Raise or lower verbosity. |
-V, --version | Print the version. |
-h, --help | Print help. |
server
Section titled “server”Runs the full server in the foreground: HTTPS API, SOCKS5 proxy and the cluster endpoint. This is what the client starts, and what the systemd units and the container image run.
kabactl server [--device <BACKEND>]| Option | Default | Description |
|---|---|---|
--device <BACKEND> | auto | Inference backend preference: auto, cuda (alias nvidia), rocm (hip, amd), wgpu (gpu, vulkan, metal) or cpu (ndarray). CPU is always the final fallback. KABA_BACKEND overrides this flag. |
-b, --bind <ADDRESS> | 0.0.0.0 | See the caution below. |
--port <PORT> | 28832 | See the caution below. |
Only wgpu and cpu are compiled into a default build. Requesting cuda or rocm on a build without them logs a warning and falls through to the next backend.
service
Section titled “service”Lifecycle controls for the same server, detached from your terminal.
| Subcommand | Description |
|---|---|
service start | Daemonize the server. Fails if one is already running. |
service stop | Send SIGTERM, wait, then SIGKILL if it has not exited. The wait is KABA_STOP_GRACE_SECS. |
service restart | Stop if running, then start. If the stop fails, no replacement is started. |
service status | Report running (with PID), stale PID file, or not running. |
service join <TICKET> | Join a cluster using a ticket. Same as cluster join. |
service install | Copy this binary to the managed location and put it on PATH. |
service uninstall | Stop the service and remove the unit, symlink and managed binary. |
service install options:
| Option | Description |
|---|---|
--system | System-wide install to /usr/local/bin. Requires root. Without it, the install is per-user to <storage>/bin with a symlink in ~/.local/bin on Linux. |
--systemd | Also write, enable and start a systemd unit. |
--force | Overwrite an existing managed binary. |
service uninstall takes --system for a system-wide install. See services for the units themselves.
cluster
Section titled “cluster”See cluster & mesh for the workflow.
| Subcommand | Description |
|---|---|
cluster invite [-n, --name <LABEL>] [--ttl <MINUTES>] | Create a join ticket. --name defaults to node, --ttl to 15. Prints a kaba_invite_… string. |
cluster join <TICKET> | Join the cluster that issued the ticket. |
cluster list | List known peers. |
cluster show <NAME_OR_NODE_ID> | Print everything known about one peer. |
cluster ping <NAME_OR_NODE_ID> | Round-trip to a peer and print the latency. |
cluster evict <NAME_OR_NODE_ID> [-y] | Remove a peer. Without -y it only prints what it would do. |
cluster repair | Re-establish keys for peers that are known but keyless. |
cluster tombstones | List evicted node IDs. An evicted node cannot rejoin while its tombstone exists. |
cluster rejoin <NODE_ID> | Clear a tombstone so that node can join again. |
Downloads and prepares base-model files into <storage>/kaba-engine. Idempotent: files already present and valid are skipped. Runs on CPU, so it works on a machine with no GPU.
kabactl get # full prep: tokenizer, safetensors, training checkpoint, GGUFkabactl get --model e2b # one GGUF for serving (fastest first install)kabactl get --only-inference # GGUFs for every variant, no training prepkabactl get --voice # speech-to-text and text-to-speech models| Option | Description |
|---|---|
--model <VARIANT> | Fetch one serving GGUF: e2b (about 2.8 GB) or e4b (about 4.4 GB). Enough for ask and for serving adapters. train still needs the bare get. |
--only-inference | Fetch the serving GGUF for every variant and skip the training conversion. |
--projections | Also fetch the multimodal projector. Image and audio input use dedicated routes, so this is only needed for KABA_VISION_MODEL=base or older projection asks. |
--voice | Fetch the voice bundle. The server also provisions this automatically at start. |
Generates an answer to a prompt and streams tokens to stdout.
kabactl ask "Summarize what I read about QUIC this week" --workspace you@example.comkabactl ask "What is 2+2?" --peer gpu-towerLocal asks drive the engine in this process and expect the server to be stopped, because the engine and the vector store want exclusive access. Remote asks go to a cluster peer over kaba/infer/v1; that peer must have accept_remote_inference = true.
| Option | Description |
|---|---|
--peer <NAME_OR_NODE_ID> | Run on a cluster peer instead of locally. |
--workspace <EMAIL> | Ground the answer on that account’s encrypted memories. Prompts for the account password, or reads KABA_WORKSPACE_PASSWORD. Without it, the base model answers with no memory grounding. |
--lora <NAME> | Apply a named adapter. On a peer, omit it to use the peer’s configured adapter, or pass --lora "" to force the peer’s base model. |
--base-model <VARIANT> | e2b or e4b. |
--max-tokens <N> | Cap on generated tokens. |
--temperature <T> | Sampling temperature. |
--repetition-penalty <P> | Repetition penalty. |
--max-context <N> | Context window in tokens. |
--gpu-layers <N> | Pin how many model layers stay on the GPU. Omitted, the node fits to free VRAM. KABA_LLAMA_GPU_LAYERS on the serving node caps the request. |
--turbo | Use the compressed working-memory cache. |
--quantize <Q> | q8 or q4. |
--gguf <PATH> | Use a specific GGUF file. |
--perf | Print performance timing. |
Trains a LoRA adapter from memories and writes it to kaba-engine/loras/<name>.{mpk,json}.
kabactl train --workspace you@example.com --last 30d --output-name octoberkabactl train --workspace you@example.com --node gpu-towerLocal training needs a GPU build and expects the server to be stopped. With --node, your corpus is streamed to a peer over kaba/train/v1, progress is shown, and the finished adapter is pulled back; that peer must have accept_remote_training = true.
Which memories
| Option | Description |
|---|---|
--workspace <EMAIL> | The account whose encrypted memories to train on. |
--since <DATE|DUR> / --until <DATE|DUR> | Window bounds: YYYY-MM-DD, an ISO instant, or a duration such as 7d or 24h. |
--last <DURATION> | Shorthand for a trailing window. Omit all three to train on everything. |
--min-records <N> | Refuse to train on fewer records. Default from config: 200. |
Where and what
| Option | Description |
|---|---|
--node <NAME_OR_NODE_ID> | Train on a cluster peer. |
--output-name <NAME> | Adapter name. |
--voice <VOICE> | How training answers are phrased: factual (default) or recall. |
--skip-summaries | Skip the summarisation pass. |
--summarise-batch <N>, --summarise-batch-trust | Batch the summarisation pass. |
--perf | Print performance timing. |
Hyperparameters (defaults come from kaba_engine_options)
| Option | Description |
|---|---|
--epochs <N> | Default 3. |
--lr <LR> | Peak learning rate. Default 1e-4. |
--lora-rank <N> | Default 16. |
--batch-size <N>, --seq-len <N>, --grad-accum <N>, --warmup-ratio <RATIO> | Batch shape and schedule. |
--no-qlora | Train without quantization. |
kabactl loras [--peer <NAME_OR_NODE_ID>] [--json]Lists adapter names in kaba-engine/loras, or on a peer. An empty list just means nothing has been trained yet.
lora-import
Section titled “lora-import”kabactl lora-import <URL> [--base-model <VARIANT>]Downloads a community adapter from a Hugging Face repo URL or a direct .gguf / .safetensors link, converts it so kabactl can run it, and writes a manifest so it shows up in kabactl loras and in the policy editor. The raw download is kept alongside as provenance. Set HF_TOKEN for private repos.
security
Section titled “security”Manages the local threat database used to screen browsing. Feeds are configured under security_options.feeds.
| Subcommand | Description |
|---|---|
security update | Rebuild the Bloom filter and index from the configured feeds. Skipped when feeds are unchanged. |
security status [--json] | Indicator count, last update time and per-source results. |
security check <VALUE> [--kind <KIND>] | Test a value. --kind defaults to domain; the other kinds are url, ip, filename and sha256. |
doctor
Section titled “doctor”kabactl doctor [--json]Reports PASS, WARN or FAIL for each resource kabactl depends on, and exits non-zero if anything failed.
| Check | Looks at |
|---|---|
config_toml | config.toml exists and parses. |
tls_cert, tls_key | Present, non-empty, and the key is not group or world readable. |
sqlite_db | kaba.db exists and is writable. |
lancedb_store | memories/kaba is readable and writable. |
engine_dir, tokenizer_json, gemma_gguf, loras_dir | Model files and adapter count. |
service_pid, service_log | Running, stale PID, or stopped. |
managed_binary, kabactl_on_path | The installed binary is executable and on PATH. |
server_port, proxy_port | Ports are free, or held by the running service. |
update_token | Whether KABA_UPDATE_TOKEN is set. |
self-update
Section titled “self-update”kabactl self-update [--check] [--to <TAG>] [--force] [--no-restart]Fetches a release, picks the asset for this platform, verifies the downloaded binary reports the expected version, and atomically swaps it over the managed binary. If the service is running it is restarted unless --no-restart is given.
| Option | Description |
|---|---|
--check | Report whether an update is available and exit. |
--to <TAG> | Install a specific release tag. |
--force | Reinstall even if already current. |
--no-restart | Leave the running service on the old binary. |
Release assets exist for linux-x64, linux-arm64, macos-arm64 and windows-x64. For a private repository set KABA_UPDATE_TOKEN to a read-scope access token; it is sent as an Authorization header, never in a URL. KABA_UPDATE_BASE_URL points at a different release host.