The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Onde Inference CLI listing page.
Command-line interface for Onde Inference.
Swift · Flutter · React Native · Rust · Website
Manage your Onde Inference account, fine-tune local models, and export them to GGUF, all from the terminal.
Install onde-cli with your favorite tool. For package docs and the full install matrix, see https://ondeinference.com/cli.
The Dart package is a thin launcher. On first run it downloads the right native binary into ~/.onde/cli, then reuses the local copy.
Download a release from GitHub Releases:
| Platform | File |
|---|---|
| macOS Apple Silicon | onde-macos-arm64 |
| macOS Intel | onde-macos-amd64 |
| Linux x64 | onde-linux-amd64 |
| Linux arm64 | onde-linux-arm64 |
| Windows x64 | onde-win-amd64.exe |
| Windows arm64 | onde-win-arm64.exe |
This opens the TUI. You can sign up or sign in right there.
| Key | What it does |
|---|---|
Tab | Move between fields |
Enter | Submit or sign out |
Ctrl+L | Go to the sign-in screen |
Ctrl+N | Go to the new account screen |
Ctrl+C | Quit |
Run onde as a Model Context Protocol server over stdio instead of the TUI:
This exposes Onde account and model-catalog operations as MCP tools — login, me, apps_list, app_create, app_rename, models_list, model_register, model_assign, hf_search — returning structured JSON. stdout is the JSON-RPC channel; tools run non-interactively and reuse the token from a TUI sign-in (or the login tool). Point any MCP client at the command onde --mcp.
onde includes a LoRA fine-tuning pipeline for Qwen2, Qwen2.5, and Qwen3 models. It runs locally: Metal on Apple Silicon, CPU elsewhere. No cloud setup. No Python environment.
The flow is straightforward: download a safetensors base model, fine-tune it with LoRA, merge the adapter back into the base weights, then export to GGUF for use in the Onde SDK.
If you want a quick refresher on what the model is actually doing at inference time, Onde has a short note on the forward pass.
Each line should be one complete conversation in Qwen's chat template:
Save the file wherever you want. The TUI lets you point to it directly.
Only safetensors models can be fine-tuned. GGUF models are already quantized, so their weights are not differentiable.
Configure the run:
| Field | Default | Notes |
|---|---|---|
| Training data | ~/.onde/finetune/train.jsonl | Path to your JSONL file |
| LoRA rank | 8 | Higher means more capacity and more memory use |
| Epochs | 3 | Full passes over the dataset |
| Learning rate | 0.0001 | AdamW default |
Press Enter to start. In a healthy run, loss usually starts dropping by epoch 2. If it stays flat, try 0.0003.
For rank 8 on a 0.6B model, the adapter is about 1.5 MB. From the fine-tune complete screen:
m to merge the adapter into the base modelg to export the merged model to GGUFThe resulting GGUF loads directly in the Onde SDK for on-device AI inference.
| Model | Size | Notes |
|---|---|---|
Qwen/Qwen3-0.6B | ~1.2 GB | Smallest and quickest to train |
Qwen/Qwen2.5-1.5B-Instruct | ~3.0 GB | Good default for instruction tuning |
Qwen/Qwen3-1.7B | ~3.4 GB | Newer small Qwen3 model |
Qwen/Qwen3-4B | ~8.0 GB | Best quality, better suited to macOS |
You can search for any of these from the Models tab with /.
Logs are written to ~/.cache/onde/debug.log.
If you installed through pub.dev, the launcher cache lives under ~/.onde/cli.
Dual-licensed under MIT and Apache 2.0.
© 2026 Splitfire AB (Onde Inference).