Skip to content

On device AI

Mattia Tadini edited this page Sep 13, 2026 · 7 revisions

On-device AI

SkillFishOS can run local LLMs on the BC-250's integrated GPU, accelerated in Vulkan — nothing leaves the machine.

The stack

  • Unsloth Studio as the engine, accelerated via Vulkan on the Cyan Skillfish GPU. It is a single native service that provides both the chat interface and an OpenAI-compatible API, listening on 127.0.0.1:8888.
  • The AI section of the SkillFishOS Control Center (skillfish-control-center, or the skillfish-ai-panel command, which now opens that same section) starts and stops it and manages everything around it.
  • The Remote Manager proxies it at /unsloth, so the same engine is reachable from another machine on the LAN without opening a second port.

It used to be Ollama plus OpenWebUI in three Docker containers with a ~6.5 GB custom image. That is gone since 26.08: one service, no Docker.

Why Vulkan matters here: measured on this hardware, the same model runs at about 210 tokens/s on the GPU against 41 on the CPU. It is not a rounding difference, it is the whole reason on-device AI is usable on this board.

The AI section

  • Engine on/off, and start-at-boot.
  • Version with an update check and a one-click Update (skillfish-unsloth-update, reruns the official installer with the Vulkan bundle forced).
  • First sign-in handled for you: Unsloth generates a random initial password at install and shuts itself down after an hour if it isn't changed, so the Control Center lets you set the Studio password and creates the API key in the same step. The key is shared with the Remote Manager.
  • Models on disk with quantization and size, plus downloading new GGUFs straight from the Hugging Face Hub by repository id (e.g. unsloth/Qwen3-4B-GGUF), with the available variants listed by size.
  • A quick chat over the OpenAI-compatible API, with measured tokens/second.
  • Memory: VRAM, GTT, RAM, swap, the model budget, and a GTT-limit slider — see Memory VRAM and GTT.
  • Network: making Studio reachable from the LAN (0.0.0.0:8888, its own login) and parallel chat slots.

The full chat, Chat with Files (RAG), the model Hub page, parallel chats and Deep Research live inside Studio's own UI at http://localhost:8888 — the Control Center handles engine control and the quick checks around it.

Tips

  • Bigger models need more memory: raise the GTT budget (and/or the UMA VRAM) before loading a large model.
  • The GPU governor's idle behaviour means the card drops to 350 MHz between prompts; under inference it ramps up. For sustained throughput pick a higher ceiling in the Tuner section's Performance preset (2100 MHz) — for pure compute the 2230 MHz point can help (with adequate voltage/cooling), but the validated-stable cap is 2200 MHz. See GPU Governor and Tuning.
  • Unlocking Compute Units (40-CU) roughly 1.8×'s the GPU's compute throughput.

Nothing here phones home — the models and the chat run entirely on your BC-250.

Community research

Two repositories by akandr are worth reading if you want to go further on this board:

  • bc250 — their own recipe for the board, on a different inference stack from ours, plus image generation with stable-diffusion.cpp. Worth reading for how it gets Vulkan working on this GPU.
  • bc250-rocm — ROCm/HIP made to work on gfx1013, with the driver patches, the recipe and the measurements.

The second one answers a question we are asked often: why Vulkan and not ROCm. Its comparison puts ROCm decode at roughly 40 to 63 percent of Vulkan across most models tested, with one model close to parity. That measurement was taken on someone else's board, which is exactly what makes it worth citing here.

It also carries a warning worth repeating. Under sustained load a GPU page fault can escalate to a GPU reset, and on this hardware the reset takes the machine with it. The captured trace shows the reset itself succeeding, and the host then stalling on a clocksource watchdog. This is not a board to leave running unattended on work that matters.

Clone this wiki locally