Two local services. One model confirmed loaded.
Your main Mac has a local coding model loaded and a second Qwen service answering. The Intel machines are not yet model hosts.
On this page: Now · Fleet · Next · Full guide · OpenCode tiers
On this page: Now · Fleet · Next · Full guide · OpenCode tiers
What is available right now.
The two services are separate. Both listen only on this Mac, so phones and other machines cannot reach them yet.
Devstral Small 2 · 24B
Loaded through Ollama. The local safe name and the full model name point to the same stored weights. OpenCode lists ollama/devstral-local-safe. Ollama reports 20 GB loaded, full GPU placement, and a 65,536-token context setting.
LOADED · LOCALQwen 3.6 · 35B-A3B
The llama.cpp service responds at 127.0.0.1:8080 and lists this model. Its model file is about 21 GB, plus its vision add-on. The server is configured for 65,536 tokens. Memory use was not isolated from the rest of the Mac.
RESPONDING · LOCALThere is 108.7 GB free. Current system readings show 36.4 GB of swap in use and 19% free memory; I cannot attribute that to the models alone. I did not remove or move anything. Use one large model at a time, and pause new downloads until you choose what to clean. The Qwen files total about 28 GB; Ollama storage uses about 14 GB.
Other model files on this Mac.
The files are present, but they are not all running. The 35B Qwen service answers requests; the amount of memory it currently occupies was not measured separately.
| Model or file | What it is for | Current state |
|---|---|---|
| Qwen 3.5 · 9B | Smaller text and image model; about 5.3 GB plus an 876 MB vision add-on. | On disk, not loaded |
| Qwen 3.6 · 35B-A3B | Larger general and vision model; about 21 GB plus an 858 MB vision add-on. | Server responds |
| Devstral Small 2 · 24B Q4 | Local coding model. Two Ollama names refer to the same model weights. | Loaded in Ollama |
| Laya checkpoint | An 804 MB model file and related project scripts are present. | Purpose and live use not confirmed |
How this connects to development.
The local route is present in your coding tool. The full coding-agent experience still needs a short, real task test before calling it dependable.
Verified: OpenCode’s model list includes the local Ollama name. Not verified today: a completed edit-test-review task through the full OpenCode agent loop. Model listing alone does not prove the cloud provider accounts are signed in.
What the rest of the fleet can do today.
Five Intel machines answered over Tailscale. None showed Ollama or llama.cpp on the normal command path, and neither common local model port answered on those machines.
Show 7 machines, 5 reachable today
Best next Intel model host. Could be tested with a small quantized model, but none is installed yet.
Good small-model and background-task candidate. No local model service found in this check.
Better for queues, archives, indexing, document preparation, and small CPU model jobs than fast chat.
Good light worker. About 7 GB RAM was available in the latest measurement; no model service found.
Older macOS and graphics. Possible CPU-only model host after checking free disk; no model service found.
Not available for a live check. Last measured at 4 GB RAM; use for routing and small automation, not a chat model.
Live inventory is pending. I did not count its shared files as proof of installed models.
No connected phone model was identified in this check. The safe plan is to use a phone as a camera, recorder, and screen for the Mac model over a private Tailscale link. On-device models depend on the phone’s exact model and memory.
Where you stand.
You already have a real local option, including a coding model. You are not yet running a distributed local fleet.
Current tools and cloud choices
- Ollama 0.34.4 is installed and running on the M4 Pro.
- llama.cpp is running a Qwen 35B service on the M4 Pro.
- OpenCode, Claude Code, and Codex CLI are installed. This check did not measure subscriptions, billing, or cloud sign-in state.
- OpenCode displays cloud model choices as well as the local Devstral option. A listed provider is not proof of usable account access.
- Kaggle remains your known free burst-compute option; it is not an always-on private server.
Next safe steps
- Keep the current models; do not download more until storage is cleaned by choice.
- Run one contained coding task through OpenCode and Devstral, then compare the result with Claude.
- When ready, test one small Q4 model on wkgd-imac. Measure speed and memory before adding it to a workflow.
- Only then expose a private model endpoint to your phones through Tailscale. Do not make either service public.
How this snapshot was checked
Checked local hardware and free disk, Ollama’s model and loaded-model lists, Ollama storage, llama.cpp health and model endpoints, OpenCode’s model list, local listening addresses, memory and swap readings, and Tailscale status. Reached five Intel machines and checked their model commands and usual ports. Phones and refused/offline machines remain unverified. This is a snapshot, not automatic monitoring.
Local evidence: ~/.ollama/models, ~/Models/gguf, 127.0.0.1:11434, and 127.0.0.1:8080.
Milo Offline brand systemOllama + OpenCode setupPrivate Tailscale service access