Your AI ecosystemEvidence-led recommendation

Spend less. Keep the judgment.

Use open models for volume. Keep Claude and ChatGPT for the work where mistakes, taste, and deep reasoning cost more than tokens.

MacPrivate daily work
KaggleFree batch compute
Frontier cloudHard final passes
Hokusai wave, part of the Milo Offline visual system
The recommendation: build a ladder, not a replacement. Start cheap. Escalate only when the work earns it.

On this page: Short answer · Tiers · Computers · Models · Fleet · Phones · Your work · Products · Plan · Sources

The short answer.

Your 48 GB M4 Pro is capable, but it is also your main production computer and its disk is already tight. One carefully chosen local model makes sense. A library of giant models does not.

Best overall system

Devstral Small 2 on the Mac for local coding and repo work. Gemma 3 12B or Qwen 3.5 9B as an optional lighter general model. Kaggle for scheduled GPU batches. Ollama Cloud or a shut-down-after-use Runpod for larger open models. Claude and ChatGPT for architecture, hard debugging, major creative direction, high-stakes research, and final review.

Capability first. Brand second.

This is the honest quality ladder. Move down the ladder for volume and privacy. Move up when the cost of a weak answer is higher than the model bill.

A

Frontier judgment

Claude and ChatGPT remain your senior partners.

  • Architecture and difficult multi-file changes
  • Ambiguous debugging and production incidents
  • Brand direction, strategy, and final creative taste
  • Current research with source checking
  • Final review before deployment
B

Large open models in the cloud

Near-frontier breadth at much lower unit cost.

  • Long document analysis
  • Large codebase first passes
  • High-volume drafting and transformation
  • Tasks too large for the Mac but not worth frontier pricing
C

Strong local workhorse

Your M4 Pro handles a focused 12B to 24B model well.

  • Routine coding, tests, refactors, and documentation
  • Private client material
  • Drafts, summaries, extraction, and structured output
  • Work that can be checked automatically
D

Small utility models

Fast, cheap, narrow, and useful across the older fleet.

  • Tagging, sorting, renaming, and classification
  • Short summaries and metadata
  • Routing work to the right model
  • Background jobs where perfection is unnecessary

Four places to run the work.

The model matters less than putting it on the right computer. Free notebooks are excellent batch workers, but they are poor always-on assistants.

This Mac

Best for private, interactive work. No per-message cost. Always available. Keep only one heavyweight model loaded.

REAL COST: DISK + BATTERY + WAIT TIME
  • Best fit: Devstral Small 2, Gemma 3 12B, Qwen 3.5 9B
  • Comfortable storage target: one 8-15 GB primary model
  • Avoid: 35B at 64K while Claude, Chrome, and builds are busy

Kaggle

Your best free burst worker. Officially offers a free P100, roughly 30 GPU hours weekly, and sessions up to 12 hours.

REAL COST: SETUP + SESSION LIMITS
  • Best for batch inference, media, audio, embeddings, and notebooks
  • Good home for 7B to 14B quantized models
  • Not a dependable 24/7 endpoint

Colab

A useful backup notebook. Free GPU access is not guaranteed and changes over time.

REAL COST: INTERRUPTIONS
  • Best for experiments and temporary jobs
  • Free sessions can end after inactivity
  • Do not design the main assistant around it

Low-cost cloud

Use Ollama Cloud for simplicity or Runpod when you need your own GPU. Shut rented machines down as soon as the job ends.

REFERENCE: A5000 24 GB FROM $0.27/HOUR, A40 48 GB FROM $0.49/HOUR
  • Ollama Cloud: easiest route to 120B-class open models
  • Runpod 24 GB: medium models and fast batches
  • Runpod 48 GB: larger quantized models without burdening the Mac

The models worth caring about.

You do not need every release. These are the useful families for your mix of development, design, research, content, documents, and media.

Devstral Small 2 24B

PRIMARY PICK

The strongest practical coding-agent choice for this Mac. Built for exploring repositories, editing files, and using tools.

Expect
Good routine coding and repo work, below Claude on difficult judgment
Footprint
About 15 GB model storage
Use for
OpenCode, tests, refactors, documentation, private code

Google Gemma 3 12B

GENERAL + VISION

Google's clean general-purpose option. It can read images and documents, and the 12B version is light enough to remain comfortable.

Expect
Good summaries, image understanding, drafting, and extraction
Footprint
About 8 GB at common quantization
Use for
Documents, screenshots, content, visual inspection

Qwen 3.5 9B

FAST ALL-ROUNDER

A quick multilingual generalist with vision and tool support. Better for speed than for final authority.

Expect
Fast first drafts and reliable structured chores
Footprint
About 7 GB
Use for
Daily Q and A, extraction, summaries, small coding tasks

GPT-OSS 20B

REASONING OPTION

OpenAI's open-weight reasoning model. Useful when thinking depth matters more than writing style.

Expect
Good reasoning and tools, sometimes verbose or uneven
Footprint
About 14 GB
Use for
Analysis, structured reasoning, agent experiments

Gemma 3 27B or Qwen 3.5 27B

HEAVIER MAC

Stronger general quality, but they compete more noticeably with your browser, builds, and cloud-agent sessions.

Expect
Better answers, slower turns, more system pressure
Footprint
About 17 GB each
Use for
Focused sessions when other heavy tools are closed

Qwen Coder, GLM, Kimi, and 120B-class models

CLOUD OPEN

The open-model ceiling. Rent them by token or by hour instead of forcing them onto local hardware.

Expect
Stronger agents and long-context work, still below frontier reliability
Footprint
Cloud only for this ecosystem
Use for
Large batches, huge repositories, cheaper second opinions

7B to 14B quantized models

KAGGLE SWEET SPOT

Small enough for free notebook GPUs and strong enough for repetitive, checkable work.

Expect
Useful output with tight instructions, weak independent judgment
Footprint
Roughly 5-10 GB
Use for
Batch summaries, tagging, extraction, rewrites, embeddings

Qwen 3.5 35B

NOT RECOMMENDED HERE

It runs, but the live test used 23.5 GB of memory and pushed this already-busy Mac toward full swap at 64K context.

Expect
Capable output with a slow cold start
Footprint
23 GB plus working memory
Use for
Rent it elsewhere if specifically needed

Your fleet is a team, not a GPU cluster.

The Intel machines are not useless. They can run small quantized models for background work, serve private endpoints, index your material, run checks, and keep queues moving while the M4 Pro stays responsive. Expect slower answers, not magic GPU speed.

christophers-macbook-pro

MAIN LOCAL MODEL

M4 Pro, 48 GB unified memory. Already runs a 35B Qwen model through llama.cpp and has a separate 9B vision model on disk. Use the 35B for focused local work; keep one large model loaded at a time. The machine has about 98 GiB free, so no model library expansion yet.

wkgd-imac

BEST INTEL MODEL HOST

32 GB RAM, Intel i5-6600, Radeon graphics, 71 GB free. A realistic candidate for 3B-7B Q4 models through llama.cpp, with background-speed expectations. Also a strong index, test, and overnight queue worker.

pmd

CPU + STORAGE WORKER

16 GB RAM, Xeon E5, 418 GB free, and older Radeon cards. Try 1B-3B Q4 CPU models for tags, extraction, and summaries. Better still for archives, document preparation, containers, and queued batch work than interactive chat.

wkgd-macbook

LIGHT MODEL WORKER

Intel i5, 16 GB RAM, 392 GB free, with about 7 GB RAM currently available. Aim at 1B-3B Q4 models and short tasks; reserve most memory for the system and other services.

wokegods-macbook-pro

INTEL MAC MODEL NODE

Intel i5, 16 GB RAM, macOS 26, 748 GB free, Metal 3. Try llama.cpp CPU inference with 3B-7B Q4; 9B may work for a focused test but will be slower. Great as a private API endpoint, model mirror, and batch worker.

wokegod-1

TINY TASK ROUTER

Intel i5, 4 GB RAM, 126 GB free. Not a chat host, but useful for a lightweight dispatcher, health checks, scheduled wakeups, and perhaps a 0.5B-1B model for very short classification.

jamess-imac-pro

PENDING SHELL ACCESS

Online with SMB storage available, but Remote Login is pending. Treat it as storage until its current processor, memory, and GPU can be verified.

chris-x-imac

INTEL MODEL CANDIDATE

Live check now succeeds: Intel i5 3.3 GHz, 4 cores, 32 GB RAM, macOS 12.7.6, Radeon 2 GB. Try CPU-only llama.cpp with 3B-7B Q4; verify available disk before downloading. Older macOS may limit the easiest app choices.

Phones can be local too.

I could not identify a connected phone model in this check, so these are safe capability bands rather than a claim about your exact device. Your phone can still be the pocket doorway to the M4 model even when it is too small to run one itself.

4 GB / older phone

Capture + ask

Use it as a camera, voice recorder, scanner, and chat screen for a model running on your Mac. A tiny 0.8B-1B offline model is an experiment, not a Claude substitute.

6-8 GB memory

Small offline helper

Try 1B-3B Q4 for rewriting, short summaries, note cleanup, and extraction. Expect short context and slower responses than a cloud app.

8 GB+ memory

Better pocket model

3B-4B is the sensible first try. A 7B quantized model may fit on some phones, but heat, battery drain, speed, and app support decide whether it is pleasant.

PhoneCapture, dictate, scan, review
Private Tailscale linkReach devices without opening a public server
M4 ProRun the stronger local model when available
Intel fleetHandle queued small-model chores in the background

Best fit for your day: use the phone as the interface and send work over Tailscale to the Mac's private model endpoint. Keep a small on-device model only for airplane mode or quick private transformations. If your iPhone supports Apple's Foundation Models, that is another no-download option for bounded writing, extraction, summaries, and structured tasks; it is separate from Ollama and cannot load your own model files.

Route the work you actually do.

Your history is dominated by app building, design, debugging, deployment, research, content, documents, media, and voice. This is where the savings come from.

App building

Local model writes routine components, tests, docs, and first-pass refactors. Cloud frontier designs architecture and resolves hard cross-system failures.

LOCAL FIRST, FRONTIER REVIEW

Design and branding

Local model organizes content and checks consistency. Claude or ChatGPT owns concept, taste, visual critique, and final copy.

FRONTIER LED

Production and deployment

Fleet machines can run checks and builds. A frontier model should review risky changes, live incidents, auth, billing, and data migrations.

AUTOMATE, THEN ESCALATE

Research and audits

Local models summarize supplied material. Cloud models with web access verify current facts, laws, pricing, product details, and sources.

LOCAL DIGEST, CLOUD VERIFY

Marketing and documents

Local handles variants, cleanup, formatting, and extraction. Frontier cloud handles positioning, final narrative, and sensitive client-facing work.

MOSTLY LOCAL

Images, video, and voice

Kaggle remains the best free batch lane for ComfyUI-style jobs, Kokoro, transcription, embeddings, and repeatable media processing.

KAGGLE FIRST

Plug it into the products.

Local models earn their keep when they take bounded, testable jobs inside products you already maintain. They should not become a new product that needs constant care.

Offline OS

The private vault and execution layer should use local inference only where a model adds real value. Routine indexing stays deterministic.

LOCAL MODEL
Draft summaries, suggest tags, explain search results, turn notes into proposed tasks, and review bounded code changes.
FLEET FIT
wkgd-imac runs indexing, queues, and tests. The M4 Pro answers only when requested.
ESCALATE
System architecture, merge rules, privacy boundaries, and difficult regressions.

Milo

Milo owns the plan and phone-first daily guidance. A small model can prepare options without becoming the final decision-maker.

LOCAL MODEL
Inbox triage, task summaries, daily brief drafts, project status explanations, and converting notes into candidate plan updates.
FLEET FIT
A background worker prepares drafts. The live product receives structured results, not an always-running chat model.
ESCALATE
Capacity logic, burnout-sensitive recommendations, product decisions, and final user-facing language.

EUANGELION

Its devotional content, Soul Audit, subscriptions, and user data need a clear line between assistance and spiritual authority.

LOCAL MODEL
Content tagging, scripture-reference extraction, accessibility drafts, content QA, test fixtures, and private editorial search.
KAGGLE
Batch audio preparation, transcription, and large content-library processing.
ESCALATE
Theology, pastoral tone, personalized spiritual guidance, safety, and final published devotionals.

MELT

High-volume agency work is the clearest cost-saving opportunity because most transformations are repeatable and reviewable.

LOCAL MODEL
Email extraction, task lists, first-draft briefs, eblast variants, social cutdowns, asset inventories, and formatting checks.
KAGGLE
Media batches, transcriptions, image operations, and large archives.
ESCALATE
Campaign concept, final client copy, strategic recommendations, and anything sent externally without another review.

FAM

Genealogy rewards careful extraction but punishes confident invention. Local AI should assist the evidence workflow, never become the source.

LOCAL MODEL
OCR cleanup, name and date extraction, caption drafts, duplicate detection, and organizing source notes.
FLEET FIT
pmd stores and processes large source collections. The M4 Pro handles visual inspection.
ESCALATE
Identity conflicts, uncertain family links, historical claims, and final published narratives.

Fantasy football

The app already has tests, backtests, deterministic projections, and two clear weekly decisions. Keep the math as code.

LOCAL MODEL
Explain projection output, summarize matchup context, generate test cases, inspect feed-shape changes, and draft ASK responses from computed data.
FLEET FIT
wkgd-macbook or pmd runs refreshes, backtests, and feed checks. The model reads the result.
ESCALATE
Projection-method changes, unexplained backtest failures, and decisions requiring current injury or news verification.

WokeGod Game

The canvas game has a small testable codebase. A coding model can handle contained implementation work if visual quality is reviewed separately.

LOCAL MODEL
Unit tests, level-data validation, balance reports, bug reproduction, documentation, and bounded JavaScript changes.
KAGGLE
Sprite and media experiments when generation needs a GPU.
ESCALATE
Game feel, art direction, progression design, novel mechanics, and difficult input/rendering bugs.

MetaOak

The glasses SDK changes and hardware constraints make current documentation more important than model size.

LOCAL MODEL
Boilerplate, tests, state-machine checks, offline docs search, and size-budget reviews against supplied references.
FLEET FIT
Use older machines for docs mirrors and repeatable builds.
ESCALATE
Current SDK/API questions, lifecycle failures, publishing decisions, and hardware-specific debugging.

Voice stack

Kokoro and Chatterbox already prove the right pattern: notebooks for burst compute, local APIs only when responsiveness matters.

LOCAL MODEL
Script cleanup, pronunciation dictionaries, chunking plans, metadata, and choosing voices from known presets.
KAGGLE
Long narration, multilingual batches, voice experiments, and transcription.
ESCALATE
Final script performance, sensitive voice cloning decisions, and creative direction.

The recommendation set.

Start lean. Prove savings for a month before adding more models or another paid service.

  1. Mac workhorseFinish Devstral Small 2 with a 64K ceiling and connect it to OpenCode. Keep only this heavyweight model locally.
  2. Light general modelAdd Gemma 3 12B only if you regularly need screenshot and document vision away from ChatGPT.
  3. Kaggle batch laneCreate one reusable notebook that accepts a job bundle, runs a 7B to 14B model, and exports results. Do not make it an always-on chat server.
  4. Intel local laneProve one 3B-7B Q4 model on wkgd-imac or wokegods-macbook-pro; use pmd and the other laptops for 1B-3B jobs, queues, indexing, OCR preparation, builds, and batch work.
  5. Cheap big-model laneTry Ollama Cloud pay-as-you-go first. Use Runpod only when you need direct control, private weights, or batch GPU hours.
  6. Frontier gateReserve Claude and ChatGPT for Tier A tasks and the final review of anything expensive, public, irreversible, or hard to test.

Evidence and limits.

Model catalogs and prices change quickly. These links are the live sources behind the recommendation.

Fleet facts were measured live through the canonical Tailscale aliases on September 27, 2026. jamess-imac-pro remains pending shell access; phone make/model and memory were not identified, so phone guidance is conditional. Disk capacity and model speed are device-specific; validate before downloading or assigning always-on work. Local model quality claims are recommendations, not a guarantee of Claude-level judgment.