Know before you download
What can your Mac run?
About 75% of your Mac's unified memory is usable for a model. Pick your memory below to see what fits. Bedrova warns you before anything would run you out of RAM.
| Model | Type | Est. peak RAM | Min recommended | Fits 16 GB? |
|---|---|---|---|---|
| Qwen3 Embedding 0.6B starter Powers document chat and search. | Embeddings | 1 GB | 8 GB | Fits |
| BGE Small EN | Embeddings | 1 GB | 8 GB | Fits |
| Kokoro 82M (TTS) starter Fast, natural speech. | Voice · text-to-speech | 1 GB | 8 GB | Fits |
| Core ML perception packs Detection, segmentation, depth, OCR and pose. Tiny and fast. | Perception | 1 GB | 8 GB | Fits |
| Qwen3 4B (4-bit) starter Runs on an 8 GB Mac. A real first local chat model. | Chat | 3 GB | 8 GB | Fits |
| Whisper large v3 turbo (STT) starter | Voice · speech-to-text | 3 GB | 8 GB | Fits |
| Qwen3-TTS (TTS) Custom-voice and voice-design profiles. | Voice · text-to-speech | 4 GB | 16 GB | Fits |
| Stable Diffusion 1.5 Bring your own weights. Light enough for an 8 GB Mac, and the ControlNet base. | Image generation | 4 GB | 8 GB | Fits |
| Qwen3.5 9B (4-bit) starter The sweet spot on 16 GB. | Chat | 7 GB | 16 GB | Fits |
| Llama 3.x 8B (4-bit) Dependable all-rounder. | Chat | 8 GB | 16 GB | Fits |
| Qwen3.5 9B (vision) starter Ask questions about screenshots and photos. | Vision | 8 GB | 16 GB | Fits |
| FLUX.2 Klein 4B starter Fast local image generation. | Image generation | 9 GB | 16 GB | Fits |
| SDXL Bring your own weights. | Image generation | 10 GB | 16 GB | Fits |
| Gemma 3 12B (4-bit) | Chat | 11 GB | 24 GB | Fits |
| FLUX.1 schnell | Image generation | 12 GB | 24 GB | Fits |
| Qwen Image | Image generation | 14 GB | 32 GB | Tight |
| Stable Diffusion 3.5 Bring your own weights. | Image generation | 16 GB | 32 GB | Too big |
| Wan 2.2 TI2V 5B Text-to-video and image-to-video, generated on your Mac. | Video generation | 16 GB | 32 GB | Too big |
| Qwen3.6 27B (4-bit) Strong reasoning without a workstation. | Chat | 18 GB | 32 GB | Too big |
| Qwen3.6 27B (vision) | Vision | 18 GB | 32 GB | Too big |
| FLUX.2 Klein 9B Higher fidelity, more memory. | Image generation | 20 GB | 32 GB | Too big |
| Qwen3.6 35B A3B (4-bit) Mixture-of-experts: big model, modest active memory. | Chat | 22 GB | 48 GB | Too big |
| Llama 70B (4-bit) | Chat | 42 GB | 64 GB | Too big |
| MiniMax M2 | Chat | 48 GB | 96 GB | Too big |
| DeepSeek V4 Flash (Q4Q8) Frontier reasoning, fully local. Chinese-origin model — runs entirely on your Mac, nothing sent to China. | Frontier chat | 92 GB | 128 GB | Too big |
| DeepSeek V4 Flash (Q9) Chinese-origin model — runs entirely on your Mac. | Frontier chat | 167 GB | 256 GB | Too big |
Estimates mirror Bedrova's model-fit-check contract; real numbers depend
on context length + other loaded models, which Bedrova accounts for live. Big models
that won't fit are blocked with a safer suggestion instead of crashing. See benchmarks.
What to run, by Mac
A good starting point for each memory tier. Bedrova recommends one automatically on first run.
16 GB
Fast everyday chat, embeddings, small images.
- FLUX.1 schnell
- Gemma 3 12B (4-bit)
- SDXL
- FLUX.2 Klein 4B
32 GB
Mid-size chat, RAG, voice, and image generation.
- Qwen3.6 35B A3B (4-bit)
- FLUX.2 Klein 9B
- Qwen3.6 27B (4-bit)
- Qwen3.6 27B (vision)
64 GB
70B-class models and vision, several at a time.
- MiniMax M2
- Llama 70B (4-bit)
- Qwen3.6 35B A3B (4-bit)
- FLUX.2 Klein 9B
128 GB+
Frontier: DeepSeek V4 Flash, fully local.
- DeepSeek V4 Flash (Q4Q8)
- MiniMax M2
- Llama 70B (4-bit)
- Qwen3.6 35B A3B (4-bit)
Across models
It's one budget.
Bedrova doesn't count models one at a time. It budgets your unified memory across everything loaded, so you can run chat, image and voice together without an out-of-memory crash.
Example · 32 GB Mac
FAQ
What AI models can my Mac run?
It depends on your unified memory. About 75% of your Mac's RAM is usable for a model. 16GB runs 7–8B chat models comfortably; 32GB handles ~13B plus image generation; 64GB+ runs 70B-class models; DeepSeek V4 Flash needs 128GB+.
Why only 75%?
macOS and your other apps need memory too, and a model's footprint grows with context length and working memory during generation. Keeping loaded models under about 75% leaves safe headroom, and the fit-check enforces it so you do not crash.
Does context length change what fits?
Yes. Longer context grows the KV cache, so the same model needs more memory with a bigger context. Bedrova accounts for this live and warns you before you run out.
Can I run several models at once?
Yes. Bedrova budgets memory across every loaded model. The totals on this page are per-model, and the app tracks the running total as you pin and load more.