Skip to content

Know before you download

What can your Mac run?

About 75% of your Mac's unified memory is usable for a model. Pick your memory below to see what fits. Bedrova warns you before anything would run you out of RAM.

Your Mac's memory
Model Type Est. peak RAM Min recommended Fits 16 GB?
Qwen3 Embedding 0.6B starter Powers document chat and search. Embeddings 1 GB 8 GB Fits
BGE Small EN Embeddings 1 GB 8 GB Fits
Kokoro 82M (TTS) starter Fast, natural speech. Voice · text-to-speech 1 GB 8 GB Fits
Core ML perception packs Detection, segmentation, depth, OCR and pose. Tiny and fast. Perception 1 GB 8 GB Fits
Qwen3 4B (4-bit) starter Runs on an 8 GB Mac. A real first local chat model. Chat 3 GB 8 GB Fits
Whisper large v3 turbo (STT) starter Voice · speech-to-text 3 GB 8 GB Fits
Qwen3-TTS (TTS) Custom-voice and voice-design profiles. Voice · text-to-speech 4 GB 16 GB Fits
Stable Diffusion 1.5 Bring your own weights. Light enough for an 8 GB Mac, and the ControlNet base. Image generation 4 GB 8 GB Fits
Qwen3.5 9B (4-bit) starter The sweet spot on 16 GB. Chat 7 GB 16 GB Fits
Llama 3.x 8B (4-bit) Dependable all-rounder. Chat 8 GB 16 GB Fits
Qwen3.5 9B (vision) starter Ask questions about screenshots and photos. Vision 8 GB 16 GB Fits
FLUX.2 Klein 4B starter Fast local image generation. Image generation 9 GB 16 GB Fits
SDXL Bring your own weights. Image generation 10 GB 16 GB Fits
Gemma 3 12B (4-bit) Chat 11 GB 24 GB Fits
FLUX.1 schnell Image generation 12 GB 24 GB Fits
Qwen Image Image generation 14 GB 32 GB Tight
Stable Diffusion 3.5 Bring your own weights. Image generation 16 GB 32 GB Too big
Wan 2.2 TI2V 5B Text-to-video and image-to-video, generated on your Mac. Video generation 16 GB 32 GB Too big
Qwen3.6 27B (4-bit) Strong reasoning without a workstation. Chat 18 GB 32 GB Too big
Qwen3.6 27B (vision) Vision 18 GB 32 GB Too big
FLUX.2 Klein 9B Higher fidelity, more memory. Image generation 20 GB 32 GB Too big
Qwen3.6 35B A3B (4-bit) Mixture-of-experts: big model, modest active memory. Chat 22 GB 48 GB Too big
Llama 70B (4-bit) Chat 42 GB 64 GB Too big
MiniMax M2 Chat 48 GB 96 GB Too big
DeepSeek V4 Flash (Q4Q8) Frontier reasoning, fully local. Chinese-origin model — runs entirely on your Mac, nothing sent to China. Frontier chat 92 GB 128 GB Too big
DeepSeek V4 Flash (Q9) Chinese-origin model — runs entirely on your Mac. Frontier chat 167 GB 256 GB Too big

Estimates mirror Bedrova's model-fit-check contract; real numbers depend on context length + other loaded models, which Bedrova accounts for live. Big models that won't fit are blocked with a safer suggestion instead of crashing. See benchmarks.

What to run, by Mac

A good starting point for each memory tier. Bedrova recommends one automatically on first run.

16 GB

Fast everyday chat, embeddings, small images.

  • FLUX.1 schnell
  • Gemma 3 12B (4-bit)
  • SDXL
  • FLUX.2 Klein 4B

32 GB

Mid-size chat, RAG, voice, and image generation.

  • Qwen3.6 35B A3B (4-bit)
  • FLUX.2 Klein 9B
  • Qwen3.6 27B (4-bit)
  • Qwen3.6 27B (vision)

64 GB

70B-class models and vision, several at a time.

  • MiniMax M2
  • Llama 70B (4-bit)
  • Qwen3.6 35B A3B (4-bit)
  • FLUX.2 Klein 9B

128 GB+

Frontier: DeepSeek V4 Flash, fully local.

  • DeepSeek V4 Flash (Q4Q8)
  • MiniMax M2
  • Llama 70B (4-bit)
  • Qwen3.6 35B A3B (4-bit)

Across models

It's one budget.

Bedrova doesn't count models one at a time. It budgets your unified memory across everything loaded, so you can run chat, image and voice together without an out-of-memory crash.

Example · 32 GB Mac

Qwen3.5 7B · 4.3 GBFLUX.1 · 11 GBWhisper · 3.1 GB 13.6 GB free

FAQ

What AI models can my Mac run?

It depends on your unified memory. About 75% of your Mac's RAM is usable for a model. 16GB runs 7–8B chat models comfortably; 32GB handles ~13B plus image generation; 64GB+ runs 70B-class models; DeepSeek V4 Flash needs 128GB+.

Why only 75%?

macOS and your other apps need memory too, and a model's footprint grows with context length and working memory during generation. Keeping loaded models under about 75% leaves safe headroom, and the fit-check enforces it so you do not crash.

Does context length change what fits?

Yes. Longer context grows the KV cache, so the same model needs more memory with a bigger context. Bedrova accounts for this live and warns you before you run out.

Can I run several models at once?

Yes. Bedrova budgets memory across every loaded model. The totals on this page are per-model, and the app tracks the running total as you pin and load more.

Coming soon no telemetry · notarized · by ideius