Skip to content

Benchmarks

How fast is local, really?

Real numbers from Bedrova's built-in benchmark across Mac tiers: load time, first token, tokens per second, peak memory. Methodology first, and no cherry-picking.

Load time seconds · at launch
First token ms · at launch
Generation tokens/sec · at launch
Peak memory GB · at launch

Bedrova ships when the product is complete; final numbers publish with a reproducible run at launch. The methodology below is locked now.

The runs

Each row is one model, quant and Mac. Filterable in the app; published in full here.

ModelQuantMacLoadFirst tokenPrompt tok/sGen tok/sPeak RAM
DeepSeek V4 Flash Q4Q8 M3 Ultra · 256 GB
DeepSeek V4 Flash Q9 M3 Ultra · 256 GB
Qwen3.6 35B A3B 4-bit M4 Max · 128 GB
Llama 70B 4-bit M4 Pro · 64 GB
Qwen3.5 7B 4-bit M4 · 32 GB
FLUX.1 schnell M4 Pro · 64 GB

6 planned runs. Image models report seconds-per-image at launch.

Methodology

  • Every number comes from Bedrova's built-in benchmark (the same metrics the app records: load time, time-to-first-token, prompt and generation tokens/sec, peak unified memory).
  • Each run states the exact model, quant, runtime-pack version, context length, and Mac (SoC + unified memory).
  • Numbers are single-machine measurements — your results vary with model, context length, thermals, and what else is loaded.
  • We publish the methodology and the raw numbers together, and update them as runtimes improve. No cherry-picking.

Limitations

  • Pre-launch: figures are placeholders until the formal public run.
  • DeepSeek V4 multi-token-prediction (MTP) and multi-request concurrency are measured separately once proven safe.
  • External-drive throughput and sustained-load thermals are reported as their own runs.

FAQ

Why are the numbers blank?

Because we publish benchmark numbers only after a formal, reproducible public run, not preliminary internal figures. The methodology and the table structure are here now so you know exactly what we will measure and can hold us to it.

How do you measure?

With Bedrova's own built-in benchmark, which records load time, time to first token, prompt and generation tokens per second, and peak unified memory. They are the same metrics the app shows you while it works.

Will my Mac match these?

Roughly, for the same model, quant and Mac. Results vary with context length, thermals, and what else is loaded. Use the numbers as a guide, then run the benchmark yourself.