Benchmarks
How fast is local, really?
Real numbers from Bedrova's built-in benchmark across Mac tiers: load time, first token, tokens per second, peak memory. Methodology first, and no cherry-picking.
Bedrova ships when the product is complete; final numbers publish with a reproducible run at launch. The methodology below is locked now.
The runs
Each row is one model, quant and Mac. Filterable in the app; published in full here.
| Model | Quant | Mac | Load | First token | Prompt tok/s | Gen tok/s | Peak RAM |
|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | Q4Q8 | M3 Ultra · 256 GB | — | — | — | — | — |
| DeepSeek V4 Flash | Q9 | M3 Ultra · 256 GB | — | — | — | — | — |
| Qwen3.6 35B A3B | 4-bit | M4 Max · 128 GB | — | — | — | — | — |
| Llama 70B | 4-bit | M4 Pro · 64 GB | — | — | — | — | — |
| Qwen3.5 7B | 4-bit | M4 · 32 GB | — | — | — | — | — |
| FLUX.1 schnell | — | M4 Pro · 64 GB | — | — | — | — | — |
6 planned runs. Image models report seconds-per-image at launch.
Methodology
- Every number comes from Bedrova's built-in benchmark (the same metrics the app records: load time, time-to-first-token, prompt and generation tokens/sec, peak unified memory).
- Each run states the exact model, quant, runtime-pack version, context length, and Mac (SoC + unified memory).
- Numbers are single-machine measurements — your results vary with model, context length, thermals, and what else is loaded.
- We publish the methodology and the raw numbers together, and update them as runtimes improve. No cherry-picking.
Limitations
- Pre-launch: figures are placeholders until the formal public run.
- DeepSeek V4 multi-token-prediction (MTP) and multi-request concurrency are measured separately once proven safe.
- External-drive throughput and sustained-load thermals are reported as their own runs.
FAQ
Why are the numbers blank?
Because we publish benchmark numbers only after a formal, reproducible public run, not preliminary internal figures. The methodology and the table structure are here now so you know exactly what we will measure and can hold us to it.
How do you measure?
With Bedrova's own built-in benchmark, which records load time, time to first token, prompt and generation tokens per second, and peak unified memory. They are the same metrics the app shows you while it works.
Will my Mac match these?
Roughly, for the same model, quant and Mac. Results vary with context length, thermals, and what else is loaded. Use the numbers as a guide, then run the benchmark yourself.