Make images
FLUX, Stable Diffusion and Qwen Image, generated on your Mac.
Native macOS · Apple Silicon · runs side by side
Run FLUX, Wan, Whisper, DeepSeek, Qwen and the latest open models on-device, side by side, through one API. MLX-first on Apple Silicon, with Core ML perception and GGUF text compatibility.
One app, every type
A chat app, an image tool, a transcription app, a video tool. Bedrova runs all of them on-device, from one library.
FLUX, Stable Diffusion and Qwen Image, generated on your Mac.
Text-to-video and image-to-video, rendered on-device. Nobody else does this locally.
Llama, Qwen, Gemma and DeepSeek V4 Flash, running locally.
Whisper transcription and natural speech, no cloud.
Ask about screenshots and photos.
OCR, detection, segmentation, depth and pose through Core ML.
Local embeddings, retrieval and citations you can click.
Point any app or agent at localhost:1337/v1.
The Loom
Most local AI apps give you one model and one thread. The Loom gives you a conversation that spans text, vision, voice and image models together. Paste a screenshot and ask about it, have the answer read aloud, generate an image from it, and keep going in the same thread.
One honest caveat: the largest reasoning models take the whole GPU, so they run a turn at a time while other lanes carry on. Bedrova shows you that rather than animating a concurrency it doesn't have.
The runtime
Your apps make one kind of call. Bedrova sends each request to the right model. Favourites stay pinned and always loaded, everything else loads on demand, and your Mac's memory is budgeted across every loaded model.
Your apps & agents
RAM budgeted across loaded models — never OOMs your Mac
Getting started
No Python, no terminal, no config files. An 8 GB Mac is enough to start.
Bedrova reads your chip, memory and free disk, then recommends a model that actually fits.
One click downloads and verifies it. Resumable, and safe to close mid-download.
Type. Add more model types later, whenever you want them.
Built for your Mac
Control + safety
Bedrova tracks memory across every loaded model and checks a new one against what's left before loading it. Either it fits, or you get told why not and what to try instead.
Free, native, and entirely on-device. One runtime for all of it.