Skip to content

Native macOS · Apple Silicon · runs side by side

Image, video, speech, vision, text. Every kind of AI model, on your Mac.

Run FLUX, Wan, Whisper, DeepSeek, Qwen and the latest open models on-device, side by side, through one API. MLX-first on Apple Silicon, with Core ML perception and GGUF text compatibility.

  • by ideius
  • no telemetry
  • notarized
  • Apple Silicon
  • free
The Bedrova Loom: a thread list on the left, a Lanes column showing Text, Text plus Vision, Image and Voice models loaded at once, and an answer with a code block and a cited source.

One app, every type

Stop running an app per model type.

A chat app, an image tool, a transcription app, a video tool. Bedrova runs all of them on-device, from one library.

Image

Make images

FLUX, Stable Diffusion and Qwen Image, generated on your Mac.

Video

Make video

Text-to-video and image-to-video, rendered on-device. Nobody else does this locally.

Chat

Chat and reasoning

Llama, Qwen, Gemma and DeepSeek V4 Flash, running locally.

Voice

Speech in and out

Whisper transcription and natural speech, no cloud.

Vision

See images

Ask about screenshots and photos.

Perception

Read and measure

OCR, detection, segmentation, depth and pose through Core ML.

Embeddings

Chat with your documents

Local embeddings, retrieval and citations you can click.

Local API

OpenAI-compatible server

Point any app or agent at localhost:1337/v1.

The Loom

One conversation. Several models at once.

Most local AI apps give you one model and one thread. The Loom gives you a conversation that spans text, vision, voice and image models together. Paste a screenshot and ask about it, have the answer read aloud, generate an image from it, and keep going in the same thread.

  • Each model type runs in its own lane, so you can see what's working.
  • Lanes that can run together do. Lanes that can't say so, plainly.
  • Tool calling and guaranteed-shape JSON, on any lane.

See how the Loom works →

The Loom's thread list and Lanes column, showing Text, Text plus Vision, Image and Voice models all loaded at the same time.
Four model types loaded at once, each in its own lane.

One honest caveat: the largest reasoning models take the whole GPU, so they run a turn at a time while other lanes carry on. Bedrova shows you that rather than animating a concurrency it doesn't have.

The runtime

One API. Every type. Side by side.

Your apps make one kind of call. Bedrova sends each request to the right model. Favourites stay pinned and always loaded, everything else loads on demand, and your Mac's memory is budgeted across every loaded model.

Your apps & agents

Chat Qwen3.5 7B always-on
Image FLUX.1 schnell always-on
Voice Whisper v3 always-on
Vision Qwen2.5-VL JIT
Embeddings Nomic Embed JIT

RAM budgeted across loaded models — never OOMs your Mac

Qwen3.5 · 4.3 GBFLUX.1 · 11 GBWhisper · 3.1 GB 13.6 GB free

See the local API →

Getting started

Running in about three minutes.

No Python, no terminal, no config files. An 8 GB Mac is enough to start.

  1. 1

    Open it

    Bedrova reads your chip, memory and free disk, then recommends a model that actually fits.

  2. 2

    Fetch a model

    One click downloads and verifies it. Resumable, and safe to close mid-download.

  3. 3

    Start talking

    Type. Add more model types later, whenever you want them.

What can my Mac run? →

Built for your Mac

Built for Apple Silicon, not ported to it.

  • Tuned for Apple Silicon with MLX, the fast path on M-series chips.
  • Runs frontier models like DeepSeek V4 Flash entirely on your Mac.
  • Real control over quants, context, sampling, mixture-of-experts and speculative decoding.
  • GGUF models work too. Broad compatibility, usually slower, and labelled that way.
  • We name the engines we build on: mlx-lm, mflux, mlx-audio, llama.cpp, Core ML.

See what runs on your Mac →

Control + safety

Run big models without nuking your Mac.

Bedrova tracks memory across every loaded model and checks a new one against what's left before loading it. Either it fits, or you get told why not and what to try instead.

Qwen3.5 · 4.3 GBFLUX.1 · 11 GBWhisper · 3.1 GB 13.6 GB free
FLUX.1 dev · 24 GB — won't fit → try FLUX.1 schnell · 11 GB

Your Mac is
the whole datacentre.

  • Private — Prompts, files and models never leave your Mac.
  • Offline — Works on a plane, in a basement, anywhere.
  • Free and unlimited — No per-token bills, no rate limits.
  • Yours — No account. Nobody's permission needed.

Every kind of AI model, running on your Mac.

Free, native, and entirely on-device. One runtime for all of it.

Running AI at work? Packwolf automates recurring work with agents, and ideius builds AI for your business.