Skip to content

Docs

Documentation

Install Bedrova, understand the runtime, and run every model type locally — chat, images, voice, vision, embeddings, and your own documents.

New here? Start with Getting started to install and run your first model, then read Concepts for the mental model the rest of the docs assume — one API, models in parallel, and RAM budgeted across everything.

Start here

  • Getting started — download to first chat in a few minutes.
  • Concepts — the runtime, parallelism, pinning vs JIT, the memory budget.
  • What your Mac can run — how the fit-check decides what loads.
  • Glossary — TTFT, quant, MoE, draft models, unified memory, and the rest.

Use a model type

Per-modality guides: Chat & reasoning, Structured output, Tool calling & MCP, Images, Voice, Vision, Documents & RAG, and Embeddings.

Tune the runtime

Go deeper on Runtime & engines, Memory & loading, Performance & speculative decoding, and Parallel requests & scheduling.

Build against it

Point your own apps at Bedrova: Server overview, the OpenAI, Anthropic, and management API references, the CLI, Integrations, and the full Settings reference.

Trust & help

Privacy & security, every network call, model trust & verification, and where your data lives. Stuck? See Troubleshooting, the FAQ, diagnostics & support bundles, or the community.