Guide · 6 min read
Running local AI on your Mac
Your Mac already has everything needed to run a serious language model — no cloud, no API keys, no monthly bill. This guide covers the three easiest ways to do it, how much RAM you actually need, and how to put the model to work in Telegram.
Why run models locally?
Three reasons, in order of how often people actually care:
- Privacy. Your prompts and chat history never leave the machine.
- It's free. No per-token pricing, no subscriptions — just electricity.
- It works offline. No internet, no problem. Airplane mode included.
The trade-off is hardware: bigger models need more RAM and run slower than a top-tier cloud API. For most everyday assistant work — reminders, summaries, drafting, answering questions — a well-chosen local model is genuinely good enough.
The three ways to run models
mlx-serve — fastest on Apple Silicon
Apple's MLX framework is built specifically for Apple Silicon's unified memory, and mlx-serve is the easiest way to get an OpenAI-compatible server from it. It's the default brain for Telebot AI — start it, point the bot at http://localhost:11234, done.
pip install mlx-serve mlx-serve
Ollama — one command, huge model library
Ollama is the most popular local runtime for good reason: install once, then pull any model from a large library with a single command. It also speaks the OpenAI-compatible API out of the box.
brew install ollama ollama pull llama3.1 ollama serve
LM Studio — point and click
If you prefer a GUI, LM Studio wraps everything in a clean app: browse models, download, chat, and run a local API server — all without touching a terminal.
How much RAM do you need?
Model size is the honest answer — a 7-billion-parameter model needs roughly 5–7 GB of memory at typical quantization, plus room for the system. Rough guide for M-series Macs:
| RAM | Comfortable model size | Good picks |
|---|---|---|
| 8 GB | 1–3B parameters | Qwen 2.5 1.5B, Llama 3.2 3B |
| 16 GB | 7–8B parameters | Llama 3.1 8B, Qwen 2.5 7B, Gemma 2 9B |
| 32 GB | 13–14B parameters | Qwen 2.5 14B, Llama 3.3 70B (quantized) |
| 64 GB+ | 30B+ parameters | Qwen 2.5 32B, Mixtral 8x7B, Llama 3.3 70B |
Rule of thumb: pick the biggest model that fits in half your RAM — that leaves the rest for the operating system, your browser, and the bot itself.
Connecting it to Telegram
Once a local server is running, any OpenAI-compatible client can use it. Telebot AI is a macOS menu bar app that bridges your Telegram bots to exactly that server — or to any cloud API — with automatic fallbacks, memory, and 30+ tools like reminders, weather and stocks.
Getting started takes about two minutes: download the app, paste your @BotFather token, and point it at localhost. Everything — history, memories, skills — stays in ~/.mlx-serve on your Mac.