Guide · 6 min read

Running local AI on your Mac

Your Mac already has everything needed to run a serious language model — no cloud, no API keys, no monthly bill. This guide covers the three easiest ways to do it, how much RAM you actually need, and how to put the model to work in Telegram.

Why run models locally?

Three reasons, in order of how often people actually care:

The trade-off is hardware: bigger models need more RAM and run slower than a top-tier cloud API. For most everyday assistant work — reminders, summaries, drafting, answering questions — a well-chosen local model is genuinely good enough.

The three ways to run models

mlx-serve — fastest on Apple Silicon

Apple's MLX framework is built specifically for Apple Silicon's unified memory, and mlx-serve is the easiest way to get an OpenAI-compatible server from it. It's the default brain for Telebot AI — start it, point the bot at http://localhost:11234, done.

pip install mlx-serve
mlx-serve

Ollama — one command, huge model library

Ollama is the most popular local runtime for good reason: install once, then pull any model from a large library with a single command. It also speaks the OpenAI-compatible API out of the box.

brew install ollama
ollama pull llama3.1
ollama serve

LM Studio — point and click

If you prefer a GUI, LM Studio wraps everything in a clean app: browse models, download, chat, and run a local API server — all without touching a terminal.

How much RAM do you need?

Model size is the honest answer — a 7-billion-parameter model needs roughly 5–7 GB of memory at typical quantization, plus room for the system. Rough guide for M-series Macs:

RAMComfortable model sizeGood picks
8 GB1–3B parametersQwen 2.5 1.5B, Llama 3.2 3B
16 GB7–8B parametersLlama 3.1 8B, Qwen 2.5 7B, Gemma 2 9B
32 GB13–14B parametersQwen 2.5 14B, Llama 3.3 70B (quantized)
64 GB+30B+ parametersQwen 2.5 32B, Mixtral 8x7B, Llama 3.3 70B

Rule of thumb: pick the biggest model that fits in half your RAM — that leaves the rest for the operating system, your browser, and the bot itself.

Connecting it to Telegram

Once a local server is running, any OpenAI-compatible client can use it. Telebot AI is a macOS menu bar app that bridges your Telegram bots to exactly that server — or to any cloud API — with automatic fallbacks, memory, and 30+ tools like reminders, weather and stocks.

Getting started takes about two minutes: download the app, paste your @BotFather token, and point it at localhost. Everything — history, memories, skills — stays in ~/.mlx-serve on your Mac.

Download Telebot AI