AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: Can I run LLMs locally instead of using APIs?

Yes. Open source models (Llama, Mistral, Qwen, Gemma, DeepSeek) can run on your own machine, and tools like Ollama and LM Studio make it a ten-minute setup: install, pull a model, chat with it offline.

One thing to be clear about: the frontier models cannot be run locally. GPT, Claude, and Gemini are proprietary. They exist only on the servers of OpenAI, Anthropic, and Google, and the API is the only way to use them. Running locally always means running open source models.

What you gain: privacy (your data never leaves your machine), zero API costs, offline access, and a deeper feel for how models behave. What you trade away: quality and convenience. Small local models are genuinely useful but noticeably behind frontier models on hard tasks, and larger local models need serious hardware, like a modern GPU or a well-equipped Mac.

One clarification, because the word "local" confuses people: companies also run open source models "locally" in the sense of on their own cloud infrastructure, so sensitive data stays inside their environment. Same idea as your laptop, at enterprise scale, and it's a real slice of AI engineering work in regulated industries.

← Back to the full FAQ