Claude Code Locally on 16 GB VRAM: From Endless Struggle to Unexpected Triumph
Introduction: The Battle with the Process
Hello, friends! I’m Mebli, and today I want to share a story that many local AI enthusiasts will recognize. We discovered that it’s possible to run a fully functional Claude Code on a laptop with just 16 GB VRAM — completely offline, without paying for Anthropic API. But before the sweet victory, we had to go through what I can only call «the battle with the process» — a long chain of errors, experiments, and late-night debugging sessions.
What Does «Battle with the Process» Mean?
It starts when you think connecting a local model to Claude Code will be straightforward. In reality, the model sometimes «forgets» tools, fails to maintain context, or outputs formats that the Claude Code parser simply doesn’t understand. That’s when the real fight begins.
First Attempts: Classic Qwen3.5
We started with the fast qwen3.5:9b-q8_0. At first glance, everything looked promising: the model recognized tools, the parser picked them up, and we moved forward. But soon serious issues appeared:
- The model returned XML wrapped in <details> and <think> tags instead of clean JSON/tool_use format.
- Claude Code lost internal state and hung on the first tool call.
- On some prompts the model ignored tools completely and sent plain text.
I tried different quantizations (q4_K_M, q5_K_M), but the problems persisted at about 70–80% success rate.
The Turning Point: Abliterated Models
In 2025 it became clear that Claude Code is not just another API-compatible tool. It expects a very specific tool-calling format and strict response structure. That’s why I turned to «abliterated» versions — models where refusal and censorship mechanisms have been removed, making them much more obedient to custom instructions.
Key Discovery: huihui_ai/qwen3.5-abliterated:9b-Claude
The model that finally worked perfectly is huihui_ai/qwen3.5-abliterated:9b-Claude. Here’s why it succeeded:
- Fine-tuned in «Claude style» with excellent tool_use compliance.
- No refusal behavior — the model follows instructions precisely.
- 9B parameters + q8_0 quantization fits comfortably into 16 GB VRAM with 32k+ context.
After switching, Claude Code immediately started reading files, executing commands, spawning sub-agents — everything worked smoothly without parser crashes.
Technical Setup That Actually Works
Here are the exact environment variables (important: use the IP of your Ollama server, not just localhost if it runs on another machine):
export ANTHROPIC_BASE_URL=http://192.168.0.108:11434
export ANTHROPIC_AUTH_TOKEN=ollama
export ANTHROPIC_API_KEY=""
Launch command:
claude --model huihui_ai/qwen3.5-abliterated:9b-Claude
If the model is not downloaded yet:
ollama pull huihui_ai/qwen3.5-abliterated:9b-Claude
Install Claude Code globally:
npm install -g @anthropic-ai/claude-code
Lessons Learned from the Battle
- Format is everything. Even small XML or <think> tags break the parser.
- Temperature matters. The sweet spot was 0.75–0.85 with top_p 0.95.
- IP connection instead of localhost was one of the trickiest hidden issues.
- Abliterated + Claude-style fine-tuning beats regular Qwen3.5 by a huge margin for tool-calling tasks.
Conclusion: Worth the Struggle
After dozens of models, GitHub issues, and late nights, we finally achieved a stable local Claude Code setup. I’ve completely ditched the paid Anthropic subscription and Cursor for most daily coding tasks. Now code is written on a remote machine or laptop without token counters or monthly fees.
This «random triumph» was actually the result of systematic problem-solving and persistence. If you have 16 GB VRAM and want to try it yourself — copy the commands above, set the correct IP, and make sure environment variables are loaded. Feel free to ask questions in the comments if you run into issues on Windows, specific GPUs, or other setups.
«The key is not just connecting the model, but making it speak the exact language Claude Code expects.»
Comments (0)