# macOS and Linux -- one command set up\ncurl -fsSL https:\/\/ollama.com\/set up.sh | sh\n\n# Confirm the model -- have to be 0.14.0+ for Claude Code compatibility\nollama model\n# Anticipated: ollama model is 0.14.x or increased\n\n# Home windows: obtain the installer from https:\/\/ollama.com\n# Native Home windows help has improved considerably in latest releases<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
After set up, Ollama begins routinely as a background service on port 11434<\/strong>. You possibly can confirm it’s working:<\/p>\n
\n# Test the Ollama server is stay\ncurl http:\/\/localhost:11434\n\n# Anticipated response:\n# Ollama is working<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Pull a coding mannequin:<\/p>\n
\n# GLM-4.7-Flash -- beneficial place to begin\n# Sturdy device calling, 128K context, matches on 8 GB VRAM\n# Apache 2.0 license\nollama pull glm-4.7-flash:newest\n\n# Qwen3-Coder -- sturdy code technology and instruction following\n# Requires 20+ GB VRAM for the total mannequin\nollama pull qwen3-coder\n\n# Devstral-Small -- particularly designed for agentic coding workflows\n# Neighborhood-tested for Claude Code compatibility\n# 24B, requires 16+ GB VRAM\nollama pull devstral-small-2:24b\n\n# Confirm the mannequin is downloaded and prepared\nollama listing\n# Reveals all pulled fashions with their sizes and modification dates<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
\/\/\u00a0<\/span>Configuring Claude Code to Use Ollama<\/h4>\nChoice 1: Shell export (present terminal session solely)<\/strong><\/p>\n
\n# Redirect Claude Code to your native Ollama server\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:11434\"\n\n# Native servers don't require actual authentication\n# Set these to any non-empty string -- Ollama ignores the worth\nexport ANTHROPIC_API_KEY=\"ollama\"\nexport ANTHROPIC_AUTH_TOKEN=\"ollama\"\n\n# Map Claude Code's mannequin tier requests to your native mannequin title\n# Claude Code internally requests sonnet\/haiku\/opus -- these variables\n# translate these tier names to no matter mannequin you could have pulled domestically\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"glm-4.7-flash:newest\"\n\n# Launch Claude Code -- it is going to now use Ollama as an alternative of the Anthropic API\nclaude<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Choice 2: ~\/.claude\/settings.json<\/code> (everlasting, applies to all classes)<\/strong><\/p>\n
This strategy survives terminal restarts and applies each time you launch Claude Code. Claude Code reads setting variables from settings.json<\/code> at startup in order that they take impact irrespective of how claude<\/code> was launched.<\/p>\n
Create or edit ~\/.claude\/settings.json<\/code>:<\/p>\n
\n{\n  \"env\": {\n    \"ANTHROPIC_BASE_URL\": \"http:\/\/localhost:11434\",\n    \"ANTHROPIC_API_KEY\": \"ollama\",\n    \"ANTHROPIC_AUTH_TOKEN\": \"ollama\",\n    \"ANTHROPIC_DEFAULT_SONNET_MODEL\": \"glm-4.7-flash:newest\",\n    \"ANTHROPIC_DEFAULT_HAIKU_MODEL\": \"glm-4.7-flash:newest\",\n    \"ANTHROPIC_DEFAULT_OPUS_MODEL\": \"glm-4.7-flash:newest\"\n  }\n}<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Choice 3: .env<\/code> file in undertaking listing (per-project override)<\/strong><\/p>\n
If you’d like a selected undertaking to make use of a special mannequin whereas holding your world settings on the Anthropic API:<\/p>\n
\n# .env in your undertaking root -- loaded routinely by Claude Code\nANTHROPIC_BASE_URL=http:\/\/localhost:11434\nANTHROPIC_API_KEY=ollama\nANTHROPIC_AUTH_TOKEN=ollama\nANTHROPIC_DEFAULT_SONNET_MODEL=qwen3-coder\nANTHROPIC_DEFAULT_HAIKU_MODEL=qwen3-coder\nANTHROPIC_DEFAULT_OPUS_MODEL=qwen3-coder<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Confirm the connection:<\/p>\n
\n# Launch Claude Code with a easy check\nclaude\n\n# Inside Claude Code, run a fundamental immediate:\n# > What mannequin are you working?\n# A neighborhood mannequin ought to reply with out making any Anthropic API calls.\n\n# To substantiate no exterior calls are being made, run with verbose logging:\nclaude --verbose\n\n# Search for traces displaying requests going to localhost:11434\n# quite than api.anthropic.com<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Full working sequence from scratch:<\/p>\n
\ncurl -fsSL https:\/\/ollama.com\/set up.sh | sh          # 1. Set up Ollama\nollama pull glm-4.7-flash:newest                       # 2. Pull mannequin (~4 GB)\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:11434\"     # 3. Redirect Claude Code\nexport ANTHROPIC_API_KEY=\"ollama\"                      # 4. Set placeholder auth\nexport ANTHROPIC_AUTH_TOKEN=\"ollama\"\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"glm-4.7-flash:newest\"\nclaude                                                  # 5. Launch<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
#\u00a0<\/span>Backend 2: LM Studio<\/h2>\n\u00a0
LM Studio is the fitting alternative if you would like a graphical interface for searching and managing fashions quite than working solely within the terminal. Since model 0.4.1, it features a native Anthropic-compatible \/v1\/messages<\/strong> endpoint \u2014 the identical path Claude Code expects \u2014 so no translation layer or proxy is required.<\/p>\n
Conditions:<\/strong><\/p>\n
\nmacOS, Home windows, or Linux\n<\/li>\n
GPU with 6+ GB VRAM beneficial (CPU-only is feasible however sluggish)\n<\/li>\n
Obtain from lmstudio.ai or use the CLI installer for headless servers\n<\/li>\n<\/ul>\nSet up and configure LM Studio:<\/p>\n
\n# On a server or VM with out a GUI -- CLI installer\ncurl -fsSL https:\/\/releases.lmstudio.ai\/cli\/set up.sh | bash\n\n# Or obtain the desktop app from https:\/\/lmstudio.ai for GUI use<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
GUI setup steps:<\/p>\n
\nOpen LM Studio and seek for a coding mannequin (search “qwen coder” or “devstral”).\n<\/li>\n
Obtain the mannequin. LM Studio handles quantization choice routinely.\n<\/li>\n
Go to the Native Server<\/strong> tab (the <><\/code> icon within the left sidebar).\n<\/li>\n
Set the context measurement. LM Studio recommends beginning with no less than 25,000 tokens and rising for higher outcomes.\n<\/li>\n
Click on Begin Server<\/strong>.\n<\/li>\n
Notice the port (default: 1234) and duplicate the mannequin title precisely as proven.\n<\/li>\n<\/ol>\n\u00a0<\/p>\n
\n\nNotice: Copy the mannequin identifier precisely. LM Studio shows the precise string it is advisable cross to ANTHROPIC_DEFAULT_SONNET_MODEL<\/code>. A mismatch right here is the commonest failure mode.\n<\/p>\n<\/blockquote>\n
\u00a0<\/p>\n
Configure Claude Code:<\/p>\n
\n# Set the bottom URL to LM Studio's native server\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:1234\"\nexport ANTHROPIC_API_KEY=\"lm-studio\"\nexport ANTHROPIC_AUTH_TOKEN=\"lm-studio\"\n\n# Substitute the mannequin title with what LM Studio exhibits to your loaded mannequin\n# Copy it precisely -- together with any model suffix or quantization tag\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"qwen2.5-coder-32b-instruct\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"qwen2.5-coder-32b-instruct\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"qwen2.5-coder-32b-instruct\"<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Or persistently in ~\/.claude\/settings.json<\/code>:<\/p>\n
\n{\n  \"env\": {\n    \"ANTHROPIC_BASE_URL\": \"http:\/\/localhost:1234\",\n    \"ANTHROPIC_API_KEY\": \"lm-studio\",\n    \"ANTHROPIC_AUTH_TOKEN\": \"lm-studio\",\n    \"ANTHROPIC_DEFAULT_SONNET_MODEL\": \"qwen2.5-coder-32b-instruct\",\n    \"ANTHROPIC_DEFAULT_HAIKU_MODEL\": \"qwen2.5-coder-32b-instruct\",\n    \"ANTHROPIC_DEFAULT_OPUS_MODEL\": \"qwen2.5-coder-32b-instruct\"\n  }\n}<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Find out how to run:<\/p>\n
\n# 1. Begin the LM Studio server from the GUI (Native Server tab > Begin Server)\n# 2. Set setting variables\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:1234\"\nexport ANTHROPIC_API_KEY=\"lm-studio\"\nexport ANTHROPIC_AUTH_TOKEN=\"lm-studio\"\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"your-model-name-here\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"your-model-name-here\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"your-model-name-here\"\n# 3. Launch\nclaude<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
#\u00a0<\/span>Backend 3: llama.cpp<\/h2>\n\u00a0
llama.cpp<\/strong> is the fitting alternative while you want direct management over inference parameters \u2014 quantization kind, KV cache configuration, batch measurement, thread rely \u2014 or if you find yourself working on a server and wish the bottom overhead. It has native Anthropic Messages API help, so no proxy or translation layer is required.<\/p>\n
Conditions:<\/strong><\/p>\n
\nA GGUF-format mannequin file (obtain from Hugging Face; seek for “GGUF” variations of any mannequin)\n<\/li>\n
CUDA-capable GPU for GPU inference, or CPU-only for slower inference\n<\/li>\n
CMake and a C++ compiler for supply builds (on Linux\/CUDA, supply is beneficial)\n<\/li>\n<\/ul>\nSet up llama.cpp:<\/p>\n
\n# macOS -- Homebrew is easiest\nbrew set up llama.cpp\n\n# Linux with CUDA -- construct from supply for greatest GPU efficiency\ngit clone https:\/\/github.com\/ggml-org\/llama.cpp\ncd llama.cpp\ncmake -B construct -DGGML_CUDA=ON          # Allow CUDA acceleration\ncmake --build construct --config Launch   # Construct\n# Binaries in .\/construct\/bin\/\n\n# Linux CPU-only construct\ncmake -B construct\ncmake --build construct --config Launch\n\n# Home windows -- pre-built binaries out there at:\n# https:\/\/github.com\/ggml-org\/llama.cpp\/releases\n# Obtain the CUDA or CPU variant matching your {hardware}<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Obtain a GGUF mannequin:<\/p>\n
\n# Set up the Hugging Face CLI in the event you shouldn't have it\npip set up huggingface-hub\n\n# Obtain GLM-4.7-Flash in Q4_K_XL quantization (~4.5 GB)\n# This quantization presents an excellent measurement\/high quality stability for coding\nhuggingface-cli obtain unsloth\/GLM-4.7-Flash-GGUF \n  GLM-4.7-Flash-UD-Q4_K_XL.gguf \n  --local-dir .\/fashions\/\n\n# Or obtain Qwen3-Coder in This autumn quantization (~15 GB for 32B)\nhuggingface-cli obtain Qwen\/Qwen3-Coder-32B-Instruct-GGUF \n  qwen3-coder-32b-instruct-q4_k_m.gguf \n  --local-dir .\/fashions\/<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Begin the llama.cpp server:<\/p>\n
\n# Begin llama-server with Anthropic API help and a 128K context window\nllama-server \n  --model .\/fashions\/GLM-4.7-Flash-UD-Q4_K_XL.gguf \n  --alias \"glm-4.7-flash\"           # This title goes in ANTHROPIC_DEFAULT_SONNET_MODEL\n  --port 8001 \n  --ctx-size 131072                 # 128K context -- necessary for big codebases\n  --flash-attn                      # Reminiscence-efficient consideration, improves pace\n  --n-gpu-layers 99                  # Offload all layers to GPU; take away for CPU-only\n\n# For CPU-only inference (no GPU):\nllama-server \n  --model .\/fashions\/GLM-4.7-Flash-UD-Q4_K_XL.gguf \n  --alias \"glm-4.7-flash\" \n  --port 8001 \n  --ctx-size 32768                  # Cut back context measurement on CPU to maintain reminiscence manageable\n  --threads 8                        # Match your CPU core rely<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Key flags defined:<\/p>\n
\n--alias<\/code>: the mannequin title string Claude Code will ship in requests. Set ANTHROPIC_DEFAULT_SONNET_MODEL<\/code> to match this precisely.\n<\/li>\n
--ctx-size<\/code>: context window in tokens. 131072 = 128K<\/strong>. Bigger is best for codebase evaluation however makes use of extra VRAM. Cut back in the event you get out-of-memory errors.\n<\/li>\n
--flash-attn<\/code>: Flash Consideration reduces peak VRAM by processing consideration in smaller blocks. Allow it at any time when your construct helps it.\n<\/li>\n
--n-gpu-layers 99<\/code>: offloads all transformer layers to the GPU. The server routinely makes use of fewer layers if VRAM is tight.\n<\/li>\n<\/ul>\nConfigure Claude Code:<\/p>\n
\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:8001\"\nexport ANTHROPIC_API_KEY=\"llama-cpp\"\nexport ANTHROPIC_AUTH_TOKEN=\"llama-cpp\"\n\n# Should match the --alias you handed to llama-server precisely\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"glm-4.7-flash\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"glm-4.7-flash\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"glm-4.7-flash\"<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Find out how to run:<\/p>\n
\n# Terminal 1: begin the llama.cpp server\nllama-server \n  --model .\/fashions\/GLM-4.7-Flash-UD-Q4_K_XL.gguf \n  --alias \"glm-4.7-flash\" \n  --port 8001 \n  --ctx-size 131072 \n  --flash-attn \n  --n-gpu-layers 99\n\n# Terminal 2: configure and launch Claude Code\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:8001\"\nexport ANTHROPIC_API_KEY=\"llama-cpp\"\nexport ANTHROPIC_AUTH_TOKEN=\"llama-cpp\"\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"glm-4.7-flash\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"glm-4.7-flash\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"glm-4.7-flash\"\nclaude<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
#\u00a0<\/span>The Full settings.json<\/code><\/h2>\n\u00a0
Atmosphere variable exports final solely so long as the terminal session. For a sturdy configuration, use ~\/.claude\/settings.json<\/code>. Claude Code reads variables from this file at startup in order that they apply irrespective of how Claude was launched \u2014 from the terminal, from a VS Code process, or from a script.<\/p>\n
Here’s a production-ready settings.json<\/code> with all variables defined:<\/p>\n
\n{\n  \"env\": {\n    \"ANTHROPIC_BASE_URL\": \"http:\/\/localhost:11434\",\n\n    \"ANTHROPIC_API_KEY\": \"ollama\",\n    \"ANTHROPIC_AUTH_TOKEN\": \"ollama\",\n\n    \"ANTHROPIC_DEFAULT_SONNET_MODEL\": \"glm-4.7-flash:newest\",\n    \"ANTHROPIC_DEFAULT_HAIKU_MODEL\": \"glm-4.7-flash:newest\",\n    \"ANTHROPIC_DEFAULT_OPUS_MODEL\": \"glm-4.7-flash:newest\",\n\n    \"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS\": \"1\"\n  }\n}<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Why CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: \"1\"<\/code> issues:<\/strong><\/p>\n
When utilizing Claude Code by way of non-Anthropic backends, Claude Code provides Anthropic-specific experimental beta flags to request headers \u2014 flags that third-party and native servers don’t acknowledge. This causes Error: Sudden worth(s) for the anthropic-beta header<\/code> on most native inference servers. Setting this variable to \"1\"<\/code> strips these headers earlier than the request goes out, which eliminates the error with out affecting any core Claude Code performance.<\/p>\n
Switching between backends:<\/strong><\/p>\n
Should you work with a number of backends \u2014 Ollama for every day use, the Anthropic API for advanced duties \u2014 the cleanest strategy is sustaining separate shell scripts quite than modifying settings.json<\/code> backwards and forwards:<\/p>\n
\n# use-local.sh -- swap to Ollama\nexport ANTHROPIC_BASE_URL=\"http:\/\/localhost:11434\"\nexport ANTHROPIC_API_KEY=\"ollama\"\nexport ANTHROPIC_AUTH_TOKEN=\"ollama\"\nexport ANTHROPIC_DEFAULT_SONNET_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_HAIKU_MODEL=\"glm-4.7-flash:newest\"\nexport ANTHROPIC_DEFAULT_OPUS_MODEL=\"glm-4.7-flash:newest\"\necho \"Claude Code \u2192 native Ollama (glm-4.7-flash)\"<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
\n# use-anthropic.sh -- swap again to the Anthropic API\nunset ANTHROPIC_BASE_URL\nunset ANTHROPIC_AUTH_TOKEN\nunset ANTHROPIC_DEFAULT_SONNET_MODEL\nunset ANTHROPIC_DEFAULT_HAIKU_MODEL\nunset ANTHROPIC_DEFAULT_OPUS_MODEL\n# ANTHROPIC_API_KEY ought to already be set to your actual key in your rc file\necho \"Claude Code \u2192 Anthropic API\"<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
Supply both script in your present session:<\/p>\n
\nsupply .\/use-local.sh\nclaude\n\n# Whenever you want the true API for a posh process:\nsupply .\/use-anthropic.sh\nclaude<\/code><\/pre>\n<\/div>\n\u00a0<\/p>\n
#\u00a0<\/span>Finest Native Fashions for Claude Code in 2026<\/h2>\n\u00a0
{Hardware} is the principle constraint. For Claude Code with native fashions to be genuinely usable for coding duties quite than only a demo, purpose for 32 GB of RAM \u2014 Apple Silicon unified reminiscence or PC RAM. 16 GB is viable with smaller quantized fashions and CPU offload, however technology pace can be noticeably slower on multi-step agentic duties.<\/p>\n
\u00a0<\/p>\n\n\n\n\n\n\n\n\n\nMannequin<\/strong><\/th>\n VRAM Wanted<\/strong><\/th>\n Context<\/strong><\/th>\n Strengths<\/strong><\/th>\n License<\/strong><\/th>\n Pull Command<\/strong><\/th>\n<\/tr>\n<\/thead>\n
glm-4.7-flash<\/a><\/td>\n 8 GB<\/td>\n 128K<\/td>\n Instrument calling, quick, low VRAM<\/td>\n Apache 2.0<\/td>\n ollama pull glm-4.7-flash<\/code><\/td>\n<\/tr>\n
devstral-small-2:24b<\/a><\/td>\n 16 GB<\/td>\n 32K<\/td>\n Agentic coding workflows<\/td>\n Apache 2.0<\/td>\n ollama pull devstral-small-2:24b<\/code><\/td>\n<\/tr>\n
qwen3-coder<\/a><\/td>\n 20 GB<\/td>\n 128K<\/td>\n Code technology, directions<\/td>\n Apache 2.0<\/td>\n ollama pull qwen3-coder<\/code><\/td>\n<\/tr>\n
qwen3.5:27b<\/a><\/td>\n 20 GB<\/td>\n 256K<\/td>\n Sturdy all-round, enormous context<\/td>\n Apache 2.0<\/td>\n ollama pull qwen3.5:27b<\/code><\/td>\n<\/tr>\n
gemma4:26b<\/a><\/td>\n 20 GB<\/td>\n 256K<\/td>\n Reasoning, 77% coding bench<\/td>\n Gemma License<\/td>\n ollama pull gemma4:26b<\/code><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n\u00a0<\/p>\n
#\u00a0<\/span>Troubleshooting Frequent Points<\/h2>\n\u00a0<\/p>\n
\nConnection refused when launching Claude Code:<\/strong> The inference server will not be working. That is the commonest difficulty and the best to diagnose.\n\n# Test if Ollama is working\ncurl http:\/\/localhost:11434\n# Anticipated: \"Ollama is working\"\n\n# Test if LM Studio server is working\ncurl http:\/\/localhost:1234\/v1\/fashions\n# Ought to return a JSON listing of loaded fashions\n\n# Test if llama-server is working\ncurl http:\/\/localhost:8001\/well being\n# Ought to return {\"standing\":\"okay\"}\n\n# If not working -- begin the server first, then launch Claude Code\nollama serve          # Ollama\n# LM Studio: use the GUI Native Server tab\n# llama.cpp: run the llama-server command from the Backend 3 part<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n
Mannequin not discovered or unknown mannequin error:<\/strong> The mannequin title in your ANTHROPIC_DEFAULT_SONNET_MODEL<\/code> doesn’t match what the server is aware of.\n\n# Checklist all fashions Ollama has out there\nollama listing\n\n# The mannequin title in ANTHROPIC_DEFAULT_SONNET_MODEL should match EXACTLY\n# together with the tag -- \"glm-4.7-flash:newest\" not \"glm-4.7-flash\"\n\n# Confirm with a direct API name to substantiate what the server sees\ncurl http:\/\/localhost:11434\/v1\/fashions<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n
Instrument calls failing or returning errors:<\/strong> For streaming device calls, which Claude Code makes use of when executing features or scripts, Ollama model 0.14.3-rc1 or later is required. Earlier variations within the 0.14.x collection had incomplete streaming device name help.\n\n# Test your Ollama model\nollama model\n\n# If beneath 0.14.3, replace Ollama\ncurl -fsSL https:\/\/ollama.com\/set up.sh | sh<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n
anthropic-beta<\/code> header error:\nYou will note: Error: Sudden worth(s) for the anthropic-beta header<\/code>. This occurs as a result of Claude Code provides Anthropic-specific experimental beta flags that native servers don’t acknowledge. Repair it by including this to your settings.json<\/code> env block:<\/p>\n
\n\"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS\": \"1\"<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n
Reverting to the Anthropic API:<\/strong>\n\n# Shell session -- unset the redirect variables\nunset ANTHROPIC_BASE_URL\nunset ANTHROPIC_AUTH_TOKEN\nunset ANTHROPIC_DEFAULT_SONNET_MODEL\nunset ANTHROPIC_DEFAULT_HAIKU_MODEL\nunset ANTHROPIC_DEFAULT_OPUS_MODEL\n\n# Then be certain that your actual API key's set\necho $ANTHROPIC_API_KEY\n# Ought to present your sk-ant-... key, not a placeholder\n\n# Should you used settings.json -- take away or remark out the env block\n# and restart Claude Code<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n
Sluggish technology pace:<\/strong> For agentic Claude Code duties, technology pace issues as a result of every device name is a spherical journey. If pace is insufficient:\n\nChange to a smaller or extra aggressively quantized mannequin (Q4_K_M as an alternative of Q8).\n<\/li>\n
Allow --flash-attn<\/code> in llama.cpp if not already set.\n<\/li>\n
Cut back context measurement (--ctx-size<\/code>); bigger contexts are slower to prefill.\n<\/li>\n
On Ollama, set OLLAMA_NUM_GPU_LAYERS=99<\/code> in your setting to pressure most GPU offload.\n<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n\u00a0<\/p>\n
#\u00a0<\/span>Conclusion<\/h2>\n\u00a0
What used to require fragile adapters and hacks is now a five-step course of. Set up the inference backend, pull a mannequin, set three setting variables, and Claude Code routes to your native machine as an alternative of Anthropic’s API. The configuration takes below 5 minutes after you have the mannequin downloaded.<\/p>\n
The sensible result’s a coding assistant that prices nothing to run after setup, has no price limits, retains your code solely in your machine, and covers the overwhelming majority of actual coding use circumstances at high quality ranges that weren’t out there in native fashions a yr in the past. Begin with Ollama and glm-4.7-flash<\/code> \u2014 it has the bottom {hardware} requirement, probably the most constant tool-calling help, and the quickest path to a working setup. As soon as that’s working, scale up the mannequin primarily based in your {hardware} and the standard degree you really want.
\u00a0
\u00a0<\/p>\n
Shittu Olumide<\/a><\/strong><\/strong><\/a> is a software program engineer and technical author enthusiastic about leveraging cutting-edge applied sciences to craft compelling narratives, with a eager eye for element and a knack for simplifying advanced ideas. It’s also possible to discover Shittu on Twitter<\/a>.<\/p>\n<\/p><\/div>\n
<\/script>
\n
<\/p>\n","protected":false},"excerpt":{"rendered":"
\u00a0 \u00a0 #\u00a0Introduction \u00a0Agentic coding classes are costly. A single Claude Code session \u2014 studying information, writing code, working checks, iterating \u2014 can burn 10\u201350x extra tokens than a plain chat dialog. At scale, that provides up quick. Add price limits that may interrupt a long-running workflow mid-session, and the dependency on a third-party API […]<\/p>\n","protected":false},"author":2,"featured_media":15702,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[458,977,1520,266,7009],"class_list":["post-15700","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-machine-learning","tag-claude","tag-code","tag-local","tag-models","tag-pairing"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=15700"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700\/revisions"}],"predecessor-version":[{"id":15701,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700\/revisions\/15701"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/15702"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=15700"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=15700"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=15700"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}

Mannequin<\/strong><\/th>\n	VRAM Wanted<\/strong><\/th>\n	Context<\/strong><\/th>\n	Strengths<\/strong><\/th>\n	License<\/strong><\/th>\n	Pull Command<\/strong><\/th>\n<\/tr>\n<\/thead>\n
glm-4.7-flash<\/a><\/td>\n	8 GB<\/td>\n	128K<\/td>\n	Instrument calling, quick, low VRAM<\/td>\n	Apache 2.0<\/td>\n	`ollama pull glm-4.7-flash<\/code><\/td>\n<\/tr>\n`
devstral-small-2:24b<\/a><\/td>\n	16 GB<\/td>\n	32K<\/td>\n	Agentic coding workflows<\/td>\n	Apache 2.0<\/td>\n	`ollama pull devstral-small-2:24b<\/code><\/td>\n<\/tr>\n`
qwen3-coder<\/a><\/td>\n	20 GB<\/td>\n	128K<\/td>\n	Code technology, directions<\/td>\n	Apache 2.0<\/td>\n	`ollama pull qwen3-coder<\/code><\/td>\n<\/tr>\n`
qwen3.5:27b<\/a><\/td>\n	20 GB<\/td>\n	256K<\/td>\n	Sturdy all-round, enormous context<\/td>\n	Apache 2.0<\/td>\n	`ollama pull qwen3.5:27b<\/code><\/td>\n<\/tr>\n`
gemma4:26b<\/a><\/td>\n	20 GB<\/td>\n	256K<\/td>\n	Reasoning, 77% coding bench<\/td>\n	Gemma License<\/td>\n	ollama pull gemma4:26b<\/code><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n\u00a0<\/p>\n #\u00a0<\/span>Troubleshooting Frequent Points<\/h2>\n\u00a0<\/p>\n \nConnection refused when launching Claude Code:<\/strong> The inference server will not be working. That is the commonest difficulty and the best to diagnose.\n\n# Test if Ollama is working \ncurl http:\/\/localhost:11434 \n# Anticipated: \"Ollama is working\" \n \n# Test if LM Studio server is working \ncurl http:\/\/localhost:1234\/v1\/fashions \n# Ought to return a JSON listing of loaded fashions \n \n# Test if llama-server is working \ncurl http:\/\/localhost:8001\/well being \n# Ought to return {\"standing\":\"okay\"} \n \n# If not working -- begin the server first, then launch Claude Code \nollama serve # Ollama \n# LM Studio: use the GUI Native Server tab \n# llama.cpp: run the llama-server command from the Backend 3 part<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n Mannequin not discovered or unknown mannequin error:<\/strong> The mannequin title in your ANTHROPIC_DEFAULT_SONNET_MODEL<\/code> doesn’t match what the server is aware of.\n\n# Checklist all fashions Ollama has out there \nollama listing \n \n# The mannequin title in ANTHROPIC_DEFAULT_SONNET_MODEL should match EXACTLY \n# together with the tag -- \"glm-4.7-flash:newest\" not \"glm-4.7-flash\" \n \n# Confirm with a direct API name to substantiate what the server sees \ncurl http:\/\/localhost:11434\/v1\/fashions<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n Instrument calls failing or returning errors:<\/strong> For streaming device calls, which Claude Code makes use of when executing features or scripts, Ollama model 0.14.3-rc1 or later is required. Earlier variations within the 0.14.x collection had incomplete streaming device name help.\n\n# Test your Ollama model \nollama model \n \n# If beneath 0.14.3, replace Ollama \ncurl -fsSL https:\/\/ollama.com\/set up.sh \| sh<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n anthropic-beta<\/code> header error:\nYou will note: Error: Sudden worth(s) for the anthropic-beta header<\/code>. This occurs as a result of Claude Code provides Anthropic-specific experimental beta flags that native servers don’t acknowledge. Repair it by including this to your settings.json<\/code> env block:<\/p>\n \n\"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS\": \"1\"<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n Reverting to the Anthropic API:<\/strong>\n\n# Shell session -- unset the redirect variables \nunset ANTHROPIC_BASE_URL \nunset ANTHROPIC_AUTH_TOKEN \nunset ANTHROPIC_DEFAULT_SONNET_MODEL \nunset ANTHROPIC_DEFAULT_HAIKU_MODEL \nunset ANTHROPIC_DEFAULT_OPUS_MODEL \n \n# Then be certain that your actual API key's set \necho $ANTHROPIC_API_KEY \n# Ought to present your sk-ant-... key, not a placeholder \n \n# Should you used settings.json -- take away or remark out the env block \n# and restart Claude Code<\/code><\/pre>\n<\/div>\n\u00a0\n<\/p>\n<\/li>\n Sluggish technology pace:<\/strong> For agentic Claude Code duties, technology pace issues as a result of every device name is a spherical journey. If pace is insufficient:\n\nChange to a smaller or extra aggressively quantized mannequin (Q4_K_M as an alternative of Q8).\n<\/li>\n Allow --flash-attn<\/code> in llama.cpp if not already set.\n<\/li>\n Cut back context measurement (--ctx-size<\/code>); bigger contexts are slower to prefill.\n<\/li>\n On Ollama, set OLLAMA_NUM_GPU_LAYERS=99<\/code> in your setting to pressure most GPU offload.\n<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n\u00a0<\/p>\n #\u00a0<\/span>Conclusion<\/h2>\n\u00a0 What used to require fragile adapters and hacks is now a five-step course of. Set up the inference backend, pull a mannequin, set three setting variables, and Claude Code routes to your native machine as an alternative of Anthropic’s API. The configuration takes below 5 minutes after you have the mannequin downloaded.<\/p>\n The sensible result’s a coding assistant that prices nothing to run after setup, has no price limits, retains your code solely in your machine, and covers the overwhelming majority of actual coding use circumstances at high quality ranges that weren’t out there in native fashions a yr in the past. Begin with Ollama and glm-4.7-flash<\/code> \u2014 it has the bottom {hardware} requirement, probably the most constant tool-calling help, and the quickest path to a working setup. As soon as that’s working, scale up the mannequin primarily based in your {hardware} and the standard degree you really want. \u00a0 \u00a0<\/p>\n Shittu Olumide<\/a><\/strong><\/strong><\/a> is a software program engineer and technical author enthusiastic about leveraging cutting-edge applied sciences to craft compelling narratives, with a eager eye for element and a knack for simplifying advanced ideas. It’s also possible to discover Shittu on Twitter<\/a>.<\/p>\n<\/p><\/div>\n <\/script> \n <\/p>\n","protected":false},"excerpt":{"rendered":" \u00a0 \u00a0 #\u00a0Introduction \u00a0Agentic coding classes are costly. A single Claude Code session \u2014 studying information, writing code, working checks, iterating \u2014 can burn 10\u201350x extra tokens than a plain chat dialog. At scale, that provides up quick. Add price limits that may interrupt a long-running workflow mid-session, and the dependency on a third-party API […]<\/p>\n","protected":false},"author":2,"featured_media":15702,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[458,977,1520,266,7009],"class_list":["post-15700","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-machine-learning","tag-claude","tag-code","tag-local","tag-models","tag-pairing"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=15700"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700\/revisions"}],"predecessor-version":[{"id":15701,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/15700\/revisions\/15701"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/15702"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=15700"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=15700"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=15700"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}