I did not want to be bothered by how many tokens I use for learning about LLM.
Thus I set up the local LLM environment where I bought GMKtec EVO-X2 AI MiniPC with 128 Gbyte memory and 2 TByte SSD
https://www.gmktec.com/products/amd-ryzen%E2%84%A2-ai-max-395-evo-x2-ai-mini-pc?variant=46826048585882.
This machine has AMD Ryzen AI-Max+ 395 (16 cores and 32 threads) and the 128 Gbyte memory can be partitioned to VRAM and memory.
I used the BIOS to set it to have VRAM for 96 GBytes so that you can run a rather big open LLM.
The size of a local LLM is big and must fit in VRAM. For example, qwen3-coder-next:latest is 51GB.
The following is for Windows 11.
Update 9/18/2026
Here is the command and the result:
> winget list python Name Id Version Available Source ---------------------------------------------------------------------- Python Launcher Python.Launcher < 3.14.7 3.14.7 winget Python 3.14.6 (64-bit) 9NCVDN91XZQP 3.14.6150.0 msstore Python 3.12.10 (64-bit) 9NCVDN91XZQP 3.12.10150.0 msstore
I had minforge3 installed python. Here is how to remove it.
> winget list python Name Id Version Available Source ----------------------------------------------------------------------------------------------- Miniforge3 25.3.0-3 (Python 3.12.10 64-bit) CondaForge.Miniforge3 25.3.0-3 26.3.2-3 winget Python Launcher Python.Launcher > 3.13.5 winget Python 3.14.6 (64-bit) 9NCVDN91XZQP 3.14.6150.0 msstore > winget uninstall --id CondaForge.Miniforge3 Found Miniforge3 25.3.0-3 (Python 3.12.10 64-bit) [CondaForge.Miniforge3] Starting package uninstall... Successfully uninstalled > winget list python Name Id Version Source ----------------------------------------------------------- Python Launcher Python.Launcher > 3.13.5 winget Python 3.14.6 (64-bit) 9NCVDN91XZQP 3.14.6150.0 msstore
Windows installs a utility called py.exe.
>py -3.14 Python 3.14.6 (tags/v3.14.6:c63aec6, Jun 10 2026, 10:26:10) [MSC v.1944 64 bit (AMD64)] on win32 >py -3.12 Python 3.12.10 (tags/v3.12.10:0cc8128, Apr 8 2025, 12:21:36) [MSC v.1943 64 bit (AMD64)] on win32
Python virtual environment creates the isolated workspace for Python interpreter, libraries and scripts. It prevents dependency conflicts. Here is how to set up the environment in the project directory
Creates the virtual environment directory, .venv, in the project directory. Using cmd
> python -m venv .venv
The directory name, .venv, can be anything and it contains include, Lib, Scripts directories and two files, .gitignore and pyvenv.cfg. Now activate the virtual environment using cmd
> .venv\Scripts\activate (.venv) (current directory)>
After you get the prompt modified with (.venv), you can install Python packages into the virtual environment without modifying the system Python directory
(.venv) (current directory)> pip install
In order to get out of the virtual environment, you issue the following command
(.venv) deactivate
This is the step.
0. Set the environmental variables (see next) 1. install ClaudeCode agent > winget install Anthropic.ClaudeCode 2. install a local LLM qwen3-coder-next:latest > ollama pull qwen3-coder-next:latest 3. launch ClaudeCode > ollama launch claude --model qwen3-coder-next
When I issue "ollama launch claude --model qwen3-coder-next:latest" After updating ClaudeCode
>winget update Anthropic.ClaudeCode
(installed v2.1.268) and it produced the yellow text warning of the following:
"qwen3-coder-next:latest" isn't described by this version's model catalog; update Claude Code, or map it with behavesAs on a modelPicker row (or modelOverrides, if it is a provider id of a model this version knows). Until then auto-compact keeps this session within 200k tokens (the context window it assumes); if the model accepts more, append [1m] to the model name for 1M, or set CLAUDE_CODE_MAX_CONTEXT_TOKENS to its real window; CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 restores the previous wait-for-the-API behavior.
I was able to eliminate the warning doing the following windows .bat file before invoking ollama launch:
rem ollama local LLM port is 11434 and openAI compatible path /v1, LM Studio 1234, Llama.cpp 8001 or 8080 set ANTHROPIC_BASE_URL=http://localhost:11434/v1 rem ollama context default to 4K for < 24 GB, 32K for 24GB-48GB, 256K for >= 48 GB (docs.ollama.com/context-length) set OLLAMA_CONTEXT_LENGTH=262144 rem qwen3-coder-next context length = 256K = 262144 set CLAUDE_CODE_MAX_CONTEXT_TOKENS=262144 rem Q4_K_M (4bit) VRAM 52GB for 256K token rem Q5_K_M (5bit) VRAM 64GB rem FP8 (8bit) VRAM 86GB for 1xH200
Check this link https://ollama.com/search.
The link is https://www.infoworld.com/article/4218328/how-to-get-better-results-from-local-llms-with-ollama.html.