Running Local LLM đ€
Running Model locally - Cursor@Home
see also
- 4-node AMD cluster - runing 512GB Model locally / needs \(\)
- 4-node Minisforum MS-S1 Max Strix Halo cluster
- tokenspeed / HN
# Models âźș
see also
- Model benchmarks
- Top Model by task
- The Best Local Agentic Coding Workflow (Complete Guide)
- Your Open Source Model Could Have a Hidden Time-Release Backdoor - demonstrate that harness context can be used to trigger a different answer in controlled condition.
| rank | Model | Size | Speed | Comments | Â |
|---|---|---|---|---|---|
| Â | Qwen3.8 27B | 15.9Go = 13Go + cont 46k | 6t/s | LM Studio 4.24 | Â |
| Â | Qwen3.6 27B | Â | Â | Â | Â |
| Â | Qwen3(?) 14B | 15.9Go = 9Go + cont 46k | 80t/s | Â | Â |
| Â | qwen3-14b-claude-4.5-opus-high-reasoning-distill | 9GB | 80 tok/sec | LM Studio 4.12 | Â |
| Â | unsloth/qwen3-coder-30b-a3b-instruct | 11GB/12.4GB | 55.46 tok/sec - 1019 tokens - 0.03s to first token | LM Studio 3.6 | Â |
| Â | llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-NVFP4-Experts-Only-GGUF | 12.5GB | 10 tok/sec - 2898 tokens - 1.16s to first token | LM Studio 4.12 | Â |
| Â | llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-NVFP4-Experts-Only-GGUF | 12.5GB | 10 tok/sec - 2898 tokens - 1.16s to first token | LM Studio 4.12 | Â |
- I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
- QWEN 3.6 27B sur 16GB VRAM : La meilleure configuration
Avec un KV Cache en Q4 tu dĂ©grades Ă©normĂ©ment la qualitĂ©. Autant passer sur un 35b a3b avec MoE (mĂȘme sans MTP) qui sera beaucoup plus rapide et de meilleure qualitĂ©, surtout si tu le prends en Q6 voire Q8. Le KV cache câest toujours Q8 minimum, ou alors turboquant, et encore (turbo4 minimum, pas turbo3).
Je sais pas comment vous faites pour du dev avec 55k de contexte, Ă moins de faire juste quelques petits scripts par ci par lĂ . Pour moi câest beaucoup trop peu (je suis dĂ©veloppeur depuis plus de 30 ans). Câest minimum 256k pour que le modĂšle ait une vue dâensemble de lâarchitecture du projet, puisse chercher dans la doc, les headers, git, ticketing, logs, etc, bref bosser dans de bonnes conditions. Jâutilise Opencode.
Par ailleurs il est indispensable dâactiver la persistence du raisonnement (preserve_thinking). Ce nâest pas fait par dĂ©faut par llama.cpp (du moins avec les modĂšles Qwen 3.6) donc le modĂšle oublie ses propres raisonnements dâun tour Ă lâautre. Pour tester, il faut lui demander de penser Ă deux nombres et de donner seulement lâun des deux. Allez voir les deux nombres dans son raisonnement. Ensuite, on lui demande de nous donner lâautre nombre auquel il pensait dans le message prĂ©cĂ©dent, et lĂ il va dire nâimporte quoi ! Parce quâil lâaura oubliĂ© ! Du coup, en activant preserve_thinking, on gagne Ă©normĂ©ment en qualitĂ© vu que le modĂšle nâoublie rien. Mais effectivement ça demande un contexte plus grand. Et mĂȘme si la qualitĂ© de la plupart des modĂšles locaux se dĂ©grade Ă partir de 128k, câest quand mĂȘme indispensable. Ăa vaut vraiment le coup. Qwen 3.6 35b a3b a pu me terminer des projets de tests quâaucun autre LLM local nâavait pu grĂące Ă ce rĂ©glage.
Pour ma part je nâai que 12 Go sur ma RTX 4070 + i5 12400F et 64 Go DDR5. Jâai Qwen 3.6 35b a3b UD-Q4_K_M, KV cache en Q8, 256k de contexte, 34 moe sur CPU, batch 16384 et ubatch 2048 (ça dĂ©pote en pp). En bench je suis Ă 70 t/s tg. En pratique entre 60 et 35 entre le dĂ©but et la fin du contexte. Jâai testĂ© 27b mais Q2 obligatoire pour le faire rentrer dans 12 GO de VRAM, donc câest pourri. Sinon jâai laissĂ© tomber MTP car dâune part ça demande de la mĂ©moire supplĂ©mentaire (donc potentiellement des couches ou KV cache Ă offloader ==> perte de performance) et de plus jâai notĂ© une baisse de qualitĂ©, mĂȘme avec un excellent taux dâacceptation.
En espérant avoir apporté quelques informations utiles !
# Qwen 3.8
Testing
- ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF 13Gb with -mtp
- Quantization: IQ3_S - Benchmark
- Qwen 3.8 27B sur 16 Go de VRAM : il a résolu le bug que DeepSeek V4 Flash a raté
-
Qwen3.8 Flash: Over 7x Faster First Token - llama.cpp vs SGLang vs FreeToken. Benchmarked! (96GB)
- Qwen 3.8 27B GSQ RCO tested - 16GB Local LLM setup - from ISTA DAS Lab Austria,
# Qwen3.6 âźș
This release delivers substantial upgrades, particularly in
-
Agentic Coding: the model now handles frontend workflows and repository-level reasoning with greater fluency and precision.
-
Thinking Preservation: weâve introduced a new option to retain reasoning context from historical messages, streamlining iterative development and reducing overhead.
see also
- Qwen 3.6 Locally: Can I use it for my daily tasks?
- 27B - better on coding task (but dense)
- 35B - faster
- Prompt Vault - a curated collection of structured coding prompts and challenges designed for testing Large Language Models (LLMs)
# Qwen3.5 âźș
Qwen3.5 is Alibabaâs new model family, including Qwen3.5-35B-A3B, 27B, 122B-A10B and 397B-A17B and the new Small series: Qwen3.5-0.8B, 2B, 4B and 9B. The multimodal hybrid reasoning LLMs deliver the strongest performances for their sizes. They support 256K context across 201 languages, have thinking + non-thinking, and excel in agentic coding, vision, chat, and long-context tasks. The 35B and 27B models work on a 22GB Mac / RAM device.
See all GGUFs here.
# Qwen3
# qwen3-14b-claude-4.5-opus-high-reasoning-distill
# Qwen3-Coder-30B-A3B-Instruct-GGUF
This is the standard instruction-tuned code model.
Typical strengths:
- generating code from prompts
- explaining code
- quick coding tasks
- IDE autocomplete / chat coding
Itâs also easier to run locally because the total model size is smaller (30B).
- use Qwen3 Coder 30B A3B Instruct. This model delivers strong coding performance and reliable tool use.
# Qwen3-Coder-Next âźș
Much larger model: ~80B parameters total / ~3B active per token
- very sparse Mixture-of-Experts
- optimized for fast repeated calls in agent loops
Designed for coding agents that run loops:
- read repo
- plan changes
- edit files
- run tests
- debug
- repeat
It was trained with environment interaction and executable coding tasks, so it learns from test results and feedback loops.
Typical strengths:
- repo-scale refactors
- debugging from logs
- multi-step tool use
- IDE agents (Cline, Claude Code, etc.)
# Quantisation
# Q4_K_M âźș
Q4_K_M is a compressed 4-bit version of an AI model that balances small size with good accuracy.
Q4 - means each model weight uses 4 bits instead of 16 bits, reducing memory usage by about 4Ă.
- smaller files
- faster inference
- lower RAM/VRAM requirements
K stands for K-block quantisation, an improved quantisation method used in GGML/GGUF models.
Instead of compressing each weight independently, it compresses blocks of weights together, which:
- preserves more information
- improves accuracy compared to older Q4 formats
M a variant of the quantisation scheme (S (small) / M (Medium) / L (Large) precision)
# GSQ-RCO
A particular quantization method/format intended to preserve quality at very low bit rates
# Model Parameters
# A3B âźș
Some newer models use a Mixture-of-Experts (MoE) architecture.
A3B â Active 3B parameters during inference.
Instead of using all parameters every time, the model:
- Contains a large total parameter count (e.g., 35B).
- Activates only a subset of them per token using a router that selects a few âexperts.â
Example
| Model name | Total parameters | Active parameters |
|---|---|---|
| Qwen3.5-35B-A3B | 35B | ~3B active |
| ERNIE-4.5-21B-A3B | 21B | ~3B active |
# -mtp
-mtp build (about 0.35 GB larger) that carries the modelâs Multi-Token Prediction head for speculative decoding in llama.cpp. The weights are otherwise identical, so quality is unchanged.
youâre adding machinery that can make decoding faster.
And the MTP speedup only happens when your inference engine actually supports/activates MTP. This repo specifically targets llama.cppâs draft-mtp speculative decoding.
# instruct âźș
This refers to a model that is specifically trained or fine-tuned to follow instructions from users in a helpful, safe, and coherent way.
# DFlash âźș
DFlash is a novel speculative decoding method that utilizes a lightweight block diffusion model for drafting. It enables efficient, high-quality parallel drafting that pushes the limits of inference speed.