E

ExLlamaV2

E
ExLlamaV2 AI v0.3.2

0.3.2

Actions: Remove sentencepiece from other workflows Signed-off-by: kingbri <8082010+kingbri1@users.noreply.github.com>

E
ExLlamaV2 AI v0.3.0

0.3.0

Add Qwen3 and Qwen3MoE support Full Changelog: v0.2.9...v0.3.0

E
ExLlamaV2 AI v0.2.9

0.2.9

Add Torch 2.7.0 wheels (big thanks to @kingbri1 for unborking the build action) Support Gemma3, text + vision Support Mistral 3.1, text + vision Support partial_rotary_factor (Phi-4 mini etc.) Support GLM4 (32B model still broken) Various fixes Full Changelog: v0.2.8...v0.2.9

E
ExLlamaV2 AI v0.2.8

0.2.8

Support Qwen2.5-VL Minor bugfixes Full Changelog: v0.2.7...v0.2.8

E
ExLlamaV2 AI v0.2.7

0.2.7

Basic video support for Qwen2-VL Support Cohere2 arch Support Granite3 arch Couple of bugfixes Full Changelog: v0.2.6...v0.2.7

E
ExLlamaV2 AI v0.2.6

0.2.6

Some small fixes, most notably for Qwen2-VL inference on Windows Full Changelog: v0.2.5...v0.2.6

E
ExLlamaV2 AI v0.2.5

0.2.5

Initial support for Qwen2-VL (images for now, no video) Some bugfixes Full Changelog: v0.2.4...v0.2.5

E
ExLlamaV2 AI v0.2.4

0.2.4

Support Pixtral Refactoring for more multimodal support Faster filter evaluation Various optimizations and bugfixes Various quality of life improvements Full Changelog: v0.2.3...v0.2.4

E
ExLlamaV2 AI v0.2.3

0.2.3

No longer use safetensors for loading weights (fix virtual memory issues on Windows especially) Disable fasttensors option (now redundant) Prioritize HF Tokenizers model when both HF and SPM models available Add XTC sampler Add YaRN support Various fixes and QoL improvements Full Changelog: v0.2.2...v0.2.3