L

LMDeploy

L
LMDeploy AI v0.15.0

v0.15.0

What's Changed 🚀 Features Support long-context and MTP prefix-cache hits by @grimoire in #4688 [Feature] Add guided decoding support for speculative decoding by @windreamer in #4559 feat(turbomind): memory allocator, object cache, and scheduler integration by @lzhangzz in #4717 feat: add AgRs all2all backend by @irexyc in #4739 DeepSeek V4 support by @grimoire in #4554 Support memdecode by @lvhan0…

L
LMDeploy AI v0.14.0

v0.14.0

What's Changed 🚀 Features FP8 kv cache quantization by @CUHKSZzxy in #4563 Support Qwen3 Omni by @CUHKSZzxy in #4411 support qwen3.5(vit) inference in turbomind backend by @irexyc in #4602 Add OpenAI Responses-compatible endpoint by @CUHKSZzxy in #4582 Add /get_ppl endpoint by @irexyc in #4679 💥 Improvements Update turbomind modeling infrastructure by @lzhangzz in #4557 refactor(turbomind): consol…

L
LMDeploy AI v0.14.0a2

v0.14.0a2

What's Changed 🚀 Features FP8 kv cache quantization by @CUHKSZzxy in #4563 Support Qwen3 Omni by @CUHKSZzxy in #4411 support qwen3.5(vit) inference in turbomind backend by @irexyc in #4602 Add OpenAI Responses-compatible endpoint by @CUHKSZzxy in #4582 💥 Improvements Update turbomind modeling infrastructure by @lzhangzz in #4557 refactor(turbomind): consolidate CUDA error handling and add manual s…

L
LMDeploy AI v0.14.0a1

v0.14.0a1

What's Changed 🚀 Features FP8 kv cache quantization by @CUHKSZzxy in #4563 💥 Improvements Update turbomind modeling infrastructure by @lzhangzz in #4557 refactor(turbomind): consolidate CUDA error handling and add manual stacktracing by @lzhangzz in #4565 Add Qwen3.5 Moe lite awq by @43758726 in #4561 [Improve]: Drain queues when sleep engine by @RunningLeon in #4577 Extend chat completions by int…

L
LMDeploy AI v0.13.0

v0.13.0

What's Changed 🚀 Features [Ascend] support qwen3.5 35BA3B by @wanfengcxz in #4485 feat: Add TurboQuant (quant_policy=42) support for KV Cache Quantization by @windreamer in #4510 [refactor] [api_server] [2/N] improve tool parsers by abstracting xml parser by @lvhan028 in #4548 feat(turbomind): integrate cublasGemmGroupedBatchedEx for Qwen3.5 MoE inference on Blackwell GPUs with memory copy optimiz…

L
LMDeploy AI v0.12.3

v0.12.3

What's Changed 🚀 Features Support video inputs by @CUHKSZzxy in #4360 feat: fully implement compressed-tensors gs32 support in TurboMind by @lapy in #4429 Draft model update params by @CUHKSZzxy in #4452 💥 Improvements support qwen3.5 on volta by @grimoire in #4405 Optimize Qwen3.5 by @lzhangzz in #4434 Builtin mrope by @grimoire in #4393 delete ray remote function return value by @grimoire in #44…

L
LMDeploy AI v0.12.2

v0.12.2

What's Changed 🚀 Features support glm5 by @grimoire in #4355 Qwen/Internlm/Llama Dense/Moe model fp8 quant online by @43758726 in #4324 Qwen3.5 by @grimoire in #4351 GLM-4.7-Flash Turbomind support by @lapy in #4362 Support router replay and ignore quant layer for qwen3.5 by @RunningLeon in #4394 [Feature] Add TurboMind support for Qwen3.5 models (dense + MoE) by @lapy in #4389 support repetition…

L
LMDeploy AI v0.12.1

v0.12.1

What's Changed 🚀 Features support glm-4.7-flash by @RunningLeon in #4320 [ascend]suppot ep by @yao-fengchen in #3696 💥 Improvements fix rotary embedding for transformers v5 by @grimoire in #4303 Improve metrics log by @CUHKSZzxy in #4297 Support ignore layers in quant config for qwen3 models by @RunningLeon in #4293 add custom noaux kernel by @grimoire in #4345 fix qwen3vl with transformers5 by @g…

L
LMDeploy AI v0.12.0

v0.12.0

What's Changed 🚀 Features Add Gloo communication to turbomind by @irexyc in #3362 [Feat] Support llm-compressor AWQ models in TurboMind by @43758726 in #4290 Router replay for gpt oss by @RunningLeon in #4298 Support llm-compressor symmetric quantized model inference in TurboMind by @43758726 in #4305 Support Intern-S1-Pro by @CUHKSZzxy in #4318 💥 Improvements Configurable max CTAs and NVLS usage…

L
LMDeploy AI v0.11.1

v0.11.1

What's Changed 🚀 Features [ascend] support dptp by @tangzhiyi11 in #4218 Support Deepseek v32 by @grimoire in #4026 💥 Improvements Improve metrics by @CUHKSZzxy in #4178 reserve blocks for dummy inputs by @grimoire in #4157 Add vision id for Qwen3-VL by @CUHKSZzxy in #4183 [Enhance]: Return routed experts when request canceled by @RunningLeon in #4197 Add mm processor args for Qwen3-VL by @CUHKSZz…

L
LMDeploy AI v0.11.0

v0.11.0

What's Changed 🚀 Features add endpoint /abort_request by @lvhan028 in #4092 Qwen3 next by @grimoire in #4039 Support Qwen3-VL by @CUHKSZzxy in #4093 Support sync weights with flattened bucket tensor by @RunningLeon in #4109 Support group router for moe models by @RunningLeon in #4120 [Feature]: return routed experts to reuse by @RunningLeon in #4090 support context parallel by @irexyc in #3951 fop…