Skip to content

Conversation

@markmc
Copy link
Member

@markmc markmc commented Oct 31, 2025

Follow on from #25712

VllmConfig is explicitly designed as a dataclass containing user-provided configuration and model metadata. It is a global configuration object that lives throughout the entire engine lifetime and is meant to be immutable after __post_init__().

KVCacheConfig is worker-specific, runtime-computed state. It has limited lifetime, and its purpose is limited to initializing the KV Cache in the model runner.

Even if we add KV cache hints to model config.json in future, this would be parsed into ModelConfig, used as input to the get_kv_cache_configs() computation, and the resulting KVCacheConfig would still be runtime state.

We are currently creating per-worker copies of VllmConfig in order to attach the runtime KVCacheConfig state. But instead we should just explicitly pass KVCacheConfig to the connector.

Make sure to handle backwards compatibility for external connector implementations (loaded via module path) that have the old style constructor signature.

Copy link
Contributor

@gemini-code-assist gemini-code-assist bot left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors KV connector instantiation to explicitly pass KVCacheConfig, which is a good improvement for separating runtime state from global configuration. The implementation of backward compatibility for external connectors is also well-handled in the factory. However, I've identified a critical issue in MultiConnector where it incorrectly instantiates its sub-connectors, which breaks backward compatibility and fails to pass the kv_cache_config. I've also found a minor issue in a test utility. Please see my comments for details and suggestions.

@markmc
Copy link
Member Author

markmc commented Oct 31, 2025

See also #25712 (comment)

Copy link

@chatgpt-codex-connector chatgpt-codex-connector bot left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@markmc
Copy link
Member Author

markmc commented Oct 31, 2025

xref #27811

/cc @KuntaiDu @njhill @heheda12345

@markmc markmc requested a review from KuntaiDu October 31, 2025 15:40
@markmc markmc force-pushed the connector-kv-cache-config branch from 4513b7c to 15ec91b Compare October 31, 2025 15:41
Follow on from vllm-project#25712

`VllmConfig` is explicitly designed as a dataclass containing
user-provided configuration and model metadata. It is a global
configuration object that lives throughout the entire engine lifetime
and is meant to be immutable after `__post_init__()`.

`KVCacheConfig` is worker-specific, runtime-computed state. It has
limited lifetime, and its purpose is limited to initializing the KV
Cache in the model runner.

Even if we add KV cache hints to model config.json in future, this
would be parsed into `ModelConfig`, used as input to the
`get_kv_cache_configs()` computation, and the resulting
`KVCacheConfig` would still be runtime state.

We are currently creating per-worker copies of VllmConfig in order
to attach the runtime `KVCacheConfig` state. But instead we should
just explicitly pass `KVCacheConfig` to the connector.

Make sure to handle backwards compatibility for external connector
implementations (loaded via module path) that have the old style
constructor signature.

Signed-off-by: Mark McLoughlin <[email protected]>
@markmc markmc force-pushed the connector-kv-cache-config branch from 15ec91b to fdc30ae Compare October 31, 2025 17:24
@markmc markmc added the ready ONLY add when PR is ready to merge/full CI is needed label Oct 31, 2025
Copy link
Collaborator

@KuntaiDu KuntaiDu left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In general LGTM. Some nits listed in comments.

f"Class {connector_name} not found in {connector_module_path}"
) from e
connector_cls = cast(type[KVConnectorBaseType], connector_cls)
if not supports_kw(connector_cls, "kv_cache_config"):
Copy link
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm: this means that we allow connector to include kv_cache_config field as init args even when it does not support hybrid allocator (not a subclass of SupportsHMA)?

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah. Since we we are currently unconditionally attaching KVCacheConfig to VllmConfig, I didn't even consider making it conditional until now

If we think that KVCacheConfig will only be used as part of implementing request_finished_all_groups() we could add init_hma() or init_kv_cache_config() to SupportsHMA ?

Copy link
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

KVCacheConfig takes effect on almost all connector methods when HMA is enabled and it is not useful at all when HMA is disabled.

That said, the current implementation looks good to me as it is simple enough.

@@ -0,0 +1,275 @@
# SPDX-License-Identifier: Apache-2.0
Copy link
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe rename this file? Like test_connector_init_with_kv_cache_config or something.

Copy link
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Obviously I don't really mind renaming it, if it helps, but my thinking is that these tests are about testing support for connectors that have not yet been updated to the new signature so it's more like "without_kv_cache_config()"

Basically, because all connectors are expected to support the new signature, we'll soon see these tests as old cruft that we need to keep around for a while

(This is different from the SupportsHMA approach - in that case, maybe only a small subset of connectors would be updated to take KVCacheConfig)

@markmc
Copy link
Member Author

markmc commented Nov 3, 2025

Just capturing some of the points from the Slack discussion, since I think it was very useful 👍

@heheda12345

I think in the long term, kv cache config will be part of vllm config. But for now, as it is not created during vllm_config initialization, my concern is that there are multiple copies of vllm_config inside vllm, and it's difficult to keep them synchronized. So I suggest @KuntaiDu to add it only as a special attribute for connector

But I don't want something like init(vllm_config, kv_cache_config) as kv_cache_config may be part of vllm_config in the future.

Even if we add KV cache hints to model config.json in future, this would be parsed into ModelConfig, used as input to the get_kv_cache_configs() computation, and the resulting KVCacheConfig would still be runtime state.

if we can do get_kv_cache_configs(model_config), then kv_cache_config can be init during config resolution and is immutable

@markmc

In future, KVCacheConfig will only ever contain immutable config from user/model, and it will no longer ever contain runtime computed state?

IMHO we should not attach runtime computed state to VllmConfig - it makes sense to have KVConnector.init(VllmConfig, KVCacheConfig) now, and work towards making KVCacheConfig only immutable config and then deprecating the KVCacheConfig argument

@heheda12345

should we keep connector interface as stable as possible?

@markmc

yes, but IMO "stability" means "backwards compatible for a (potentially long) deprecation period"

Copy link
Collaborator

@NickLucche NickLucche left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm, let's keep compat code under control though

self,
vllm_config: "VllmConfig",
role: KVConnectorRole,
kv_cache_config: "KVCacheConfig",
Copy link
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@KuntaiDu do you need to change LMCache connector? kv_cache_config is not available in vllm config now.

Copy link
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I need to change LMCache PR correspondingly, but we can merge this PR first and I can change it in LMCache afterwards.

Copy link
Collaborator

@heheda12345 heheda12345 left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM! Leave this PR to @KuntaiDu for compatibility with LMCache + HMA

@KuntaiDu
Copy link
Collaborator

KuntaiDu commented Nov 4, 2025

LGTM! Leave this PR to @KuntaiDu for compatibility with LMCache + HMA

Thanks for reminder! I will fix LMCache compatibility after this PR is merged (still a lot of work on LMCache side before turning on HMA by default for LMCache).

@KuntaiDu KuntaiDu merged commit 58279c6 into vllm-project:main Nov 4, 2025
54 checks passed
omerpaz95 pushed a commit to omerpaz95/vllm that referenced this pull request Nov 4, 2025
juliendenize pushed a commit to juliendenize/vllm that referenced this pull request Nov 6, 2025
zWaNg3 added a commit to fangyuchu/vllm that referenced this pull request Nov 7, 2025
* add fault_report_addr in FaultToleranceConfig

* add handle fault&get_fault_info api

Signed-off-by: w00689259 <[email protected]>

* remove fault_report_address in CoreEngineActorManager __init__

Signed-off-by: a798347923 <[email protected]>

* ruff format

Signed-off-by: a798347923 <[email protected]>

* add handle fault&get_fault_info api

Signed-off-by: w00689259 <[email protected]>

* fix one bug.

Signed-off-by: fangyuchu <[email protected]>

* add fault_report_port in FaultToleranceConfig

Signed-off-by: a798347923 <[email protected]>

* add zmq_addr concatenate with fault_report_addr and fault_report_port

Signed-off-by: a798347923 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fix some bug

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* fault reporter bug fix

Signed-off-by: w00689259 <[email protected]>

* remove fault_report_addr in FaultToleranceConfig

Signed-off-by: a798347923 <[email protected]>

* refactor: relocate method serialization functions to serial_util.py

Signed-off-by: fangyuchu <[email protected]>

* fix actor bug

* fix actor bug

* add engine_core_cmd_addr in FaultToleranceConfig

Signed-off-by: a798347923 <[email protected]>

* add and use _stop_worker_execution in EngineCoreGuard

Signed-off-by: a798347923 <[email protected]>

* add and use run in WorkerGuard

Signed-off-by: a798347923 <[email protected]>

* fix actor bug

* fix bug

* fix sentinel

* fix bug vllm/v1/engine/core.py:847: error: Missing positional argument "tp_size" in call to "EngineCoreGuard"

Signed-off-by: a798347923 <[email protected]>

* fix bug error: Missing positional arguments "length", "byteorder" in call to "to_bytes" of "int"

Signed-off-by: a798347923 <[email protected]>

* fix bug in fault tolerance mode

Signed-off-by: w00689259 <[email protected]>

* fix bug in fault tolerance mode

Signed-off-by: w00689259 <[email protected]>

* change fault_report_port to internal_fault_report_port
add external_fault_notify_port

Signed-off-by: a798347923 <[email protected]>

* change fault_report_port to internal_fault_report_port
add external_fault_notify_port

Signed-off-by: a798347923 <[email protected]>

* add _recv_cmd func
use deserialize_method_call and run_method in run func

Signed-off-by: a798347923 <[email protected]>

* Update core.py

fix bug error: Need type annotation for "kwargs" (hint: "kwargs: dict[<type>, <type>] = ...")

Signed-off-by: a798347923 <[email protected]>

* add self.ctx.term() in shutdown()

Signed-off-by: a798347923 <[email protected]>

* changed import deserialize_method_call,serialize_method_call

Signed-off-by: a798347923 <[email protected]>

* changed init worker_guard in init_device

Signed-off-by: a798347923 <[email protected]>

* Update core.py

add import serialize_method_call

Signed-off-by: a798347923 <[email protected]>

* Update gpu_worker.py

changed init WorkerGuard in init_device

Signed-off-by: a798347923 <[email protected]>

* Update gpu_worker.py

FIX BUG self.worker_guard: WorkerGuard|None = None

Signed-off-by: a798347923 <[email protected]>

* Update gpu_worker.py

fix bug error: Argument 1 to "deserialize_method_call" has incompatible type "str | None"; expected "str"  [arg-type]

Signed-off-by: a798347923 <[email protected]>

* Update gpu_worker.py

ruff format

Signed-off-by: a798347923 <[email protected]>

* Update core.py

ruff-format

Signed-off-by: a798347923 <[email protected]>

* actively send exception information

Signed-off-by: w00689259 <[email protected]>

* actively send exception information

Signed-off-by: w00689259 <[email protected]>

* actively send exception information

Signed-off-by: w00689259 <[email protected]>

* change engine_core_cmd_addr(str) to engine_core_cmd_addrs(list[str]) in EngineZmqAddresses

Signed-off-by: a798347923 <[email protected]>

* change engine_core_cmd_addr(str) to engine_core_cmd_addrs(list[str]) in EngineZmqAddresses

Signed-off-by: a798347923 <[email protected]>

* Update utils.py

delete engine_core_cmd_addr in EngineZmqAddresses

Signed-off-by: a798347923 <[email protected]>

* Remove redundant configuration: fault-pub-port

Signed-off-by: fangyuchu <[email protected]>

* Send pause instructions after receiving fault info in ClientGuard

Signed-off-by: fangyuchu <[email protected]>

* change engine_core_guard_identities from dict[int, bytes] to list[bytes]

Signed-off-by: a798347923 <[email protected]>

* fix bug "only the worker guard of engine core 0 can receive messages sent from engine core guard

Signed-off-by: a798347923 <[email protected]>

* change local_rank to rank_in_group in WorkerGuard

Signed-off-by: a798347923 <[email protected]>

* changed del self.client_cmd_registry[int(unhealthy_engine.engine_id)]

Signed-off-by: a798347923 <[email protected]>

* add gloo communication timeout

* fix some bug

* add  stateless_process_group gloo_comm_timeout

* reconstruct fault receiver&fault handler

Signed-off-by: w00689259 <[email protected]>

* fix some bug

* reconstruct fault receiver&fault handler

Signed-off-by: w00689259 <[email protected]>

* reconstruct fault receiver&fault handler

Signed-off-by: w00689259 <[email protected]>

* fix return format

Signed-off-by: w00689259 <[email protected]>

* fix return format

Signed-off-by: w00689259 <[email protected]>

* fix return format

Signed-off-by: w00689259 <[email protected]>

* add abort request

* fix some bug

* fix some bug

* fix some bug

* add dt for client guard

Signed-off-by: w00689259 <[email protected]>

* add dt for client guard

Signed-off-by: w00689259 <[email protected]>

* add dt for client guard

Signed-off-by: w00689259 <[email protected]>

* Implementation of two types of pause: a soft one by using flag signals and a hard one by aborting nccl communicators.

Signed-off-by: fangyuchu <[email protected]>

* Refine certain log forms and fix a minor bug in pause function.

Signed-off-by: fangyuchu <[email protected]>

* Refactor and abstract the recv_msg logic in CG,ECG,WG.

Signed-off-by: fangyuchu <[email protected]>

* [Frontend] Align finish_reason when tool is called with OpenAI (vllm-project#25054)

Signed-off-by: Sungyoon Jeong <[email protected]>
Co-authored-by: Chauncey <[email protected]>

* [Hybrid] Pass kernel block size to builders (vllm-project#27753)

Signed-off-by: Thomas Parnell <[email protected]>

* [Bugfix] Padded Eagle Specdec with Chunked Prefill (vllm-project#26263)

Signed-off-by: Rémi Delacourt <[email protected]>
Signed-off-by: Rémi Delacourt <[email protected]>
Signed-off-by: remi <[email protected]>
Co-authored-by: Benjamin Chislett <[email protected]>

* [XPU]Refine Dockerfile.xpu, avoid oneccl dependency issue (vllm-project#27964)

Signed-off-by: Kunshang Ji <[email protected]>

* Add and check method uuid when sending commands and receiving results.

Signed-off-by: fangyuchu <[email protected]>

* Add ORCA endpoint load metrics support (vllm-project#24905)

Signed-off-by: Misha Efimov <[email protected]>

* [CI/Build] Remove the flaky gpt-oss lora test (vllm-project#27966)

Signed-off-by: Jee Jee Li <[email protected]>

* Abstract the logic of sending instructions and waiting responses from FaultHandler

Signed-off-by: fangyuchu <[email protected]>

* [Model] Add PaddleOCR-VL Model Support  (vllm-project#27758)

Signed-off-by: zhangyue <[email protected]>
Signed-off-by: Roger Wang <[email protected]>
Signed-off-by: Isotr0py <[email protected]>
Signed-off-by: zhangyue66 <[email protected]>
Co-authored-by: Roger Wang <[email protected]>
Co-authored-by: Isotr0py <[email protected]>

* Add options in EngineCoreGuard to recv execution results from WorkerGuard

Signed-off-by: fangyuchu <[email protected]>

* Early exit for MoE LoRA kernels (vllm-project#27131)

Signed-off-by: gnovack <[email protected]>
Co-authored-by: Jee Jee Li <[email protected]>

* [Bugfix] Skip gs:// model paths for speculator detection (vllm-project#27846)

Signed-off-by: Peter Schuurman <[email protected]>

* [BUG] Make 'binary' default option for saving torch compile artifacts when using standalone_compile (vllm-project#27616)

Signed-off-by: ahao-anyscale <[email protected]>

* [CI/Testing] Add basic single node dual batch overlap test (vllm-project#27235)

Signed-off-by: Lucas Wilkinson <[email protected]>

* [Spec Decode] Integrate Suffix Decoding from Arctic Inference (vllm-project#25784)

Co-authored-by: Aurick Qiao <[email protected]>

* [Feature][Benchmarks] Support `inf` burstiness (vllm-project#26941)

Signed-off-by: Sophie du Couédic <[email protected]>

* [Bugfix][Qwen][Multimodal] Move Qwen2_5_vl sdpa to custom op and reenable compile (vllm-project#27764)

Signed-off-by: Lucas Kabela <[email protected]>

* [Bugfix] change FlashMLA reorder_batch_threshold (vllm-project#27777)

Signed-off-by: Matthew Bonanni <[email protected]>

* [Docs] add runai_streamer_sharded to LoadConfig (vllm-project#27937)

Signed-off-by: Andy Xie <[email protected]>

* Add TP parameter to attention tests (vllm-project#27683)

Signed-off-by: Matthew Bonanni <[email protected]>

* [Bugfix][plugin] fla crash on plugin (vllm-project#27322)

* [Bugfix] Fix MoE Routing Simulation (vllm-project#28002)

Signed-off-by: Tyler Michael Smith <[email protected]>

* Remove the tpu docker image nightly build. (vllm-project#27997)

Signed-off-by: Qiliang Cui <[email protected]>

* [Bugfix][ROCm] Fix ViT rotary embeddings for torch.compile compatibility on ROCm (vllm-project#27748)

Signed-off-by: vllmellm <[email protected]>

* [LoRA] Lora shrink swizzle (vllm-project#27694)

Signed-off-by: li2haipeng <[email protected]>
Signed-off-by: Haipeng Li <[email protected]>
Co-authored-by: Jee Jee Li <[email protected]>

* [Refactor] Lazy import tool_parser (vllm-project#27974)

Signed-off-by: chaunceyjiang <[email protected]>

* [NIXL][XPU] Pin NIXL version to 0.7.0 (vllm-project#27849)

Signed-off-by: zhenwei-intel <[email protected]>

* [Metrics] Enable sleep state metric outside of dev mode (vllm-project#27867)

Signed-off-by: Mark McLoughlin <[email protected]>

* [Bug] Batch invariant: Fix flash attn MLA `RuntimeError: scheduler_metadata must have shape (metadata_size)` (vllm-project#27884)

* [CPU]Improve dynamic 4bit moe performance (vllm-project#27240)

Signed-off-by: Zhang Xiangze <[email protected]>

* [CI/Build] Update LM Eval Version in AMD CI (vllm-project#27944)

Signed-off-by: zhewenli <[email protected]>

* [KV Connector] Make KVCacheConfig an explicit constructor argument (vllm-project#27887)

Signed-off-by: Mark McLoughlin <[email protected]>

* [Model] fix ernie45 reasoning_parser (vllm-project#27973)

Signed-off-by: wangyafeng <[email protected]>

* [CI/Build] Fix OpenAI API correctness on AMD CI (vllm-project#28022)

Signed-off-by: zhewenli <[email protected]>

* [BugFix][Performance] Restore flashinfer autotuning for all scenarios (vllm-project#27904)

* Support worker reinitialization after hard pause; add task queue in FaultHandler to ensure sequential task execution

Signed-off-by: fangyuchu <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* Load tuned fused_moe_lora shrink and expand kernel configs separately (vllm-project#27435)

Signed-off-by: Yu Gong <[email protected]>
Co-authored-by: Jee Jee Li <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* resolve conflicts

Signed-off-by: w00689259 <[email protected]>

* Support using Int4PreshuffledTensor after loading (vllm-project#26066)

Signed-off-by: Jerry Zhang <[email protected]>

* [Core] Enable StatLogger in LLMEngine (vllm-project#28020)

Signed-off-by: Zhuohan Li <[email protected]>

* [Model][Bugfix] fix pipeline parallelism support for NemotronH (vllm-project#27968)

Signed-off-by: Tomer Asida <[email protected]>

* [Model] add optimal triton fused moe configs for NemotronH MoE (vllm-project#27967)

Signed-off-by: Tomer Asida <[email protected]>

* [Kernels] Isolate modular kernel code from FusedMoEMethodBase subclasses. (vllm-project#27123)

* [BugFix] Fix incorrect preallocated sampled_token_ids tensor size (vllm-project#28025)

Signed-off-by: Nick Hill <[email protected]>

* [Perf] SM100 - add swap AB optimization to CUTLASS FP8 GEMM (vllm-project#27284)

Signed-off-by: Faqin Zhong <[email protected]>
Co-authored-by: Faqin Zhong <[email protected]>
Co-authored-by: Michael Goin <[email protected]>

* [PERF] Decouple projections from GDN custom op (vllm-project#27512)

Signed-off-by: Vadim Gimpelson <[email protected]>

* [model] Add support for openPangu_Ultra_MoE (vllm-project#27521)

Signed-off-by: yuantao <[email protected]>
Signed-off-by: yt0428 <[email protected]>
Co-authored-by: Jee Jee Li <[email protected]>

* [PerfFix] Avoid separate thread for MP executor shm spin (vllm-project#28012)

Signed-off-by: Nick Hill <[email protected]>

* [AsyncScheduling] Don't schedule past request max_tokens (vllm-project#27922)

Signed-off-by: Nick Hill <[email protected]>

* Remove deprecated `--rope-scaling` and `--rope-theta` (vllm-project#28006)

Signed-off-by: Harry Mellor <[email protected]>

* [ROCm][Perf] New design on ROCm AITER MHA backend Implementation (vllm-project#25763)

Signed-off-by: ganyi <[email protected]>

* Added disable rule to track files under benchmarks/lib (vllm-project#28048)

Signed-off-by: Nadav Kluger <[email protected]>

* [Multimodal] Make MediaConnector extensible. (vllm-project#27759)

Signed-off-by: Chenheli Hua <[email protected]>

* [ROCm] gemm_a16w16 upstreaming (vllm-project#26969)

Signed-off-by: Aleksandr Malyshev <[email protected]>
Co-authored-by: Aleksandr Malyshev <[email protected]>

* Revert "[PERF] Decouple projections from GDN custom op" (vllm-project#28080)

Signed-off-by: Vadim Gimpelson <[email protected]>

* add engine core ut

Signed-off-by: w00689259 <[email protected]>

* add engine core ut

Signed-off-by: w00689259 <[email protected]>

* [Qwen3-Next] MOE configs for A100-SXM4-80GB TP4 TP8 (vllm-project#27740)

* [XPU] Add gpt-oss model support for Intel GPU (vllm-project#27786)

Signed-off-by: Kunshang Ji <[email protected]>

* [CI/Build] Enable some fixed tests in AMD CI (vllm-project#28078)

Signed-off-by: zhewenli <[email protected]>

* [V0 deprecation] Remove VLLM_USE_V1 usage in most modules (vllm-project#27955)

Signed-off-by: wangxiyuan <[email protected]>

* [Bugfix] Fix encoder-only model support for transformers backend (vllm-project#28021)

Signed-off-by: Isotr0py <[email protected]>
Signed-off-by: Harry Mellor <[email protected]>
Co-authored-by: Harry Mellor <[email protected]>

* [BugFix] Fix DCP Assert (AssertionError: DCP not support reorder_batch_threshold > 1 now.) (vllm-project#28100)

Signed-off-by: Lucas Wilkinson <[email protected]>

* [Model, Core] Support Granite Speech & LoRA for STT (vllm-project#24455)

* [Refactor] Lazy-loaded reasoning_parser (vllm-project#28092)

Signed-off-by: chaunceyjiang <[email protected]>

* [Refactor] to simplify and extract the shared logic between chat completion and responses (vllm-project#27961)

Signed-off-by: chaunceyjiang <[email protected]>

* [bugfix] fix wrong `dcp_local_seq_lens` calc (vllm-project#27518)

Signed-off-by: Qiu <[email protected]>

* [Hybrid allocator + kv connector] revert connector test changes related to hybrid allocator (vllm-project#28011)

Signed-off-by: KuntaiDu <[email protected]>

* [Misc] fix import error for DeepSeekR1ReasoningParser (vllm-project#28114)

Signed-off-by: chaunceyjiang <[email protected]>

* Fix excessive logging noise by reducing the log level of the MinimaxM2ToolParser import success message (vllm-project#27635)

Signed-off-by: minatoaquaMK2 <[email protected]>

* Bugfix: Cutlass FP8 FusedMoE bad scaling factors (vllm-project#27255)

Signed-off-by: Amir Klein <[email protected]>
Co-authored-by: Tyler Michael Smith <[email protected]>
Co-authored-by: Michael Goin <[email protected]>

* [Graph Partition][Cache] Use inductor partition ops config (vllm-project#27702)

Signed-off-by: Boyuan Feng <[email protected]>

* [XPU] Enable custom routing functions in IPEX for Llama4 (vllm-project#28004)

Signed-off-by: frost-intel <[email protected]>

* add kimi reasoning parser (vllm-project#28128)

Signed-off-by: wangzhengtao <[email protected]>
Co-authored-by: wangzhengtao <[email protected]>

* [DCP] check return_lse for all layers in dcp (vllm-project#27929)

Signed-off-by: Chen Zhang <[email protected]>

* [BugFix] Support EP/DP + EPLB with MTP (vllm-project#25311)

Signed-off-by: ilmarkov <[email protected]>
Signed-off-by: Sage Moore <[email protected]>
Co-authored-by: Sage Moore <[email protected]>
Co-authored-by: Tyler Michael Smith <[email protected]>
Co-authored-by: Lucas Wilkinson <[email protected]>

* Enabling cooperative multi-gpu tests on multi-gpu nodes (vllm-project#27986)

Signed-off-by: Alexei V. Ivanov <[email protected]>

* [ROCm][MLA] Support block-size > 1 for AITER MLA backend  (vllm-project#27224)

Signed-off-by: ganyi <[email protected]>
Co-authored-by: wuhuikx <[email protected]>

* [Bugfix] Validate custom logits processor xargs for online serving (vllm-project#27560)

Signed-off-by: Isotr0py <[email protected]>

* [misc] add vLLM Beijing Meetup (vllm-project#28127)

Signed-off-by: Jiaju Zhang <[email protected]>

* [Kernel] Fuse computation of g and beta for Gated Delta Net (vllm-project#28095)

Signed-off-by: zjy0516 <[email protected]>

* [Core] add support for reasoning parser plugins (vllm-project#28075)

Signed-off-by: walter beller-morales <[email protected]>

* [Bugfix] vLLM should check Inductor config for compile cache enablement status (vllm-project#27637)

Signed-off-by: Yanan Cao <[email protected]>

* [FlashInfer] Avoid FlashInfer block_size 16 + head_size 256 on blackwell (vllm-project#27994)

Signed-off-by: Chen Zhang <[email protected]>

* [CI]: Add LMCacheConnector Unit Tests (vllm-project#27852)

Signed-off-by: Samuel Shen <[email protected]>
Co-authored-by: Samuel Shen <[email protected]>
Co-authored-by: Yihua Cheng <[email protected]>

* [Feature] Extend batch invariant torch.compile to B200 (vllm-project#27856)

Signed-off-by: PaulZhang12 <[email protected]>

* [Bugfix] Fix Qwen3-Reranker-8B load (vllm-project#28117)

Signed-off-by: wang.yuqi <[email protected]>

* [Docs] Clean up README_TUNING.md (vllm-project#28088)

Signed-off-by: windsonsea <[email protected]>

* [Hardware][IBM Z] Optimize s390x Dockerfile (vllm-project#28023)

Signed-off-by: Rehan Khan <[email protected]>

* [Chore] Remove Nemotron-Nano-VL config copy (vllm-project#28126)

Signed-off-by: Isotr0py <[email protected]>

* [Docs] Add guide to debugging vLLM-torch.compile integration (vllm-project#28094)

Signed-off-by: Richard Zou <[email protected]>

* [Feature]: Add corrupted request metric to V1 metrics system. (vllm-project#27306)

Signed-off-by: atalhens <[email protected]>

* [CI/Build] Update checking logic in cutlass_group_gemm_supported  (vllm-project#27948)

Signed-off-by: zhewenli <[email protected]>

* [CI/Build] Fix `test_defaults_with_usage_context` in AMD CI (vllm-project#27926)

Signed-off-by: zhewenli <[email protected]>

* [Core][Hybrid allocator + connector 2/n] Unify `remove_skipped_blocks` by `get_last_useful_token` (vllm-project#25431)

Signed-off-by: KuntaiDu <[email protected]>

* [Debugging] Add annotation for easier trace analysis (vllm-project#22496)

* [PERF] Decouple projections from GDN custom op. Attempt 2 (vllm-project#28083)

Signed-off-by: Vadim Gimpelson <[email protected]>

* [Bug] Fix cpu disable shared_experts `VLLM_DISABLE_SHARED_EXPERTS_STREAM` (vllm-project#28157)

Signed-off-by: yewentao256 <[email protected]>

* [Bug] Fix env string `"0"` same to `True` (vllm-project#28159)

Signed-off-by: yewentao256 <[email protected]>

* Ensure WorkerGuard command execution returns result; fix missing set_device when TP>1

Signed-off-by: fangyuchu <[email protected]>

* [Feature] Enable TP + EP `shared_experts` overlap with router, 3.7% E2E performance improvement (vllm-project#28164)

Signed-off-by: yewentao256 <[email protected]>

* [CI Failure] `nm-testing/Qwen2-0.5B-Instruct-FP8-SkipQKV` was removed from HF. Skip it in tests (vllm-project#28170)

Signed-off-by: Vadim Gimpelson <[email protected]>

* [Misc] Remove the duplicate code (vllm-project#28111)

Signed-off-by: chaunceyjiang <[email protected]>

* rename& format logger

Signed-off-by: w00689259 <[email protected]>

* rename& format logger

Signed-off-by: w00689259 <[email protected]>

* feat(nccl): enable non-blocking NCCL communicators to support ncclCommAbort

Signed-off-by: fangyuchu <[email protected]>

---------

Signed-off-by: w00689259 <[email protected]>
Signed-off-by: a798347923 <[email protected]>
Signed-off-by: fangyuchu <[email protected]>
Signed-off-by: zWaNg3 <[email protected]>
Signed-off-by: a798347923 <[email protected]>
Signed-off-by: Sungyoon Jeong <[email protected]>
Signed-off-by: Thomas Parnell <[email protected]>
Signed-off-by: Rémi Delacourt <[email protected]>
Signed-off-by: Rémi Delacourt <[email protected]>
Signed-off-by: remi <[email protected]>
Signed-off-by: Kunshang Ji <[email protected]>
Signed-off-by: Misha Efimov <[email protected]>
Signed-off-by: Jee Jee Li <[email protected]>
Signed-off-by: zhangyue <[email protected]>
Signed-off-by: Roger Wang <[email protected]>
Signed-off-by: Isotr0py <[email protected]>
Signed-off-by: zhangyue66 <[email protected]>
Signed-off-by: gnovack <[email protected]>
Signed-off-by: Peter Schuurman <[email protected]>
Signed-off-by: ahao-anyscale <[email protected]>
Signed-off-by: Lucas Wilkinson <[email protected]>
Signed-off-by: Sophie du Couédic <[email protected]>
Signed-off-by: Lucas Kabela <[email protected]>
Signed-off-by: Matthew Bonanni <[email protected]>
Signed-off-by: Andy Xie <[email protected]>
Signed-off-by: Tyler Michael Smith <[email protected]>
Signed-off-by: Qiliang Cui <[email protected]>
Signed-off-by: vllmellm <[email protected]>
Signed-off-by: li2haipeng <[email protected]>
Signed-off-by: Haipeng Li <[email protected]>
Signed-off-by: chaunceyjiang <[email protected]>
Signed-off-by: zhenwei-intel <[email protected]>
Signed-off-by: Mark McLoughlin <[email protected]>
Signed-off-by: Zhang Xiangze <[email protected]>
Signed-off-by: zhewenli <[email protected]>
Signed-off-by: wangyafeng <[email protected]>
Signed-off-by: Yu Gong <[email protected]>
Signed-off-by: Jerry Zhang <[email protected]>
Signed-off-by: Zhuohan Li <[email protected]>
Signed-off-by: Tomer Asida <[email protected]>
Signed-off-by: Nick Hill <[email protected]>
Signed-off-by: Faqin Zhong <[email protected]>
Signed-off-by: Vadim Gimpelson <[email protected]>
Signed-off-by: yuantao <[email protected]>
Signed-off-by: yt0428 <[email protected]>
Signed-off-by: Harry Mellor <[email protected]>
Signed-off-by: ganyi <[email protected]>
Signed-off-by: Nadav Kluger <[email protected]>
Signed-off-by: Chenheli Hua <[email protected]>
Signed-off-by: Aleksandr Malyshev <[email protected]>
Signed-off-by: wangxiyuan <[email protected]>
Signed-off-by: Qiu <[email protected]>
Signed-off-by: KuntaiDu <[email protected]>
Signed-off-by: minatoaquaMK2 <[email protected]>
Signed-off-by: Amir Klein <[email protected]>
Signed-off-by: Boyuan Feng <[email protected]>
Signed-off-by: frost-intel <[email protected]>
Signed-off-by: wangzhengtao <[email protected]>
Signed-off-by: Chen Zhang <[email protected]>
Signed-off-by: ilmarkov <[email protected]>
Signed-off-by: Sage Moore <[email protected]>
Signed-off-by: Alexei V. Ivanov <[email protected]>
Signed-off-by: Jiaju Zhang <[email protected]>
Signed-off-by: zjy0516 <[email protected]>
Signed-off-by: walter beller-morales <[email protected]>
Signed-off-by: Yanan Cao <[email protected]>
Signed-off-by: Samuel Shen <[email protected]>
Signed-off-by: PaulZhang12 <[email protected]>
Signed-off-by: wang.yuqi <[email protected]>
Signed-off-by: windsonsea <[email protected]>
Signed-off-by: Rehan Khan <[email protected]>
Signed-off-by: Richard Zou <[email protected]>
Signed-off-by: atalhens <[email protected]>
Signed-off-by: yewentao256 <[email protected]>
Co-authored-by: fangyuchu <[email protected]>
Co-authored-by: a798347923 <[email protected]>
Co-authored-by: w00689259 <[email protected]>
Co-authored-by: fangyuchu <[email protected]>
Co-authored-by: TianZhuo <[email protected]>
Co-authored-by: a798347923 <[email protected]>
Co-authored-by: Sungyoon Jeong <[email protected]>
Co-authored-by: Chauncey <[email protected]>
Co-authored-by: Thomas Parnell <[email protected]>
Co-authored-by: Rémi Delacourt <[email protected]>
Co-authored-by: Benjamin Chislett <[email protected]>
Co-authored-by: Kunshang Ji <[email protected]>
Co-authored-by: Misha Efimov <[email protected]>
Co-authored-by: Jee Jee Li <[email protected]>
Co-authored-by: zhang-prog <[email protected]>
Co-authored-by: Roger Wang <[email protected]>
Co-authored-by: Isotr0py <[email protected]>
Co-authored-by: gnovack <[email protected]>
Co-authored-by: pwschuurman <[email protected]>
Co-authored-by: ahao-anyscale <[email protected]>
Co-authored-by: Lucas Wilkinson <[email protected]>
Co-authored-by: Aurick Qiao <[email protected]>
Co-authored-by: Aurick Qiao <[email protected]>
Co-authored-by: Sophie du Couédic <[email protected]>
Co-authored-by: Lucas Kabela <[email protected]>
Co-authored-by: Matthew Bonanni <[email protected]>
Co-authored-by: Ning Xie <[email protected]>
Co-authored-by: Hank_ <[email protected]>
Co-authored-by: Tyler Michael Smith <[email protected]>
Co-authored-by: QiliangCui <[email protected]>
Co-authored-by: vllmellm <[email protected]>
Co-authored-by: li2haipeng <[email protected]>
Co-authored-by: liuzhenwei <[email protected]>
Co-authored-by: Mark McLoughlin <[email protected]>
Co-authored-by: Wentao Ye <[email protected]>
Co-authored-by: xiangze-arm <[email protected]>
Co-authored-by: Zhewen Li <[email protected]>
Co-authored-by: CSWYF3634076 <[email protected]>
Co-authored-by: Varun Sundar Rabindranath <[email protected]>
Co-authored-by: yugong333 <[email protected]>
Co-authored-by: Jerry Zhang <[email protected]>
Co-authored-by: Zhuohan Li <[email protected]>
Co-authored-by: tomeras91 <[email protected]>
Co-authored-by: bnellnm <[email protected]>
Co-authored-by: Nick Hill <[email protected]>
Co-authored-by: lyrisz <[email protected]>
Co-authored-by: Faqin Zhong <[email protected]>
Co-authored-by: Michael Goin <[email protected]>
Co-authored-by: Vadim Gimpelson <[email protected]>
Co-authored-by: yt0428 <[email protected]>
Co-authored-by: Harry Mellor <[email protected]>
Co-authored-by: Pleaplusone <[email protected]>
Co-authored-by: nadavkluger <[email protected]>
Co-authored-by: Chenheli Hua <[email protected]>
Co-authored-by: Aleksandr Malyshev <[email protected]>
Co-authored-by: Aleksandr Malyshev <[email protected]>
Co-authored-by: tou <[email protected]>
Co-authored-by: wangxiyuan <[email protected]>
Co-authored-by: Alex Brooks <[email protected]>
Co-authored-by: Qiu <[email protected]>
Co-authored-by: Kuntai Du <[email protected]>
Co-authored-by: Eric Yue <[email protected]>
Co-authored-by: amirkl94 <[email protected]>
Co-authored-by: Boyuan Feng <[email protected]>
Co-authored-by: Frost Mitchell <[email protected]>
Co-authored-by: bigmoyan <[email protected]>
Co-authored-by: wangzhengtao <[email protected]>
Co-authored-by: Chen Zhang <[email protected]>
Co-authored-by: Ilya Markov <[email protected]>
Co-authored-by: Sage Moore <[email protected]>
Co-authored-by: Alexei-V-Ivanov-AMD <[email protected]>
Co-authored-by: wuhuikx <[email protected]>
Co-authored-by: Jiaju Zhang <[email protected]>
Co-authored-by: Jiangyun Zhu <[email protected]>
Co-authored-by: Walter Beller-Morales <[email protected]>
Co-authored-by: gmagogsfm <[email protected]>
Co-authored-by: Samuel Shen <[email protected]>
Co-authored-by: Samuel Shen <[email protected]>
Co-authored-by: Yihua Cheng <[email protected]>
Co-authored-by: Paul Zhang <[email protected]>
Co-authored-by: wang.yuqi <[email protected]>
Co-authored-by: Michael Yao <[email protected]>
Co-authored-by: R3hankhan <[email protected]>
Co-authored-by: Richard Zou <[email protected]>
Co-authored-by: Snehlata <[email protected]>
Co-authored-by: Dayeol Lee <[email protected]>
ZhengHongming888 pushed a commit to ZhengHongming888/vllm that referenced this pull request Nov 8, 2025
xuebwang-amd pushed a commit to xuebwang-amd/vllm that referenced this pull request Nov 13, 2025
devpatelio pushed a commit to SumanthRH/vllm that referenced this pull request Nov 29, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kv-connector ready ONLY add when PR is ready to merge/full CI is needed v1

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants