[Bug]: EPLB load statistics problem

### Your current environment

<details>
<summary>The output of <code>python collect_env.py</code></summary>

```text
Your output of `python collect_env.py` here
```

</details>

A bug in load statistics: load statistics are collected during expert selection. However, only the top k experts for the current token that are routed to the local node are counted. The load of experts routed to other nodes is not counted.

### 🐛 Describe the bug

https://github.com/vllm-project/vllm/blob/555e7225bcb9cdf9b037ce064e48987dbc3e13a0/vllm/model_executor/layers/fused_moe/layer.py#L1336

```
            if expert_map is not None:
                topk_ids_local = expert_map[topk_ids]
                topk_ids_flatten = topk_ids_local.flatten()
            else:
                topk_ids_flatten = topk_ids.flatten()

            # Should be equivalent to:
            # ```
            # topk_ids_masked = topk_ids_local[topk_ids_local >= 0]
            # expert_load_view += topk_ids_masked.bincount(
            #     minlength=expert_load_view.shape[0])
            # ```
            # We use `scatter_add_` since `bincount` cannot be compiled

            # Performance optimization:
            # `masked_fill` is significantly faster than `masked_select`
            invalid_mask = topk_ids_flatten < 0
            # Replace invalid expert ids with 0 (just a dummy position)
            # to avoid out-of-bounds errors in scatter_add_
            index = topk_ids_flatten.masked_fill_(invalid_mask, 0)
            # `src` is the valid mask, which is 1 for valid and 0 for invalid
            src = ~invalid_mask

            expert_load_view.scatter_add_(dim=0,
                                          index=index.long(),
                                          src=src.to(expert_load_view))
```

### Before submitting a new issue...

- [x] Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the [documentation page](https://docs.vllm.ai/en/latest/), which can answer lots of frequently asked questions.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Uh oh!

[Bug]: EPLB load statistics problem #21883

Your current environment

🐛 Describe the bug

Before submitting a new issue...

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Uh oh!

[Bug]: EPLB load statistics problem #21883

Description

Your current environment

🐛 Describe the bug

Before submitting a new issue...

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions