Add paged MLA/KDA serving for Ling-3.0 (Bailing V3) - #607
Conversation
2ab7e6c to
3c149a0
Compare
d13c00e to
c2efa4f
Compare
Groundwork for the Nemotron-H stack in #644, PR 1 of 4. The hybrid runtime's geometry is currently ten undeclared attributes that lifecycle mutates onto the runner, and `cache_policy` sizes the linear side as `num_layers` minus the SDPA count. That arithmetic breaks as soon as a model has a third kind of layer. On the 52-layer Nemotron checkpoint it reports 46 recurrent layers when only 23 are. Because of that, this PR moves the geometry behind one owner, a typed layer plan built by `families/gdn.py`, in the same shape as the encoder pooling loaders from #608. Lifecycle shrinks to a four-line route. The loader table comes with the next PR, so lifecycle imports the GDN builder directly for now. For evidence, the main one is an interleaved A/B on `Qwen/Qwen3.5-0.8B`. Greedy token ids are bit-identical and `per_block_bytes` is identical. `num_blocks` wobbles a few blocks on both arms since overhead is measured at profile time. The full suite gives 1921 passed with 15 new tests, and error messages on malformed configs got stricter, each pinned by a full-string test. #598 and #607 both touch these hunks, and #607 passes constructor kwargs this PR deletes, so whichever lands second rebases. Happy to be the one who does. The `MambaSpec` shape and the attention-side wiring stay put until the table lands. --------- Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com> Signed-off-by: Yuan Lik Xun <lxyuan0420@gmail.com> Co-authored-by: Yuan Lik Xun <lxyuan0420@gmail.com>
6f5fafd to
658d77a
Compare
There was a problem hiding this comment.
The Ling-3.0 part looks well scoped. My concern is what's grown around it:
Eight of the twenty-four files aren't Bailing at all. they're mlx-lm compatibility fixes. cache.state becomes cache.cache in contiguous_cache and prefix caching, plus adjustments in gguf, LoRA, TurboQuant, pooling and xlm_roberta. That rename hits every model through model_runner.py, not just Ling-3.0. Could you split the bump and those fixes into their own PR and rebase Ling-3.0 on top? Same shape as #501 and #645.
Thanks, agreed. I’ll split the MLX/MLX-LM bump and the related compatibility fixes into a separate PR, then rebase this PR on top so it only contains the Ling-3.0 MLA/KDA support and related tests and documentation. |
658d77a to
377ca5e
Compare
## Summary - Bump the exact MLX runtime pin from 0.32.0 to 0.32.1. - Advance the MLX-LM pin from `254d153f` to `9e6acca6`. - Update contiguous-cache handling for MLX-LM's `ArraysCache.state` to `ArraysCache.cache` rename. - Adapt GGUF, LoRA, TurboQuant, pooling, and XLM-RoBERTa code to the updated dependency APIs and type information. ## Motivation This repository-wide dependency and compatibility update was split from #607 so that the Ling-3.0 MLA/KDA implementation can remain model-scoped. PR #607 will be rebased on top of this change. The pinned MLX-LM revision requires MLX 0.32.1 and contains APIs not present in the latest MLX-LM release. No other open PR was found for the same MLX 0.32.1 or `ArraysCache` compatibility update. ## Validation Environment: Apple M3 Pro with 18 GB unified memory, macOS 14.4.1, Python 3.12.12, MLX 0.32.1, and MLX-LM at `9e6acca691e64d6d8bb808c328fcdea459099cca`. ```bash VLLM_METAL_BUILD_FROM_SOURCE=1 \ .venv/bin/python -m pytest -p no:cacheprovider \ -m "not slow" tests -q .venv/bin/python -m ruff check . .venv/bin/python -m ruff format --check . .venv/bin/python -m mypy vllm_metal git diff --check origin/main...HEAD ``` Results: - Pytest: `1997 passed, 18 skipped, 40 deselected`. - Ruff: all checks passed; 291 files already formatted. - Mypy: no issues in 137 source files. - Git whitespace check: passed. ## AI assistance AI assistance was used to inspect the dependency changes, separate the compatibility update from #607, review the diff, and run the reported tests. The submitter reviewed the changed code and remains responsible for the contribution. Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
feabb91 to
f7469e0
Compare
f7469e0 to
00e2f1c
Compare
Hi @ricky-chaoju, #723 has been merged, and I’ve rebased this PR onto the latest |
|
Sorry, I missed this one. Scope looks much better at 15 files. It needs a merge from main before it can go anywhere though. It is 45 commits behind, and #778 dropped Probably a deletion rather than a port. The MLA routing in cache_policy.py is the part that will actually need reworking. |
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The cross-cutting MLA/KDA serving changes require final human review.
Review effort: Lite
Findings: None
What changed in this PR
Adds paged MLA/KDA serving support for Ling-3.0 Bailing V3 hybrid models.
Changes:
- Adds Bailing topology, KDA state handling, and hybrid runtime integration.
- Extends MLA cache and projection support.
- Adds tests and documents supported checkpoints.
| File | Summary |
|---|---|
vllm_metal/v1/model_lifecycle.py |
Preserves model architectures. |
vllm_metal/v1/cache_policy.py |
Configures hybrid MLA caching. |
vllm_metal/attention/runtime/mla.py |
Rebinds MLA caches. |
vllm_metal/attention/runtime/hybrid.py |
Supports hybrid MLA/KDA execution. |
vllm_metal/attention/runtime/families/bailing.py |
Defines Bailing topology and state family. |
vllm_metal/attention/runtime/factory.py |
Registers Bailing runtime plans. |
vllm_metal/attention/impls/mla.py |
Adds Bailing MLA compatibility. |
vllm_metal/attention/impls/kda.py |
Implements paged KDA state execution. |
tests/test_model_lifecycle.py |
Tests configuration resolution and dimensions. |
tests/test_mla_paged_backend.py |
Tests paged MLA parity. |
tests/test_kv_cache_spec_capacity.py |
Tests hybrid cache specifications. |
tests/test_kda_paged_backend.py |
Tests KDA state parity. |
tests/stub_runner.py |
Adds Bailing test helpers. |
tests/attention/test_hybrid_runtime_plan.py |
Tests runtime planning and patching. |
docs/supported_models.md |
Documents Ling-3.0 support limits. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Copilot review overview
Review effort: Lite
Findings: 3
Open (4)
This allocates a newArraysCache(size=4)for every request on every call, and performs a… · New This assumesdtypeshas at least 2 entries; if it doesn’t, it will raise anIndexErrorat… · Newsetdefault("architectures", ...)won’t populate HFarchitecturesif_extract_model_args(...)… · Newlayer_idxis accepted but never used insideKDAPagedAttentionWrapper. If the argument is… · New
|
Agreed with @ricky-chaoju above: needs a rebase onto main (45 commits behind), and |
00e2f1c to
aba9c92
Compare
|
Hi @ricky-chaoju @ericcurtin, I rebuilt this PR on the latest |
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Copilot review overview
Review effort: Lite
Findings: 2
Open (3)
Resolved since last review (4)
setdefault("architectures", ...)won’t populate HFarchitecturesif_extract_model_args(...)… This assumesdtypeshas at least 2 entries; if it doesn’t, it will raise anIndexErrorat… This allocates a newArraysCache(size=4)for every request on every call, and performs a…layer_idxis accepted but never used insideKDAPagedAttentionWrapper. If the argument is…
| def __call__( | ||
| self, | ||
| x: mx.array, | ||
| mask: mx.array | None = None, | ||
| cache: nn.Module | None = None, | ||
| **kwargs: Any, | ||
| ) -> mx.array: | ||
| ctx = get_context() | ||
| if ctx is None: | ||
| return self._inner(x, mask=mask, cache=cache) | ||
|
|
||
| step = self._prepare_step(ctx) | ||
| self._kda_state_cache.apply_pending_conv_state(self._kda_cache_idx) | ||
| self._kda_state_cache.apply_pending_recurrent_state(self._kda_cache_idx) | ||
| if step.num_decode_requests == step.num_requests: | ||
| return self._run_decode(x, step) | ||
| return self._run_requests(x, step) | ||
|
|
||
| def _prepare_step(self, ctx: PagedAttentionContext) -> _KDAStep: | ||
| cu_seqlens = ctx.cu_seqlens | ||
| if cu_seqlens is None or len(cu_seqlens) < 2: | ||
| raise RuntimeError("KDA wrapper requires cu_seqlens") | ||
| num_requests = len(cu_seqlens) - 1 | ||
| return _KDAStep( | ||
| cu_seqlens=cu_seqlens, | ||
| slot_ids=self._kda_state_cache.step_slot_ids( | ||
| ctx, self._kda_cache_idx, num_requests | ||
| ), | ||
| num_requests=num_requests, | ||
| num_decode_requests=ctx.num_decode_requests, | ||
| ) | ||
|
|
||
| def _run_decode(self, x: mx.array, step: _KDAStep) -> mx.array: | ||
| rows = x.reshape(step.num_requests, 1, x.shape[-1]) | ||
| cache = self._slot_cache(mx.array(step.slot_ids, dtype=mx.int32)) | ||
| output = self._inner(rows, mask=None, cache=cache) | ||
| cache.flush() | ||
| return output.reshape(1, step.num_requests, -1) | ||
|
|
||
| def _run_requests(self, x: mx.array, step: _KDAStep) -> mx.array: | ||
| outputs = [] | ||
| for req_idx, slot in enumerate(step.slot_ids): | ||
| start = step.cu_seqlens[req_idx] | ||
| end = step.cu_seqlens[req_idx + 1] | ||
| cache = self._slot_cache(mx.array([slot], dtype=mx.int32)) | ||
| outputs.append(self._inner(x[:, start:end, :], mask=None, cache=cache)) | ||
| cache.flush() | ||
| return mx.concatenate(outputs, axis=1) |
|
Thanks! MLA latent writes look like they will not persist into the shared storage views (see the TODO in |
aba9c92 to
f9562f2
Compare
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
Preserve upstream cache aliases with paged row scatter and fall back from the single-pass MLA kernel for padded layouts. Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
18cc9ce to
6fbb74c
Compare
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
|
Hi @ericcurtin, I've addressed the shared MLA cache writes (including a read-back test), hybrid MLA TurboQuant rejection, and the shared KDA/Mamba2 slot-cache adapter. The PR is open again; local regression and end-to-end tests passed. The macOS CI test is still queued. Could you take another look once it completes? Thanks! |
| cache.latent_caches = storage.views(latent_tensors) | ||
| cache.dtype = cache.latent_caches[0].dtype | ||
| cache.has_dense_pages = all(tensor.is_contiguous() for tensor in latent_tensors) |
|
Merged, thanks! Follow-on: (1) reject MLA layers spanning more than one scheduler group, since the wrapper reads ctx group 0. (2) share the step/decode/request loops between |



Summary
Dependencies
main, including Bump MLX to 0.32.1 and update the mlx-lm pin #723 (mlx==0.32.1and the MLX-LM pin at9e6acca691e64d6d8bb808c328fcdea459099cca).Validation
Local host: Apple M3 Pro with 18 GB unified memory, macOS 14.4.1, and Python 3.12.12. Current
mainrequires macOS 15 or later, so native artifacts were source-built for this local host; supported-platform CI remains authoritative.Full non-slow test suite:
VLLM_METAL_BUILD_FROM_SOURCE=1 \ VLLM_METAL_DISABLE_NAX=1 \ .venv/bin/python -m pytest -p no:cacheprovider \ -m "not slow" tests -qResult:
2715 passed, 14 skipped, 53 deselected.ruff check,ruff format --check, targetedmypy, andgit diff --checkalso passed.Full-model serving used the 24-layer Ling-3.0 Tiny checkpoint converted from the official BF16 checkpoint to MLX-native MXFP8:
VLLM_METAL_BUILD_FROM_SOURCE=1 \ .venv/bin/vllm serve /path/to/ling_v3_tiny_mxfp8_selective \ --host 127.0.0.1 \ --port 8000 \ --trust-remote-code \ --max-model-len 2048 \ --max-num-seqs 2 \ --default-chat-template-kwargs '{"enable_thinking": false}'The engine resolved
BailingMoeV3ForCausalLM, initialized the shared MLA/KDA cache in align mode, and served the 6 MLA / 18 KDA topology./healthand/v1/chat/completionsreturned HTTP 200. A discount prompt returned$480; two concurrent multiplication requests returned391and1702.Scope
Duplicate-work check
Searches for open PRs using
Ling-3.0 Bailing MLA KDAandBailing hybridfound no duplicate implementation other than this PR.AI assistance
AI assistance was used to inspect upstream implementations, implement and review the changes, and run the reported tests. The submitter reviewed the changed code and remains responsible for the contribution.