Skip to content

Add paged MLA/KDA serving for Ling-3.0 (Bailing V3) - #607

Merged
ericcurtin merged 4 commits into
vllm-project:mainfrom
FENP:bailing_v3_support
Sep 30, 2026
Merged

ericcurtin merged 4 commits into
vllm-project:mainfrom
FENP:bailing_v3_support

Conversation

@FENP

@FENP FENP commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add paged serving for Ling-3.0's hybrid MLA/KDA architecture.
  • Register KDA as a state family backed by the shared scheduler-managed recurrent state cache.
  • Extend the generic hybrid runtime to bind MLA layers to vLLM's shared cache allocation.
  • Extend the existing MLA wrapper for Bailing V3's query shape and gated output projection.
  • Add focused topology, cache-layout, MLA parity, KDA prefill/decode parity, and documentation coverage.

Dependencies

Validation

Local host: Apple M3 Pro with 18 GB unified memory, macOS 14.4.1, and Python 3.12.12. Current main requires macOS 15 or later, so native artifacts were source-built for this local host; supported-platform CI remains authoritative.

Full non-slow test suite:

VLLM_METAL_BUILD_FROM_SOURCE=1 \
VLLM_METAL_DISABLE_NAX=1 \
  .venv/bin/python -m pytest -p no:cacheprovider \
  -m "not slow" tests -q

Result: 2715 passed, 14 skipped, 53 deselected.

ruff check, ruff format --check, targeted mypy, and git diff --check also passed.

Full-model serving used the 24-layer Ling-3.0 Tiny checkpoint converted from the official BF16 checkpoint to MLX-native MXFP8:

VLLM_METAL_BUILD_FROM_SOURCE=1 \
  .venv/bin/vllm serve /path/to/ling_v3_tiny_mxfp8_selective \
  --host 127.0.0.1 \
  --port 8000 \
  --trust-remote-code \
  --max-model-len 2048 \
  --max-num-seqs 2 \
  --default-chat-template-kwargs '{"enable_thinking": false}'

The engine resolved BailingMoeV3ForCausalLM, initialized the shared MLA/KDA cache in align mode, and served the 6 MLA / 18 KDA topology. /health and /v1/chat/completions returned HTTP 200. A discount prompt returned $480; two concurrent multiplication requests returned 391 and 1702.

Scope

  • Supported weights: the official BF16 checkpoint and MLX-native MXFP8 converted from BF16.
  • Direct loading of the official serialized block-FP8 checkpoint is not supported.
  • Ling-3.0 Tiny was validated end to end. Ling-3.0 Flash was validated only at the configuration and direct-Q model-structure level; its full checkpoint was not loaded.
  • MTP weights may be ignored during loading; this change does not execute MTP decoding.

Duplicate-work check

Searches for open PRs using Ling-3.0 Bailing MLA KDA and Bailing hybrid found no duplicate implementation other than this PR.

AI assistance

AI assistance was used to inspect upstream implementations, implement and review the changes, and run the reported tests. The submitter reviewed the changed code and remains responsible for the contribution.

@FENP FENP changed the title Add paged MLA/KDA serving for Ling-3.0 Bailing V3 Add paged MLA/KDA serving for Ling-3.0 Aug 14, 2026
@FENP
FENP force-pushed the bailing_v3_support branch from 2ab7e6c to 3c149a0 Compare August 25, 2026 12:27
@FENP
FENP force-pushed the bailing_v3_support branch 2 times, most recently from d13c00e to c2efa4f Compare September 2, 2026 03:58
@FENP
FENP marked this pull request as ready for review September 2, 2026 03:58
ricky-chaoju added a commit that referenced this pull request Sep 3, 2026
Groundwork for the Nemotron-H stack in #644, PR 1 of 4. The hybrid
runtime's geometry is currently ten undeclared attributes that lifecycle
mutates onto the runner, and `cache_policy` sizes the linear side as
`num_layers` minus the SDPA count. That arithmetic breaks as soon as a
model has a third kind of layer. On the 52-layer Nemotron checkpoint it
reports 46 recurrent layers when only 23 are.

Because of that, this PR moves the geometry behind one owner, a typed
layer plan built by `families/gdn.py`, in the same shape as the encoder
pooling loaders from #608. Lifecycle shrinks to a four-line route. The
loader table comes with the next PR, so lifecycle imports the GDN
builder directly for now.

For evidence, the main one is an interleaved A/B on `Qwen/Qwen3.5-0.8B`.
Greedy token ids are bit-identical and `per_block_bytes` is identical.
`num_blocks` wobbles a few blocks on both arms since overhead is
measured at profile time. The full suite gives 1921 passed with 15 new
tests, and error messages on malformed configs got stricter, each pinned
by a full-string test.

#598 and #607 both touch these hunks, and #607 passes constructor kwargs
this PR deletes, so whichever lands second rebases. Happy to be the one
who does. The `MambaSpec` shape and the attention-side wiring stay put
until the table lands.

---------

Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com>
Signed-off-by: Yuan Lik Xun <lxyuan0420@gmail.com>
Co-authored-by: Yuan Lik Xun <lxyuan0420@gmail.com>
@FENP
FENP force-pushed the bailing_v3_support branch 3 times, most recently from 6f5fafd to 658d77a Compare September 7, 2026 06:53

@ricky-chaoju ricky-chaoju left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Ling-3.0 part looks well scoped. My concern is what's grown around it:

Eight of the twenty-four files aren't Bailing at all. they're mlx-lm compatibility fixes. cache.state becomes cache.cache in contiguous_cache and prefix caching, plus adjustments in gguf, LoRA, TurboQuant, pooling and xlm_roberta. That rename hits every model through model_runner.py, not just Ling-3.0. Could you split the bump and those fixes into their own PR and rebase Ling-3.0 on top? Same shape as #501 and #645.

@FENP FENP changed the title Add paged MLA/KDA serving for Ling-3.0 Add paged MLA/KDA serving for Ling-3.0 (Bailing V3) Sep 7, 2026
@FENP

FENP commented Sep 7, 2026

Copy link
Copy Markdown
Contributor Author

The Ling-3.0 part looks well scoped. My concern is what's grown around it:

Eight of the twenty-four files aren't Bailing at all. they're mlx-lm compatibility fixes. cache.state becomes cache.cache in contiguous_cache and prefix caching, plus adjustments in gguf, LoRA, TurboQuant, pooling and xlm_roberta. That rename hits every model through model_runner.py, not just Ling-3.0. Could you split the bump and those fixes into their own PR and rebase Ling-3.0 on top? Same shape as #501 and #645.

Thanks, agreed. I’ll split the MLX/MLX-LM bump and the related compatibility fixes into a separate PR, then rebase this PR on top so it only contains the Ling-3.0 MLA/KDA support and related tests and documentation.

@FENP
FENP marked this pull request as draft September 7, 2026 07:36
@FENP
FENP force-pushed the bailing_v3_support branch from 658d77a to 377ca5e Compare September 7, 2026 07:57
ricky-chaoju pushed a commit that referenced this pull request Sep 7, 2026
## Summary

- Bump the exact MLX runtime pin from 0.32.0 to 0.32.1.
- Advance the MLX-LM pin from `254d153f` to `9e6acca6`.
- Update contiguous-cache handling for MLX-LM's `ArraysCache.state` to
`ArraysCache.cache` rename.
- Adapt GGUF, LoRA, TurboQuant, pooling, and XLM-RoBERTa code to the
updated dependency APIs and type information.

## Motivation

This repository-wide dependency and compatibility update was split from
#607 so that the Ling-3.0 MLA/KDA implementation can remain
model-scoped. PR #607 will be rebased on top of this change.

The pinned MLX-LM revision requires MLX 0.32.1 and contains APIs not
present in the latest MLX-LM release. No other open PR was found for the
same MLX 0.32.1 or `ArraysCache` compatibility update.

## Validation

Environment: Apple M3 Pro with 18 GB unified memory, macOS 14.4.1,
Python 3.12.12, MLX 0.32.1, and MLX-LM at
`9e6acca691e64d6d8bb808c328fcdea459099cca`.

```bash
VLLM_METAL_BUILD_FROM_SOURCE=1 \
  .venv/bin/python -m pytest -p no:cacheprovider \
  -m "not slow" tests -q

.venv/bin/python -m ruff check .
.venv/bin/python -m ruff format --check .
.venv/bin/python -m mypy vllm_metal
git diff --check origin/main...HEAD
```

Results:

- Pytest: `1997 passed, 18 skipped, 40 deselected`.
- Ruff: all checks passed; 291 files already formatted.
- Mypy: no issues in 137 source files.
- Git whitespace check: passed.

## AI assistance

AI assistance was used to inspect the dependency changes, separate the
compatibility update from #607, review the diff, and run the reported
tests. The submitter reviewed the changed code and remains responsible
for the contribution.

Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
@FENP
FENP force-pushed the bailing_v3_support branch 2 times, most recently from feabb91 to f7469e0 Compare September 8, 2026 02:58
@FENP
FENP marked this pull request as ready for review September 8, 2026 02:58
@FENP
FENP force-pushed the bailing_v3_support branch from f7469e0 to 00e2f1c Compare September 10, 2026 08:17
@FENP

FENP commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

The Ling-3.0 part looks well scoped. My concern is what's grown around it:

Eight of the twenty-four files aren't Bailing at all. they're mlx-lm compatibility fixes. cache.state becomes cache.cache in contiguous_cache and prefix caching, plus adjustments in gguf, LoRA, TurboQuant, pooling and xlm_roberta. That rename hits every model through model_runner.py, not just Ling-3.0. Could you split the bump and those fixes into their own PR and rebase Ling-3.0 on top? Same shape as #501 and #645.

Hi @ricky-chaoju, #723 has been merged, and I’ve rebased this PR onto the latest main. It now contains only the Ling-3.0 support and related tests/docs. Could you please take another look? Thanks!

@ricky-chaoju

Copy link
Copy Markdown
Collaborator

Sorry, I missed this one.

Scope looks much better at 15 files. It needs a merge from main before it can go anywhere though. It is 45 commits behind, and #778 dropped create_state_cache from StateFamilySpec along with the StateCacheFactory protocol, so BAILING_FAMILY does not construct any more: TypeError: StateFamilySpec.__init__() got an unexpected keyword argument 'create_state_cache'.

Probably a deletion rather than a port. _create_kda_state_cache is line for line the old create_gdn_state_cache with a different error string, and that one went away in #778 too. hybrid.py now builds one shared PagedStateCache from storage.state_views() for every family, so the factory and the kwarg can just go and the geometry still reaches it through the plan.

The MLA routing in cache_policy.py is the part that will actually need reworking.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The cross-cutting MLA/KDA serving changes require final human review.

Review effort: Lite
Findings: None

What changed in this PR

Adds paged MLA/KDA serving support for Ling-3.0 Bailing V3 hybrid models.

Changes:

  • Adds Bailing topology, KDA state handling, and hybrid runtime integration.
  • Extends MLA cache and projection support.
  • Adds tests and documents supported checkpoints.
File Summary
vllm_metal/​v1/​model_lifecycle.py Preserves model architectures.
vllm_metal/​v1/​cache_policy.py Configures hybrid MLA caching.
vllm_metal/​attention/​runtime/​mla.py Rebinds MLA caches.
vllm_metal/​attention/​runtime/​hybrid.py Supports hybrid MLA/KDA execution.
vllm_metal/​attention/​runtime/​families/​bailing.py Defines Bailing topology and state family.
vllm_metal/​attention/​runtime/​factory.py Registers Bailing runtime plans.
vllm_metal/​attention/​impls/​mla.py Adds Bailing MLA compatibility.
vllm_metal/​attention/​impls/​kda.py Implements paged KDA state execution.
tests/​test_model_lifecycle.py Tests configuration resolution and dimensions.
tests/​test_mla_paged_backend.py Tests paged MLA parity.
tests/​test_kv_cache_spec_capacity.py Tests hybrid cache specifications.
tests/​test_kda_paged_backend.py Tests KDA state parity.
tests/​stub_runner.py Adds Bailing test helpers.
tests/​attention/​test_hybrid_runtime_plan.py Tests runtime planning and patching.
docs/​supported_models.md Documents Ling-3.0 support limits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@ericcurtin
ericcurtin requested a lite review from Copilot September 25, 2026 19:00

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Warning

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

Copilot review overview

Review effort: Lite
Findings: 3 Medium severity · 1 Low severity

Open (4)

Comment thread vllm_metal/attention/impls/kda.py Outdated
Comment thread vllm_metal/attention/runtime/families/bailing.py Outdated
Comment thread vllm_metal/v1/model_lifecycle.py Outdated
Comment thread vllm_metal/attention/impls/kda.py
@ericcurtin

Copy link
Copy Markdown
Collaborator

Agreed with @ricky-chaoju above: needs a rebase onto main (45 commits behind), and _create_kda_state_cache/the StateCacheFactory kwarg should be deleted rather than ported since #778 already dropped create_state_cache — hybrid.py builds the shared PagedStateCache for every family now. cache_policy.py's MLA routing will need reworking to match.

@ericcurtin
ericcurtin marked this pull request as draft September 25, 2026 20:18
@FENP
FENP force-pushed the bailing_v3_support branch from 00e2f1c to aba9c92 Compare September 29, 2026 04:17
@FENP
FENP marked this pull request as ready for review September 29, 2026 04:18
@FENP

FENP commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Hi @ricky-chaoju @ericcurtin, I rebuilt this PR on the latest main: the obsolete state-cache factory is gone, MLA routing now uses the shared hybrid cache path, and macOS 15 CI is green. Could you please take another look? Thanks!

@ericcurtin
ericcurtin requested a lite review from Copilot September 29, 2026 07:00

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment on lines +153 to +200
def __call__(
self,
x: mx.array,
mask: mx.array | None = None,
cache: nn.Module | None = None,
**kwargs: Any,
) -> mx.array:
ctx = get_context()
if ctx is None:
return self._inner(x, mask=mask, cache=cache)

step = self._prepare_step(ctx)
self._kda_state_cache.apply_pending_conv_state(self._kda_cache_idx)
self._kda_state_cache.apply_pending_recurrent_state(self._kda_cache_idx)
if step.num_decode_requests == step.num_requests:
return self._run_decode(x, step)
return self._run_requests(x, step)

def _prepare_step(self, ctx: PagedAttentionContext) -> _KDAStep:
cu_seqlens = ctx.cu_seqlens
if cu_seqlens is None or len(cu_seqlens) < 2:
raise RuntimeError("KDA wrapper requires cu_seqlens")
num_requests = len(cu_seqlens) - 1
return _KDAStep(
cu_seqlens=cu_seqlens,
slot_ids=self._kda_state_cache.step_slot_ids(
ctx, self._kda_cache_idx, num_requests
),
num_requests=num_requests,
num_decode_requests=ctx.num_decode_requests,
)

def _run_decode(self, x: mx.array, step: _KDAStep) -> mx.array:
rows = x.reshape(step.num_requests, 1, x.shape[-1])
cache = self._slot_cache(mx.array(step.slot_ids, dtype=mx.int32))
output = self._inner(rows, mask=None, cache=cache)
cache.flush()
return output.reshape(1, step.num_requests, -1)

def _run_requests(self, x: mx.array, step: _KDAStep) -> mx.array:
outputs = []
for req_idx, slot in enumerate(step.slot_ids):
start = step.cu_seqlens[req_idx]
end = step.cu_seqlens[req_idx + 1]
cache = self._slot_cache(mx.array([slot], dtype=mx.int32))
outputs.append(self._inner(x[:, start:end, :], mask=None, cache=cache))
cache.flush()
return mx.concatenate(outputs, axis=1)
Comment thread vllm_metal/attention/runtime/hybrid.py Outdated
Comment thread vllm_metal/attention/caches/mla_cache.py Outdated
@ericcurtin

Copy link
Copy Markdown
Collaborator

Thanks! MLA latent writes look like they will not persist into the shared storage views (see the TODO in runtime/mla.py), so please add a read-back test through from_upstream. Also reject TurboQuant on hybrid MLA, and share the slot-cache wrapper with mamba2.py instead of copying it in kda.py. Setting to draft, please reopen when ready for re-review.

@ericcurtin
ericcurtin marked this pull request as draft September 29, 2026 08:23
@FENP
FENP force-pushed the bailing_v3_support branch from aba9c92 to f9562f2 Compare September 29, 2026 09:49
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
Preserve upstream cache aliases with paged row scatter and fall back from the single-pass MLA kernel for padded layouts.

Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
@FENP
FENP force-pushed the bailing_v3_support branch from 18cc9ce to 6fbb74c Compare September 30, 2026 09:38
Signed-off-by: FENP <yuanyongjie.yyj@antgroup.com>
@FENP
FENP marked this pull request as ready for review September 30, 2026 11:26
@FENP

FENP commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Hi @ericcurtin, I've addressed the shared MLA cache writes (including a read-back test), hybrid MLA TurboQuant rejection, and the shared KDA/Mamba2 slot-cache adapter. The PR is open again; local regression and end-to-end tests passed. The macOS CI test is still queued. Could you take another look once it completes? Thanks!

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

MLA cache binding may use the wrong scheduler group and corrupt cache addressing; preserve the MLA group identity before approval.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity

Open (2)
Resolved since last review (2)

Comment on lines +58 to +60
cache.latent_caches = storage.views(latent_tensors)
cache.dtype = cache.latent_caches[0].dtype
cache.has_dense_pages = all(tensor.is_contiguous() for tensor in latent_tensors)
@ericcurtin
ericcurtin merged commit cfd1cb6 into vllm-project:main Sep 30, 2026
3 checks passed
@ericcurtin

Copy link
Copy Markdown
Collaborator

Merged, thanks! Follow-on: (1) reject MLA layers spanning more than one scheduler group, since the wrapper reads ctx group 0. (2) share the step/decode/request loops between KDAPagedAttentionWrapper and Mamba2PagedStateWrapper.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants