Skip to content

vllm-project/afd-plugin

Repository files navigation

afd-plugin

Overview

afd-plugin is a vLLM external plugin for Attention-FFN Disaggregation (AFD). It provides plugin-owned worker classes, model runners, model wrappers, connectors, configuration validation, compatibility shims, and hardware-gated integration tests for GPU and Ascend NPU deployments.

Note

This project is still experimental and needs more large-scale testing across different hardware backends.

The target runtime is vLLM v0.19.1. The plugin does not modify the vLLM source tree. AFD behavior is installed through the vllm.general_plugins entry point, --additional-config, automatically selected role workers, plugin-owned model wrappers, and narrow version-scoped compatibility shims.

Architecture

afd-plugin architecture

Current Status

Core runtime support:

  • vLLM plugin registration, AFD configuration, and runtime validation.
  • Attention/FFN workers, model runners, model wrappers, and connector-driven execution for CUDA and Ascend NPU.
  • Eager and FULL_DECODE_ONLY graph execution, plus backend-specific profiling support.

Model support:

Model family Registered architectures Plugin model wrappers Notes
DeepSeekV2 / DeepSeekV3 / DeepSeekV3.2 DeepseekForCausalLM, DeepseekV2ForCausalLM, DeepseekV3ForCausalLM, DeepseekV32ForCausalLM AFDDeepseekForCausalLM, AFDDeepseekV2ForCausalLM, AFDDeepseekV3ForCausalLM DeepSeekV3.2 uses AFDDeepseekV3ForCausalLM. Each AFD role constructs and loads only its role-required model components, while shared embedding, normalization, and output components remain available where required by the model lifecycle.

Connector support:

See the recipe index for deployment and benchmark examples.

Connector Platform Recommend Stage Sync or Async Graph Support Notes
P2pNcclAFDConnector CUDA Decode Sync FULL_DECODE_ONLY CUDA graph FFN ranks are ordered before Attention ranks. num_attention_ranks must be greater than or equal to num_ffn_ranks and divisible by it. See the DeepSeek V2 Lite recipe.
CAMP2pAFDConnector Ascend NPU Decode Sync FULL_DECODE_ONLY ACL graph Uses HCCL/CAMP2P custom ops. Ascend ops build by default on NPU platforms. See the synchronous DeepSeek V3.2 recipe.
CAMAsyncAFDConnector Ascend NPU Prefill Async Not supported Uses CAM async-DP custom ops and requires async=true with the Ascend NPU workers. See the asynchronous DeepSeek V3.2 recipe.

Connector implementations are grouped by backend package: afd_plugin.connectors.gpu for GPU-only connectors, afd_plugin.connectors.npu for NPU-only connectors.

Known gaps:

  • vLLM versions other than 0.19.1 are not claimed as supported.
  • vLLM/vLLM-Ascend model runner v2 is not supported.
  • GPU and NPU E2E tests are opt-in and require real hardware plus model weights.
  • GPU CUDA graph support is limited to FULL_DECODE_ONLY.
  • GPU DBO plus CUDA graph is limited to exactly two ubatches.

Install

Requires Python 3.10-3.13 and uv.

git clone https://github.com/vllm-project/afd-plugin.git
cd afd-plugin
uv sync --group dev

vllm is an optional runtime extra so CPU-only or macOS development environments can still run import/config tests without a CUDA wheel.

GPU installation

For Linux / CUDA-capable environments only; Ascend users should skip this command:

uv sync --group dev --extra vllm

The optional extra pins vllm==0.19.1.

Ascend NPU installation

AFD's Ascend path is validated on openEuler 22.03 (aarch64) with Ascend 910C / Atlas A3. Install a compatible driver and firmware, and confirm the devices with npu-smi info. Use this compatible release baseline:

Component Version
Python 3.10 or 3.11
vLLM 0.19.1
vLLM-Ascend 0.19.1rc1
CANN / NNAL 8.5.1
torch 2.9.0
torch-npu 2.9.0

Environment

The following A3/openEuler environment has been validated:

docker pull quay.io/ascend/vllm-ascend:v0.19.1rc1-a3-openeuler

Use the fixed vLLM-Ascend installation guide to start the image with the device and driver configuration for your host. The image includes the matched CANN and NNAL environment. Run the remaining commands inside the container from the AFD repository root.

Install AFD

From the AFD repository root:

AFD_BUILD_ASCEND_OPS=1 \
SOC_VERSION=ascend910_9391 \
python -m pip install -v --no-build-isolation --no-deps -e .

--no-deps preserves the matched NPU runtime, and --no-build-isolation uses its CANN/torch-npu toolchain. The command above forces the Ascend op build for reproducibility. Otherwise, leave AFD_BUILD_ASCEND_OPS unset to use the auto-detection described in Development, set it to 1 if detection misses an Ascend environment, or set it to 0 to skip the ops.

Verify

python - <<'PY'
import torch
import torch_npu
import vllm_ascend

from afd_plugin.compat.npu import ensure_afd_ascend_ops_loaded

assert torch.npu.is_available(), "torch-npu cannot see an Ascend device"
ensure_afd_ascend_ops_loaded()
print("AFD_OPS_OK")
PY

After AFD_OPS_OK, the environment is ready to run the NPU examples and E2E tests. See the synchronous NPU recipe. For implementation details, see the Attention runtime design and FFN runtime design.

Using the Plugin

Install or sync the distribution as vllm-afd-plugin. Python imports use the afd_plugin package name.

AFD is configured through vLLM --additional-config. There is no separate --afd-config flag. When AFD is configured and --worker-cls is omitted, the plugin automatically selects the Attention or FFN worker for the active CUDA or standard Ascend NPU platform. Explicit AFD worker paths remain accepted for compatibility with existing commands, but are not required or stable launch interfaces.

GPU Attention-side shape:

vllm serve /path/to/DeepSeek-V2-Lite \
  --served-model-name deepseek-v2-lite-afd-attention \
  --data-parallel-size 1 \
  --tensor-parallel-size 1 \
  --enable-expert-parallel \
  --enforce-eager \
  --host 127.0.0.1 \
  --port 18000 \
  --additional-config '{"afd":{"role":"attention","connector":"P2pNcclAFDConnector","host":"127.0.0.1","port":6239,"num_attention_ranks":1,"num_ffn_ranks":1}}'

GPU FFN-side shape:

vllm serve /path/to/DeepSeek-V2-Lite \
  --served-model-name deepseek-v2-lite-afd-ffn \
  --data-parallel-size 1 \
  --tensor-parallel-size 1 \
  --enable-expert-parallel \
  --enforce-eager \
  --host 127.0.0.1 \
  --port 18001 \
  --additional-config '{"afd":{"role":"ffn","connector":"P2pNcclAFDConnector","host":"127.0.0.1","port":6239,"num_attention_ranks":1,"num_ffn_ranks":1}}'

NPU uses the same config channel with CAMP2pAFDConnector; the plugin selects the NPU worker automatically:

vllm serve /path/to/DeepSeek-V2-Lite \
  --served-model-name deepseek-v2-lite-afd-attention \
  --data-parallel-size 1 \
  --tensor-parallel-size 1 \
  --enable-expert-parallel \
  --enforce-eager \
  --host 127.0.0.1 \
  --port 18000 \
  --additional-config '{"afd":{"role":"attention","connector":"CAMP2pAFDConnector","host":"127.0.0.1","port":6239,"num_attention_ranks":1,"num_ffn_ranks":1}}'

Attention and FFN may be started in either order. Send requests only to the Attention API server. FFN workers are connector-driven; scheduler-driven FFN execute_model() calls fail fast.

For repeatable local smoke testing, prefer the bundled runner:

uv run python tests/e2e/runner.py \
  --model /path/to/DeepSeek-V2-Lite \
  --device-backend gpu \
  --num-attention-ranks 1 \
  --num-ffn-ranks 1 \
  --attention-gpus 0 \
  --ffn-gpus 1 \
  --api-port-base 18000 \
  --afd-port 6239 \
  --common-vllm-arg=--trust-remote-code

For NPU, use --device-backend npu; the runner maps the same device arguments to ASCEND_RT_VISIBLE_DEVICES and selects CAMP2pAFDConnector.

AFD Config

The canonical config shape is:

{
  "afd": {
    "role": "attention",
    "connector": "P2pNcclAFDConnector",
    "host": "127.0.0.1",
    "port": 1239,
    "num_attention_ranks": 2,
    "num_ffn_ranks": 1,
    "afd_role_rank": 0,
    "compute_gate_on_attention": false,
    "connector_extra_config": {}
  }
}

role must be attention or ffn. connector must be P2pNcclAFDConnector, CAMP2pAFDConnector, or CAMAsyncAFDConnector. AFD is active when additional_config["afd"] is present and passes common AFD config validation; omit additional_config["afd"] to disable AFD. Connector-owned connector_extra_config is strictly validated by the selected connector parser when the connector/runtime is constructed. The plugin also accepts selected compatibility aliases such as afd_role, afd_connector, afd_host, and afd_port.

Development

Run the default CPU-safe checks:

uv run pytest
uv run ruff check .

Native C/C++ sources are grouped by backend under csrc/: Ascend/CANN sources live in csrc/npu, including the a2e and e2a ACLNN operators, and csrc/gpu is reserved for GPU native sources.

Ascend custom ops are built automatically only when the build environment looks like an Ascend NPU platform, for example when torch_npu, CANN environment variables, or the default Ascend toolkit path are present. GPU builds skip Ascend ops by default. Set AFD_BUILD_ASCEND_OPS=1 or AFD_BUILD_ASCEND_OPS=0 to override the auto-detection.

E2E Test

To run E2E tests, use the run-e2e skill.

License

afd-plugin is licensed under the Apache License 2.0.

Cite

If you find afd-plugin helpful in your research or projects, please consider citing it:

@misc{afdplugin2026,
  title={afd-plugin: Attention-FFN Disaggregation for vLLM},
  author={AFD Plugin Contributors},
  year={2026},
  howpublished={\url{https://github.com/vllm-project/afd-plugin}},
}

About

vLLM plugin for attention-ffn disaggregation support

Resources

Stars

Watchers

Forks

Releases

Packages

Used by

Contributors

Languages