[Refactor] Integrated some lightllm kernels into token-attention (#4946)

* add some req for inference * clean codes * add codes * add some lightllm deps * clean codes * hello * delete rms files * add some comments * add comments * add doc * add lightllm deps * add lightllm cahtglm2 kernels * add lightllm cahtglm2 kernels * replace rotary embedding with lightllm kernel * add some commnets * add some comments * add some comments * add * replace fwd kernel att1 * fix a arg * add * add * fix token attention * add some comments * clean codes * modify comments * fix readme * fix bug * fix bug --------- Co-authored-by: cuiqing.li <lixx336@gmail.com> Co-authored-by: CjhHa1 <cjh18671720497@outlook.com>
2025-09-11 13:59:08 +00:00 · 2023-10-19 22:22:47 +08:00
parent 11009103be
commit 3a41e8304e
20 changed files with 160 additions and 1555 deletions
--- a/colossalai/kernel/triton/self_attention_nofusion.py
+++ b/colossalai/kernel/triton/self_attention_nofusion.py
@@ -12,6 +12,7 @@ if HAS_TRITON:
    from .qkv_matmul_kernel import qkv_gemm_4d_kernel
    from .softmax import softmax_kernel

+    # adpeted from https://github.com/microsoft/DeepSpeed/blob/master/deepspeed/ops/transformer/inference/triton/triton_matmul_kernel.py#L312
    def self_attention_forward_without_fusion(
        q: torch.Tensor, k: torch.Tensor, v: torch.Tensor, input_mask: torch.Tensor, scale: float
    ):
@@ -141,6 +142,7 @@ if HAS_TRITON:
        )
        return output.view(batches, -1, d_model)

+    # modified from https://github.com/microsoft/DeepSpeed/blob/master/deepspeed/ops/transformer/inference/triton/attention.py#L212
    def self_attention_compute_using_triton(
        qkv, input_mask, layer_past, alibi, scale, head_size, triangular=False, use_flash=False
    ):