ColossalAI/colossalai/lazy
Hongxin Liu 2b415e5999
[shardformer] support ep for deepseek v3 (#6185)
* [feature] support ep for deepseek v3

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* fix test

* [shardformer] fix deepseek v3 init

* [lazy] fit lora for lazy init

* [example] support npu for deepseek v3

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-11 16:10:25 +08:00
..
__init__.py
construction.py
lazy_init.py [shardformer] support ep for deepseek v3 (#6185) 2025-02-11 16:10:25 +08:00
pretrained.py [Feature] Zigzag Ring attention (#5905) 2024-08-16 13:56:38 +08:00