flybird11111
28cf1e2c57
fix
2025-04-09 15:20:14 +08:00
flybird11111
397875e640
Update build_on_pr.yml
2025-04-09 15:14:17 +08:00
flybird11111
ca914147eb
Update test_fp16_torch.py
2025-04-09 14:01:47 +08:00
pre-commit-ci[bot]
3491a9f7e3
[pre-commit.ci] auto fixes from pre-commit.com hooks
...
for more information, see https://pre-commit.ci
2025-04-01 07:34:49 +00:00
flybird11111
4b8b67ae23
fix
2025-04-01 15:32:11 +08:00
pre-commit-ci[bot]
822556a8ca
[pre-commit.ci] auto fixes from pre-commit.com hooks
...
for more information, see https://pre-commit.ci
2025-03-31 08:17:17 +00:00
flybird11111
621cb93bb1
fix
2025-03-31 16:16:15 +08:00
flybird11111
8c66b7c3e9
fix
2025-03-31 15:39:37 +08:00
flybird11111
837a503f50
fix
2025-03-31 15:32:51 +08:00
flybird11111
43885a4317
fix
2025-03-31 15:17:30 +08:00
flybird11111
6c728df3e3
fix
2025-03-31 11:22:59 +08:00
flybird11111
0b81be7f7f
add ci machine
2025-03-28 18:04:03 +08:00
flybird11111
40cf89d66e
Merge branch 'hpcaitech:main' into upgrade-transformers
2025-03-27 18:11:39 +08:00
flybird11111
3ecb5000e3
test for upgrading transformers
2025-03-27 18:08:37 +08:00
duanjunwen
44d4053fec
[HotFix] update load lora model Readme; ( #6240 )
...
* [fix] update load lora model Readme;
* [fix] update lora infer readme
* [fix] remove useless comments
2025-03-07 14:14:26 +08:00
Hongxin Liu
6d676ee0e9
[release] update version ( #6236 )
2025-03-03 16:15:09 +08:00
Hongxin Liu
56fe130b15
[hotfix] fix lora load ( #6231 )
...
* [hotfix] fix lora load
* [hotfix] fix hp load
* accelerate deepseek loading
2025-03-01 19:04:14 +08:00
Hongxin Liu
f32861ccc5
[misc] update torch version ( #6206 )
...
* [misc] update torch version
* fix test
* fix test
* fix test
* fix test
2025-02-24 14:35:48 +08:00
YeAnbang
b9e60559b8
Merge pull request #6208 from hpcaitech/grpo_dev
...
[Chat] fix colossalchat bugs
2025-02-20 21:23:16 +08:00
pre-commit-ci[bot]
7595c453a5
[pre-commit.ci] auto fixes from pre-commit.com hooks
...
for more information, see https://pre-commit.ci
2025-02-20 10:25:19 +00:00
YeAnbang
53834b74b9
fix num_train_step update
2025-02-20 18:24:04 +08:00
YeAnbang
0171884664
fix inference rebatching bug
2025-02-20 17:28:49 +08:00
Hongxin Liu
9379cbd668
[release] update version ( #6195 )
...
* [release] update version
* fix test
* fix test
2025-02-20 11:36:18 +08:00
binmakeswell
24dee8f0b7
[doc] DeepSeek V3/R1 news ( #6199 )
...
* [doc] DeepSeek V3/R1 news
* [doc] DeepSeek V3/R1 news
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-19 15:07:29 +08:00
Hongxin Liu
f73ae55394
[application] add lora sft example data ( #6198 )
2025-02-18 20:18:18 +08:00
Tong Li
f8b9e88484
[application] Update README ( #6196 )
...
* remove unused ray
* remove unused readme
* update readme
* update readme
* update
* update
* add link
* update readme
* update readme
* fix link
* update code
* update cititaion
* update
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* update readme
* update project
* add images
* update link
* update note
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-18 20:17:56 +08:00
Hongxin Liu
d54642a263
[application] add lora sft example ( #6192 )
...
* [application] add lora sft example
* update requirements
* update readme
* update comment
* update ci
2025-02-18 13:06:38 +08:00
YeAnbang
d20c8ffd97
Add GRPO and Support RLVR for PPO ( #6186 )
...
* add grpo, support rlvr
* add grpo, support rlvr
* tested deepseek r1 pipeline
* add ci
* verify grpo r1
* verify grpo r1
* update readme, remove unused code
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* remove path
* clean code
* fix circular import
* fix ci OOM
* fix ci OOM
* skip kto tp, fix qwen generation
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-18 09:43:36 +08:00
flybird11111
ce0ec40811
[checkpointio] fix for async io ( #6189 )
2025-02-14 17:34:13 +08:00
flybird11111
510ff7bec2
Merge branch 'hpcaitech:main' into main
2025-02-14 15:24:15 +08:00
Hongxin Liu
5ff5323538
[hotfix] fix zero optim save ( #6191 )
2025-02-14 15:09:50 +08:00
Hongxin Liu
014837e725
[shardformer] support pipeline for deepseek v3 and optimize lora save ( #6188 )
...
* [shardformer] support pipeline for deepseek v3
* [checkpointio] fix lora save
* [devops] update ci env
* [booster] optimize lora
* fix test
* fix test
2025-02-14 14:48:54 +08:00
flybird11111
6fc6a059a0
fix for async io
2025-02-13 14:06:57 +08:00
Wenxuan Tan
ec73f1b5e2
[CI] Cleanup Dist Optim tests with shared helper funcs ( #6125 )
...
* Refractor and cleanup using common helper funcs. Tests passed
* Update comments
* Fix relative import
* Fix param fetching bug
2025-02-12 13:42:34 +08:00
flybird11111
5c09d726a6
[checkpointio] fix checkpoint for 3d ( #6187 )
...
* fix checkpoint io for 3d
* fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Update hybrid_parallel_checkpoint_io.py
* fix
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-12 11:54:55 +08:00
Hongxin Liu
2b415e5999
[shardformer] support ep for deepseek v3 ( #6185 )
...
* [feature] support ep for deepseek v3
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix test
* [shardformer] fix deepseek v3 init
* [lazy] fit lora for lazy init
* [example] support npu for deepseek v3
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-11 16:10:25 +08:00
flybird11111
17062c83b9
[hotfix] fix hybrid checkpointio for sp+dp ( #6184 )
...
* Update hybrid_parallel_plugin.py
* Update hybrid_parallel_plugin.py
* Update hybrid_parallel_plugin.py
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Update build_on_pr.yml
* Update test_zerobubble_pp.py
* fix
* fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-02-06 17:21:04 +08:00
Wenxuan Tan
ca0aa2365d
[Issue template] Add checkbox asking for details to reproduce error ( #6104 )
...
* Add checkbox asking about reproducing error
* update
* Update
* Update checkbox
2025-01-24 14:36:25 +08:00
Lemon Qin
97e60cbbcb
[checkpointio] gather tensor before unpad it if the tensor is both padded and distributed ( #6168 )
2025-01-21 10:23:15 +08:00
Guangyao Zhang
5b094a836b
[Inference]Fix example in readme ( #6178 )
2025-01-08 11:51:50 +08:00
Hongxin Liu
ee81366cac
[checkpointio] support load-pin overlap ( #6177 )
...
* [checkpointio] support load-pin overlap
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* [test] add conftest
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-01-07 16:16:04 +08:00
Hongxin Liu
479067e9bc
[release] update version ( #6174 )
...
* [release] update version
* [devops] fix test pypi ci
* [devops] fix test pypi ci
2025-01-03 11:52:23 +08:00
pre-commit-ci[bot]
7fdef9fd6b
[pre-commit.ci] pre-commit autoupdate ( #6113 )
...
updates:
- [github.com/pre-commit/mirrors-clang-format: v19.1.2 → v19.1.5](https://github.com/pre-commit/mirrors-clang-format/compare/v19.1.2...v19.1.5 )
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-01-02 10:23:20 +08:00
duanjunwen
a9bedc7a43
[Sharderformer] Support zbv in Sharderformer Policy ( #6150 )
...
* [feat] Sharderformer support zbv
* [feat] support chatglm2, command, deepseek for zbv
* [feat] support zbv in shardformer policy:
falcon,gptj,mistral,opt,qwen2,t5, vit, whisper
* [feat] support GPT2FusedLinearConv1D
* [feat] support GPT2FusedLinear (without tp)
* [fix] debug FusedConvLinear
* [shardfromer] support gpt2 policy for zbv, support GPT2FusedLinearConv
Col and Row.
* [Shardformer] support FusedLinear1D base for zbv
* [shardformer] support zbv in FusedLinear1D base, Col, Row
* [shardformer] support zbv in blip2 and sam policy
* [shardformer] fix bug incorrect number of gradients; add fusedLinear
base testcase;
* [fix] fix incorrect number of gradients ;
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* [Shardformer] add en doc for zbv;
* [fix] fix typo in Model compatibility table
* [fix] fix API Reference typo
* [Shardformer] add zh-Han doc for zbv
* [fix] fix Linear name; update en & zh doc
* [fix] fix shardformer doc import err
* [fix] fix shardconfig import in doc
* [fix] fix shardformer doc
* [fix] fix shardconfig doc
* [fix] fix config
* [fix] remove shardconfig
* [fix] fix doc
* [feat] add zbv doc string
* [fix] rm doc
* [fix] fix doc
* [fix] empty zbv doc
* [fix] ifx torch version
* [fix] fix torch version
* [fix] fix torch versions
* [fix] fix torch versions
* [fix] fix pyramid versions
* [fix] fix pyramid, zope version
* [fix] try fix workflow
* [fix] try import ShardConfig in yml
* [fix] fix workflow
* [fix] fix workflow
* [fix] fix workflow
* [fix] fix workflow
* [fix] fix ci
* [fix] fix zbv doc
* [fix] fix param for qkv linear, gpt2fused linear; fix requirments;
* [fix] fix policy use fused_linear
* [fix] fix weight grad none, err caused by weight ptr change
* [fix] fix comm in WeightGradStore
* [fix] fix WeightGradStore pop param
* [fix] remove useless param in doc; fix gpt2 qkv test;
* [shardformer] simplify execute_w_pass_grad_accum;
* [fix] rm useless comments
* [shardformer] simplify execute_w_pass_grad_accum & execute_w_pass
* [shardformer] Run meaningful doc test
* [shadformer] fix doc test cmd;
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2025-01-02 10:22:26 +08:00
Hongxin Liu
af06d162cf
[checkpointio] support non blocking pin load ( #6172 )
...
* [checkpointio] support non blocking pin load
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2024-12-25 17:03:25 +08:00
binmakeswell
836992438f
[news] release colossalai for sora ( #6166 )
...
* [news] release colossalai for sora
* [news] release colossalai for sora
* [news] release colossalai for sora
* [news] release colossalai for sora
2024-12-23 21:59:39 +08:00
Hongxin Liu
8b0ed61490
[hotfix] improve compatibility ( #6165 )
2024-12-23 18:57:08 +08:00
binmakeswell
5f82bfa636
[doc] add bonus event ( #6164 )
...
* [doc] add bonus event
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2024-12-23 17:41:59 +08:00
duanjunwen
fa9d0318e4
[Hotfix] hotfix normalization ( #6163 )
...
* [fix] hotfix normalization
* [hotfix] force doc ci test
* [hotfix] fallback doc
2024-12-23 16:29:48 +08:00
flybird11111
130229fdcb
[checkpointio]support asyncio for 3d ( #6152 )
...
* fix
* fix
* fix
* fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
* Update utils.py
* fix
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2024-12-23 10:24:22 +08:00