Yishuo Wang
|
2929eb262e
|
support npu glm4 (#11539)
|
2024-07-09 15:46:49 +08:00 |
|
binbin Deng
|
9274282ef7
|
Support pipeline parallel for glm-4-9b-chat (#11463)
|
2024-07-03 14:25:28 +08:00 |
|
Yishuo Wang
|
e8dd8e97ef
|
fix chatglm lookahead on ARC (#11320)
|
2024-06-14 16:26:11 +08:00 |
|
Yishuo Wang
|
7f65836cb9
|
fix chatglm2/3-32k/128k fp16 (#11311)
|
2024-06-14 09:58:07 +08:00 |
|
Xin Qiu
|
1b0c4c8cb8
|
use new rotary two in chatglm4 (#11312)
* use new rotary two in chatglm4
* rempve
|
2024-06-13 19:02:18 +08:00 |
|
Xin Qiu
|
f1410d6823
|
refactor chatglm4 (#11301)
* glm4
* remove useless code
* stype
* add rope_ratio
* update
* fix fp16
* fix style
|
2024-06-13 18:06:04 +08:00 |
|
Xin Qiu
|
592f7aa61e
|
Refine glm1-4 sdp (#11276)
* chatglm
* update
* update
* change chatglm
* update sdpa
* update
* fix style
* fix
* fix glm
* update glm2-32k
* update glm2-32k
* fix cpu
* update
* change lower_bound
|
2024-06-12 17:11:56 +08:00 |
|
Xin Qiu
|
dbc3c2d72d
|
glm4 sdp (#11253)
* glm4 sdp
* fix style
* update comment
|
2024-06-07 15:42:23 +08:00 |
|
Xin Qiu
|
2f809116e2
|
optimize Chatglm4 (#11239)
* chatglm4
* update
* update
* add rms norm
* chatglm4
|
2024-06-06 18:25:20 +08:00 |
|