ipex-llm

Author	SHA1	Message	Date
Yina Chen	fc7f8feb83	Support compress kv (#11642 ) * mistral snapkv * update * mtl update * update * update * update * add comments * style fix * fix style * support llama * llama use compress kv * support mistral 4.40 * fix style * support diff transformers versions * move snapkv util to kv * fix style * meet comments & small fix * revert all in one * fix indent --------- Co-authored-by: leonardozcm <leonardo1997zcm@gmail.com>	2024-07-26 16:02:00 +08:00
Yishuo Wang	6bcdc6cc8f	fix qwen2 cpu (#11663 )	2024-07-26 13:41:51 +08:00
Wang, Jian4	23681fbf5c	Support codegeex4-9b for lightweight-serving (#11648 ) * add options, support prompt and not return end_token * enable openai parameter * set do_sample None and update style	2024-07-26 09:41:03 +08:00
Guancheng Fu	a4d30a8211	Change logic for detecting if vllm is available (#11657 ) * fix * fix	2024-07-25 15:24:19 +08:00
Xiangyu Tian	4499d25c26	LLM: Fix ParallelLMHead convert in vLLM cpu (#11654 )	2024-07-25 13:07:19 +08:00
binbin Deng	777e61d8c8	Fix qwen2 & int4 on NPU (#11646 )	2024-07-24 13:14:39 +08:00
Yishuo Wang	1b3b46e54d	fix chatglm new model (#11639 )	2024-07-23 13:44:56 +08:00
Xiangyu Tian	060792a648	LLM: Refine Pipeline Parallel FastAPI (#11587 ) Refine Pipeline Parallel FastAPI	2024-07-22 15:52:05 +08:00
Wang, Jian4	1eed0635f2	Add lightweight serving and support tgi parameter (#11600 ) * init tgi request * update openai api * update for pp * update and add readme * add to docker * add start bash * update * update * update	2024-07-19 13:15:56 +08:00
Xiangyu Tian	d27a8cd08c	Fix Pipeline Parallel dtype (#11623 )	2024-07-19 13:07:40 +08:00
Yishuo Wang	d020ad6397	add save_low_bit support for DiskEmbedding (#11621 )	2024-07-19 10:34:53 +08:00
Guoqiong Song	380717f50d	fix gemma for 4.41 (#11531 ) * fix gemma for 4.41	2024-07-18 15:02:50 -07:00
Guoqiong Song	5a6211fd56	fix minicpm for transformers>=4.39 (#11533 ) * fix minicpm for transformers>=4.39	2024-07-18 15:01:57 -07:00
Yishuo Wang	0209427cf4	Add disk_embedding parameter to support put Embedding layer on CPU (#11617 )	2024-07-18 17:06:06 +08:00
Xiangyu Tian	4594a3dd6c	LLM: Fix DummyLayer.weight device in Pipeline Parallel (#11612 )	2024-07-18 13:39:34 +08:00
Yishuo Wang	f4077fa905	fix llama3-8b npu long input stuck (#11613 )	2024-07-18 11:08:17 +08:00
Zhao Changmin	e5c0058c0e	fix baichuan (#11606 )	2024-07-18 09:43:36 +08:00
Guoqiong Song	d64711900a	Fix cohere model on transformers>=4.41 (#11575 ) * fix cohere model for 4-41	2024-07-17 17:18:59 -07:00
Wang, Jian4	9c15abf825	Refactor fastapi-serving and add one card serving(#11581 ) * init fastapi-serving one card * mv api code to source * update worker * update for style-check * add worker * update bash * update * update worker name and add readme * rename update * rename to fastapi	2024-07-17 11:12:43 +08:00
Yishuo Wang	5837bc0014	fix chatglm3 npu output (#11590 )	2024-07-16 18:16:30 +08:00
Guancheng Fu	06930ab258	Enable ipex-llm optimization for lm head (#11589 ) * basic * Modify convert.py * fix	2024-07-16 16:48:44 +08:00
Yina Chen	99c22745b2	fix qwen 14b fp6 abnormal output (#11583 )	2024-07-16 10:59:00 +08:00
Yishuo Wang	c279849d27	add disk embedding api (#11585 )	2024-07-16 10:43:39 +08:00
Xiangyu Tian	79c742dfd5	LLM: Add XPU Memory Optimizations for Pipeline Parallel (#11567 ) Add XPU Memory Optimizations for Pipeline Parallel	2024-07-16 09:44:50 +08:00
Yishuo Wang	019da6c0ab	use mlp silu_mul fusion in qwen2 to optimize memory usage (#11574 )	2024-07-13 16:32:54 +08:00
Yishuo Wang	a945500a98	fix internlm xcomposser stream chat (#11564 )	2024-07-11 18:21:17 +08:00
Zhao Changmin	b9c66994a5	add npu sdp (#11562 )	2024-07-11 16:57:35 +08:00
binbin Deng	2b8ad8731e	Support pipeline parallel for glm-4v (#11545 )	2024-07-11 16:06:06 +08:00
Xiangyu Tian	7f5111a998	LLM: Refine start script for Pipeline Parallel Serving (#11557 ) Refine start script and readme for Pipeline Parallel Serving	2024-07-11 15:45:27 +08:00
Zhao Changmin	105e124752	optimize phi3-v encoder npu performance and add multimodal example (#11553 ) * phi3-v * readme	2024-07-11 13:59:14 +08:00
Cengguang Zhang	70ab1a6f1a	LLM: unify memory optimization env variables. (#11549 ) * LLM: unify memory optimization env variables. * fix comments.	2024-07-11 11:01:28 +08:00
Yishuo Wang	994e49a510	optimize internlm xcomposser performance again (#11551 )	2024-07-10 17:08:56 +08:00
Yishuo Wang	82f9514303	optimize internlm xcomposer2 performance (#11550 )	2024-07-10 15:57:04 +08:00
Zhao Changmin	3c16c9f725	Optimize baichuan on NPU (#11548 ) * baichuan_npu	2024-07-10 13:18:48 +08:00
Yishuo Wang	7dc6756d86	add disk embedding (#11543 )	2024-07-09 17:38:40 +08:00
Yishuo Wang	99b2802d3b	optimize qewn2 memory (#11535 )	2024-07-09 17:14:01 +08:00
Yishuo Wang	2929eb262e	support npu glm4 (#11539 )	2024-07-09 15:46:49 +08:00
Xiangyu Tian	a1cede926d	Fix update_kv_cache in Pipeline-Parallel-Serving for glm4-9b model (#11537 )	2024-07-09 14:08:04 +08:00
Yishuo Wang	c26651f91f	add mistral npu support (#11523 )	2024-07-08 13:17:15 +08:00
binbin Deng	252426793b	Fix setting of `use_quantize_kv_cache` on different GPU in pipeline parallel (#11516 )	2024-07-08 09:27:01 +08:00
Yishuo Wang	7cb09a8eac	optimize qwen2 memory usage again (#11520 )	2024-07-05 17:32:34 +08:00
Zhao Changmin	f7e957aaf9	Clean npu dtype branch (#11515 ) * clean branch * create_npu_kernels	2024-07-05 15:45:26 +08:00
Yishuo Wang	14ce058004	add chatglm3 npu support (#11518 )	2024-07-05 15:31:27 +08:00
Xin Qiu	a31f2cbe13	update minicpm.py (#11517 ) * update minicpm * meet code review	2024-07-05 15:25:44 +08:00
Zhao Changmin	24de13fc45	Optimize stablelm on NPU (#11512 ) * stablelm_optimize	2024-07-05 14:21:57 +08:00
Xiangyu Tian	7d8bc83415	LLM: Partial Prefilling for Pipeline Parallel Serving (#11457 ) LLM: Partial Prefilling for Pipeline Parallel Serving	2024-07-05 13:10:35 +08:00
binbin Deng	60de428b37	Support pipeline parallel for qwen-vl (#11503 )	2024-07-04 18:03:57 +08:00
Zhao Changmin	57b8adb189	[WIP] Support npu load_low_bit method (#11502 ) * npu_load_low_bit	2024-07-04 17:15:34 +08:00
Yishuo Wang	1a8bab172e	add minicpm 1B/2B npu support (#11507 )	2024-07-04 16:31:04 +08:00
Yishuo Wang	bb0a84044b	add qwen2 npu support (#11504 )	2024-07-04 11:01:25 +08:00

1 2 3 4 5 ...

312 commits