ipex-llm

Author	SHA1	Message	Date
Yishuo Wang	7cb09a8eac	optimize qwen2 memory usage again (#11520 )	2024-07-05 17:32:34 +08:00
Yuwen Hu	8f376e5192	Change igpu perf to mainly test int4+fp16 (#11513 )	2024-07-05 17:12:33 +08:00
Jun Wang	1efb6ebe93	[ADD] add `transformer_int4_fp16_loadlowbit_gpu_win` api (#11511 ) * [ADD] add transformer_int4_fp16_loadlowbit_gpu_win api * [UPDATE] add int4_fp16_lowbit config and description * [FIX] fix run.py mistake * [FIX] fix run.py mistake * [FIX] fix indent; change dtype=float16 to model.half()	2024-07-05 16:38:41 +08:00
Zhao Changmin	f7e957aaf9	Clean npu dtype branch (#11515 ) * clean branch * create_npu_kernels	2024-07-05 15:45:26 +08:00
Yishuo Wang	14ce058004	add chatglm3 npu support (#11518 )	2024-07-05 15:31:27 +08:00
Xin Qiu	a31f2cbe13	update minicpm.py (#11517 ) * update minicpm * meet code review	2024-07-05 15:25:44 +08:00
Zhao Changmin	24de13fc45	Optimize stablelm on NPU (#11512 ) * stablelm_optimize	2024-07-05 14:21:57 +08:00
Xiangyu Tian	7d8bc83415	LLM: Partial Prefilling for Pipeline Parallel Serving (#11457 ) LLM: Partial Prefilling for Pipeline Parallel Serving	2024-07-05 13:10:35 +08:00
binbin Deng	60de428b37	Support pipeline parallel for qwen-vl (#11503 )	2024-07-04 18:03:57 +08:00
Zhao Changmin	57b8adb189	[WIP] Support npu load_low_bit method (#11502 ) * npu_load_low_bit	2024-07-04 17:15:34 +08:00
Jun Wang	f07937945f	[REMOVE] remove all useless repo-id in benchmark/igpu-perf (#11508 )	2024-07-04 16:38:34 +08:00
Yishuo Wang	1a8bab172e	add minicpm 1B/2B npu support (#11507 )	2024-07-04 16:31:04 +08:00
Yishuo Wang	bb0a84044b	add qwen2 npu support (#11504 )	2024-07-04 11:01:25 +08:00
Xin Qiu	f84ca99b9f	optimize gemma2 rmsnorm (#11500 )	2024-07-03 15:21:03 +08:00
Wang, Jian4	61c36ba085	Add pp_serving verified models (#11498 ) * add verified models * update * verify large model * update commend	2024-07-03 14:57:09 +08:00
binbin Deng	9274282ef7	Support pipeline parallel for glm-4-9b-chat (#11463 )	2024-07-03 14:25:28 +08:00
Yishuo Wang	d97c2664ce	use new fuse rope in stablelm family (#11497 )	2024-07-03 11:08:26 +08:00
Xu, Shuo	52519e07df	remove models we no longer need in benchmark. (#11492 ) Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-02 17:20:48 +08:00
Zhao Changmin	6a0134a9b2	support q4_0_rtn (#11477 ) * q4_0_rtn	2024-07-02 16:57:02 +08:00
Yishuo Wang	5e967205ac	remove the code converts input to fp16 before calling batch forward kernel (#11489 )	2024-07-02 16:23:53 +08:00
Wang, Jian4	4390e7dc49	Fix codegeex2 transformers version (#11487 )	2024-07-02 15:09:28 +08:00
Yishuo Wang	ec3a912ab6	optimize npu llama long context performance (#11478 )	2024-07-01 16:49:23 +08:00
Heyang Sun	913e750b01	fix non-string deepseed config path bug (#11476 ) * fix non-string deepseed config path bug * Update lora_finetune_chatglm.py	2024-07-01 15:53:50 +08:00
binbin Deng	48ad482d3d	Fix import error caused by pydantic on cpu (#11474 )	2024-07-01 15:49:49 +08:00
Yishuo Wang	39bcb33a67	add sdp support for stablelm 3b (#11473 )	2024-07-01 14:56:15 +08:00
Zhao Changmin	cf8eb7b128	Init NPU quantize method and support q8_0_rtn (#11452 ) * q8_0_rtn * fix float point	2024-07-01 13:45:07 +08:00
Yishuo Wang	319a3b36b2	fix npu llama2 (#11471 )	2024-07-01 10:14:11 +08:00
Heyang Sun	07362ffffc	ChatGLM3-6B LoRA Fine-tuning Demo (#11450 ) * ChatGLM3-6B LoRA Fine-tuning Demo * refine * refine * add 2-card deepspeed * refine format * add mpi4py and deepspeed install	2024-07-01 09:18:39 +08:00
Xiangyu Tian	fd933c92d8	Fix: Correct num_requests in benchmark for Pipeline Parallel Serving (#11462 )	2024-06-28 16:10:51 +08:00
SONG Ge	a414e3ff8a	add pipeline parallel support with load_low_bit (#11414 )	2024-06-28 10:17:56 +08:00
Cengguang Zhang	d0b801d7bc	LLM: change write mode in all-in-one benchmark. (#11444 ) * LLM: change write mode in all-in-one benchmark. * update output style.	2024-06-27 19:36:38 +08:00
binbin Deng	987017ef47	Update pipeline parallel serving for more model support (#11428 )	2024-06-27 18:21:01 +08:00
Yishuo Wang	029ff15d28	optimize npu llama2 first token performance (#11451 )	2024-06-27 17:37:33 +08:00
Qiyuan Gong	4e4ecd5095	Control sys.modules ipex duplicate check with BIGDL_CHECK_DUPLICATE_IMPORT (#11453 ) * Control sys.modules ipex duplicate check with BIGDL_CHECK_DUPLICATE_IMPORT。	2024-06-27 17:21:45 +08:00
Yishuo Wang	c6e5ad668d	fix internlm xcomposser meta-instruction typo (#11448 )	2024-06-27 15:29:43 +08:00
Yishuo Wang	f89ca23748	optimize npu llama2 perf again (#11445 )	2024-06-27 15:13:42 +08:00
Yishuo Wang	cf0f5c4322	change npu document (#11446 )	2024-06-27 13:59:59 +08:00
binbin Deng	508c364a79	Add precision option in PP inference examples (#11440 )	2024-06-27 09:24:27 +08:00
Yishuo Wang	2a0f8087e3	optimize qwen2 gpu memory usage again (#11435 )	2024-06-26 16:52:29 +08:00
Shaojun Liu	ab9f7f3ac5	FIX: Qwen1.5-GPTQ-Int4 inference error (#11432 ) * merge_qkv if quant_method is 'gptq' * fix python style checks * refactor * update GPU example	2024-06-26 15:36:22 +08:00
Guancheng Fu	99cd16ef9f	Fix error while using pipeline parallism (#11434 )	2024-06-26 15:33:47 +08:00
Jiao Wang	40fa23560e	Fix LLAVA example on CPU (#11271 ) * update * update * update * update	2024-06-25 20:04:59 -07:00
Yishuo Wang	ca0e69c3a7	optimize npu llama perf again (#11431 )	2024-06-26 10:52:54 +08:00
Yishuo Wang	9f6e5b4fba	optimize llama npu perf (#11426 )	2024-06-25 17:43:20 +08:00
binbin Deng	e473b8d946	Add more qwen1.5 and qwen2 support for pipeline parallel inference (#11423 )	2024-06-25 15:49:32 +08:00
binbin Deng	aacc1fd8c0	Fix shape error when run qwen1.5-14b using deepspeed autotp (#11420 )	2024-06-25 13:48:37 +08:00
Yishuo Wang	3b23de684a	update npu examples (#11422 )	2024-06-25 13:32:53 +08:00
Xiangyu Tian	8ddae22cfb	LLM: Refactor Pipeline-Parallel-FastAPI example (#11319 ) Initially Refactor for Pipeline-Parallel-FastAPI example	2024-06-25 13:30:36 +08:00
SONG Ge	34c15d3a10	update pp document (#11421 )	2024-06-25 10:17:20 +08:00
Xin Qiu	9e4ee61737	rename BIGDL_OPTIMIZE_LM_HEAD to IPEX_LLM_LAST_LM_HEAD and add qwen2 (#11418 )	2024-06-24 18:42:37 +08:00

1 2 3 4 5 ...

1542 commits