ipex-llm

Author	SHA1	Message	Date
Yishuo Wang	2929eb262e	support npu glm4 (#11539 )	2024-07-09 15:46:49 +08:00
Xiangyu Tian	a1cede926d	Fix update_kv_cache in Pipeline-Parallel-Serving for glm4-9b model (#11537 )	2024-07-09 14:08:04 +08:00
Cengguang Zhang	fa81dbefd3	LLM: update multi gpu write csv in all-in-one benchmark. (#11538 )	2024-07-09 11:14:17 +08:00
Xin Qiu	69701b3ec8	fix typo in python/llm/scripts/README.md (#11536 )	2024-07-09 09:53:14 +08:00
Jason Dai	099486afb7	Update README.md (#11530 )	2024-07-08 20:18:41 +08:00
binbin Deng	66f6ffe4b2	Update GPU HF-Transformers example structure (#11526 )	2024-07-08 17:58:06 +08:00
Xu, Shuo	f9a199900d	add model RWKV/v5-Eagle-7B-HF to igpu benchmark (#11528 ) Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-08 15:50:16 +08:00
Shaojun Liu	9b37ca6027	remove (#11527 )	2024-07-08 15:49:52 +08:00
Yishuo Wang	c26651f91f	add mistral npu support (#11523 )	2024-07-08 13:17:15 +08:00
Jun Wang	5a57e54400	[ADD] add 5 new models for igpu-perf (#11524 )	2024-07-08 11:12:15 +08:00
Xu, Shuo	64cfed602d	Add new models to benchmark (#11505 ) * Add new models to benchmark * remove Qwen/Qwen-VL-Chat to pass the validation --------- Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-08 10:35:55 +08:00
binbin Deng	252426793b	Fix setting of `use_quantize_kv_cache` on different GPU in pipeline parallel (#11516 )	2024-07-08 09:27:01 +08:00
Yishuo Wang	7cb09a8eac	optimize qwen2 memory usage again (#11520 )	2024-07-05 17:32:34 +08:00
Yuwen Hu	8f376e5192	Change igpu perf to mainly test int4+fp16 (#11513 )	2024-07-05 17:12:33 +08:00
Jun Wang	1efb6ebe93	[ADD] add `transformer_int4_fp16_loadlowbit_gpu_win` api (#11511 ) * [ADD] add transformer_int4_fp16_loadlowbit_gpu_win api * [UPDATE] add int4_fp16_lowbit config and description * [FIX] fix run.py mistake * [FIX] fix run.py mistake * [FIX] fix indent; change dtype=float16 to model.half()	2024-07-05 16:38:41 +08:00
Zhao Changmin	f7e957aaf9	Clean npu dtype branch (#11515 ) * clean branch * create_npu_kernels	2024-07-05 15:45:26 +08:00
Yishuo Wang	14ce058004	add chatglm3 npu support (#11518 )	2024-07-05 15:31:27 +08:00
Xin Qiu	a31f2cbe13	update minicpm.py (#11517 ) * update minicpm * meet code review	2024-07-05 15:25:44 +08:00
Zhao Changmin	24de13fc45	Optimize stablelm on NPU (#11512 ) * stablelm_optimize	2024-07-05 14:21:57 +08:00
Xiangyu Tian	7d8bc83415	LLM: Partial Prefilling for Pipeline Parallel Serving (#11457 ) LLM: Partial Prefilling for Pipeline Parallel Serving	2024-07-05 13:10:35 +08:00
Shaojun Liu	72b4efaad4	Enhanced XPU Dockerfiles: Optimized Environment Variables and Documentation (#11506 ) * Added SYCL_CACHE_PERSISTENT=1 to xpu Dockerfile * Update the document to add explanations for environment variables. * update quickstart	2024-07-04 20:18:38 +08:00
binbin Deng	60de428b37	Support pipeline parallel for qwen-vl (#11503 )	2024-07-04 18:03:57 +08:00
Zhao Changmin	57b8adb189	[WIP] Support npu load_low_bit method (#11502 ) * npu_load_low_bit	2024-07-04 17:15:34 +08:00
Jun Wang	f07937945f	[REMOVE] remove all useless repo-id in benchmark/igpu-perf (#11508 )	2024-07-04 16:38:34 +08:00
Yishuo Wang	1a8bab172e	add minicpm 1B/2B npu support (#11507 )	2024-07-04 16:31:04 +08:00
Yishuo Wang	bb0a84044b	add qwen2 npu support (#11504 )	2024-07-04 11:01:25 +08:00
Shaojun Liu	932ef78131	Update Workflow Inputs, Runner, and PR Validation Process (#11501 ) * update check-artifact runner label to Shire * update github.event.inputs to inputs * update PR template	2024-07-03 16:49:54 +08:00
Xin Qiu	f84ca99b9f	optimize gemma2 rmsnorm (#11500 )	2024-07-03 15:21:03 +08:00
Wang, Jian4	61c36ba085	Add pp_serving verified models (#11498 ) * add verified models * update * verify large model * update commend	2024-07-03 14:57:09 +08:00
binbin Deng	9274282ef7	Support pipeline parallel for glm-4-9b-chat (#11463 )	2024-07-03 14:25:28 +08:00
Shaojun Liu	e7ab93b55c	Update pull_request_template.md (#11484 ) * Update pull_request_template.md * refine	2024-07-03 11:13:16 +08:00
Yishuo Wang	d97c2664ce	use new fuse rope in stablelm family (#11497 )	2024-07-03 11:08:26 +08:00
Jun Wang	18c973dc3e	Wang jun/ipex llm workflow (#11499 ) * [update] merge manually build for testing function to manualy build * [FIX] change public type to string * [FIX] change public type to string * [FIX] remove github.event prefix for inputs	2024-07-03 10:13:42 +08:00
Yuwen Hu	e53bd4401c	Small typo fixes in binary build workflow (#11494 )	2024-07-02 19:11:43 +08:00
Yuwen Hu	4e32c92979	Further fix for triggering perf test from commit (#11493 ) * Further fix for triggering perf test from commit * Small fix	2024-07-02 18:56:53 +08:00
Xu, Shuo	52519e07df	remove models we no longer need in benchmark. (#11492 ) Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-02 17:20:48 +08:00
Zhao Changmin	6a0134a9b2	support q4_0_rtn (#11477 ) * q4_0_rtn	2024-07-02 16:57:02 +08:00
Jun Wang	6352c718f3	[update] merge manually build for testing function to manualy build (#11491 )	2024-07-02 16:28:15 +08:00
Yishuo Wang	5e967205ac	remove the code converts input to fp16 before calling batch forward kernel (#11489 )	2024-07-02 16:23:53 +08:00
Yuwen Hu	1638573f56	Update llama cpp quickstart regarding windows prerequisites to avoid misleading (#11490 )	2024-07-02 16:15:47 +08:00
Yuwen Hu	986b10e397	Further fix for performance tests triggered by pr (#11488 )	2024-07-02 15:29:42 +08:00
Yuwen Hu	bb6953c19e	Support pr validate perf test (#11486 ) * Support triggering performance tests through commits * Small fix * Small fix * Small fixes	2024-07-02 15:20:42 +08:00
Wang, Jian4	4390e7dc49	Fix codegeex2 transformers version (#11487 )	2024-07-02 15:09:28 +08:00
Guancheng Fu	4fbb0d33ae	Pin compute runtime version for xpu images (#11479 ) * pin compute runtime version * fix done	2024-07-01 21:41:02 +08:00
Shaojun Liu	a1164e45b6	Enable Release Pypi workflow to be called in another repo (#11483 )	2024-07-01 19:48:21 +08:00
Yuwen Hu	fb4774b076	Update pull request template for manually-ttriggered Unit tests (#11482 )	2024-07-01 19:06:29 +08:00
Yuwen Hu	ca24794dd0	Fixes for performance test triggering (#11481 )	2024-07-01 18:39:54 +08:00
Yuwen Hu	6bdc562f4c	Enable triggering nightly tests/performance tests from another repo (#11480 ) * Enable triggering from another workflow for nightly tests and example tests * Enable triggering from another workflow for nightly performance tests	2024-07-01 17:45:42 +08:00
Yishuo Wang	ec3a912ab6	optimize npu llama long context performance (#11478 )	2024-07-01 16:49:23 +08:00
Heyang Sun	913e750b01	fix non-string deepseed config path bug (#11476 ) * fix non-string deepseed config path bug * Update lora_finetune_chatglm.py	2024-07-01 15:53:50 +08:00

... 2 3 4 5 6 ...

3278 commits