ipex-llm

Author	SHA1	Message	Date
Chu,Youcheng	28d1c972da	add mixed_precision argument on ppl wikitext evaluation (#11813 ) * fix: delete ipex extension import in ppl wikitext evaluation * feat: add mixed_precision argument on ppl wikitext evaluation * fix: delete mix_precision command in perplex evaluation for wikitext * fix: remove fp16 mixed-presicion argument * fix: Add a space. --------- Co-authored-by: Jinhe Tang <jin.tang1337@gmail.com>	2024-08-15 17:58:53 +08:00
Chu,Youcheng	3ac83f8396	fix: delete ipex extension import in ppl wikitext evaluation (#11806 ) Co-authored-by: Jinhe Tang <jin.tang1337@gmail.com>	2024-08-15 13:40:01 +08:00
Yuwen Hu	356281cb80	Further all-in-one benchmark update `continuation` task (#11784 ) * Further update prompt for continuation task, and disable lookup candidate update strategy on MTL * style fix	2024-08-14 14:39:34 +08:00
Yuwen Hu	81824ff8c9	Fix stdout in all-in-one benchmark to utf-8 (#11772 )	2024-08-13 10:51:08 +08:00
Yuwen Hu	f97a77ea4e	Update all-in-one benchmark for `continuation` task input preparation (#11760 ) * All use 8192.txt for prompt preparation for now * Small fix * Fix text encoding mode to utf-8 * Small update	2024-08-12 17:49:45 +08:00
Jin, Qiao	05989ad0f9	Update npu example and all in one benckmark (#11766 )	2024-08-12 16:46:46 +08:00
Ruonan Wang	66fe2ee464	initial support of `IPEX_LLM_PERFORMANCE_MODE` (#11754 ) * add perf mode * update * fix style	2024-08-09 19:04:09 +08:00
Zijie Li	e7f7141781	Add benchmark util for transformers 4.42 (#11725 ) * add new benchmark_util.py Add new benchmark_util.py for transformers>=4.43.1. The old one renamed to benchmark_util_prev.py. * Small fix to import code * Update __init__.py * fix file names * Update lint-python Update lint-python to exclude benchmark_util_4_29.py benchmark_util_4_43.py * Update benchmark_util_4_43.py * add benchmark_util for transformers 4.42	2024-08-07 08:48:07 +08:00
Zijie Li	8fb36b9f4a	add new benchmark_util.py (#11713 ) * add new benchmark_util.py	2024-08-05 16:18:48 +08:00
RyuKosei	1da1f1dd0e	Combine two versions of run_wikitext.py (#11597 ) * Combine two versions of run_wikitext.py * Update run_wikitext.py * Update run_wikitext.py * aligned the format * update error display * simplified argument parser --------- Co-authored-by: jenniew <jenniewang123@gmail.com>	2024-07-29 15:56:16 +08:00
Yishuo Wang	7f88ce23cd	add more gemma2 optimization (#11673 )	2024-07-29 11:13:00 +08:00
Yishuo Wang	3e8819734b	add basic gemma2 optimization (#11672 )	2024-07-29 10:46:51 +08:00
Qiyuan Gong	0c6e0b86c0	Refine continuation get input_str (#11652 ) * Remove duplicate code in continuation get input_str. * Avoid infinite loop in all-in-one due to test_length not in the list.	2024-07-25 14:41:19 +08:00
Xu, Shuo	7f80db95eb	Change run.py in benchmark to support phi-3-vision in arc-perf (#11638 ) Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-23 09:51:36 +08:00
Zhao Changmin	06745e5742	Add npu benchmark all-in-one script (#11571 ) * npu benchmark	2024-07-15 10:42:37 +08:00
Xu, Shuo	1355b2ce06	Add model Qwen-VL-Chat to iGPU-perf (#11558 ) * Add model Qwen-VL-Chat to iGPU-perf * small fix --------- Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-11 15:39:02 +08:00
Xu, Shuo	028ad4f63c	Add model phi-3-vision-128k-instruct to iGPU-perf benchmark (#11554 ) * try to improve MIniCPM performance * Add model phi-3-vision-128k-instruct to iGPU-perf benchmark --------- Co-authored-by: ATMxsp01 <shou.xu@intel.com>	2024-07-10 17:26:30 +08:00
Cengguang Zhang	fa81dbefd3	LLM: update multi gpu write csv in all-in-one benchmark. (#11538 )	2024-07-09 11:14:17 +08:00
Shaojun Liu	9b37ca6027	remove (#11527 )	2024-07-08 15:49:52 +08:00
Jun Wang	1efb6ebe93	[ADD] add `transformer_int4_fp16_loadlowbit_gpu_win` api (#11511 ) * [ADD] add transformer_int4_fp16_loadlowbit_gpu_win api * [UPDATE] add int4_fp16_lowbit config and description * [FIX] fix run.py mistake * [FIX] fix run.py mistake * [FIX] fix indent; change dtype=float16 to model.half()	2024-07-05 16:38:41 +08:00
Cengguang Zhang	d0b801d7bc	LLM: change write mode in all-in-one benchmark. (#11444 ) * LLM: change write mode in all-in-one benchmark. * update output style.	2024-06-27 19:36:38 +08:00
RyuKosei	05a8d051f6	Fix run.py run_ipex_fp16_gpu (#11361 ) * fix a bug on run.py * Update run.py fixed the format problem --------- Co-authored-by: sgwhat <ge.song@intel.com>	2024-06-20 10:29:32 +08:00
hxsz1997	44f22cba70	add config and default value (#11344 ) * add config and default value * add config in taml * remove lookahead and max_matching_ngram_size in config * remove streaming and use_fp16_torch_dtype in test yaml * update task in readme * update commit of task	2024-06-18 15:28:57 +08:00
hxsz1997	99b309928b	Add lookahead in test_api: transformer_int4_fp16_gpu (#11337 ) * add lookahead in test_api:transformer_int4_fp16_gpu * change the short prompt of summarize * change short prompt to cnn_64 * change short prompt of summarize	2024-06-17 17:41:41 +08:00
binbin Deng	6ea1e71af0	Update PP inference benchmark script (#11323 )	2024-06-17 09:59:36 +08:00
binbin Deng	f97cce2642	Fix import error of ds autotp (#11307 )	2024-06-13 16:22:52 +08:00
Ruonan Wang	986af21896	fix perf test(#11295 )	2024-06-13 10:35:48 +08:00
Ruonan Wang	14b1e6b699	Fix gguf_q4k (#11293 ) * udpate embedding parameter * update benchmark	2024-06-12 20:43:08 +08:00
Yuwen Hu	fac49f15e3	Remove manual importing ipex in all-in-one benchmark (#11272 )	2024-06-11 09:32:13 +08:00
Shaojun Liu	85df5e7699	fix nightly perf test (#11251 )	2024-06-07 09:33:14 +08:00
hxsz1997	b6234eb4e2	Add task in allinone (#11226 ) * add task * update prompt * modify typos * add more cases in summarize * Make the summarize & QA prompt preprocessing as a util function	2024-06-06 17:22:40 +08:00
Wenjing Margaret Mao	231b968aba	Modify the check_results.py to support batch 2&4 (#11133 ) * add batch 2&4 and exclude to perf_test * modify the perf-test&437 yaml * modify llm_performance_test.yml * remove batch 4 * modify check_results.py to support batch 2&4 * change the batch_size format * remove genxir * add str(batch_size) * change actual_test_casese in check_results file to support batch_size * change html highlight * less models to test html and html_path * delete the moe model * split batch html * split * use installing from pypi * use installing from pypi - batch2 * revert cpp * revert cpp * merge two jobs into one, test batch_size in one job * merge two jobs into one, test batch_size in one job * change file directory in workflow * try catch deal with odd file without batch_size * modify pandas version * change the dir * organize the code * organize the code * remove Qwen-MOE * modify based on feedback * modify based on feedback * modify based on second round of feedback * modify based on second round of feedback + change run-arc.sh mode * modify based on second round of feedback + revert config * modify based on second round of feedback + revert config * modify based on second round of feedback + remove comments * modify based on second round of feedback + remove comments * modify based on second round of feedback + revert arc-perf-test * modify based on third round of feedback * change error type * change error type * modify check_results.html * split batch into two folders * add all models * move csv_name * revert pr test * revert pr test --------- Co-authored-by: Yishuo Wang <yishuo.wang@intel.com>	2024-06-05 15:04:55 +08:00
Kai Huang	f93664147c	Update config.yaml (#11208 ) * update config.yaml * fix * minor * style	2024-06-04 19:58:18 +08:00
Yina Chen	711fa0199e	Fix fp6k phi3 ppl core dump (#11204 )	2024-06-04 16:44:27 +08:00
Cengguang Zhang	3eb13ccd8c	LLM: fix input length condition in deepspeed all-in-one benchmark. (#11185 )	2024-06-03 10:05:43 +08:00
ZehuaCao	4127b99ed6	Fix null pointer dereferences error. (#11125 ) * delete unused function on tgi_server * update * update * fix style	2024-05-30 16:16:10 +08:00
hxsz1997	62b2d8af6b	Add lookahead in all-in-one (#11142 ) * add lookahead in allinone * delete save to csv in run_transformer_int4_gpu * change lookup to lookahead * fix the error of add model.peak_memory * Set transformer_int4_gpu as the default option * add comment of transformer_int4_fp16_lookahead_gpu	2024-05-28 15:39:58 +08:00
Zhao Changmin	15d906a97b	Update linux igpu run script (#11098 ) * update run script	2024-05-22 17:18:07 +08:00
Kai Huang	f63172ef63	Align ppl with llama.cpp (#11055 ) * update script * remove * add header * update readme	2024-05-22 16:43:11 +08:00
Wang, Jian4	74950a152a	Fix tgi_api_server error file name (#11075 )	2024-05-20 16:48:40 +08:00
Wang, Jian4	d9f71f1f53	Update benchmark util for example using (#11027 ) * mv benchmark_util.py to utils/ * remove * update	2024-05-15 14:16:35 +08:00
binbin Deng	4053a6ef94	Update environment variable setting in AutoTP with arc (#11018 )	2024-05-15 10:23:58 +08:00
Shaojun Liu	7f8c5b410b	Quickstart: Run PyTorch Inference on Intel GPU using Docker (on Linux or WSL) (#10970 ) * add entrypoint.sh * add quickstart * remove entrypoint * update * Install related library of benchmarking * update * print out results * update docs * minor update * update * update quickstart * update * update * update * update * update * update * add chat & example section * add more details * minor update * rename quickstart * update * minor update * update * update config.yaml * update readme * use --gpu * add tips * minor update * update	2024-05-14 12:58:31 +08:00
ZehuaCao	99255fe36e	fix ppl (#10996 )	2024-05-13 13:57:19 +08:00
Xin Qiu	dfa3147278	update (#10944 )	2024-05-08 14:28:05 +08:00
Cengguang Zhang	0edef1f94c	LLM: add min_new_tokens to all in one benchmark. (#10911 )	2024-05-06 09:32:59 +08:00
Yuwen Hu	1a8a93d5e0	Further fix nightly perf (#10901 )	2024-04-28 10:18:58 +08:00
Yuwen Hu	ddfdaec137	Fix nightly perf (#10899 ) * Fix nightly perf by adding default value in benchmark for use_fp16_torch_dtype * further fixes	2024-04-28 09:39:29 +08:00
binbin Deng	f51bf018eb	Add benchmark script for pipeline parallel inference (#10873 )	2024-04-26 15:28:11 +08:00
Cengguang Zhang	cd369c2715	LLM: add device id to benchmark utils. (#10877 )	2024-04-25 14:01:51 +08:00
Cengguang Zhang	eb39c61607	LLM: add min new token to perf test. (#10869 )	2024-04-24 14:32:02 +08:00
yb-peng	c9dee6cd0e	Update 8192.txt (#10824 ) * Update 8192.txt * Update 8192.txt with original text	2024-04-23 14:02:09 +08:00
Wang, Jian4	23c6a52fb0	LLM: Fix ipex torchscript=True error (#10832 ) * remove * update * remove torchscript	2024-04-22 15:53:09 +08:00
Kai Huang	053ec30737	Transformers ppl evaluation on wikitext (#10784 ) * tranformers code * cache	2024-04-18 15:27:18 +08:00
hxsz1997	0d518aab8d	Merge pull request #10697 from MargarettMao/ceval combine english and chinese, remove nan	2024-04-12 14:37:47 +08:00
jenniew	cdbb1de972	Mark Color Modification	2024-04-12 14:00:50 +08:00
yb-peng	2685c41318	Modify all-in-one benchmark (#10726 ) * Update 8192 prompt in all-in-one * Add cpu_embedding param for linux api * Update run.py * Update README.md	2024-04-11 13:38:50 +08:00
Wenjing Margaret Mao	289cc99cd6	Update README.md (#10700 ) Edit "summarize the results"	2024-04-09 16:01:12 +08:00
Wenjing Margaret Mao	d3116de0db	Update README.md (#10701 ) edit "summarize the results"	2024-04-09 15:50:25 +08:00
Chen, Zhentao	d59e0cce5c	Migrate harness to ipexllm (#10703 ) * migrate to ipexlm * fix workflow * fix run_multi * fix precision map * rename ipexlm to ipexllm * rename bigdl to ipex in comments	2024-04-09 15:48:53 +08:00
jenniew	591bae092c	combine english and chinese, remove nan	2024-04-08 19:37:51 +08:00
yb-peng	2d88bb9b4b	add test api transformer_int4_fp16_gpu (#10627 ) * add test api transformer_int4_fp16_gpu * update config.yaml and README.md in all-in-one * modify run.py in all-in-one * re-order test-api * re-order test-api in config * modify README.md in all-in-one * modify README.md in all-in-one * modify config.yaml --------- Co-authored-by: pengyb2001 <arda@arda-arc21.sh.intel.com> Co-authored-by: ivy-lv11 <zhicunlv@gmail.com>	2024-04-07 15:47:17 +08:00
Wang, Jian4	9ad4b29697	LLM: CPU benchmark using tcmalloc (#10675 )	2024-04-07 14:17:01 +08:00
binbin Deng	d9a1153b4e	LLM: upgrade deepspeed in AutoTP on GPU (#10647 )	2024-04-07 14:05:19 +08:00
binbin Deng	27be448920	LLM: add `cpu_embedding` and peak memory record for deepspeed autotp script (#10621 )	2024-04-02 17:32:50 +08:00
Shaojun Liu	a10f5a1b8d	add python style check (#10620 ) * add python style check * fix style checks * update runner * add ipex-llm-finetune-qlora-cpu-k8s to manually_build workflow * update tag to 2.1.0-SNAPSHOT	2024-04-02 16:17:56 +08:00
Ruonan Wang	d6af4877dd	LLM: remove ipex.optimize for gpt-j (#10606 ) * remove ipex.optimize * fix * fix	2024-04-01 12:21:49 +08:00
WeiguangHan	fbeb10c796	LLM: Set different env based on different Linux kernels (#10566 )	2024-03-27 17:56:33 +08:00
Ruonan Wang	ea4bc450c4	LLM: add esimd sdp for pvc (#10543 ) * add esimd sdp for pvc * update * fix * fix batch	2024-03-26 19:04:40 +08:00
Shaojun Liu	c563b41491	add nightly_build workflow (#10533 ) * add nightly_build workflow * add create-job-status-badge action * update * update * update * update setup.py * release * revert	2024-03-26 12:47:38 +08:00
Wang, Jian4	16b2ef49c6	Update_document by heyang (#30 )	2024-03-25 10:06:02 +08:00
Wang, Jian4	9df70d95eb	Refactor bigdl.llm to ipex_llm (#24 ) * Rename bigdl/llm to ipex_llm * rm python/llm/src/bigdl * from bigdl.llm to from ipex_llm	2024-03-22 15:41:21 +08:00
binbin Deng	85ef3f1d99	LLM: add empty cache in deepspeed autotp benchmark script (#10488 )	2024-03-21 10:51:23 +08:00
Xiangyu Tian	5a5fd5af5b	LLM: Add speculative benchmark on CPU/XPU (#10464 ) Add speculative benchmark on CPU/XPU.	2024-03-21 09:51:06 +08:00
Xiangyu Tian	cbe24cc7e6	LLM: Enable BigDL IPEX Int8 (#10480 ) Enable BigDL IPEX Int8	2024-03-20 15:59:54 +08:00
Jin Qiao	e41d556436	LLM: change fp16 benchmark to model.half (#10477 ) * LLM: change fp16 benchmark to model.half * fix	2024-03-20 13:38:39 +08:00
Jin Qiao	e9055c32f9	LLM: fix fp16 mem record in benchmark (#10461 ) * LLM: fix fp16 mem record in benchmark * change style	2024-03-19 16:17:23 +08:00
Jin Qiao	0451103a43	LLM: add int4+fp16 benchmark script for windows benchmarking (#10449 ) * LLM: add fp16 for benchmark script * remove transformer_int4_fp16_loadlowbit_gpu_win	2024-03-19 11:11:25 +08:00
Yuxuan Xia	f36224aac4	Fix ceval run.sh (#10410 )	2024-03-14 10:57:25 +08:00
Wang, Jian4	0193f29411	LLM : Enable gguf float16 and Yuan2 model (#10372 ) * enable float16 * add yun files * enable yun * enable set low_bit on yuan2 * update * update license * update generate * update readme * update python style * update	2024-03-13 10:19:18 +08:00
Xiangyu Tian	0ded0b4b13	LLM: Enable BigDL IPEX optimization for int4 (#10319 ) Enable BigDL IPEX optimization for int4	2024-03-12 17:08:50 +08:00
Lilac09	5809a3f5fe	Add run-hbm.sh & add user guide for spr and hbm (#10357 ) * add run-hbm.sh * add spr and hbm guide * only support quad mode * only support quad mode * update special cases * update special cases	2024-03-12 16:15:27 +08:00
binbin Deng	5d996a5caf	LLM: add benchmark script for deepspeed autotp on gpu (#10380 )	2024-03-12 15:19:57 +08:00
WeiguangHan	17bdb1a60b	LLM: add whisper models into nightly test (#10193 ) * LLM: add whisper models into nightly test * small fix * small fix * add more whisper models * test all cases * test specific cases * collect the csv * store the resut * to html * small fix * small test * test all cases * modify whisper_csv_to_html	2024-03-11 20:00:47 +08:00
Yuxuan Xia	0c8d3c9830	Add C-Eval HTML report (#10294 ) * Add C-Eval HTML report * Fix C-Eval workflow pr trigger path * Fix C-Eval workflow typos * Add permissions to C-Eval workflow * Fix C-Eval workflow typo * Add pandas dependency * Fix C-Eval workflow typo	2024-03-07 16:44:49 +08:00
Shaojun Liu	178eea5009	upload bigdl-llm wheel to sourceforge for backup (#10321 ) * test: upload to sourceforge * update scripts * revert	2024-03-05 16:36:01 +08:00
WeiguangHan	fd81d66047	LLM: Compress some models to save space (#10315 ) * LLM: compress some models to save space * add deleted comments	2024-03-04 17:53:03 +08:00
Yuwen Hu	27d9a14989	[LLM] all-on-one update: memory optimize and streaming output (#10302 ) * Memory saving for continous in-out pair run and add support for streaming output on MTL iGPU * Small fix * Small fix * Add things back	2024-03-01 18:02:30 +08:00
Keyan (Kyrie) Zhang	59861f73e5	Add Deepseek-6.7B (#9991 ) * Add new example Deepseek * Add new example Deepseek * Add new example Deepseek * Add new example Deepseek * Add new example Deepseek * modify deepseek * modify deepseek * Add verified model in README * Turn cpu_embedding=True in Deepseek example --------- Co-authored-by: Shengsheng Huang <shengsheng.huang@intel.com>	2024-02-28 11:36:39 +08:00
hxsz1997	cba61a2909	Add html report of ppl (#10218 ) * remove include and language option, select the corresponding dataset based on the model name in Run * change the nightly test time * change the nightly test time of harness and ppl * save the ppl result to json file * generate csv file and print table result * generate html * modify the way to get parent folder * update html in parent folder * add llm-ppl-summary and llm-ppl-summary-html * modify echo single result * remove download fp16.csv * change model name of PR * move ppl nightly related files to llm/test folder * reformat * seperate make_table from make_table_and_csv.py * separate make_csv from make_table_and_csv.py * update llm-ppl-html * remove comment * add Download fp16.results	2024-02-27 17:37:08 +08:00
Chen, Zhentao	213ef06691	fix readme	2024-02-24 00:38:08 +08:00
Chen, Zhentao	6fe5344fa6	separate make_csv from the file	2024-02-23 16:33:38 +08:00
Chen, Zhentao	bfa98666a6	fall back to make_table.py	2024-02-23 16:33:38 +08:00
Chen, Zhentao	f315c7f93a	Move harness nightly related files to llm/test folder (#10209 ) * move harness nightly files to test folder * change workflow file path accordingly * use arc01 when pr * fix path * fix fp16 csv path	2024-02-23 11:12:36 +08:00
Yuwen Hu	21de2613ce	[LLM] Add model loading time record for all-in-one benchmark (#10201 ) * Add model loading time record in csv for all-in-one benchmark * Small fix * Small fix to number after .	2024-02-22 13:57:18 +08:00
Yuxuan Xia	7cbc2429a6	Fix C-Eval ChatGLM loading issue (#10206 ) * Add c-eval workflow and modify running files * Modify the chatglm evaluator file * Modify the ceval workflow for triggering test * Modify the ceval workflow file * Modify the ceval workflow file * Modify ceval workflow * Adjust the ceval dataset download * Add ceval workflow dependencies * Modify ceval workflow dataset download * Add ceval test dependencies * Add ceval test dependencies * Correct the result print * Fix the nightly test trigger time * Fix ChatGLM loading issue	2024-02-22 10:00:43 +08:00
yb-peng	b1a97b71a9	Harness eval: Add is_last parameter and fix logical operator in highlight_vals (#10192 ) * Add is_last parameter and fix logical operator in highlight_vals * Add script to update HTML files in parent folder * Add running update_html_in_parent_folder.py in summarize step * Add licence info * Remove update_html_in_parent_folder.py in Summarize the results for pull request	2024-02-21 14:45:32 +08:00
Chen, Zhentao	39d37bd042	upgrade harness package version in workflow (#10188 ) * upgrade harness * update readme	2024-02-21 11:21:30 +08:00
Yuwen Hu	001c13243e	[LLM] Add support for `low_low_bit` benchmark on Windows GPU (#10167 ) * Add support for low_low_bit performance test on Windows GPU * Small fix * Small fix * Save memory during converting model process * Drop the results for first time when loading in low bit on mtl igpu for better performance * Small fix	2024-02-21 10:51:52 +08:00
yb-peng	de3dc609ee	Modify harness evaluation workflow (#10174 ) * Modify table head in harness * Specify the file path of fp16.csv * change run to run nightly and run pr to debug * Modify the way to get fp16.csv to downloading from github * Change the method to calculate diff in html table * Change the method to calculate diff in html table * Re-arrange job order * Re-arrange job order * Change limit * Change fp16.csv path * Change highlight rules * Change limit	2024-02-20 18:55:43 +08:00

1 2 3 4 5 ...

262 commits