ipex-llm

Author	SHA1	Message	Date
Cengguang Zhang	51a133de56	LLM: add fuse rope and norm optimization for Baichuan. (#9166 ) * add fuse rope optimization. * add rms norm optimization.	2023-10-13 17:36:52 +08:00
Jin Qiao	db7f938fdc	LLM: add replit and starcoder to gpu pytorch model example (#9154 )	2023-10-13 15:44:17 +08:00
Jin Qiao	797b156a0d	LLM: add dolly-v1 and dolly-v2 to gpu pytorch model example (#9153 )	2023-10-13 15:43:35 +08:00
Yishuo Wang	259cbb4126	[LLM] add initial bigdl-llm-init (#9150 )	2023-10-13 15:31:45 +08:00
Cengguang Zhang	433f408081	LLM: Add fuse rope and norm optimization for Aquila. (#9161 ) * add fuse norm optimization. * add fuse rope optimization	2023-10-13 14:18:37 +08:00
SONG Ge	e7aa67e141	[LLM] Add rope optimization for internlm (#9159 ) * add rope and norm optimization for internlm and gptneox * revert gptneox back and split with pr#9155 # * add norm_forward * style fix * update * update	2023-10-13 14:18:28 +08:00
Jin Qiao	f754ab3e60	LLM: add baichuan and baichuan2 to gpu pytorch model example (#9152 )	2023-10-13 13:44:31 +08:00
Ruonan Wang	b8aee7bb1b	LLM: Fix Qwen kv_cache optimization (#9148 ) * first commit * ut pass * accelerate rotate half by using common util function * fix style	2023-10-12 15:49:42 +08:00
binbin Deng	69942d3826	LLM: fix model check before attention optimization (#9149 )	2023-10-12 15:21:51 +08:00
JIN Qiao	1a1ddc4144	LLM: Add Replit CPU and GPU example (#9028 )	2023-10-12 13:42:14 +08:00
JIN Qiao	d74834ff4c	LLM: add gpu pytorch-models example llama2 and chatglm2 (#9142 )	2023-10-12 13:41:48 +08:00
Ruonan Wang	4f34557224	LLM: support num_beams in all-in-one benchmark (#9141 ) * support num_beams * fix	2023-10-12 13:35:12 +08:00
Ruonan Wang	62ac7ae444	LLM: fix inaccurate input / output tokens of current all-in-one benchmark (#9137 ) * first fix * fix all apis * fix	2023-10-11 17:13:34 +08:00
Lilac09	e02fbb40cc	add bigdl-llm-tutorial into llm-inference-cpu image (#9139 ) * add bigdl-llm-tutorial into llm-inference-cpu image * modify Dockerfile * modify Dockerfile	2023-10-11 16:41:04 +08:00
ZehuaCao	65dd73b62e	Update manually_build.yml (#9138 ) * Update manually_build.yml fix llm-serving-tdx image build dir * Update manually_build.yml	2023-10-11 15:07:09 +08:00
binbin Deng	eb3fb18eb4	LLM: improve PyTorch API doc (#9128 )	2023-10-11 15:03:39 +08:00
Ziteng Zhang	4a0a3c376a	Add stand-alone mode on cpu for finetuning (#9127 ) * Added steps for finetune on CPU in stand-alone mode * Add stand-alone mode to bigdl-lora-finetuing-entrypoint.sh * delete redundant docker commands * Update README.md Turn to intelanalytics/bigdl-llm-finetune-cpu:2.4.0-SNAPSHOT and append example outputs to allow users to check the running * Update bigdl-lora-finetuing-entrypoint.sh Add some tunable parameters * Add parameters --cpus and -e WORKER_COUNT_DOCKER * Modified the cpu number range parameters * Set -ppn to CCL_WORKER_COUNT * Add related configuration suggestions in README.md	2023-10-11 15:01:21 +08:00
binbin Deng	995b0f119f	LLM: update some gpu examples (#9136 )	2023-10-11 14:23:56 +08:00
Ruonan Wang	1c8d5da362	LLM: fix llama tokenizer for all-in-one benchmark (#9129 ) * fix tokenizer for gpu benchmark * fix ipex fp16 * meet code review * fix	2023-10-11 13:39:39 +08:00
binbin Deng	2ad67a18b1	LLM: add mistral examples (#9121 )	2023-10-11 13:38:15 +08:00
Ruonan Wang	1363e666fc	LLM: update benchmark_util.py for beam search (#9126 ) * update reorder_cache * fix	2023-10-11 09:41:53 +08:00
Guoqiong Song	e8c5645067	add LLM example of aquila on GPU (#9056 ) * aquila, dolly-v1, dolly-v2, vacuna	2023-10-10 17:01:35 -07:00
Yuwen Hu	dc70fc7b00	Update performance tests for dependency of bigdl-core-xe-esimd (#9124 )	2023-10-10 19:32:17 +08:00
Lilac09	30e3c196f3	Merge pull request #9108 from Zhengjin-Wang/main Add instruction for chat.py in bigdl-llm-cpu	2023-10-10 16:40:52 +08:00
Lilac09	1e78b0ac40	Optimize LoRA Docker by Shrinking Image Size (#9110 ) * modify dockerfile * modify dockerfile	2023-10-10 15:53:17 +08:00
Ruonan Wang	388f688ef3	LLM: update setup.py to add `bigdl-core-xe` package (#9122 )	2023-10-10 15:02:48 +08:00
Zhao Changmin	1709beba5b	LLM: Explicitly close pickle file pointer before removing temporary directory (#9120 ) * fp close	2023-10-10 14:57:23 +08:00
Yuwen Hu	0e09dd926b	[LLM] Fix example test (#9118 ) * Update llm example test link due to example layout change * Add better change detect	2023-10-10 13:24:18 +08:00
Ruonan Wang	ad7d9231f5	LLM: add benchmark script for Max gpu and ipex fp16 gpu (#9112 ) * add pvc bash * meet code review * rename to run-max-gpu.sh	2023-10-10 10:18:41 +08:00
Lilac09	6264381f2e	Merge pull request #9117 from Zhengjin-Wang/manually_build add llm-serving-xpu on github action	2023-10-10 10:09:06 +08:00
Zhengjin Wang	0dbb3a283e	amend manually_build	2023-10-10 10:03:23 +08:00
Zhengjin Wang	bb3bb46400	add llm-serving-xpu on github action	2023-10-10 09:48:58 +08:00
binbin Deng	e4d1457a70	LLM: improve transformers style API doc (#9113 )	2023-10-10 09:31:00 +08:00
Yuwen Hu	65212451cc	[LLM] Small update to performance tests (#9106 ) * small updates to llm performance tests regarding model handling * Small fix	2023-10-09 16:55:25 +08:00
Zhao Changmin	edccfb2ed3	LLM: Check model device type (#9092 ) * check model device	2023-10-09 15:49:15 +08:00
Heyang Sun	2c0c9fecd0	refine LLM containers (#9109 )	2023-10-09 15:45:30 +08:00
binbin Deng	5e9962b60e	LLM: update example layout (#9046 )	2023-10-09 15:36:39 +08:00
Yina Chen	4c4f8d1663	[LLM]Fix Arc falcon abnormal output issue (#9096 ) * update * update * fix error & style * fix style * update train * to input_seq_size	2023-10-09 15:09:37 +08:00
Wang	a1aefdb8f4	modify README	2023-10-09 13:36:29 +08:00
Wang	3814abf95a	add instruction for chat.py	2023-10-09 12:57:28 +08:00
Zhao Changmin	548e4dd5fe	LLM: Adapt transformers models for `optimize model` SL (#9022 ) * LLM: Adapt transformers model for SL	2023-10-09 11:13:44 +08:00
Ruonan Wang	f64257a093	LLM: basic api support for esimd fp16 (#9067 ) * basic api support for fp16 * fix style * fix * fix error and style * fix style * meet code review * update based on comments	2023-10-09 11:05:17 +08:00
Wang	a42c25436e	Merge remote-tracking branch 'upstream/main'	2023-10-09 10:55:18 +08:00
JIN Qiao	65373d2a8b	LLM: adjust portable zip content (#9054 ) * LLM: adjust portable zip content * LLM: adjust portable zip README	2023-10-09 10:51:19 +08:00
Guancheng Fu	df8df751c4	Modify readme for bigdl-llm-serving-cpu (#9105 )	2023-10-09 09:56:09 +08:00
Heyang Sun	2756f9c20d	XPU QLoRA Container (#9082 ) * XPU QLoRA Container * fix apt issue * refine	2023-10-08 11:04:20 +08:00
ZehuaCao	aad68100ae	Add trusted-bigdl-llm-serving-tdx image. (#9093 ) * add entrypoint in cpu serving * kubernetes support for fastchat cpu serving * Update Readme * add image to manually_build action * update manually_build.yml * update README.md * update manually_build.yaml * update attestation_cli.py * update manually_build.yml * update Dockerfile * rename * update trusted-bigdl-llm-serving-tdx Dockerfile	2023-10-08 10:13:51 +08:00
Xin Qiu	b3e94a32d4	change log4error import (#9098 )	2023-10-08 09:23:28 +08:00
Kai Huang	78ea7ddb1c	Combine apply_rotary_pos_emb for gpt-neox (#9074 )	2023-10-07 16:27:46 +08:00
Heyang Sun	0b40ef8261	separate trusted and native llm cpu finetune from lora (#9050 ) * seperate trusted-llm and bigdl from lora finetuning * add k8s for trusted llm finetune * refine * refine * rename cpu to tdx in trusted llm * solve conflict * fix typo * resolving conflict * Delete docker/llm/finetune/lora/README.md * fix --------- Co-authored-by: Uxito-Ada <seusunheyang@foxmail.com> Co-authored-by: leonardozcm <leonardo1997zcm@gmail.com>	2023-10-07 15:26:59 +08:00

1 2 3 4 5 ...

1516 commits