Install AWQ
Setup environment (please refer to this link for more details):Chat with the CLI
Benchmark
- Through 4-bit weight quantization, AWQ helps to run larger language models within the device memory restriction and prominently accelerates token generation. All benchmarks are done with group_size 128.
-
Benchmark on NVIDIA RTX A6000:
-
NVIDIA RTX 4090:
-
NVIDIA Jetson Orin: