Skip to main content

Install AWQ

Setup environment (please refer to this link for more details):

Chat with the CLI

Benchmark

  • Through 4-bit weight quantization, AWQ helps to run larger language models within the device memory restriction and prominently accelerates token generation. All benchmarks are done with group_size 128.
  • Benchmark on NVIDIA RTX A6000:
  • NVIDIA RTX 4090:
  • NVIDIA Jetson Orin: