Skip to main content
  1. Window user: use the old-cuda branch.
  2. Linux user: recommend the fastest-inference-4bit branch.

Install

Setup environment:
Chat with the CLI:
Start model worker:

Benchmark