Files
llama.cpp/docs/models.md
Georgi Gerganov 6b36c23056 readme : refresh (#26280)
* docs : center badges and links, remove Hot topics

- Use <div align="center"> for GitHub-compatible centering
- Add dev branches and compile times links
- Remove Hot topics section

Assisted-by: llama.cpp:Qwen3.6-27B

* readme : remove sections

* docs : center badges, remove Hot topics, extract sections, remove tools

- Use <div align="center"> for GitHub-compatible centering
- Add dev branches and compile times links
- Add lib llama API and llama-server REST API links
- Remove Hot topics section
- Remove Recent API changes section
- Extract XCFramework section into docs/xcframework.md
- Extract Completions section into docs/completions.md
- Extract Obtaining and quantizing models into docs/models.md
- Remove tools usage sections (llama-cli, llama-server, etc.)
- Move Contributing section to the end

Assisted-by: llama.cpp:Qwen3.6-27B

* cont : arrange links

* cont : fix ws

* cont : remove seminal papers

* cont : change sample model

* cont : trim-down contributing section

* cont : sort backends alphabetically

* cont : words

* cont : add fig captions

* docs : models words

* readme : shorter caption

* cont : fix typo

* cont : add window frame to screenshot
2026-07-30 16:14:37 +03:00

1.9 KiB

Obtaining and quantizing models

The Hugging Face platform hosts thousands of models compatible with llama.cpp:

You can use any llama.cpp-compatible model from Hugging Face using this CLI argument: -hf <user>/<model>[:quant]. For example:

llama cli -hf ggml-org/gemma-3-1b-it-GGUF

You can use the same CLI invocation to download from other sites, by pointing the MODEL_ENDPOINT environment variable to an endpoint compatible with the Hugging Face API. llama.cpp can also run models you have downloaded locally to your filesystem.

After downloading a model, use the CLI tools to run it locally - see below.

llama.cpp requires the model to be stored in the GGUF file format. Models in other data formats can be converted to GGUF using the convert_*.py Python scripts in this repo. To learn more about model quantization, read this documentation

The Hugging Face platform provides a variety of online tools for converting, quantizing and hosting models with llama.cpp: