mirror of
https://github.com/ollama/ollama.git
synced 2026-07-23 09:10:53 -05:00
Add a laguna-v8 renderer/parser matching the Laguna XS 2.1 template, and fix v2 handling of embedded thinking and structured tool arguments. Prevent FP16 overflow in Metal's quantized routed-MoE prefill path by scaling the linear branch and folding the inverse into the routing scale. Other backends and token-generation paths are unchanged. Add comprehensive v2/v8 Jinja parity and parser tests.
17 KiB
17 KiB