mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-07 02:20:48 -05:00
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> elopment environment, causing every line to show as changed in diffs against master. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> imit bailout: claims a ring slot via atomicAdd (single-GPU host atomics work on RTX 5090), writes fields, fences, sets completion flag, then all threads exit - Watchdog thread simply polls ring head counters every 1ms and prints any new complete records — no CUDA event queries, no mutex, no queue - Zero overhead on the dispatch path (no queue posting, no memset) - Watchdog shutdown returns within ~1ms (atomic bool, no drain) - On bailout the kernel skips Phase 3 entirely and exits cleanly Verified: 20/20 prefill soak test clean at ~1112 t/s, no hangs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> P32, tensors <= 256 KB. Notes in NOTES-allreduce.md. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
8 lines
349 B
Plaintext
8 lines
349 B
Plaintext
# Force LF line endings everywhere (prevent Windows CRLF conversion).
|
|
* text=auto eol=lf
|
|
|
|
# Treat the generated single-file WebUI build as binary for diff purposes.
|
|
# Git's pack-file delta compression still works (byte-level), but this prevents
|
|
# git diff from printing the entire minified file on every change.
|
|
tools/server/public/index.html -diff
|