* docs: add AGENTS.md with verified Tesla T4 (16GB) inference notes
Documents two things found while verifying inference on a real T4:
1. Cold-cache first /v1/audio/speech call can hit the 300s
OMNIVOICE_GENERATE_TIMEOUT_S because the checkpoint download happens
inside that budget — workaround via existing POST /models/install or
raising the timeout, no code change needed.
2. The OpenAI-compatible endpoint silently ignores num_step/guidance_scale
(schema doesn't declare them) — use native /generate for those.
Also documents the T4 acceleration checklist (dtype/attention/int8/CUDA
graphs) and measured VRAM (peak 2.05GB). No code changes.
* fix(docs): make /models/install workaround command actually executable
Addresses Greptile review: the instruction omitted the required
repo_id body field (InstallModelRequest rejects an empty body).
* fix(docs): correct port in /models/install example (3900, not 8000)
The app serves on port 3900 (confirmed: /health returns 200 there,
connection refused on 8000). Verified the exact corrected curl command
returns 200 {"status":"install_started",...}.
* move T4 notes to docs/hardware-notes-tesla-t4.md — AGENTS.md is the auto-loaded agent-instructions filename
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>