Compare commits
572
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
38c2405fa1 | ||
|
|
97b153df3a | ||
|
|
4e37e21745 | ||
|
|
3bfdeeddba | ||
|
|
ac35351000 | ||
|
|
764fed09bc | ||
|
|
4fdde15079 | ||
|
|
14a8a9cd46 | ||
|
|
1af1ecc5bd | ||
|
|
e0879027d5 | ||
|
|
5298c8e5e1 | ||
|
|
1de4fedad7 | ||
|
|
c5f32d020a | ||
|
|
f5ea363724 | ||
|
|
ae68dd5bfd | ||
|
|
2fd4cb3caf | ||
|
|
e532c9d8dd | ||
|
|
5516750a4f | ||
|
|
c079b721ed | ||
|
|
664c9ea4b5 | ||
|
|
faa3d39836 | ||
|
|
b0cae840c4 | ||
|
|
566aad7bba | ||
|
|
e9965b2cb4 | ||
|
|
52887ae44d | ||
|
|
b0f1558db0 | ||
|
|
dea884a22d | ||
|
|
bee9fa25ff | ||
|
|
06c15ce37f | ||
|
|
bb9149ce1f | ||
|
|
4c10e802ab | ||
|
|
c021fac1d8 | ||
|
|
24e0737e0d | ||
|
|
06f7c4af7b | ||
|
|
c9a587ea75 | ||
|
|
cef2b73e02 | ||
|
|
50b1184df3 | ||
|
|
c48e8c4ff5 | ||
|
|
0c9f1bd7c4 | ||
|
|
b6a0f96e90 | ||
|
|
93771a0201 | ||
|
|
a4f87bad6e | ||
|
|
703c665c6c | ||
|
|
f301db1b02 | ||
|
|
c90d32845a | ||
|
|
9a7f5fd7ed | ||
|
|
25393b4654 | ||
|
|
e11270990b | ||
|
|
f21ad9a27d | ||
|
|
fc45ec3b5b | ||
|
|
1eda03827d | ||
|
|
555210bbdb | ||
|
|
78d6dbe802 | ||
|
|
e3efe4a39a | ||
|
|
36a9515320 | ||
|
|
9618f26f6e | ||
|
|
9e12c7b181 | ||
|
|
baaef34a4a | ||
|
|
5b842280c8 | ||
|
|
9d71c7cedd | ||
|
|
770fe35da9 | ||
|
|
db48c7851c | ||
|
|
512c2fa7d6 | ||
|
|
55e6d9a07b | ||
|
|
192c1f5127 | ||
|
|
7610a0d86e | ||
|
|
2afea02c94 | ||
|
|
46b1fb2f57 | ||
|
|
ac427fb821 | ||
|
|
924172e012 | ||
|
|
14d6b90836 | ||
|
|
2cde7663a1 | ||
|
|
2bf922d9e5 | ||
|
|
d52a4c89a7 | ||
|
|
6be80a4b3d | ||
|
|
f67e5da289 | ||
|
|
b945ddd163 | ||
|
|
525a2809eb | ||
|
|
1b7fcf92c3 | ||
|
|
78a6c489c9 | ||
|
|
5e4f863c4d | ||
|
|
12794f3dae | ||
|
|
0ce5cca1b5 | ||
|
|
744a0550d3 | ||
|
|
56e9e42ac2 | ||
|
|
c2f27c39d2 | ||
|
|
60b696f37f | ||
|
|
5237f7849a | ||
|
|
a5f4d61eda | ||
|
|
c7d0e314d9 | ||
|
|
c4f214858b | ||
|
|
7aa2333823 | ||
|
|
fe012ad3a5 | ||
|
|
1b7e858886 | ||
|
|
fd74865f7c | ||
|
|
0f237a9b25 | ||
|
|
ae25a6a594 | ||
|
|
aa984075a7 | ||
|
|
80e0b91ce8 | ||
|
|
d1e4c847cb | ||
|
|
91515310f0 | ||
|
|
6087085d43 | ||
|
|
2f66e95958 | ||
|
|
f905230fd8 | ||
|
|
b6ed598770 | ||
|
|
ac09da3c24 | ||
|
|
ef8636ee78 | ||
|
|
72249a8d8e | ||
|
|
8e2928ca56 | ||
|
|
de962d164a | ||
|
|
3d76f1ca26 | ||
|
|
6203fe41e9 | ||
|
|
8e87755845 | ||
|
|
3013cdc241 | ||
|
|
939ed4b12d | ||
|
|
9e768c9552 | ||
|
|
18d2d1e076 | ||
|
|
8d0b826c92 | ||
|
|
939bdc7248 | ||
|
|
71d4114354 | ||
|
|
5f7df6ac52 | ||
|
|
12216ecc9a | ||
|
|
1339fa6814 | ||
|
|
4b5b19ba47 | ||
|
|
7d1f44f8bd | ||
|
|
647ddc842b | ||
|
|
638392994a | ||
|
|
dba3074376 | ||
|
|
d0422b8913 | ||
|
|
a31b50c5cb | ||
|
|
cb4aaa1580 | ||
|
|
d0ab15c376 | ||
|
|
212a1afb69 | ||
|
|
58a4674316 | ||
|
|
bcc47882e9 | ||
|
|
355f7465c8 | ||
|
|
9426006d3c | ||
|
|
f82ec3f579 | ||
|
|
b4fc4cb53f | ||
|
|
c6d7a19f42 | ||
|
|
c902a42f6b | ||
|
|
7a24b459bc | ||
|
|
0862b3efed | ||
|
|
ac63eff33d | ||
|
|
f92e33f8e3 | ||
|
|
70aecf7ad0 | ||
|
|
164534016a | ||
|
|
7b7b491ba0 | ||
|
|
0a019b5d7c | ||
|
|
4da62e664d | ||
|
|
508943ae71 | ||
|
|
42b4bf856f | ||
|
|
c119afb023 | ||
|
|
0ba91e6d82 | ||
|
|
28d9d2fbe8 | ||
|
|
5069872752 | ||
|
|
030436a549 | ||
|
|
b9324b84b3 | ||
|
|
353cc40fdc | ||
|
|
a4cb4afe37 | ||
|
|
075ab7b68b | ||
|
|
e5a0a553e9 | ||
|
|
b7eb5428c7 | ||
|
|
bb4e99a336 | ||
|
|
accef57865 | ||
|
|
d07fa1ffff | ||
|
|
d313045b13 | ||
|
|
6a44733297 | ||
|
|
fd7d2d94dd | ||
|
|
f82f35ee97 | ||
|
|
1c37c479be | ||
|
|
b55098f044 | ||
|
|
854609e459 | ||
|
|
7d46563ea8 | ||
|
|
b6143cce08 | ||
|
|
512aad5c48 | ||
|
|
8526ff56f7 | ||
|
|
a68f85c51b | ||
|
|
df7dab43b6 | ||
|
|
f48a33edcf | ||
|
|
791a46d9dd | ||
|
|
61eb27dbba | ||
|
|
19d371f9d6 | ||
|
|
c760347339 | ||
|
|
ee8a8fb928 | ||
|
|
d7177be66c | ||
|
|
88e59326fd | ||
|
|
f02880cc5c | ||
|
|
f9ff185a22 | ||
|
|
cc96de27b8 | ||
|
|
e8af4fae12 | ||
|
|
e0309be21f | ||
|
|
2e4465a3f4 | ||
|
|
9ffc9d52d7 | ||
|
|
cea6678ea7 | ||
|
|
40199e39cf | ||
|
|
c21696b9ae | ||
|
|
67817b89bf | ||
|
|
20bc24c312 | ||
|
|
810abf9187 | ||
|
|
cfb318c50b | ||
|
|
2e1e52f26e | ||
|
|
2f663011e5 | ||
|
|
773c4728c8 | ||
|
|
6b68c4d9fe | ||
|
|
5f5bb04f41 | ||
|
|
cea00a4889 | ||
|
|
6268eacf35 | ||
|
|
a2653824ce | ||
|
|
71056bcdb4 | ||
|
|
b5441bea36 | ||
|
|
d7751ad7ce | ||
|
|
2b027cb54a | ||
|
|
d2fb292217 | ||
|
|
7afb82bb43 | ||
|
|
d630db29f5 | ||
|
|
f49d31d5b3 | ||
|
|
4f4d21ff5b | ||
|
|
b82205fbab | ||
|
|
7601f0f1a8 | ||
|
|
5c7f8a8e7a | ||
|
|
7dcb88452c | ||
|
|
db50111c54 | ||
|
|
6b7c2bac8f | ||
|
|
48d1b22fd9 | ||
|
|
6a6dd83efa | ||
|
|
09a260bb26 | ||
|
|
a754faec04 | ||
|
|
8fc003ad7c | ||
|
|
99a534882c | ||
|
|
81f889f3ff | ||
|
|
693a43b26c | ||
|
|
d848c2eb3b | ||
|
|
d5d28ddabf | ||
|
|
3622102ace | ||
|
|
31d7b4d251 | ||
|
|
5e3af17d34 | ||
|
|
67ef5e22eb | ||
|
|
a84dc9e8ef | ||
|
|
b15fa1209a | ||
|
|
8a2c9a4828 | ||
|
|
adebe98e0f | ||
|
|
97e0e70bf2 | ||
|
|
66f7ef8cfe | ||
|
|
d01fb5e7cc | ||
|
|
bc22276fec | ||
|
|
db9e9d7fa1 | ||
|
|
e098d1280c | ||
|
|
2127aa7716 | ||
|
|
0d3fb07c1f | ||
|
|
3ad31feccf | ||
|
|
18f56a7940 | ||
|
|
870b6f93ba | ||
|
|
4c0bf2e010 | ||
|
|
7d6f5c5b0c | ||
|
|
4f4dc7e068 | ||
|
|
56a420a898 | ||
|
|
d5c0fa6f75 | ||
|
|
fe53e02693 | ||
|
|
56f8b616cc | ||
|
|
f1016cdedd | ||
|
|
68194c004d | ||
|
|
23dd8728d1 | ||
|
|
22ab31e0c7 | ||
|
|
fc237d3929 | ||
|
|
260519d4ee | ||
|
|
9b1f2149ec | ||
|
|
311ff43de0 | ||
|
|
8f140e550d | ||
|
|
25c41e630c | ||
|
|
ffc5c07949 | ||
|
|
503407272e | ||
|
|
6b2683eb32 | ||
|
|
4999bac7a8 | ||
|
|
132fdf313e | ||
|
|
9bfa3ec751 | ||
|
|
c637166f02 | ||
|
|
658269d42f | ||
|
|
ae85dd36e4 | ||
|
|
17fda0d79f | ||
|
|
c6fbedbecb | ||
|
|
e2b5b3d6ed | ||
|
|
15c6cc698d | ||
|
|
3beebc6d57 | ||
|
|
a9d90936ce | ||
|
|
4b601ee6c1 | ||
|
|
03c1396b14 | ||
|
|
276bf626da | ||
|
|
09a582c65a | ||
|
|
adfe7b8588 | ||
|
|
d3eb502395 | ||
|
|
a88ad00bdf | ||
|
|
ceef6d66bd | ||
|
|
035add4c5c | ||
|
|
31abce74b6 | ||
|
|
7668d045fa | ||
|
|
7157fd0fbe | ||
|
|
fb98bbfaf6 | ||
|
|
a935c08d6a | ||
|
|
8d60929236 | ||
|
|
6e651ecbbd | ||
|
|
4a53fd2668 | ||
|
|
955c62e671 | ||
|
|
89c5d46660 | ||
|
|
70091bf97c | ||
|
|
0189b2a84f | ||
|
|
0d7625ee1a | ||
|
|
5d99271865 | ||
|
|
b5879343a0 | ||
|
|
5e9538a049 | ||
|
|
ece08bd790 | ||
|
|
e3e8e80888 | ||
|
|
4ad5aa2831 | ||
|
|
1c344bf9ec | ||
|
|
37aea507a4 | ||
|
|
9979f1a84e | ||
|
|
af216d8411 | ||
|
|
65ea0dc376 | ||
|
|
3105647afb | ||
|
|
50be1944b4 | ||
|
|
e014d63253 | ||
|
|
55bcfbdb6a | ||
|
|
12b398b662 | ||
|
|
9790d28922 | ||
|
|
7736e9cc60 | ||
|
|
fdeb8a40fc | ||
|
|
ebbe74b14b | ||
|
|
6a733a4c52 | ||
|
|
5cde34c113 | ||
|
|
bd84169ff2 | ||
|
|
4486b479a1 | ||
|
|
74aa444485 | ||
|
|
e1e96ae5a2 | ||
|
|
0d39a5b283 | ||
|
|
33ce88e465 | ||
|
|
10a0fc27f0 | ||
|
|
63ade08b6f | ||
|
|
fc7021e2b4 | ||
|
|
1038a487a8 | ||
|
|
19b9b4a136 | ||
|
|
d9bf251665 | ||
|
|
5bc9f7856a | ||
|
|
8428517d95 | ||
|
|
6caabf3b0b | ||
|
|
6b802beaf9 | ||
|
|
e059c4f330 | ||
|
|
8998faa98b | ||
|
|
2ed38c476e | ||
|
|
078bc06428 | ||
|
|
99f93164be | ||
|
|
af245a45ee | ||
|
|
dbbad3d58e | ||
|
|
8667b34ca5 | ||
|
|
bcd2e53175 | ||
|
|
f18a7adddf | ||
|
|
a968bc57b3 | ||
|
|
7ac5e5ab3d | ||
|
|
f2cfdaa430 | ||
|
|
a768e84ffe | ||
|
|
4648c23256 | ||
|
|
1e96748402 | ||
|
|
169e36d5b6 | ||
|
|
135dd0a3a7 | ||
|
|
ef51c6b75d | ||
|
|
24252712d2 | ||
|
|
b49d2954d5 | ||
|
|
bf8da941d1 | ||
|
|
ca3bf8367f | ||
|
|
0c255dfabd | ||
|
|
4008499ea5 | ||
|
|
e32bc94248 | ||
|
|
fc8db259a9 | ||
|
|
0921292319 | ||
|
|
9afc9db287 | ||
|
|
3b64692eed | ||
|
|
30307ae883 | ||
|
|
9dbf45ca2c | ||
|
|
0e7ee3d912 | ||
|
|
140e9262b9 | ||
|
|
18a0c11c23 | ||
|
|
a0b84a6b61 | ||
|
|
8c67b10942 | ||
|
|
760e9e5ac7 | ||
|
|
25498dd17b | ||
|
|
16020c2178 | ||
|
|
e4471008b4 | ||
|
|
2823d6a8ea | ||
|
|
57150ede64 | ||
|
|
df46b71be2 | ||
|
|
a41e66e0f0 | ||
|
|
2b2bfe1a7f | ||
|
|
f728071819 | ||
|
|
59a487e742 | ||
|
|
ef9bda2e17 | ||
|
|
e8a6950898 | ||
|
|
43d47f4d36 | ||
|
|
93bbfd64ea | ||
|
|
c973b1c983 | ||
|
|
c4530b4919 | ||
|
|
b4b10c1431 | ||
|
|
99218e2c34 | ||
|
|
b759ae7749 | ||
|
|
b2f5de328e | ||
|
|
f2dd6c95db | ||
|
|
ee3e147401 | ||
|
|
af7f1ceb3d | ||
|
|
0cc22ad803 | ||
|
|
62b0e05c1d | ||
|
|
aed50121b4 | ||
|
|
079be2fbe5 | ||
|
|
f463276ca5 | ||
|
|
5a1ada0757 | ||
|
|
6396cc2b3f | ||
|
|
997ba9e673 | ||
|
|
a438f5db0e | ||
|
|
f72ab3a3b2 | ||
|
|
1d30a50c1c | ||
|
|
c36df0ddf0 | ||
|
|
85cb7242ea | ||
|
|
c547da351c | ||
|
|
602be02a17 | ||
|
|
f395c90aa1 | ||
|
|
01cb810c65 | ||
|
|
f4e5d7b9ab | ||
|
|
caedcde0c6 | ||
|
|
2a36a24612 | ||
|
|
610303ec11 | ||
|
|
3a114d62dd | ||
|
|
65e6c275c6 | ||
|
|
9a90fe41be | ||
|
|
d55838542e | ||
|
|
9c0a5469cc | ||
|
|
c7a8d24af6 | ||
|
|
7d963cf297 | ||
|
|
b9cb6c204f | ||
|
|
6157c1717b | ||
|
|
c3b54393c4 | ||
|
|
f0bb3fb8b2 | ||
|
|
157f2987dc | ||
|
|
426bd1ecea | ||
|
|
d34490fd15 | ||
|
|
9b9d4d6579 | ||
|
|
eee35ab771 | ||
|
|
0095c4a089 | ||
|
|
ef30dbece7 | ||
|
|
a9062c2fa1 | ||
|
|
b7ebd6e54d | ||
|
|
cd58bbded0 | ||
|
|
b9183bb2a6 | ||
|
|
e9495beb8b | ||
|
|
409dd016ce | ||
|
|
db10e4b70f | ||
|
|
3613b3d39b | ||
|
|
4277859a44 | ||
|
|
eefc45ec49 | ||
|
|
2f5a52f9eb | ||
|
|
729cad8912 | ||
|
|
86f9eda65f | ||
|
|
32c96eeee4 | ||
|
|
f5d52e479e | ||
|
|
eafd94d0a6 | ||
|
|
2b5e3cdbc6 | ||
|
|
2229a68cfc | ||
|
|
0581fb69fd | ||
|
|
85a368a08b | ||
|
|
a8c57c816b | ||
|
|
360d1d29c8 | ||
|
|
7366555c95 | ||
|
|
b8ff2721a5 | ||
|
|
da26a240e2 | ||
|
|
fb3d9b7139 | ||
|
|
6c47ae7b21 | ||
|
|
fda73384f0 | ||
|
|
84add8b279 | ||
|
|
8b510d73db | ||
|
|
b1194c6223 | ||
|
|
75f8924222 | ||
|
|
2a6d089f4a | ||
|
|
5d134f22ec | ||
|
|
08af0a971e | ||
|
|
ff0ce4d37d | ||
|
|
1ab03067cf | ||
|
|
7a2f86066b | ||
|
|
d91beef0fd | ||
|
|
574b634688 | ||
|
|
62ab62fc57 | ||
|
|
ecbd152c43 | ||
|
|
fec2e7b5f3 | ||
|
|
04bc9cedac | ||
|
|
e397e10d64 | ||
|
|
8e1da3c22f | ||
|
|
b92cc3bb79 | ||
|
|
d757f160fa | ||
|
|
a9ac8b1b46 | ||
|
|
f88bfecabc | ||
|
|
753891adc9 | ||
|
|
4efde9ce4e | ||
|
|
ae2fa83b02 | ||
|
|
a671004cb7 | ||
|
|
ed6fca46a9 | ||
|
|
65d37288ce | ||
|
|
293b0812f4 | ||
|
|
b441a486cc | ||
|
|
d4bf1fe9a9 | ||
|
|
5574cbbc16 | ||
|
|
13eae6ff02 | ||
|
|
fd5e78cdab | ||
|
|
67c316e81b | ||
|
|
81fb585557 | ||
|
|
3675d750d6 | ||
|
|
bbbf185cf3 | ||
|
|
0c4ed0546f | ||
|
|
17df9210cd | ||
|
|
aedca15f7e | ||
|
|
7a14c31a8e | ||
|
|
53ff367c1f | ||
|
|
ed746bab57 | ||
|
|
847ca6d44c | ||
|
|
784494dc3c | ||
|
|
6c42b50fe3 | ||
|
|
2594abebf8 | ||
|
|
e4f00bc564 | ||
|
|
6a2ada2f11 | ||
|
|
9ac1045917 | ||
|
|
31d90def83 | ||
|
|
d5c33eda27 | ||
|
|
36a38bdbb9 | ||
|
|
02c3a651ac | ||
|
|
fa148ecb55 | ||
|
|
f172d0c3be | ||
|
|
fcac8e1bae | ||
|
|
4e6c36848c | ||
|
|
f2302e8c95 | ||
|
|
8b4e4ebf56 | ||
|
|
7f7a4c5f83 | ||
|
|
e5916acc01 | ||
|
|
06e69d6d6b | ||
|
|
e4ce065864 | ||
|
|
f95e73b710 | ||
|
|
b20cd6cb95 | ||
|
|
1515b46adb | ||
|
|
999345de41 | ||
|
|
ac287c612f | ||
|
|
d80286562e | ||
|
|
9d30774ee2 | ||
|
|
65236774aa | ||
|
|
0ed7d2ec22 | ||
|
|
a95041f1e6 | ||
|
|
4053397921 | ||
|
|
e2446c3e61 | ||
|
|
ccd984f324 | ||
|
|
89f0a25082 | ||
|
|
b351dc5b1a | ||
|
|
026410fbc6 | ||
|
|
b6a1ee50ff | ||
|
|
3c3c37615c | ||
|
|
c44f2fc042 | ||
|
|
62604ec9bb | ||
|
|
08569397d3 | ||
|
|
c9a468e500 | ||
|
|
aec4694628 | ||
|
|
5df1dc973c | ||
|
|
8cc978588d | ||
|
|
6f80110a42 | ||
|
|
c5ac548100 | ||
|
|
26d7a333e8 | ||
|
|
75eb7c6099 | ||
|
|
1d06c6f079 | ||
|
|
267ded0e79 | ||
|
|
549a56fc2d | ||
|
|
6e47b15e82 | ||
|
|
5e08dde5e0 |
@@ -85,14 +85,20 @@ names: there is no `desktop=prod` (note the **hyphen** in `desktop-prod`).
|
||||
|
||||
Requires [Rust](https://rustup.rs/) and platform-specific Tauri dependencies — see the [Tauri prerequisites](https://v2.tauri.app/start/prerequisites/).
|
||||
|
||||
After installing Rust with rustup on macOS/Linux, either open a new terminal or
|
||||
load Cargo into the current one before starting the desktop app:
|
||||
After installing Rust with rustup (or `uv` with its installer), a terminal that
|
||||
was already open still has the old `PATH`. The desktop launchers (`bun desktop`,
|
||||
`bun desktop-prod`, `bun desktop-fresh`) detect this and add `~/.cargo/bin` /
|
||||
`~/.local/bin` for that run, printing a one-line note; to make it permanent,
|
||||
open a new terminal, or on macOS/Linux load Cargo into the current one:
|
||||
|
||||
```bash
|
||||
source "$HOME/.cargo/env"
|
||||
bun desktop
|
||||
```
|
||||
|
||||
If Rust is genuinely not installed, the launchers stop up front with the
|
||||
install command instead of failing later inside `cargo metadata`.
|
||||
|
||||
On Linux, errors such as `Package gdk-3.0 was not found`, `pango.pc` missing,
|
||||
or `javascriptcoregtk-4.1` missing mean the native packages above were not
|
||||
installed; changing `PKG_CONFIG_PATH` does not fix libraries that are absent.
|
||||
|
||||
@@ -11,6 +11,11 @@ on:
|
||||
push:
|
||||
branches: [main]
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
windows_wix_diagnostic:
|
||||
description: Run only the tiny nonpublishing Windows MSI authoring diagnostic
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
@@ -21,6 +26,7 @@ env:
|
||||
|
||||
jobs:
|
||||
test:
|
||||
if: ${{ !inputs.windows_wix_diagnostic }}
|
||||
name: Tests (backend + frontend)
|
||||
runs-on: ubuntu-22.04
|
||||
env:
|
||||
@@ -189,6 +195,7 @@ jobs:
|
||||
# and `cargo test --lib` runs the shell's unit tests natively on each OS.
|
||||
# Full bundling stays in release.yml on tag push.
|
||||
tauri-cross-platform:
|
||||
if: ${{ !inputs.windows_wix_diagnostic }}
|
||||
name: Tauri shell check (${{ matrix.label }})
|
||||
needs: test
|
||||
strategy:
|
||||
@@ -289,6 +296,7 @@ jobs:
|
||||
# job above misses. Narrow scope (tests/smoke/ only) — full pytest stays
|
||||
# on Linux until Phase 1's INST-01 lands setuptools for WhisperX.
|
||||
smoke-matrix:
|
||||
if: ${{ !inputs.windows_wix_diagnostic }}
|
||||
name: Smoke (${{ matrix.label }})
|
||||
needs: test
|
||||
strategy:
|
||||
@@ -428,17 +436,68 @@ jobs:
|
||||
PY
|
||||
|
||||
- name: Run smoke tests
|
||||
# Exercise credential paths on native Windows as well as POSIX hosts.
|
||||
if: matrix.backend_supported
|
||||
run: uv run --no-sync pytest tests/smoke/ -q --tb=short
|
||||
run: uv run --no-sync pytest tests/smoke/ tests/test_hf_token_cache_paths.py -q --tb=short
|
||||
env:
|
||||
HF_HUB_OFFLINE: "1" # same no-silent-downloads guard as the main pytest job
|
||||
HF_HUB_CACHE: ${{ runner.temp }}/pockettts-empty-hf-cache
|
||||
|
||||
# The isolated backend session, on Windows. The `test` job runs it on
|
||||
# Linux only, which is how four tests that CANNOT pass on Windows shipped
|
||||
# unnoticed: two reach for os.WNOHANG and os.waitid (POSIX-only, an
|
||||
# AttributeError before the first assertion), one asserts a RuntimeError
|
||||
# that `backend_drain_fd` returns None instead of raising off POSIX, and
|
||||
# one raced the OS reaping a crashed child — a race Linux won and Windows
|
||||
# lost every time. All four were invisible to CI and hit every Windows
|
||||
# contributor on their first `pytest` run. Forty seconds closes the class.
|
||||
- name: Isolated backend session (Windows)
|
||||
if: runner.os == 'Windows' && matrix.backend_supported
|
||||
run: uv run --no-sync pytest backend/tests/ -q --tb=short
|
||||
env:
|
||||
HF_HUB_OFFLINE: "1"
|
||||
|
||||
# Artifact commits depend on native Windows rename/replace semantics;
|
||||
# Linux emulation cannot exercise sharing rules or path parsing.
|
||||
# test_worker_task_store and test_worker_inbound_transport joined this
|
||||
# step after a Windows run found a real portability bug the Linux-only
|
||||
# `test` job could not see: a staged input's artifact id was built with
|
||||
# os.path.join, so a Windows control plane persisted and shipped
|
||||
# `inputs\<sha>.wav` — which a Linux worker cannot resolve. These suites
|
||||
# need no ffmpeg, so they cost seconds here.
|
||||
- name: Remote-worker artifact paths (Windows)
|
||||
if: runner.os == 'Windows' && matrix.backend_supported
|
||||
run: uv run --no-sync pytest tests/test_worker_upload_server.py tests/test_worker_server_integrity.py -q --tb=short
|
||||
run: >-
|
||||
uv run --no-sync pytest
|
||||
tests/test_worker_upload_server.py
|
||||
tests/test_worker_server_integrity.py
|
||||
tests/test_worker_task_store.py
|
||||
tests/test_worker_inbound_transport.py
|
||||
-q --tb=short
|
||||
env:
|
||||
HF_HUB_OFFLINE: "1"
|
||||
HF_HUB_CACHE: ${{ runner.temp }}/worker-artifact-empty-hf-cache
|
||||
|
||||
windows-wix-diagnostic:
|
||||
name: Windows MSI authoring (no publishing)
|
||||
needs: test
|
||||
if: ${{ !cancelled() && (inputs.windows_wix_diagnostic || needs.test.result == 'success') }}
|
||||
runs-on: windows-2022
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: oven-sh/setup-bun@v1
|
||||
- name: Bundle canonical system and per-user templates with a tiny payload
|
||||
shell: pwsh
|
||||
run: ./scripts/diagnose-windows-wix.ps1
|
||||
- name: Restore hosted Installer policy after failed standard-user installation
|
||||
shell: powershell
|
||||
run: ./scripts/test-msi-policy-cleanup.ps1
|
||||
- name: Preserve verbose linker output and rendered authoring
|
||||
if: always()
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: windows-wix-diagnostic
|
||||
path: wix-diagnostic-artifacts/
|
||||
if-no-files-found: warn
|
||||
retention-days: 3
|
||||
|
||||
@@ -148,15 +148,20 @@ jobs:
|
||||
preview-gate:
|
||||
name: Preview gate
|
||||
runs-on: ubuntu-22.04
|
||||
permissions:
|
||||
contents: read
|
||||
outputs:
|
||||
is_preview: ${{ steps.decide.outputs.is_preview }}
|
||||
proceed: ${{ steps.decide.outputs.proceed }}
|
||||
stable_tag: ${{ steps.decide.outputs.stable_tag }}
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
with:
|
||||
fetch-depth: 50
|
||||
- id: decide
|
||||
shell: bash
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
event="${{ github.event_name }}"
|
||||
@@ -171,6 +176,13 @@ jobs:
|
||||
exit 1
|
||||
fi
|
||||
echo "is_preview=true" >> "$GITHUB_OUTPUT"
|
||||
# Resolve once before the matrix starts so every platform stamps
|
||||
# against the same immutable Stable-channel snapshot.
|
||||
STABLE_TAG=$(gh release view --repo "$GITHUB_REPOSITORY" --json tagName --jq .tagName)
|
||||
[[ "$STABLE_TAG" =~ ^v[0-9]+\.[0-9]+\.[0-9]+$ ]] || {
|
||||
echo "::error::latest stable release has an invalid tag"; exit 1;
|
||||
}
|
||||
echo "stable_tag=$STABLE_TAG" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "is_preview=false" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
@@ -497,35 +509,22 @@ jobs:
|
||||
echo "APPLE_TEAM_ID=$TID"
|
||||
} >> "$GITHUB_ENV"
|
||||
|
||||
# Stamp each preview build with a unique, monotonically increasing semver
|
||||
# PRERELEASE so the updater actually offers it (a rolling preview that
|
||||
# always reported the static 0.3.0 never looked "newer", so no update was
|
||||
# ever delivered). Ephemeral, CI-only — never committed. Tauri reads the
|
||||
# bundle + updater version from tauri.conf.json, so rewriting it here
|
||||
# stamps the artifacts + latest.json. Under the versioning hard rule
|
||||
# (owner-set 2026-06-11) main is always last-release + 1, so BASE-N is a
|
||||
# prerelease of the NEXT version and semver-sorts ABOVE the last stable
|
||||
# (0.3.6-N > 0.3.5) — preview users naturally upgrade past stable, and
|
||||
# the Windows MSI ProductVersion (which strips the prerelease → 0.3.6)
|
||||
# is also correctly above the last stable.
|
||||
# Stamp each preview with a numeric prerelease that is strictly above the
|
||||
# latest stable release. Main may intentionally retain the released
|
||||
# version while AUTO_VERSION_BUMP is disabled; in that case the helper
|
||||
# advances the preview base by one patch so stable users can still opt in
|
||||
# and receive it. The edit is ephemeral and never committed.
|
||||
- name: Stamp preview version
|
||||
if: needs.preview-gate.outputs.is_preview == 'true'
|
||||
shell: bash
|
||||
env:
|
||||
STABLE_TAG: ${{ needs.preview-gate.outputs.stable_tag }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# package.json is the single source of truth; tauri.conf.json reads its
|
||||
# version from it ("version": "../package.json"), so stamping
|
||||
# package.json restamps the whole bundle.
|
||||
CONF=frontend/package.json
|
||||
BASE=$(jq -r .version "$CONF")
|
||||
# MSI/WiX requires the semver pre-release identifier to be numeric-only
|
||||
# (and <= 65535). "preview.N" hard-fails the Windows bundler, so the
|
||||
# preview stamp is BASE-N — still sorts below the stable BASE for the
|
||||
# updater, still unique per run.
|
||||
PREVIEW_VERSION="${BASE}-${{ github.run_number }}"
|
||||
tmp=$(mktemp)
|
||||
jq --arg v "$PREVIEW_VERSION" '.version = $v' "$CONF" > "$tmp"
|
||||
mv "$tmp" "$CONF"
|
||||
PREVIEW_VERSION=$(python scripts/stamp-preview-version.py \
|
||||
--package-json frontend/package.json \
|
||||
--stable-tag "$STABLE_TAG" \
|
||||
--run-number "${{ github.run_number }}")
|
||||
echo "Stamped preview version: $PREVIEW_VERSION"
|
||||
|
||||
# The rolling `preview` release is REUSED every night, and macOS updater
|
||||
@@ -605,6 +604,21 @@ jobs:
|
||||
fi
|
||||
done < /tmp/stale.txt
|
||||
|
||||
# A retried job reuses its version and can collide with installers it
|
||||
# uploaded before a later step failed. Keep other versions/arches intact;
|
||||
# macOS versionless updater archives are scoped by release tag and arch.
|
||||
- name: Clear this target's installer assets on retry
|
||||
if: github.run_attempt > 1
|
||||
shell: bash
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
RELEASE_TAG: ${{ (needs.preview-gate.outputs.is_preview == 'true') && 'preview' || github.ref_name }}
|
||||
RELEASE_TARGET: ${{ matrix.rust_target }}
|
||||
run: |
|
||||
VERSION=$(python -c 'import json; print(json.load(open("frontend/package.json"))["version"])')
|
||||
python scripts/clear-release-rerun-assets.py \
|
||||
--tag "$RELEASE_TAG" --version "$VERSION" --target "$RELEASE_TARGET"
|
||||
|
||||
- name: Build + release (Tauri)
|
||||
uses: tauri-apps/tauri-action@v0
|
||||
env:
|
||||
@@ -660,6 +674,7 @@ jobs:
|
||||
set -euo pipefail
|
||||
python ../scripts/render-per-user-wix.py \
|
||||
--source src-tauri/wix/main.wxs \
|
||||
--system-wxs src-tauri/target/${{ matrix.rust_target }}/release/wix/x64/main.wxs \
|
||||
--output src-tauri/target/wix-per-user/main.wxs
|
||||
bunx tauri build --target ${{ matrix.rust_target }} --bundles msi \
|
||||
--config src-tauri/tauri.per-user.conf.json
|
||||
@@ -780,7 +795,7 @@ jobs:
|
||||
set -euo pipefail
|
||||
MSI=$(find frontend/src-tauri/target/${{ matrix.rust_target }}/release/bundle/msi -name '*Current*User*.msi' | head -1)
|
||||
powershell.exe -NoProfile -ExecutionPolicy Bypass \
|
||||
-File scripts/smoke-per-user-msi.ps1 -MsiPath "$(cygpath -w "$MSI")"
|
||||
-File scripts/smoke-per-user-msi.ps1 -MsiPath "$(cygpath -w "$MSI")" -PrepareHostedRunner
|
||||
|
||||
# linuxdeploy re-links .DirIcon as an ABSOLUTE symlink into the build
|
||||
# machine AFTER tauri's files-map has placed the real icon bytes — the
|
||||
|
||||
+10
@@ -170,3 +170,13 @@ bin/omnivoice-tts-linux-aarch64
|
||||
# committed (they ship with the app); the per-language source WAVs are just the
|
||||
# inputs scripts/render_dub_demo_audio.py hands to scripts/build_dub_demo.sh.
|
||||
backend/assets/samples/demo/dubbing/*.src.wav
|
||||
|
||||
# Stray sqlite session artifacts (`<db-path>.ses`). An in-memory DB yields the
|
||||
# literal name `:memory:.ses`, and a path containing `:` cannot be checked out
|
||||
# on Windows at all — committing one fails every Windows CI job at the git
|
||||
# checkout step, before a single test runs. Guarded by
|
||||
# tests/test_no_windows_hostile_paths.py.
|
||||
*.ses
|
||||
|
||||
# Generated Windows MSI diagnostic logs and installer payloads
|
||||
/wix-diagnostic-artifacts/
|
||||
|
||||
@@ -25,6 +25,8 @@ regexes = [
|
||||
'''^hf_QWERTYUIOPasdfghjklZXCVBNM0123456789xyzAB$''',
|
||||
# NLLB generation length argument, not the value of a credential.
|
||||
'''^max_length=400$''',
|
||||
# Dubbing pane split-position localStorage key, not a credential.
|
||||
'''^omnivoice\.dubSplit\.v1$''',
|
||||
# cryptography's Ed25519 private-key type name, not key material.
|
||||
'''^Ed25519PrivateKey$''',
|
||||
]
|
||||
|
||||
@@ -33,6 +33,11 @@ Binding for every AI agent (Claude, Codex, Cursor, review bots, …). CLAUDE.md
|
||||
- `frontend/package.json` dep changes require regenerating root `bun.lock` (Docker runs `--frozen-lockfile`).
|
||||
- Issues: absorb or decline — never defer to a future version. Check the open-PR queue before implementing community-reported fixes.
|
||||
|
||||
## Shared select controls
|
||||
|
||||
- Use `frontend/src/components/SearchableSelect.jsx` for all new or redesigned select boxes. Reuse `VoiceSelector` for voice choices. Do not introduce native `<select>` controls.
|
||||
- Provide a localized `ariaLabel`; use `menuPortal` inside scrolling or clipping containers. Preserve keyboard selection and disabled states.
|
||||
|
||||
## Agent skills
|
||||
|
||||
Project development skills are pinned in `skills-lock.json` and installed under
|
||||
|
||||
+183
-2
@@ -8,26 +8,207 @@ the frozen-backend fallback mirror it for their toolchains.
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
## [0.5.2] — 2026-09-10
|
||||
|
||||
**Highlights**
|
||||
|
||||
- Supertonic-3 and PocketTTS show their license Accept button again, so they can be enabled (#2017)
|
||||
- An engine that can't run on your platform says so, instead of telling you to install it (#2018)
|
||||
- MOSS-TTS-v1.5, Confucius4-TTS, dots.tts, Supertonic-3 and PocketTTS install in one click, each in its own environment, so switching engines and back never breaks a working one (#2015, #2016)
|
||||
- A pronunciation entry that is stored but not applied yet says so, instead of looking like it did not match (#1949)
|
||||
- A bare 500 report now names the backend error class, so two unrelated faults stop filing the same issue (#1773)
|
||||
- A rejected dubbing source language now names the code it rejected (#1960)
|
||||
- The first-run install log is kept on disk instead of vanishing with the setup screen (#1847)
|
||||
- `bun run desktop` reclaims port 3900 from a backend the app itself left running, instead of refusing to start (#1974)
|
||||
- A dictation shortcut another app already owns now says so, instead of silently doing nothing (#1858)
|
||||
- Quitting on Windows is no longer reported as a crash on the next launch (#1898)
|
||||
- A Reduce motion switch in Settings, for calm without changing your whole system (#1857)
|
||||
- A light theme, and System Auto now follows a light-mode OS instead of staying dark (#1973) — thanks @CoDe-ReDz!
|
||||
- Generating from a one-character input now says the input was too short, instead of quoting a convolution error (#1826)
|
||||
- First run asks about text size before the install, not after it (#1849)
|
||||
- Cloning without a reference clip now says so, instead of naming library parameters you cannot set (#1879)
|
||||
- Upgrading torch for an RTX 50-series card no longer trades one startup crash for another, and the upgrade is documented (#1931)
|
||||
- A generation timeout now points at the compute-time budget in Settings rather than an environment variable (#1808)
|
||||
- An engine you have not installed now says so, instead of reporting a failed check (#1866)
|
||||
- The Accessibility prompt no longer floats over first-run setup and every other app until you grant it (#1845, #1886)
|
||||
- The last onboarding step offers to install a speech-to-text model instead of failing three times when none is installed (#1856)
|
||||
- A download that fails because the folder sits behind a mount point Windows will not cross now says so, and where to move it (#1957)
|
||||
- A GPU that is merely short on free memory is no longer told to reinstall its drivers (#1812) — thanks @michaelhuamanflores!
|
||||
- An error thrown by a browser extension is filtered on Safari and the macOS app too, not only on Chromium (#1901) — thanks @Chang-Jin-Lee!
|
||||
- Choosing the China mirror no longer re-races the network on every dependency step, which cost seconds per step on blocked connections (#1892) — thanks @yuezheng2006!
|
||||
- The backend log panel reports a log it cannot read instead of quietly showing less (#1847) — thanks @Chang-Jin-Lee!
|
||||
- The floating dictation bubble adds pause, resume, stop, close, and a multiline preview (#1952)
|
||||
- Transcriptions checks model readiness and offers an inline download and shortcut hints (#1952)
|
||||
- Transcriptions' missing-model prompt lists every dictation model by accuracy vs latency, languages and size, so you install the one that fits — or switch to one already on disk (#1952)
|
||||
- The Engines menu's Transcription tab picks the dictation model under Sherpa-ONNX, and that choice now also drives Sherpa transcription (#1952)
|
||||
- A failure with no stage attached no longer borrows another stage's advice, so a text-to-speech error stops telling you the video server dropped the download (#1943)
|
||||
- A generation failure that the app cannot classify now names the backend error class, so two unrelated faults stop arriving as the same untriageable report (#1800)
|
||||
- Transcriptions dictation wakes the desktop recorder, presents one contextual start action, and centers its microphone icon with the label (#1902)
|
||||
- Colab transcription and dubbing now include an explicit ASR model setup step (#1922) — thanks @nidhi-singh02!
|
||||
- Apple Silicon now shows one canonical OmniVoice choice in the engine picker while retaining its automatic crash-isolated sidecar runtime (#1913)
|
||||
- Validate current-user Windows installers under a standard account on hosted runners (#1883)
|
||||
- Model downloads survive a flaky connection instead of restarting from zero (#1940)
|
||||
- `bun run dev` recovers on Windows instead of demanding Task Manager (#1941)
|
||||
- The desktop app builds and opens from a fresh clone again (#1818) — thanks @flutterkage2k!
|
||||
- GPUs with less VRAM than the engine needs no longer get half the compute-time budget a CPU gets (#1806) — thanks @VishvakR!
|
||||
- Gallery voice previews play again — the quality guard was rejecting good renders as silent (#1819) — thanks @flutterkage2k!
|
||||
- Tilde-separated number ranges are spoken clearly without running their endpoints together (#1821) — thanks @flutterkage2k!
|
||||
- Voice modes use themed tabs, with Synthesize and Convert pinned below their scrolling forms (#1823)
|
||||
- Fix current-user Windows installer validation and nested resource cleanup (#1873)
|
||||
- Keep generated frontend assets available while building the current-user Windows installer (#1881)
|
||||
- Voice cloning now starts with a clear upload-or-record choice, reveals recording and reference details only when needed, and keeps sampling controls under Production Overrides (#1817)
|
||||
- The first-run welcome line uses an instruction accepted by OmniVoice and VoiceDesign engines (#1861) — thanks @psiberfunk!
|
||||
- audio.cpp joins the engine lineup as an opt-in CPU backend for Breeze-TTS-2 (English + Chinese, clone + voice design, explicit Model Catalogue install, no Python venv) (#1891)
|
||||
- audio.cpp uses installed native CUDA, HIP, Metal, and Vulkan providers and preserves device routing across remote workers (#1926)
|
||||
- Show estimated and measured model, dependency, cache, and temporary disk costs in the engine catalogue (#1718)
|
||||
- CosyVoice setup guidance now separates downloaded model files from the runtime that makes the engine available.
|
||||
- Preview builds now stay newer than Stable even when automatic post-release version bumps are disabled (#1762)
|
||||
- CosyVoice setup guidance now separates downloaded model files from the runtime that makes the engine available (#1761)
|
||||
- MCP tools can now keep audio out of agent context by returning files and accepting base-path-confined file inputs (#1760) — thanks @agudmund!
|
||||
- Hear a dub line as you type it — an opt-in live preview streams TTS for the edited segment (#1769) — thanks @mvanhorn!
|
||||
- Studio gains a Convert method: re-say any clip in one of your saved voices, speech to speech, fully local (#1765) — thanks @mvanhorn!
|
||||
- Hardsub video export gains an opt-in karaoke word-highlight caption style (#1764) — thanks @mvanhorn!
|
||||
- The batch queue can now watch a folder: new videos dropped into it are dubbed automatically (#1768) — thanks @mvanhorn!
|
||||
- The audiobook player now shows the chapter text and highlights the word being narrated (#1766) — thanks @mvanhorn!
|
||||
- The dub editor gains a casting board: drag voice chips onto speakers, dropdowns stay in sync (#1767) — thanks @mvanhorn!
|
||||
|
||||
### Changed
|
||||
|
||||
- Tauri 2.11.5 with refreshed plugins (dialog, updater, log, opener, positioner, single-instance), React 19.3, TanStack Query 5.102, lucide 1.43, posthog-js 1.428, and the rest of the npm workspace on current minors; jsdom 30, jest-dom 7, concurrently 10, taze 21 (#1952)
|
||||
- eslint ignores `src-tauri/`, so a local Tauri build no longer floods `lint:hooks` with parse errors from generated assets (#1952)
|
||||
- Casting uses responsive SVG voice cards and searchable speaker menus that stay above surrounding panels (#1823)
|
||||
- Dubbing aligns output settings, brings review status forward, and simplifies transcript and glossary editing; Launchpad files and voices reflow into responsive grids (#1823)
|
||||
- Transcript segments use three readable rows for text, timing/status and voice controls, with heights that adapt to wrapping (#1823)
|
||||
- Dragging the waveform pans horizontally while a click still seeks, keeping the timed transcript aligned (#1823)
|
||||
- Bulk segment editing uses searchable voice and language menus, readable language names and a responsive selection toolbar (#1823)
|
||||
- Dubbing overlays playback controls on video, combines waveform and transcript in a compact timeline, and removes header/action background fills (#1823)
|
||||
- Dubbing uses compact casting, translation and output controls with responsive rows to leave more room for editing (#1823)
|
||||
- Export uses grouped format settings, themed track menus and switches, with a pinned filename summary and download action (#1823)
|
||||
- Dubbing output settings use icon-labelled switches, themed track and speaker menus, and clearer timing/transcript controls (#1823)
|
||||
- Casting voice menus use searchable themed options with SVG preset icons instead of native dropdowns (#1823)
|
||||
- Dubbing groups casting and translation controls with readable labels, SVG icons, searchable menus, and compact timeline spacing (#1823)
|
||||
- Production Overrides use readable icon-labelled controls and accessible Denoise/Postprocess switches (#1823)
|
||||
- Expanded navigation uses a theme-accent tint with subtle static wave gradients (#1823)
|
||||
- Convert groups source audio, target voice, and timing options into clearer controls; design choices include theme-matched SVG icons (#1823)
|
||||
- The expandable sidebar reveals workspace labels with restrained active states; language menus adapt to multiple columns on wider screens (#1823)
|
||||
- Voice design and recording use themed, keyboard-accessible selectors with clearer spacing and labels (#1823)
|
||||
- Voice tabs and upload/record controls have subtle SVG motion; Text adds clipboard paste and the upload area fills available height (#1823)
|
||||
- The title-bar label cycles through active speech, transcription, and LLM engines; bundled model labels correctly say OmniVoice (#1823)
|
||||
- The top-bar Engines panel groups Speech, Transcription, and LLM choices into tabs, with compact memory controls and no duplicate pickers (#1823)
|
||||
- Voice Design simplified: the 12-row fine-grained block collapses to one summary line with a five-field editor, English accent and Chinese dialect merge into a single field, and the starting-point chips now show 5 with an overflow toggle (#1793)
|
||||
|
||||
### Added
|
||||
|
||||
- The audiobook result is now a synced-lyrics player: chapter text follows playback with the current word highlighted and click-to-seek, timed from the render's own chapter durations with a karaoke-style even split — no ASR pass, fully local (#1766) — thanks @mvanhorn!
|
||||
- The dub CAST strip expands into a project-level casting board: drag voice chips (clone profiles, design presets, Default) onto speaker rows — or pick from a keyboard listbox — writing the same per-speaker cast fields as the existing dropdowns (#1767) — thanks @mvanhorn!
|
||||
- Studio's new Convert method turns a dropped or recorded clip into an existing voice profile's voice, with optional source-duration matching (#1765) — thanks @mvanhorn!
|
||||
- Opt-in watch folder on the batch queue: pick a directory once and new videos are auto-enqueued with your last Add-to-queue settings, with pause/stop controls and copy-in-progress protection — files upload as bytes, paths never leave the app (#1768) — thanks @mvanhorn!
|
||||
- Hardsub export can now burn karaoke word-highlight captions: an opt-in Line | Karaoke control renders a word-timed ASS sweep from timings persisted at transcription, with an even-split fallback for older jobs and translated tracks, plus a `GET /dub/ass/{job_id}` sidecar (#1764) — thanks @mvanhorn!
|
||||
- Windows releases now include an independently updatable per-user MSI that installs and uninstalls without elevation (#1713)
|
||||
- Dub segments can now stream live TTS while you edit a translated line — opt-in toggle, existing `/ws/tts` socket, shared generation admission, exports still render at full quality (#1769) — thanks @mvanhorn!
|
||||
- Engine status and diagnostic bundles now record loaded execution provider, device, precision, fallback stage, accelerator identity, runtime versions, and parent-process memory visibility (#1717)
|
||||
|
||||
### Docs
|
||||
|
||||
- The CosyVoice guide now states that packaged builds have no one-click runtime installer and records the exact readiness checks exposed by [Discussion 1631](https://github.com/debpalash/VoiceStudio/discussions/1631).
|
||||
- PowerShell Docker setup now generates the administrator key without requiring Python on the host (#1993) — thanks @yangfan-yf-yf!
|
||||
- The torch upgrade an RTX 50-series card needs is written down, with the second pin file the resolver checks and the command that proves the kernels are there (#1931)
|
||||
- Docker quick starts now explain the AMD64-only images and direct Apple Silicon users to the native macOS app (#1921) — thanks @yangfan-yf-yf!
|
||||
- audio.cpp (Breeze-TTS-2) is now a documented opt-in engine: prebuilt binary install, explicit GGUF download, voice modes, and the weights' research/non-commercial terms (#1891)
|
||||
- `docs/STRUCTURE.md` describes the tree as it is today, and a test now keeps its counts honest (#1981) — thanks @Dawcraft!
|
||||
- Local gigastt is now documented as a supported OpenAI-compatible ASR endpoint, with loopback privacy distinguished from remote servers (#1736) — thanks @ekhodzitsky!
|
||||
- The CosyVoice guide now states that packaged builds have no one-click runtime installer and records the exact readiness checks exposed by [Discussion 1631](https://github.com/debpalash/VoiceStudio/discussions/1631) (#1761)
|
||||
- A production private-API guide now covers pinned containers, root credentials, network isolation, streaming proxies, health checks, upgrades, and benchmark evidence (#1720)
|
||||
- RX 6700 XT/gfx1031 over WSL2 ROCDXG is now explicitly unverified until a published end-to-end GPU workload proves the mapped path (#1716)
|
||||
|
||||
### Fixed
|
||||
|
||||
- One-click engine installs no longer inherit VoiceStudio's own PyTorch pin, which made MOSS-TTS-v1.5 and Confucius4 impossible to install (#2024)
|
||||
- Uninstalling a translation engine no longer removes a package VoiceStudio or another engine still needs (#2019)
|
||||
- Closing the dictation pill on Windows removes it from the screen: an empty dark rectangle used to stay there, always on top, until the app was quit (#2009)
|
||||
- The dictation pill on Windows no longer sits inside a bordered card wider than the pill itself (#2009)
|
||||
- Dictation uses the model you picked instead of one remembered from before the backend started, so it stops reporting no speech-to-text model while one is installed — and when none is, the main window offers the download (#2012)
|
||||
- The remote-worker loop-responsiveness tests no longer turn a build red over milliseconds of scheduling noise on shared CI hardware (#1990)
|
||||
- Remote GPU workers work when the machine running VoiceStudio is on Windows: a staged input is now identified the same way on every operating system, instead of with a path only Windows can read (#2005)
|
||||
- The pronunciation list badges an IPA or CMU entry as not applied yet, so you can see it without running a test (#1949) — thanks @utkarsha741!
|
||||
- A remote-worker test no longer fails at random on Windows CI: it waited for a background thread by spinning the event loop that thread's work needed (#1990)
|
||||
- The isolated backend test session passes on a stock Windows checkout, and CI now runs it there so it stays that way (#1990)
|
||||
- Windows contributors can run the test suite without Developer Mode: tests that create a symlink now skip instead of failing with `WinError 1314` (#1990)
|
||||
- The crash details dialog now says what the exit code means and what to try, instead of showing a raw number and a log (#1927)
|
||||
- A crash report now carries the backend's actual last words: the log tail is captured after the dying process's final output lands, not the instant it exits (#1850)
|
||||
- The first-run setup screen no longer mislabels a step when the bootstrap restarts itself: Rust now says which attempt each stage and log line belongs to, instead of the screen guessing from a once-a-second poll (#1900)
|
||||
- A port-3900 conflict now names who is actually holding it, and gives the command that ends an orphaned backend, instead of telling you to quit an app that has no window (#1933) — thanks @Chang-Jin-Lee!
|
||||
- Windows desktop launches no longer freeze at "Loading ML runtime (PyTorch)": the parent-liveness watchdog polls the stdin pipe instead of leaving a read pending, which deadlocked numpy's OpenBLAS initializer (#1952, #1955)
|
||||
- `bun desktop-prod` and `bun desktop-fresh` find Rust and uv from a terminal opened before they were installed, as `bun desktop` already did; a missing Rust toolchain fails up front with the install steps (#1952)
|
||||
- Voice synthesis progress no longer races to a fabricated 95%; it stays indeterminate until the active generation path reports real progress (#1907) — thanks @psiberfunk!
|
||||
- The Backend log tab keeps showing history across a log rollover, instead of going nearly empty until new lines arrive (#1920)
|
||||
- Clearing the logs now empties the rotated log files too, so it frees the space it appears to (#1920)
|
||||
- An error thrown by a browser extension no longer offers to file itself as a VoiceStudio bug (#1901)
|
||||
- Clearing the desktop logs no longer wipes the backend's stderr, which is the only record a native crash leaves behind and is meant to survive a respawn (#1510)
|
||||
- Long audiobook chapters now use the same device- and text-length-aware synthesis timeout as other TTS routes (#1910) — thanks @psiberfunk!
|
||||
- Interrupted audiobook renders can resume cached chapters after tab navigation, and their chapter cache is available from the recovery card (#1911) — thanks @psiberfunk!
|
||||
- System-check details and storage paths beginning with a number or a slash no longer render with their leading text moved to the end of the line (#1848) — thanks @psiberfunk!
|
||||
- An unavailable engine's row now links to that engine's guide, so the generic "check installation and configuration" message has somewhere to send you (#1866) — thanks @psiberfunk!
|
||||
- The backend log now records which engine failed a health check and whether its probe raised, instead of a line that identified neither (#1866) — thanks @psiberfunk!
|
||||
- The first-run Activity log counts every line instead of freezing at 200 while the install is still running, and Copy now hands back the whole run rather than the last 200 lines (#1847) — thanks @psiberfunk!
|
||||
- A first-run failure that happened early in a long install keeps its specific advice, instead of falling back to the generic retry hint once the log scrolled past 200 lines (#1847) — thanks @psiberfunk!
|
||||
- Opening the log panel no longer clips the Launchpad's heading and slides the feature cards up over it — the page scrolls instead of squashing itself (#1859) — thanks @psiberfunk!
|
||||
- Segmented model downloads split files into 16 MB ranges instead of one range per connection, so a dropped connection refetches one range rather than restarting the file (#1940)
|
||||
- The download accelerator is kept across retries after a transient network failure and resumes from its manifest, instead of falling back to a from-zero `snapshot_download` (#1940)
|
||||
- `dev-backend.mjs` stops the backend by process tree on Windows, so an orphaned uvicorn no longer holds port 3900 and turns a source reload into three phantom crashes (#1941)
|
||||
- `clear-dev-ports.mjs` can free a stuck development port on Windows again, bound to the inspected process instance so a recycled pid is never terminated (#1941)
|
||||
- Checkout-ownership matching no longer resolves POSIX paths with the host's separator, which made the guard's own test fail on Windows (#1941)
|
||||
- Install documentation help now prints correctly on Windows consoles using legacy encodings (#1815) — thanks @dajiaohuang!
|
||||
- Saved transcriptions with missing or invalid timestamps now remain readable (#1799) — thanks @yunaremaia and @tvbht!
|
||||
- Transcribing with an engine that reports no segment end no longer fails with a server error; the null timing is passed through the way the segment list already expects (#1904) — thanks @aeroglu!
|
||||
- Copying a saved transcription now uses the shared clipboard helper and reports failed copies accurately (#1803) — thanks @tvbht!
|
||||
- Voice reference preparation reclaims allocator memory before one bounded retry, then reports persistent GPU out-of-memory failures (#1811)
|
||||
- `bun run desktop` now opens on a fresh clone: the Vite alias for `@tauri-apps/plugin-dialog` no longer assumes a nested `frontend/node_modules`, which bun's workspace hoisting leaves empty (#1818) — thanks @flutterkage2k!
|
||||
- Slow backend startups remain running with progress updates, and Retry interrupts startup without stale timeout failures (#1809)
|
||||
- Backend connection errors report crashes only when recorded evidence exists, and diagnostic waits honor cancellation (#1810)
|
||||
- A CUDA or ROCm GPU with less VRAM than the engine needs now gets the CPU compute-time budget instead of the shorter accelerated one, since it pages to system RAM and renders slower than the CPU would — applied to local generation, voice conversion, and remote worker deadlines alike (#1806) — thanks @VishvakR!
|
||||
- Gallery previews no longer fail with "the voice engine returned no audible audio" on perfectly good renders: the degenerate-buzz guard measured spectral flatness over the whole clip (so the value tracked clip length) against a threshold calibrated on a synthetic signal, and rejected real speech in every language tested (#1819) — thanks @flutterkage2k!
|
||||
- Speak tilde separators in integer, signed, and decimal ranges in English, Korean, Japanese, and Chinese (#1821) — thanks @flutterkage2k!
|
||||
- Keep recording and conversion work safe while switching methods, synchronize dubbing language controls, and localize timeline controls and timing warnings (#1841)
|
||||
- Audiobook is now a Write → Cast → Produce tab workspace matching the voice workspace, with the warnings/progress/result rail pinned below (#1841)
|
||||
- Gallery uses a workspace header with zone tabs, hairline section dividers, theme-token cards, and borderless import rows (#1841)
|
||||
- Gallery cards reset native button faces, cluster icon actions in the header so Use voice never wraps, and use a roomier grid floor (#1841)
|
||||
- Gallery filters gain name search, removable iconified pills with clear-all, and dimension icons on every facet (#1841)
|
||||
- Dubbing playback starts before waveform decoding, automatic cast names are readable, and transcript timestamps have more room (#1823)
|
||||
- The title-bar engine button stays compact and stable while cycling labels, with engine names aligned right (#1823)
|
||||
- Long dubbing segment errors wrap in a bounded scrollable notice instead of widening the editor (#1823)
|
||||
- Voice dropdowns match their field width, use theme accents, and show recent voices only once (#1823)
|
||||
- Language menus no longer show a pale frame around their search header (#1823)
|
||||
- The notification count stays inside the title bar instead of clipping above the bell (#1823)
|
||||
- The workspace engine menu opens beside its button instead of at the opposite edge of the page (#1823)
|
||||
- Cloning reuses the dubbing language picker with flags, search, and single selection, opening above the pinned synthesis controls (#1823)
|
||||
- The first-run welcome line uses an instruction accepted by OmniVoice and VoiceDesign engines (#1861) — thanks @psiberfunk!
|
||||
- The header status dot now honors OS Reduce Motion instead of pulsing regardless (#1862) — thanks @psiberfunk!
|
||||
- Onboarding reads Hugging Face tokens locally, preserves Windows CLI logins, and requires successful discovery before replacing saved credentials (#1852) — thanks @psiberfunk!
|
||||
- The logs panel no longer reports “All clear” before log retrieval succeeds or while logs contain warnings or errors (#1870) — thanks @motodriver!
|
||||
- MOSS accelerator routing and status match runtime selection, with CPU fallback when device probing fails (#1830) — thanks @li-lizhe!
|
||||
- Confucius accelerator routing tolerates failed device probes, and dots.tts keeps safe default precision on non-CUDA hosts (#1831) — thanks @li-lizhe!
|
||||
- On macOS, the header status dot and kicker no longer render underneath the overlaid traffic lights (#1863) — thanks @psiberfunk!
|
||||
- The capture widget can hide after recording and recover from being left visible while idle (#1865) — thanks @psiberfunk!
|
||||
- macOS retains the shared desktop window sizing, resize limits, and file-drop behavior when native chrome is applied (#1865) — thanks @psiberfunk!
|
||||
- On macOS, the header no longer shows Windows-style minimize/maximize/close buttons alongside the native traffic lights (#1865) — thanks @psiberfunk!
|
||||
- Release retries replace their own partially uploaded installers without colliding with existing assets (#1871)
|
||||
- Timed-out voice engines finish process cleanup before retrying, and old timeout callbacks cannot kill replacement engines (#1872)
|
||||
- Fast macOS process exits no longer turn a completed shutdown into a permission error (#1809)
|
||||
- The bootstrap splash no longer shows fabricated first-run install steps on a warm start or repair sync — a step now renders done only once it was actually observed (#1894)
|
||||
- A deliberate, clean quit killed by the desktop shell's short shutdown grace no longer gets reported as a crash on next launch — the run sentinel now clears before the slower shutdown steps instead of after (#1895)
|
||||
- Model Catalogue engine rows stack into one column on narrow shells instead of clipping actions off-screen (#1891)
|
||||
- Simplified Chinese locale completed: all 486 missing keys translated and the parity ratchet tightened to zero (#1877) — thanks @yearth!
|
||||
- The generation compute-time budget is now a Settings control (Performance & Device) instead of an env-var-only setting the timeout error recommended with no UI path — the error copy points there too, and long CPU/MPS renders get an upfront heads-up before they start (#1787)
|
||||
- Windows: the backend can now start when the install path contains non-English characters (e.g. a CJK username) on a non-UTF-8 system code page — a new or broken Python environment now builds at an ASCII-safe path automatically (a healthy existing one is never relocated), and a specific error message names the cause and a working fix if the interpreter still crashes in `site` (#1783)
|
||||
- Exports and other native-picker actions no longer 403 with "Invalid or expired desktop authorization" when the desktop app and backend resolve different data directories, e.g. dev mode or a custom data folder (#1781)
|
||||
- Voice Design no longer lets you pick a Chinese dialect and an English accent together — the picker keeps them mutually exclusive instead of round-tripping a 400 (#1771)
|
||||
- The desktop app no longer attaches to an already-running backend on version string alone: it now verifies the backend's actual code fingerprint too, so an orphaned or manually started backend reporting the current version but running older code (e.g. a stale `destination_path` export 422) gets replaced instead of adopted (#1770)
|
||||
- Korean locale overhauled: 231 mistranslations corrected and all 493 missing keys translated (#1776) — thanks @j30231!
|
||||
- Japanese "Cleaning…" clone status now reads as denoising instead of housekeeping (#1775) — thanks @j30231!
|
||||
- The batch dubbing queue now has a UI entry point — a quiet link on the Dub landing (it was previously unreachable: the app switched on a mode nothing ever set) (#1768) — thanks @mvanhorn!
|
||||
- OpenAI-compatible ASR now requires HTTPS outside loopback and refuses redirects so audio stays on the configured origin (#1736)
|
||||
- Windows isolated engines now retain direct Job ownership without an extra Python supervisor process that can deadlock the child loader (#1734)
|
||||
- The setup splash now waits through the backend's full startup budget instead of reporting slow Windows CUDA initialization as stuck after two minutes (#1749)
|
||||
- Dubbing jobs can now reuse every source-language code produced by automatic ASR detection without a 400 error on the next upload (#1737)
|
||||
- Incomplete Sherpa-ONNX model snapshots now self-repair before recognizer startup instead of failing on a missing ONNX file (#1733)
|
||||
- OmniVoice subprocess startup now allows slow packaged Windows Python runtimes to signal readiness before termination (#1711)
|
||||
- SRT files selected during source analysis now wait for speaker cloning, then replace transcript text without losing voices (#1709)
|
||||
|
||||
+10
-4
@@ -10,10 +10,10 @@ Copyright 2024-present Palash Debnath and VoiceStudio contributors.
|
||||
|
||||
VoiceStudio is **free and open-source software, licensed under the GNU
|
||||
Affero General Public License, Version 3 (AGPL-3.0)**. You are free to use,
|
||||
copy, modify, and redistribute it — and that **includes commercial and internal
|
||||
business use**: run the app, use its outputs commercially, sell the audio you
|
||||
produce with it, provide professional/client services with it, and deploy it
|
||||
within your organization.
|
||||
copy, modify, and redistribute it. That **includes commercial and internal
|
||||
business use** of the application itself. Model weights, tokenizers, and other
|
||||
third-party assets retain their own terms; this application license does not
|
||||
grant or summarize rights under those separate terms.
|
||||
|
||||
Because this is the **Affero** GPL, one additional obligation applies: if you
|
||||
modify VoiceStudio and make that modified version available to others over
|
||||
@@ -41,6 +41,12 @@ is **separately licensed under Apache License 2.0** by its upstream authors and
|
||||
is not relicensed here. Apache License 2.0 is compatible with, and may be
|
||||
combined under, the GNU AGPL-3.0. See `pyproject.toml`.
|
||||
|
||||
Downloaded model weights are not relicensed by VoiceStudio. The default
|
||||
`k2-fsa/OmniVoice` model card identifies its code as Apache-2.0 and pretrained
|
||||
weights as CC-BY-NC. Its `audio_tokenizer/LICENSE` contains separate Boson
|
||||
Higgs Audio 2 and Meta Llama community terms. A commercial license for
|
||||
VoiceStudio-owned code does not replace any of those terms.
|
||||
|
||||
Third-party dependencies retain their own licenses. See `Cargo.lock`,
|
||||
`bun.lock`, and `uv.lock` for the resolved set.
|
||||
|
||||
|
||||
@@ -1,26 +1,30 @@
|
||||
<div align="center">
|
||||
<a href="https://trendshift.io/repositories/28176?utm_source=repository-badge&utm_medium=badge&utm_campaign=badge-repository-28176" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/28176" alt="debpalash%2FVoiceStudio | Trendshift" width="250" height="55" /></a>
|
||||
|
||||
<img src="docs/logo.png" alt="VoiceStudio logo" width="120" height="120" />
|
||||
<p><img src="docs/logo.png" alt="VoiceStudio logo" width="120" height="120" /></p>
|
||||
<h1>VoiceStudio</h1>
|
||||
<p>
|
||||
<a href="https://trendshift.io/repositories/28176?utm_source=repository-badge&utm_medium=badge&utm_campaign=badge-repository-28176" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/repositories/28176" alt="VoiceStudio ranking on Trendshift" width="220" height="48" /></a>
|
||||
</p>
|
||||
<p><sub>Previously OmniVoice-Studio</sub></p>
|
||||
<h3>Local voice cloning, dubbing, dictation, and long-form audio.</h3>
|
||||
<p>16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, and Linux</p>
|
||||
<p><strong>Local-first.</strong> No account, API key, subscription, or usage meter for the core workflow.</p>
|
||||
<h3>Clone voices, dub video, dictate, and produce long-form audio on your own hardware.</h3>
|
||||
<p>16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker</p>
|
||||
<p>No account, API key, subscription, or usage meter for the local workflow.</p>
|
||||
|
||||
<p>
|
||||
<a href="#install">Install</a> ·
|
||||
<a href="#features">Features</a> ·
|
||||
<a href="#comparison">Compare</a> ·
|
||||
<a href="#requirements">Requirements</a> ·
|
||||
<a href="#hardware-recommendations">Hardware</a> ·
|
||||
<a href="#engines">Engines</a> ·
|
||||
<a href="#architecture">Architecture</a> ·
|
||||
<a href="#api">API</a> ·
|
||||
<a href="#documentation">Docs</a> ·
|
||||
<a href="#faq">FAQ</a> ·
|
||||
<a href="README_CN.md"><strong>简体中文</strong></a>
|
||||
</p>
|
||||
|
||||
<p>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/debpalash/VoiceStudio/ci.yml?branch=main&style=flat-square&label=CI" alt="CI status" /></a>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="GitHub stars" /></a>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/releases"><img src="https://img.shields.io/github/downloads/debpalash/VoiceStudio/total?style=flat-square&color=8b5cf6&label=downloads" alt="Total downloads" /></a>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="Latest release" /></a>
|
||||
@@ -38,7 +42,7 @@
|
||||
</div>
|
||||
|
||||
> [!WARNING]
|
||||
> **Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work or `main` for current fixes. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues).
|
||||
> **Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work. `main` contains the newest fixes and may change between releases. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues).
|
||||
|
||||
## At a glance
|
||||
|
||||
@@ -51,33 +55,70 @@
|
||||
| **Compute** | CUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers |
|
||||
| **Interfaces** | Desktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server |
|
||||
| **Storage** | Voices, projects, settings, and outputs stay on the machine by default |
|
||||
| **License** | AGPL-3.0; optional engines keep their own model licenses |
|
||||
| **License** | AGPL-3.0 application; downloaded models keep their upstream terms |
|
||||
|
||||
The Voice workspace starts with three tabs: **From audio** for cloning, **By design** for creating a voice, and **Convert** for speech-to-speech conversion. Each tab displays its own workflow, with Synthesize Audio or Convert pinned below the scrolling form. The top-bar **Engines** panel combines engine selection, loaded models, and unload/flush controls; <kbd>Ctrl</kbd>/<kbd>Cmd</kbd>+<kbd>E</kbd> opens it. The searchable language picker shares Dubbing’s flags and language list layout, selects one output language, and retains Auto and the full cloning catalogue. Language options flow into multiple columns when space allows. Expand **Workspaces** in the sidebar to reveal navigation labels; Escape collapses it.
|
||||
|
||||
Dubbing starts with file upload or URL import and nearby language choices. Its **Projects** panel lists previous dubs so they can be reopened by clicking anywhere on a card; action buttons operate independently. Advanced import options include captions and optional YouTube sign-in. Dubbing places playback controls over the video with background blur and combines the waveform and timed transcript in one compact editing surface. Drag the zoomed waveform left or right to pan; click to seek. Translation language and ISO-code controls stay synchronized; Auto clears any previous language code and dialect. Transcript items group editable text, timing and status, and voice controls into three readable rows that wrap with the panel width. Output Options stays compact with the active settings shown in its summary; expand it to change output, timing, or voice matching. Transcript, glossary, and paste controls share a toolbar above the segment editor. Project details, workflow steps, and Generate/Verify/Export actions use an unfilled header.
|
||||
|
||||
The Audiobook Script editor fills the available workspace beneath its markup toolbar; Voices and Book settings stay in their own tabs.
|
||||
|
||||
Output settings use aligned rows; review status appears before the collapsible transcript and glossary. Glossary terms have labelled entry fields and an explicit edit action. Launchpad arranges recent files and saved voices side by side when space allows, with responsive card grids and visible Open actions.
|
||||
|
||||
The casting board shows icon-based voice cards and searchable selectors for each speaker. Drag a card onto a speaker or choose a voice from that speaker’s menu.
|
||||
|
||||
<a id="install"></a>
|
||||
|
||||
## Install
|
||||
|
||||
Download a package from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest), then follow the platform guide.
|
||||
|
||||
| Platform | Package | Guide |
|
||||
|---|---|---|
|
||||
| macOS 13.3+ | DMG, Apple Silicon | [Install on macOS](docs/install/macos.md) |
|
||||
| Windows 10/11 | MSI, x64 | [Install on Windows](docs/install/windows.md) |
|
||||
| macOS 13.3+ | Apple Silicon DMG | [Install on macOS](docs/install/macos.md) |
|
||||
| Windows 10/11 | x64 MSI; choose the current-user build when listed to install without admin access | [Install on Windows](docs/install/windows.md#install-pre-built-msi) |
|
||||
| Linux | AppImage, x86_64 with glibc 2.39+ | [Install on Linux](docs/install/linux.md) |
|
||||
| Docker | CUDA, ROCm, or CPU; worker-only GPU profiles | [Run with Docker](docs/install/docker.md) |
|
||||
| Docker | Linux/AMD64 images; CUDA, ROCm, CPU, and worker-only GPU profiles | [Run with Docker](docs/install/docker.md) |
|
||||
|
||||
Download packages from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest). First launch creates a managed Python environment and downloads the default model. Later launches reuse both.
|
||||
First launch creates a managed Python environment and downloads the default model. Later launches reuse both.
|
||||
|
||||
> [!NOTE]
|
||||
> On macOS, first launch needs a one-time right-click → **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](docs/install/macos.md) instead.
|
||||
> On macOS, first launch needs a one-time right-click, then **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](docs/install/macos.md) instead.
|
||||
|
||||
### Quick Docker run
|
||||
|
||||
The published images are **`linux/amd64` only**. On Apple Silicon, use the
|
||||
[native macOS app](docs/install/macos.md) for GPU acceleration. ARM64 hosts
|
||||
should read the [architecture requirements](docs/install/docker.md#architecture)
|
||||
before pulling an image.
|
||||
|
||||
```bash
|
||||
docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable
|
||||
```
|
||||
|
||||
### First voice
|
||||
|
||||
1. Launch VoiceStudio and open **Voice Cloning**.
|
||||
2. Add a clean voice sample. Three seconds works; 5–15 seconds usually gives a better prompt.
|
||||
2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt.
|
||||
3. Enter text, choose a language, then select **Generate**.
|
||||
|
||||
> [!TIP]
|
||||
> **Try without installing:** Run VoiceStudio in the cloud via the [Google Colab notebook](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb). Explore audio quality comparisons in [benchmarks](docs/benchmarks.md) and prompt design tips in [expressive speech](docs/expressive-speech.md).
|
||||
|
||||
### Audio samples
|
||||
|
||||
Listen to sample outputs produced locally with VoiceStudio:
|
||||
|
||||
| Workflow | Prompt / Reference Audio | Generated Audio |
|
||||
|---|---|---|
|
||||
| **Voice Cloning** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) |
|
||||
| **Voice Design** (US News Anchor) | *"Clear, authoritative American broadcast tone"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) |
|
||||
| **Voice Design** (UK Audiobook) | *"Warm, expressive British storytelling voice"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) |
|
||||
| **Video Dubbing** (Multilingual) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [Spanish](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [French](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [Japanese](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [Chinese](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) |
|
||||
|
||||
### Run from source
|
||||
|
||||
Install the [development prerequisites](.github/CONTRIBUTING.md#development-setup), then:
|
||||
Install the [development prerequisites](.github/CONTRIBUTING.md#development-setup) (Node 20+/Bun and Python 3.11+), then:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/debpalash/VoiceStudio.git
|
||||
@@ -86,7 +127,7 @@ bun install
|
||||
bun run desktop
|
||||
```
|
||||
|
||||
Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md) for services, tests, and platform packages.
|
||||
The desktop launcher configures Python dependencies on first run via `uv` automatically. Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md) for services, tests, and platform packages.
|
||||
|
||||
### If setup fails
|
||||
|
||||
@@ -101,22 +142,22 @@ Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md
|
||||
|
||||
| Area | Included |
|
||||
|---|---|
|
||||
| **Voice Cloning** | Zero-shot synthesis from a short reference clip |
|
||||
| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions |
|
||||
| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video |
|
||||
| **Voice Cloning** | Zero-shot synthesis from a short reference clip ([guide](docs/engines/README.md)) |
|
||||
| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions ([expressive speech](docs/expressive-speech.md)) |
|
||||
| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video; compact translation settings include track selection, and completed dubs flag timing issues for review ([export guide](docs/dubbing/export.md)) |
|
||||
| **Stories and audiobooks** | Multi-voice scripts · EPUB/PDF import · chapter rendering · `.m4b` export |
|
||||
| **[Dictation Widget](docs/features/dictation.md)** | System-wide shortcut, live transcription, optional local-LLM cleanup |
|
||||
| **Vocal Isolation** | Demucs speech/background separation |
|
||||
| **Speaker Diarization** | Pyannote and WhisperX speaker assignment |
|
||||
| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress |
|
||||
| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models |
|
||||
| **Remote Model Downloads** | Install models on enrolled remote workers with live progress |
|
||||
| **GPU Auto-Detect** | CUDA, MPS, ROCm, and CPU routing with per-engine checks |
|
||||
| **Speaker Diarization** | Pyannote and WhisperX speaker assignment ([guide](docs/features/diarization.md)) |
|
||||
| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos |
|
||||
| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models ([catalogue](docs/engines/README.md)) |
|
||||
| **Remote Model Downloads** | Install models on enrolled remote workers with live progress ([guide](docs/downloading-models.md)) |
|
||||
| **GPU Auto-Detect** | CUDA, MPS, ROCm, and CPU routing with per-engine checks ([performance](docs/performance.md)) |
|
||||
| **AI Watermark** | AudioSeal embedding and detection |
|
||||
| **MCP Server** | Synthesis and transcription tools for MCP clients |
|
||||
| **Diagnostics** | Self-checks, error journal, logs, and scrubbed support bundles |
|
||||
| **MCP Server** | Synthesis and transcription tools for MCP clients ([guide](docs/mcp.md)) |
|
||||
| **Diagnostics** | Self-checks, error journal, logs, and scrubbed support bundles ([troubleshooting](docs/install/troubleshooting.md)) |
|
||||
| **Local-first** | Core creation stays local; network-backed features are explicit opt-ins |
|
||||
| **Extensible** | Registry-based TTS, ASR, and plugin interfaces |
|
||||
| **Extensible** | Registry-based TTS, ASR, and plugin interfaces ([acceptance](docs/engine-acceptance.md)) |
|
||||
|
||||
<table>
|
||||
<tr>
|
||||
@@ -159,10 +200,20 @@ Requirements vary by engine. These values cover the default local workflow.
|
||||
| **Disk** | 10 GB free | 20 GB+ SSD |
|
||||
| **GPU** | Optional; CPU mode is supported | NVIDIA CUDA or Apple Silicon |
|
||||
| **VRAM** | 4 GB when using a GPU | 8 GB+; large optional engines need more |
|
||||
| **Python from source** | 3.11+ | 3.11–3.12 |
|
||||
| **Python from source** | 3.11+ | 3.11 or 3.12 |
|
||||
|
||||
ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See [performance](docs/performance.md), [benchmarks](docs/benchmarks.md), and [engine disk usage](docs/engines/disk-usage.md).
|
||||
|
||||
<a id="hardware-recommendations"></a>
|
||||
|
||||
### Recommended stack by hardware
|
||||
|
||||
| Hardware | Recommended TTS | Recommended ASR | Why |
|
||||
|---|---|---|---|
|
||||
| **Apple Silicon (M1–M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | Native unified memory, lowest latency on macOS |
|
||||
| **NVIDIA GPU (8 GB+ VRAM)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | High-fidelity zero-shot cloning, word timestamps, diarization |
|
||||
| **Low VRAM / CPU-only** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | Low memory footprint, optimized CPU inference |
|
||||
|
||||
<a id="engines"></a>
|
||||
|
||||
## Engines
|
||||
@@ -175,22 +226,22 @@ Engine support is capability-specific. Check cloning, language, platform, memory
|
||||
|
||||
| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License |
|
||||
|---|:---:|:---:|:---:|:---:|:---:|:---:|---|
|
||||
| **VoiceStudio** (default, powered by k2-fsa/OmniVoice) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
|
||||
| **CosyVoice 3** | 9 + 18 dialects | Yes | Yes | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| **GPT-SoVITS** | 5 | Yes | — | CUDA/CPU | — | CUDA/CPU | MIT |
|
||||
| **VoxCPM2** | 30 | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 |
|
||||
| **MOSS-TTS-Nano** | 20 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| **KittenTTS** | English | — | — | CPU | CPU | CPU | MIT |
|
||||
| **MLX-Audio** | Model-dependent | Varies | Varies | — | MLX | — | Varies |
|
||||
| **Sherpa-ONNX** | 20+ | — | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| **IndexTTS 2.5** ⚡ | ZH · EN · JA · ES · AR | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Bilibili model license¹ |
|
||||
| **OmniVoice GGUF** ⚡ | 600+ | Yes | Yes | CUDA/CPU | MPS/CPU | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
|
||||
| **OmniVoice (subprocess)** ⚡ | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0](LICENSE-NOTICE.md) model |
|
||||
| **PocketTTS** ⚡ | EN · FR · DE · PT · IT · ES | Yes | — | CPU | CPU | CPU | CC-BY-4.0, gated² |
|
||||
| **Supertonic 3** ⚡ | 31 | — | — | CPU | CPU | CPU | OpenRAIL-M |
|
||||
| **MOSS-TTS-v1.5** ⚡ | 31 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| **dots.tts** ⚡ | 24 | Yes | — | CUDA/CPU | CPU | — | Apache-2.0 |
|
||||
| **Confucius4-TTS** ⚡ | 14 | Yes | — | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| [**VoiceStudio** (default, powered by k2-fsa/OmniVoice)](docs/engines/omnivoice.md) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ |
|
||||
| [**CosyVoice 3**](docs/engines/cosyvoice.md) | 9 + 18 dialects | Yes | Yes | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| [**GPT-SoVITS**](docs/engines/gpt-sovits.md) | 5 | Yes | No | CUDA/CPU | No | CUDA/CPU | MIT |
|
||||
| [**VoxCPM2**](docs/engines/voxcpm2.md) | 30 | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 |
|
||||
| [**MOSS-TTS-Nano**](docs/engines/moss-tts-nano.md) | 20 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| [**KittenTTS**](docs/engines/kittentts.md) | English | No | No | CPU | CPU | CPU | MIT |
|
||||
| [**MLX-Audio**](docs/engines/mlx-audio.md) | Model-dependent | Varies | Varies | No | MLX | No | Varies |
|
||||
| [**Sherpa-ONNX**](docs/engines/sherpa-onnx.md) | 20+ | No | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| [**IndexTTS 2.5** ⚡](docs/engines/indextts.md) | ZH · EN · JA · ES · AR | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Bilibili model license¹ |
|
||||
| [**OmniVoice GGUF** ⚡](docs/engines/omnivoice-gguf.md) | 600+ | Yes | Yes | CUDA/CPU | MPS/CPU | CUDA/CPU | [AGPL-3.0](LICENSE) app · [review the derivative model terms](https://huggingface.co/Serveurperso/OmniVoice-GGUF#license)³ |
|
||||
| [**OmniVoice (subprocess; opt-in off MPS)** ⚡](docs/engines/omnivoice-subprocess.md) | 600+ | Yes | Yes | CUDA/CPU | MPS via default OmniVoice | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ |
|
||||
| [**PocketTTS** ⚡](docs/engines/pockettts.md) | EN · FR · DE · PT · IT · ES | Yes | No | CPU | CPU | CPU | CC-BY-4.0, gated² |
|
||||
| [**Supertonic 3** ⚡](docs/engines/supertonic3.md) | 31 | No | No | CPU | CPU | CPU | OpenRAIL-M |
|
||||
| [**MOSS-TTS-v1.5** ⚡](docs/engines/moss-tts-v15.md) | 31 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
| [**dots.tts** ⚡](docs/engines/dots-tts.md) | 24 | Yes | No | CUDA/CPU | CPU | No | Apache-2.0 |
|
||||
| [**Confucius4-TTS** ⚡](docs/engines/confucius4-tts.md) | 14 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 |
|
||||
|
||||
⚡ Installed or registered on demand.
|
||||
|
||||
@@ -198,6 +249,8 @@ Engine support is capability-specific. Check cloning, language, platform, memory
|
||||
|
||||
² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use.
|
||||
|
||||
³ The OmniVoice snapshot also includes an audio tokenizer under separate [Boson Higgs Audio 2 and Meta Llama community terms](https://huggingface.co/k2-fsa/OmniVoice/blob/main/audio_tokenizer/LICENSE). VoiceStudio's application license does not replace model or tokenizer terms.
|
||||
|
||||
Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first.
|
||||
|
||||
<a id="asr-engines"></a>
|
||||
@@ -206,17 +259,17 @@ Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voic
|
||||
|
||||
| Engine | ID | Languages | Best fit |
|
||||
|---|---|:---:|---|
|
||||
| **WhisperX** (default) | `whisperx` | ~100 | Dubbing, subtitles, word-level timing |
|
||||
| **Faster-Whisper** | `faster-whisper` | ~100 | General cross-platform transcription |
|
||||
| **Faster-Whisper (isolated)** | `faster-whisper-isolated` | ~100 | Crash-isolated batch transcription |
|
||||
| **MLX Whisper** | `mlx-whisper` | ~100 | Apple Silicon |
|
||||
| **PyTorch Whisper** | `pytorch-whisper` | ~100 | CUDA, MPS, and CPU fallback |
|
||||
| **Parakeet TDT** | `nemo-parakeet` | English + 25 EU | Fast CPU/CUDA transcription |
|
||||
| **Parakeet TDT v3 (MLX)** | `parakeet-mlx` | 25 EU | Apple Silicon dictation and word timestamps |
|
||||
| **Moonshine** | `moonshine` | English | Low-power, low-latency ONNX |
|
||||
| **FunASR** | `funasr` | 50+ | VAD and inline diarization |
|
||||
| **sherpa-onnx** (live dictation) | `sherpa-onnx-asr` | Model-dependent | Streaming CPU dictation |
|
||||
| **OpenAI-compatible** ⚠️ remote | `openai-compat-asr` | Server-dependent | Qwen3-ASR or another compatible endpoint; audio leaves the machine |
|
||||
| [**WhisperX** (default)](docs/engines/whisperx.md) | `whisperx` | ~100 | Dubbing, subtitles, word-level timing |
|
||||
| [**Faster-Whisper**](docs/engines/faster-whisper.md) | `faster-whisper` | ~100 | General cross-platform transcription |
|
||||
| [**Faster-Whisper (isolated)**](docs/engines/faster-whisper-isolated.md) | `faster-whisper-isolated` | ~100 | Crash-isolated batch transcription |
|
||||
| [**MLX Whisper**](docs/engines/mlx-whisper.md) | `mlx-whisper` | ~100 | Apple Silicon |
|
||||
| [**PyTorch Whisper**](docs/engines/pytorch-whisper.md) | `pytorch-whisper` | ~100 | CUDA, MPS, and CPU fallback |
|
||||
| [**Parakeet TDT**](docs/engines/nemo-parakeet.md) | `nemo-parakeet` | English + 25 EU | Fast CPU/CUDA transcription |
|
||||
| [**Parakeet TDT v3 (MLX)**](docs/engines/parakeet-mlx.md) | `parakeet-mlx` | 25 EU | Apple Silicon dictation and word timestamps |
|
||||
| [**Moonshine**](docs/engines/moonshine.md) | `moonshine` | English | Low-power, low-latency ONNX |
|
||||
| [**FunASR**](docs/engines/funasr.md) | `funasr` | 50+ | VAD and inline diarization |
|
||||
| [**sherpa-onnx** (live dictation)](docs/engines/sherpa-onnx-asr.md) | `sherpa-onnx-asr` | Model-dependent | Streaming CPU dictation |
|
||||
| [**OpenAI-compatible** ⚠️ configured server](docs/engines/openai-compatible-asr.md) | `openai-compat-asr` | Server-dependent | Local gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server |
|
||||
|
||||
WhisperX and Faster-Whisper retry with `int8` when efficient `float16` is unavailable. Pin `ASR_COMPUTE_TYPE=int8` or `float32` only if automatic selection still fails.
|
||||
|
||||
@@ -251,8 +304,8 @@ FastAPI backend
|
||||
|
||||
- The desktop talks to a loopback-only backend on `localhost:3900`.
|
||||
- Loopback API calls need no server key. Remote access requires a share PIN or API key.
|
||||
- Remote workers and OpenAI-compatible ASR are opt-in. The UI identifies when audio leaves the machine.
|
||||
- Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata—not text, audio, file names, or projects.
|
||||
- Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed.
|
||||
- Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects.
|
||||
|
||||
<a id="api"></a>
|
||||
|
||||
@@ -287,12 +340,19 @@ with client.audio.speech.with_streaming_response.create(
|
||||
response.stream_to_file("speech.wav")
|
||||
```
|
||||
|
||||
The bundled Rust control sidecar also lets Herdr, coding agents, VS Code,
|
||||
desktop apps, and TUIs trigger the existing system-wide dictation flow or reuse
|
||||
its safe native insertion. See the [speech platform guide](docs/speech-platform.md).
|
||||
The full API reference is in **Settings → OpenAPI Reference**. For LAN,
|
||||
Tailscale, or proxy access, read [API authentication](docs/api-auth.md) before
|
||||
exposing the backend.
|
||||
```bash
|
||||
# Quick test via cURL
|
||||
curl http://localhost:3900/v1/audio/speech \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"model": "tts-1", "input": "Made on my own hardware.", "voice": "default", "response_format": "wav"}' \
|
||||
--output speech.wav
|
||||
```
|
||||
|
||||
The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps,
|
||||
and TUIs trigger the system-wide dictation flow or reuse its native text
|
||||
insertion. See the [speech platform guide](docs/speech-platform.md). The full API
|
||||
reference is in **Settings → OpenAPI Reference**. For LAN, Tailscale, or proxy
|
||||
access, read [API authentication](docs/api-auth.md) before exposing the backend.
|
||||
|
||||
### Agent skills
|
||||
|
||||
@@ -305,6 +365,36 @@ npx skills add debpalash/VoiceStudio
|
||||
- `omnivoice`: synthesize speech and transcribe audio through local VoiceStudio.
|
||||
- `oss-maintainer`: the repository's open-source maintenance workflow.
|
||||
|
||||
### Model Context Protocol (MCP)
|
||||
|
||||
VoiceStudio mounts an MCP server at `http://localhost:3900/mcp` for Claude Desktop, Cursor, and AI agents:
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"voicestudio": {
|
||||
"url": "http://localhost:3900/mcp"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For clients requiring stdio transport, use the bundled local shim (`docs/mcp.json`):
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"voicestudio": {
|
||||
"command": "python",
|
||||
"args": ["-m", "backend.mcp_shim"],
|
||||
"cwd": "/path/to/VoiceStudio"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
See the [MCP guide](docs/mcp.md) for tools (`generate_speech`, `clone_voice`, `transcribe`), file streaming modes, and client bindings.
|
||||
|
||||
### Google Colab
|
||||
|
||||
[](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb)
|
||||
@@ -326,6 +416,8 @@ The [notebook](notebooks/OmniVoice_Studio_Colab.ipynb) runs the app and web UI o
|
||||
| Track changes | [Changelog](CHANGELOG.md) · [roadmap](docs/ROADMAP.md) · [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) |
|
||||
| Remove everything | [Uninstall guide](docs/install/uninstall.md) |
|
||||
|
||||
<a id="faq"></a>
|
||||
|
||||
## FAQ
|
||||
|
||||
<details>
|
||||
@@ -337,19 +429,19 @@ Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the l
|
||||
<details>
|
||||
<summary><strong>How much VRAM do I need?</strong></summary>
|
||||
|
||||
A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12–16 GB or more. Check the [benchmarks](docs/benchmarks.md) and engine guide.
|
||||
A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the [benchmarks](docs/benchmarks.md) and engine guide.
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Why does a longer reference clip not always improve the clone?</strong></summary>
|
||||
|
||||
Cloning is zero-shot: the clip is a prompt, not training data. Use 5–15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see [data preparation](docs/data_preparation.md) and [training](docs/training.md).
|
||||
Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see [data preparation](docs/data_preparation.md) and [training](docs/training.md).
|
||||
</details>
|
||||
|
||||
<details>
|
||||
<summary><strong>Can I use generated audio commercially?</strong></summary>
|
||||
|
||||
Yes under VoiceStudio's AGPL-3.0 terms. Optional engines and model weights may use different licenses; review the selected engine's license before commercial use.
|
||||
VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use.
|
||||
</details>
|
||||
|
||||
<details>
|
||||
@@ -371,17 +463,30 @@ Use `scripts/uninstall.sh` on macOS/Linux or `scripts\uninstall.ps1` on Windows.
|
||||
- [Good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue) for a scoped starting point.
|
||||
- [Contributing guide](.github/CONTRIBUTING.md) for setup, tests, and pull requests.
|
||||
|
||||
<p align="center">
|
||||
<a href="https://star-history.com/#debpalash/VoiceStudio&Date">
|
||||
<img src="https://api.star-history.com/svg?repos=debpalash/VoiceStudio&type=Date" alt="Star History Chart" width="100%" />
|
||||
</a>
|
||||
</p>
|
||||
|
||||
## Support development
|
||||
|
||||
VoiceStudio is free and has no paid tier. Donations fund development and infrastructure.
|
||||
|
||||
[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsorship details](SPONSORS.md)
|
||||
|
||||
## Responsible use and safety
|
||||
|
||||
VoiceStudio enables zero-shot voice cloning and speech generation on personal hardware. Please use it responsibly:
|
||||
- **Consent:** Only clone or synthesize voices with explicit permission from the speaker.
|
||||
- **Audio provenance:** VoiceStudio integrates [AudioSeal](https://github.com/facebookresearch/audioseal) imperceptible watermarking by default to detect and identify synthetic speech without altering sound quality.
|
||||
- **Local privacy:** For the default local workflow, audio recordings, transcripts, voices, and projects remain strictly on your local disk; data leaves your device only when you explicitly configure remote workers or external ASR endpoints.
|
||||
|
||||
## License
|
||||
|
||||
VoiceStudio is licensed under [AGPL-3.0](LICENSE). You may run it, modify it, use it internally, and sell generated audio. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license is available for proprietary embedding; contact **VoiceStudio@palash.dev**. See [LICENSE-NOTICE.md](LICENSE-NOTICE.md) for the plain-language scope.
|
||||
VoiceStudio is licensed under [AGPL-3.0](LICENSE). You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact **VoiceStudio@palash.dev**. See [LICENSE-NOTICE.md](LICENSE-NOTICE.md) for the plain-language scope.
|
||||
|
||||
Optional engines and downloaded models retain their own licenses. The bundled `omnivoice/` model remains Apache-2.0 upstream.
|
||||
Optional engines and downloaded models retain their own licenses. The bundled `omnivoice/` Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms.
|
||||
|
||||
## Acknowledgments
|
||||
|
||||
|
||||
+67
-3
@@ -20,6 +20,7 @@
|
||||
</p>
|
||||
|
||||
<p>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/debpalash/VoiceStudio/ci.yml?branch=main&style=flat-square&label=CI" alt="CI 状态" /></a>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/stargazers"><img src="https://img.shields.io/github/stars/debpalash/VoiceStudio?style=flat-square&color=f59e0b" alt="Star 数" /></a>
|
||||
<a href="https://github.com/debpalash/VoiceStudio/releases/latest"><img src="https://img.shields.io/github/v/release/debpalash/VoiceStudio?style=flat-square&color=10b981" alt="版本" /></a>
|
||||
<a href="LICENSE"><img src="https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square" alt="许可证" /></a>
|
||||
@@ -65,11 +66,27 @@
|
||||
- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md)
|
||||
- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio)
|
||||
|
||||
```bash
|
||||
# Docker 快速运行 (CPU / 本地环回模式)
|
||||
docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable
|
||||
```
|
||||
|
||||
**三步克隆出你的第一个声音:**
|
||||
|
||||
1. **安装并启动。** 首次启动会自动搭建 Python 运行环境并下载模型权重——启动画面会逐步显示进度(仅首次,需要几分钟;之后即开即用)。
|
||||
2. 从启动台打开**语音克隆**,拖入任意声音的 **3 秒音频**。
|
||||
3. **输入一句话,点击生成。** 音频完全属于你——在你的设备上生成和保存,支持 646 种语言。
|
||||
3. **输入一句话,点击生成。** 音频在你的设备上生成并保存,支持 646 种语言(商业使用前请审阅所选模型与分词器的许可条款)。
|
||||
|
||||
### 🎧 音频示例
|
||||
|
||||
在线试听 VoiceStudio 本地生成的实际音频样例:
|
||||
|
||||
| 工作流 | 提示词 / 参考音频 | 生成音频 |
|
||||
|---|---|---|
|
||||
| **声音克隆** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) |
|
||||
| **声音设计** (美语新闻主播) | *"清晰、权威的美国广播级音色"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) |
|
||||
| **声音设计** (英式有声书) | *"温暖生动的英式故事讲述音色"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) |
|
||||
| **视频配音** (多语种) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [西班牙语](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [法语](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [日语](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [中文](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) |
|
||||
|
||||
觉得慢?[docs/performance.md](docs/performance.md) 讲清了生成时间到底花在哪里、有哪些调优开关,以及“它变慢了”的三个经典原因。各引擎/设备的实测数据见 [docs/benchmarks.md](docs/benchmarks.md)。
|
||||
|
||||
@@ -217,6 +234,16 @@ Hugging Face Token 的配置见
|
||||
> [!IMPORTANT]
|
||||
> **macOS Intel(x86_64)不支持本地后端:** 应用 UI 可以安装,但 Python 后端无法运行,因为 PyTorch 已不再发布 Intel Mac 轮子([#889](https://github.com/debpalash/VoiceStudio/issues/889))。Intel Mac 用户仍可让 UI 指向另一台机器上的远程后端——参见 [docs/install/macos.md](docs/install/macos.md)。
|
||||
|
||||
<a id="hardware-recommendations"></a>
|
||||
|
||||
### 💡 按硬件推荐引擎配置
|
||||
|
||||
| 硬件配置 | 推荐 TTS 引擎 | 推荐 ASR 语音识别 | 优势 |
|
||||
|---|---|---|---|
|
||||
| **Apple Silicon (M1–M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | 原生统一内存,macOS 上延迟最低、性能最强 |
|
||||
| **NVIDIA 显卡 (8 GB+ 显存)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | 极致零样本克隆品质、字级时间戳对齐与说话人分离 |
|
||||
| **低显存 / 仅 CPU 设备** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | 超低内存占用,针对 CPU 指令集深度优化 |
|
||||
|
||||
<a id="tts-engines"></a>
|
||||
|
||||
### 🗣️ TTS 引擎
|
||||
@@ -338,9 +365,9 @@ print(result.text)
|
||||
|
||||
### 📓 在 Google Colab 上运行
|
||||
|
||||
[](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/VoiceStudio_Studio_Colab.ipynb)
|
||||
[](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb)
|
||||
|
||||
没有本地 GPU?官方笔记本([notebooks/VoiceStudio_Studio_Colab.ipynb](notebooks/VoiceStudio_Studio_Colab.ipynb))可在免费的 Colab T4 上启动完整应用(包含 Web 界面):在笔记本内直接构建前端,用 uv 安装后端(复用 Colab 预装的 CUDA PyTorch),并通过 Colab 内置端口代理打开界面。无需第三方隧道,也无需任何 API 密钥。随后还有一套覆盖全部主要功能的 API 导览,全部可在笔记本内直接播放:多语言 TTS、声音克隆与声音设计、已保存的声音档案、语音转写、AI 水印检测、OpenAI 兼容 API、多角色故事、带章节的 m4b 有声书,以及一个附带人声分离音轨的迷你视频配音。
|
||||
没有本地 GPU?官方笔记本([notebooks/OmniVoice_Studio_Colab.ipynb](notebooks/OmniVoice_Studio_Colab.ipynb))可在免费的 Colab T4 上启动完整应用(包含 Web 界面):在笔记本内直接构建前端,用 uv 安装后端(复用 Colab 预装的 CUDA PyTorch),并通过 Colab 内置端口代理打开界面。无需第三方隧道,也无需任何 API 密钥。随后还有一套覆盖全部主要功能的 API 导览,全部可在笔记本内直接播放:多语言 TTS、声音克隆与声音设计、已保存的声音档案、语音转写、AI 水印检测、OpenAI 兼容 API、多角色故事、带章节的 m4b 有声书,以及一个附带人声分离音轨的迷你视频配音。
|
||||
|
||||
### 🤝 智能体技能(Agent Skills)
|
||||
|
||||
@@ -352,6 +379,36 @@ npx skills add debpalash/omnivoice-studio
|
||||
|
||||
内含两个 [skills](https://skills.sh):**`omnivoice`**——让任何智能体通过你的本地安装进行语音合成与转录(包括你克隆的声音),免费且离线;以及 **`oss-maintainer`**——本项目所遵循的维护者方法论,适合任何用智能体运营自己开源项目的人。
|
||||
|
||||
### 🔌 模型上下文协议(MCP 服务器)
|
||||
|
||||
VoiceStudio 在 `http://localhost:3900/mcp` 挂载了 MCP 服务,可供 Claude Desktop、Cursor 与自主智能体调用:
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"voicestudio": {
|
||||
"url": "http://localhost:3900/mcp"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
对于需要 stdio 管道传输的客户端,请使用内置的本地桥接脚本(`docs/mcp.json`):
|
||||
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"voicestudio": {
|
||||
"command": "python",
|
||||
"args": ["-m", "backend.mcp_shim"],
|
||||
"cwd": "/path/to/VoiceStudio"
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
支持 `generate_speech`、`clone_voice`、`transcribe` 等工具与流式文件输出模式,详见 [docs/mcp.md](docs/mcp.md)。
|
||||
|
||||
---
|
||||
|
||||
## 🗺️ 路线图
|
||||
@@ -542,6 +599,13 @@ VoiceStudio **免费**且采用 **AGPL-3.0** 许可——没有付费版,没
|
||||
VoiceStudio 完全本地运行——卸载就是删除应用及其写入的文件夹(模型缓存、Python 环境、你的声音/项目、配置)。运行 <code>scripts/uninstall.sh</code>(macOS/Linux)或 <code>scripts\uninstall.ps1</code>(Windows)——它会先以干跑方式列出每个文件夹及其大小,加 <code>--yes</code> 才会真正删除。完整的各平台路径列表和应用移除步骤见 <a href="docs/install/uninstall.md"><b>docs/install/uninstall.md</b></a>。
|
||||
</details>
|
||||
|
||||
## 🛡️ 负责任使用与安全
|
||||
|
||||
VoiceStudio 在个人硬件上提供零样本语音克隆与语音创作能力。我们提倡负责任的技术使用:
|
||||
- **明确授权:** 严禁在未经说话人本人知情并明确授权的情况下克隆其声音。
|
||||
- **AI 溯源:** VoiceStudio 默认集成 [AudioSeal](https://github.com/facebookresearch/audioseal) 不可见神经音频水印,在完全不影响听感音质的前提下精准标记合成语音。
|
||||
- **本地隐私:** 默认本地工作流下,所有音频、声音档案、项目与转录文本始终保存在你的本地设备上;仅当你主动配置远程工作节点或第三方 ASR 端点时,相应数据才会传输到对应服务。
|
||||
|
||||
---
|
||||
|
||||
<a id="license"></a>
|
||||
|
||||
@@ -32,17 +32,129 @@ def _public_routing_reason(status: object, diagnostic: object) -> str:
|
||||
return _ROUTING_BY_STATUS.get(status, _ROUTING_UNAVAILABLE)
|
||||
|
||||
|
||||
# Categories for WHY an engine is unavailable. The probe's own sentence cannot
|
||||
# cross the boundary — it carries exception text, local paths and sometimes
|
||||
# credentials — but "Engine unavailable. Check installation and configuration."
|
||||
# told the user nothing at all, and "Last error: A previous engine check
|
||||
# failed." reads like a crash rather than "you have not installed this yet"
|
||||
# (#1866). Classifying the private diagnostic into an owned sentence keeps the
|
||||
# boundary intact and still names the kind of problem and the place to fix it.
|
||||
_UNAVAILABLE_NOT_INSTALLED = (
|
||||
"This engine's package isn't installed yet. Install it from "
|
||||
"Model Catalogue → Engines."
|
||||
)
|
||||
# An engine gated behind an in-app license review (Supertonic-3, PocketTTS).
|
||||
# The Model Catalogue shows its Accept button only when the reason matches
|
||||
# /license not accepted/i (EngineCompatibilityMatrix.reasonMentionsLicense), so
|
||||
# this sentence must keep those words: collapsing it into the generic line hid
|
||||
# the only way to enable those engines.
|
||||
_UNAVAILABLE_LICENSE = (
|
||||
"License not accepted yet. Review and accept it in "
|
||||
"Model Catalogue → Engines to enable this engine."
|
||||
)
|
||||
# An engine that cannot run on this machine at all: Apple-Silicon-only MLX,
|
||||
# PyTorch with no Intel Mac build. "Isn't installed yet" or "check
|
||||
# installation" sent people after an install that could never work.
|
||||
_UNAVAILABLE_PLATFORM = (
|
||||
"This engine doesn't run on this computer's platform. Its guide lists "
|
||||
"the platforms it supports."
|
||||
)
|
||||
# Apple Silicon whose PyTorch cannot use the GPU (MPS): the platform is
|
||||
# right, the installation is not. MLX-Audio / MLX-Whisper need MPS (#390).
|
||||
_UNAVAILABLE_NO_MPS = (
|
||||
"This engine needs Apple's GPU (MPS), and this installation's PyTorch "
|
||||
"can't use it. Updating macOS or reinstalling VoiceStudio usually "
|
||||
"restores it."
|
||||
)
|
||||
_UNAVAILABLE_NEEDS_CONFIG = (
|
||||
"This engine needs to be configured before it can run. Open "
|
||||
"Model Catalogue → Engines to finish setting it up."
|
||||
)
|
||||
_UNAVAILABLE_FILE_MISSING = (
|
||||
"A file this engine needs is missing or unreadable. Reinstall it from "
|
||||
"Model Catalogue → Engines."
|
||||
)
|
||||
|
||||
# The same two cases for an engine the app cannot install for you. "Install it
|
||||
# from Model Catalogue → Engines" sent people to a page with no Install button
|
||||
# for that engine — most of the catalogue — which reads as the app being
|
||||
# broken. The row's own guide link (``docs_url``) is the real next step.
|
||||
_UNAVAILABLE_NOT_INSTALLED_MANUAL = (
|
||||
"This engine isn't installed yet, and it has no one-click install. "
|
||||
"Its guide lists the install steps."
|
||||
)
|
||||
_UNAVAILABLE_FILE_MISSING_MANUAL = (
|
||||
"A file this engine needs is missing or unreadable. Its guide lists the "
|
||||
"install steps."
|
||||
)
|
||||
_MANUAL_INSTALL_VARIANT = {
|
||||
_UNAVAILABLE_NOT_INSTALLED: _UNAVAILABLE_NOT_INSTALLED_MANUAL,
|
||||
_UNAVAILABLE_FILE_MISSING: _UNAVAILABLE_FILE_MISSING_MANUAL,
|
||||
}
|
||||
|
||||
# Matched against the lowered probe text. Ordered most specific first: a
|
||||
# missing file often also says "not installed", and the file case has the more
|
||||
# useful remedy of the two.
|
||||
_UNAVAILABLE_SIGNATURES = (
|
||||
# First: its probe text also says "Open Model Catalogue", and the
|
||||
# license is the one gap only the user can close.
|
||||
(_UNAVAILABLE_LICENSE, ("license not accepted",)),
|
||||
# Before the install and file checks: a platform reason often also says
|
||||
# "unavailable" or names a missing wheel, and no install can fix it. Not
|
||||
# "apple silicon only": mlx-audio says that on an M-series Mac too, when
|
||||
# the package is merely missing and installing does help.
|
||||
(_UNAVAILABLE_PLATFORM, (
|
||||
"requires apple silicon", "not supported on this platform",
|
||||
"unavailable on intel macs", "no macos x86_64 wheel",
|
||||
"no windows install", "not supported on windows",
|
||||
)),
|
||||
(_UNAVAILABLE_NO_MPS, ("torch mps unavailable",)),
|
||||
(_UNAVAILABLE_FILE_MISSING, (
|
||||
"file is missing", "file is empty", "file is unreadable",
|
||||
"script missing", "binary", "not found at",
|
||||
)),
|
||||
(_UNAVAILABLE_NEEDS_CONFIG, (
|
||||
"environment variable", "configure a server endpoint", "api key",
|
||||
"unconfigured", "set the", "base url",
|
||||
)),
|
||||
(_UNAVAILABLE_NOT_INSTALLED, (
|
||||
"not installed", "package missing", "not available", "no module named",
|
||||
"import ", "unavailable:", "failed to load",
|
||||
)),
|
||||
)
|
||||
|
||||
|
||||
def _public_unavailable_reason(diagnostic: object) -> str:
|
||||
"""Map a private availability probe to an accurate stable category."""
|
||||
private = diagnostic.lower() if isinstance(diagnostic, str) else ""
|
||||
for public, markers in _UNAVAILABLE_SIGNATURES:
|
||||
if any(marker in private for marker in markers):
|
||||
return public
|
||||
return _UNAVAILABLE
|
||||
|
||||
|
||||
def public_backends(entries: list[dict]) -> list[dict]:
|
||||
"""Copy registry entries while replacing service diagnostics.
|
||||
|
||||
Availability probes may contain exception text, local paths, tracebacks, or
|
||||
credentials. Installation hints are registry-authored and remain intact.
|
||||
credentials. Registry-authored fields are not probe output and remain
|
||||
intact: ``install_hint``, ``setup_snippet`` and ``docs_url`` are all
|
||||
VoiceStudio-owned constants keyed on the engine id, so an unavailable row
|
||||
still has something actionable to show and somewhere to send the user
|
||||
(#1866) even though ``reason``/``last_error`` are replaced here.
|
||||
"""
|
||||
safe: list[dict] = []
|
||||
for entry in entries:
|
||||
item = dict(entry)
|
||||
if item.get("reason") is not None:
|
||||
item["reason"] = _UNAVAILABLE
|
||||
reason = _public_unavailable_reason(item["reason"])
|
||||
# Only a row that explicitly says it has NO one-click install gets
|
||||
# the manual wording. Rows without the field (ASR, LLM,
|
||||
# translation — some of which have installers of their own) keep
|
||||
# the line that points at Model Catalogue.
|
||||
if item.get("one_click_install") is False:
|
||||
reason = _MANUAL_INSTALL_VARIANT.get(reason, reason)
|
||||
item["reason"] = reason
|
||||
if item.get("last_error") is not None:
|
||||
item["last_error"] = _PREVIOUS_FAILURE
|
||||
if item.get("routing_reason") is not None:
|
||||
|
||||
@@ -57,11 +57,12 @@ _PREVIEW_SEED = 42
|
||||
# 32 reliably converges to speech across the gallery's instruct/script space
|
||||
# at a one-time (cached) render cost.
|
||||
_PREVIEW_NUM_STEP = 32
|
||||
# Spectral-flatness floor below which a render is a degenerate tonal artifact
|
||||
# rather than speech. Real, mastered speech sits ~0.04–0.07; a tonal buzz
|
||||
# collapses to <0.005. 0.015 separates the two with wide margin and sits well
|
||||
# below even breathy/whisper voices (which are broadband → high flatness).
|
||||
_DEGENERATE_FLATNESS = 0.015
|
||||
# Reject near-pure tonal artifacts using mean framed spectral flatness.
|
||||
# Calibrated against the tracked speech demos exercised by
|
||||
# test_archetype_preview_quality.py: the quietest (Mandarin dubbing, 44.1 kHz)
|
||||
# measures ~7.7e-6, while the worst tested tonal buzz measures ~3.3e-9.
|
||||
# 1e-7 leaves >10x margin on both sides without rejecting low-flatness speech.
|
||||
_DEGENERATE_FLATNESS = 1e-7
|
||||
|
||||
|
||||
def _preview_key(a: dict) -> str:
|
||||
@@ -248,24 +249,46 @@ def _is_blank_audio(audio_tensor) -> bool:
|
||||
return False
|
||||
|
||||
|
||||
_FLATNESS_FRAME = 1024
|
||||
_FLATNESS_HOP = 512
|
||||
#: Frames quieter than this fraction of the loudest frame's energy are the gaps
|
||||
#: between words, not speech; their spectrum is the noise floor and averaging it
|
||||
#: in drags the measurement toward the value of whatever silence sounds like.
|
||||
_FLATNESS_FRAME_FLOOR = 1e-4
|
||||
|
||||
|
||||
def _spectral_flatness(audio_tensor) -> Optional[float]:
|
||||
"""Geometric-mean / arithmetic-mean of the power spectrum.
|
||||
"""Mean per-frame geometric-mean / arithmetic-mean of the power spectrum.
|
||||
|
||||
~1.0 for broadband noise, →0 for a pure tone. The degenerate diffusion
|
||||
renders this guards against are near-pure tonal buzzes (flatness <0.005),
|
||||
distinct from both silence (caught by ``_is_blank_audio``) and real speech
|
||||
(~0.04+). Returns ``None`` if it can't be computed so callers don't act on
|
||||
a bad measurement.
|
||||
renders this guards against are near-pure tonal buzzes, distinct from both
|
||||
silence (caught by ``_is_blank_audio``) and real speech. Returns ``None``
|
||||
if it can't be computed so callers don't act on a bad measurement.
|
||||
|
||||
Measured over short frames and averaged — the standard definition. A single
|
||||
FFT of the whole clip (what this used to do) is not the same quantity: its
|
||||
frequency resolution grows with clip length, so speech harmonics carve
|
||||
ever-deeper nulls into the spectrum and the geometric mean collapses. That
|
||||
made the result depend on how long the clip was rather than on what it
|
||||
sounded like, and put real speech below the rejection threshold.
|
||||
"""
|
||||
try:
|
||||
import torch
|
||||
|
||||
t = audio_tensor if isinstance(audio_tensor, torch.Tensor) else torch.as_tensor(audio_tensor)
|
||||
t = t.detach().to("cpu", dtype=torch.float32).flatten()
|
||||
if t.numel() < 1024 or not torch.isfinite(t).all():
|
||||
t = t.detach().to("cpu", dtype=torch.float32)
|
||||
if t.ndim > 1:
|
||||
t = t.mean(dim=0)
|
||||
t = t.flatten()
|
||||
if t.numel() < _FLATNESS_FRAME or not torch.isfinite(t).all():
|
||||
return None
|
||||
spec = torch.fft.rfft(t * torch.hann_window(t.numel())).abs().pow(2) + 1e-12
|
||||
return float(torch.exp(torch.mean(torch.log(spec))) / torch.mean(spec))
|
||||
frames = t.unfold(0, _FLATNESS_FRAME, _FLATNESS_HOP)
|
||||
spec = torch.fft.rfft(frames * torch.hann_window(_FLATNESS_FRAME)).abs().pow(2) + 1e-12
|
||||
energy = spec.sum(dim=1)
|
||||
spec = spec[energy > energy.max() * _FLATNESS_FRAME_FLOOR]
|
||||
if spec.shape[0] == 0:
|
||||
return None
|
||||
return float((torch.exp(spec.log().mean(dim=1)) / spec.mean(dim=1)).mean())
|
||||
except Exception: # never let the checker itself block a render
|
||||
return None
|
||||
|
||||
|
||||
@@ -718,6 +718,10 @@ def _remote_chapter_call(chapter, *, engine_id, default_voice, voice_map,
|
||||
"expressive": opts.to_manifest(), "watermark": bool(watermark_enabled()),
|
||||
}
|
||||
signature = hashlib.sha256(json.dumps(params, sort_keys=True, default=str).encode()).hexdigest()
|
||||
# The worker synthesizes from ``spans``, but the gateway and scheduler read
|
||||
# top-level ``text`` to scale the remote execution deadline. Add this after
|
||||
# the signature so existing content-addressed remote cache keys still hit.
|
||||
params["text"] = "\n".join(row["text"] for row in rows)
|
||||
wav_path = os.path.join(cache_dir, f"remote-{signature}.wav")
|
||||
|
||||
def decode(result):
|
||||
@@ -739,7 +743,7 @@ async def _run_chapter(chapter, *, operation="audiobook", decision, job, default
|
||||
voice_map, lexicon, cache_dir):
|
||||
"""Run one chapter through the gateway; local preparation stays lazy."""
|
||||
from services import gpu_gateway
|
||||
from services.tts_backend import active_backend_id
|
||||
from services.tts_backend import active_backend_id, get_backend_class
|
||||
|
||||
engine_id = active_backend_id()
|
||||
remote, remote_cache = _remote_chapter_call(
|
||||
@@ -753,15 +757,27 @@ async def _run_chapter(chapter, *, operation="audiobook", decision, job, default
|
||||
return remote_cache, float(info.duration), True, None
|
||||
|
||||
async def prepare_local():
|
||||
from services.model_manager import generate_timeout_s
|
||||
|
||||
synth, sr, resolve, local_engine = await _prepare_synth(
|
||||
default_voice, language=language, opts=opts, voice_map=voice_map
|
||||
)
|
||||
try:
|
||||
timeout_engine = get_backend_class(local_engine)
|
||||
except ValueError:
|
||||
# Tests and third-party integrations may inject a synth under a
|
||||
# non-catalogue id. Keep the canonical host/text policy available;
|
||||
# registered production engines still add their routing metadata.
|
||||
timeout_engine = None
|
||||
return gpu_gateway.LocalCall(
|
||||
fn=lambda: _render_chapter_cached(
|
||||
chapter, synth, sr, local_engine, resolve, cache_dir, lexicon,
|
||||
language, opts, voice_map,
|
||||
),
|
||||
what="Audiobook chapter",
|
||||
timeout=generate_timeout_s(
|
||||
remote.params["text"], engine=timeout_engine
|
||||
),
|
||||
)
|
||||
|
||||
return await gpu_gateway.run(
|
||||
|
||||
@@ -111,6 +111,24 @@ BATCH_WIDTH_ENV = "OMNIVOICE_DUB_BATCH_WIDTH"
|
||||
#: that costs more than the saving.
|
||||
_MAX_BATCH_WIDTH = 16
|
||||
|
||||
# Bound each allocation while persisting multipart uploads. Video inputs can
|
||||
# be many gigabytes; `await UploadFile.read()` with no size used to mirror the
|
||||
# entire file in process memory before writing it back out.
|
||||
_UPLOAD_CHUNK_BYTES = 1024 * 1024
|
||||
|
||||
|
||||
async def _save_upload(upload: UploadFile, destination: str) -> None:
|
||||
try:
|
||||
with open(destination, "wb") as output:
|
||||
while chunk := await upload.read(_UPLOAD_CHUNK_BYTES):
|
||||
output.write(chunk)
|
||||
except BaseException:
|
||||
try:
|
||||
unlink_if_present(destination)
|
||||
except FileCleanupError:
|
||||
logger.warning("Could not remove incomplete batch upload", exc_info=True)
|
||||
raise
|
||||
|
||||
|
||||
def _native_batch_width(backend) -> int:
|
||||
"""How many segments to render in one native batch on THIS host.
|
||||
@@ -690,9 +708,7 @@ async def enqueue_batch_job(
|
||||
ext = os.path.splitext(video.filename or "video.mp4")[1] or ".mp4"
|
||||
video_path = os.path.join(batch_dir, f"{job_id}{ext}")
|
||||
|
||||
with open(video_path, "wb") as f:
|
||||
content = await video.read()
|
||||
f.write(content)
|
||||
await _save_upload(video, video_path)
|
||||
|
||||
job = {
|
||||
"id": job_id,
|
||||
|
||||
@@ -28,6 +28,17 @@ router = APIRouter()
|
||||
logger = logging.getLogger("omnivoice.capture")
|
||||
|
||||
|
||||
def _timing(value):
|
||||
"""A segment timing, or ``None`` when the engine could not determine one.
|
||||
|
||||
``dict.get(key, 0)`` hands back a stored ``None`` rather than the default,
|
||||
because the key is present — so rounding it raised and took a transcript
|
||||
that was otherwise fine down with it (#1904). Pass the null through instead:
|
||||
the segment list renders whichever half of the range is known.
|
||||
"""
|
||||
return round(value, 2) if isinstance(value, (int, float)) else None
|
||||
|
||||
|
||||
def _truthy(value: Optional[str]) -> bool:
|
||||
"""Parse a multipart form flag. Treats '1'/'true'/'yes'/'on'/'auto'
|
||||
(any case) as on; everything else — including None — as off."""
|
||||
@@ -162,10 +173,15 @@ async def transcribe_audio(
|
||||
from services.text_polish import polish_text
|
||||
full_text = polish_text(full_text)
|
||||
|
||||
# Calculate audio duration from segments if available
|
||||
# Calculate audio duration from segments if available. A segment whose
|
||||
# timing the engine could not determine carries end=None (sherpa's
|
||||
# _sherpa_result when the sample rate yields no duration, and every
|
||||
# plain-text OpenAI-compatible response), so measure only the ones that
|
||||
# have a number and keep 0.0 when none do.
|
||||
duration = 0.0
|
||||
if segments:
|
||||
duration = max(s.get("end", 0) for s in segments)
|
||||
ends = [e for e in (s.get("end") for s in segments) if isinstance(e, (int, float))]
|
||||
duration = max(ends) if ends else 0.0
|
||||
|
||||
detected_lang = result.get("language", language or "unknown")
|
||||
|
||||
@@ -194,8 +210,8 @@ async def transcribe_audio(
|
||||
"text": full_text,
|
||||
"segments": [
|
||||
{
|
||||
"start": round(s.get("start", 0), 2),
|
||||
"end": round(s.get("end", 0), 2),
|
||||
"start": _timing(s.get("start", 0)),
|
||||
"end": _timing(s.get("end", 0)),
|
||||
"text": s.get("text", "").strip(),
|
||||
}
|
||||
for s in segments
|
||||
|
||||
@@ -55,6 +55,17 @@ from services.text_polish import polish_text
|
||||
router = APIRouter()
|
||||
logger = logging.getLogger("omnivoice.capture_ws")
|
||||
|
||||
|
||||
def _timing(value):
|
||||
"""A segment timing, or ``None`` when the engine could not determine one.
|
||||
|
||||
``dict.get(key, 0)`` returns a stored ``None`` rather than the default, so
|
||||
rounding it raised (#1904). The null is the honest answer here — this module
|
||||
emits it deliberately for un-endpointed utterances — and the segment list
|
||||
renders whichever half of the range is known.
|
||||
"""
|
||||
return round(value, 2) if isinstance(value, (int, float)) else None
|
||||
|
||||
SPEECH_PROTOCOL = "voicestudio.speech.v1"
|
||||
PLATFORM_STREAM_PATH = "/v1/audio/transcriptions/stream"
|
||||
|
||||
@@ -349,6 +360,7 @@ async def ws_transcribe(websocket: WebSocket):
|
||||
audio_chunks: list[bytes] = []
|
||||
total_bytes = 0
|
||||
last_audio_time = time.monotonic()
|
||||
paused = False
|
||||
running = True
|
||||
partial_text = ""
|
||||
# Track whether the client initiated the disconnect. When True the
|
||||
@@ -366,7 +378,7 @@ async def ws_transcribe(websocket: WebSocket):
|
||||
message as the authoritative result and skip the duplicate HTTP
|
||||
POST that used to run on every dictation.
|
||||
"""
|
||||
nonlocal total_bytes, last_audio_time, running, client_disconnected
|
||||
nonlocal total_bytes, last_audio_time, running, client_disconnected, paused
|
||||
try:
|
||||
while running:
|
||||
msg = await websocket.receive()
|
||||
@@ -397,6 +409,10 @@ async def ws_transcribe(websocket: WebSocket):
|
||||
total_bytes += len(data)
|
||||
last_audio_time = time.monotonic()
|
||||
continue
|
||||
if msg.get("text") in ("PAUSE", "RESUME"):
|
||||
paused = msg["text"] == "PAUSE"
|
||||
last_audio_time = time.monotonic()
|
||||
continue
|
||||
if _is_end_control(msg.get("text")):
|
||||
# Client signals end-of-audio but stays connected for `final`.
|
||||
running = False
|
||||
@@ -429,6 +445,9 @@ async def ws_transcribe(websocket: WebSocket):
|
||||
if not running:
|
||||
break
|
||||
|
||||
if paused:
|
||||
continue
|
||||
|
||||
# Check silence timeout
|
||||
if time.monotonic() - last_audio_time > SILENCE_TIMEOUT_S and total_bytes > MIN_BUFFER_BYTES:
|
||||
running = False
|
||||
@@ -1187,13 +1206,19 @@ async def _transcribe_buffer_full(
|
||||
from services.refinement import collapse_repetitive_artifacts
|
||||
full_text = collapse_repetitive_artifacts(full_text)
|
||||
|
||||
duration = max((s.get("end", 0) for s in segments), default=0.0)
|
||||
# end=None means the engine could not determine the timing — this
|
||||
# module writes exactly that in its own streaming payloads, and
|
||||
# sherpa's _sherpa_result does too when the sample rate yields no
|
||||
# duration. Measure only real numbers, and pass the nulls through
|
||||
# rather than rounding them (#1904).
|
||||
ends = [e for e in (s.get("end") for s in segments) if isinstance(e, (int, float))]
|
||||
duration = max(ends) if ends else 0.0
|
||||
|
||||
return {
|
||||
"text": full_text,
|
||||
"segments": [
|
||||
{"start": round(s.get("start", 0), 2),
|
||||
"end": round(s.get("end", 0), 2),
|
||||
{"start": _timing(s.get("start", 0)),
|
||||
"end": _timing(s.get("end", 0)),
|
||||
"text": s.get("text", "").strip()}
|
||||
for s in segments
|
||||
],
|
||||
|
||||
@@ -86,6 +86,18 @@ def list_dictation_models():
|
||||
}
|
||||
|
||||
|
||||
@router.get("/dictation/readiness", dependencies=[Depends(require_local)])
|
||||
def dictation_readiness(model_id: str | None = None) -> dict:
|
||||
"""Check capture's model selection without loading or downloading weights."""
|
||||
from services.asr_backend import asr_model_missing_error
|
||||
|
||||
missing = asr_model_missing_error(
|
||||
purpose="dictation",
|
||||
sherpa_model_id=model_id or _read_prefs()["model_id"],
|
||||
)
|
||||
return {"ready": missing is None, "missing": missing}
|
||||
|
||||
|
||||
@router.get("/dictation/prefs", dependencies=[Depends(require_local)])
|
||||
def get_dictation_prefs():
|
||||
return _read_prefs()
|
||||
|
||||
@@ -518,9 +518,12 @@ _ingest_gen = dub_pipeline.ingest_pipeline
|
||||
#: container so a mislabelled video can't slip past the video-skipping branch.
|
||||
_AUDIO_EXTS = {".wav", ".mp3", ".m4a", ".aac", ".flac", ".ogg", ".opus", ".wma"}
|
||||
|
||||
# Source-language choices exposed by the first-party dub UI. Keeping this an
|
||||
# allow-list rejects language names and private-use BCP-47 tags before they are
|
||||
# persisted as ASR overrides. Values are normalized to lowercase below.
|
||||
# Source-language choices exposed by the first-party dub UI, plus every
|
||||
# language code Whisper can write back after auto-detection. A restored job
|
||||
# may reuse that detected value as the next upload's override, so rejecting our
|
||||
# own persisted codes strands otherwise valid dubbing sessions (#1737).
|
||||
# Keeping this an allow-list still rejects language names and private-use
|
||||
# BCP-47 tags. Values are normalized to lowercase below.
|
||||
_DUB_SOURCE_LANG_CODES = frozenset({
|
||||
"af", "sq", "am", "ar", "hy", "az", "eu", "be", "bn", "bs", "bg",
|
||||
"my", "ca", "cmn-hans", "cmn-hant", "hr", "cs", "da", "nl", "en",
|
||||
@@ -531,19 +534,48 @@ _DUB_SOURCE_LANG_CODES = frozenset({
|
||||
"ru", "sm", "gd", "sr", "sn", "sd", "si", "sk", "sl", "so", "es",
|
||||
"su", "sw", "sv", "tg", "ta", "te", "th", "tr", "uk", "ur", "uz",
|
||||
"vi", "cy", "xh", "yi", "yo", "zu",
|
||||
"as", "ba", "bo", "br", "fo", "lb", "ln", "mg", "nn", "oc", "sa",
|
||||
"tk", "tl", "tt", "yue", "zh",
|
||||
})
|
||||
|
||||
|
||||
def _source_lang_override(value: str | None) -> str | None:
|
||||
"""Normalize a user-selected source language; auto/und means detect."""
|
||||
"""Normalize a user-selected source language; auto/und means detect.
|
||||
|
||||
A rejection NAMES the code it rejected. "Invalid source language code" on
|
||||
its own cannot be acted on or reported usefully: it does not say which of
|
||||
the ninety-odd codes was wrong, so neither the user nor a maintainer
|
||||
reading the auto-filed issue can tell whether the picker offered something
|
||||
the backend does not accept, or a stale preference from an older build is
|
||||
still being sent (#1960).
|
||||
|
||||
The value is a language code the user chose from a menu — not private
|
||||
data — and the neighbouring engine validator already echoes its input the
|
||||
same way.
|
||||
"""
|
||||
code = (value or "").strip().lower()
|
||||
if code in {"", "auto", "und"}:
|
||||
return None
|
||||
if code not in _DUB_SOURCE_LANG_CODES:
|
||||
raise HTTPException(status_code=400, detail="Invalid source language code")
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail=(
|
||||
f"Invalid source language code: {code!r}. Pick a language from "
|
||||
"the Dubbing source-language menu, or leave it on auto-detect."
|
||||
),
|
||||
)
|
||||
return code
|
||||
|
||||
|
||||
def _detected_source_lang(value: str | None) -> str:
|
||||
"""Normalize an ASR language without truncating valid three-letter codes."""
|
||||
code = (value or "en").split("_", 1)[0].strip().lower()
|
||||
if code in _DUB_SOURCE_LANG_CODES:
|
||||
return code
|
||||
short = code[:2]
|
||||
return short if short in _DUB_SOURCE_LANG_CODES else "en"
|
||||
|
||||
|
||||
@router.post("/dub/upload")
|
||||
async def dub_upload(
|
||||
video: UploadFile = File(...),
|
||||
@@ -1809,9 +1841,9 @@ async def dub_transcribe_stream(
|
||||
except Exception as e:
|
||||
logger.warning("speaker_clone extraction skipped: %s", e)
|
||||
|
||||
job["source_lang"] = job.get("source_lang_override") or (
|
||||
(detected_lang or "en").split("_")[0][:2] or "en"
|
||||
).lower()
|
||||
job["source_lang"] = job.get("source_lang_override") or _detected_source_lang(
|
||||
detected_lang
|
||||
)
|
||||
job["full_transcript"] = " ".join(s.get("text", "") for s in final_segs)
|
||||
_save_job(job_id, job)
|
||||
|
||||
@@ -2008,9 +2040,9 @@ async def dub_transcribe(job_id: str, num_speakers: Optional[int] = None):
|
||||
except Exception as e:
|
||||
logger.warning("Failed to unload ASR backend: %s", e)
|
||||
|
||||
job["source_lang"] = job.get("source_lang_override") or (
|
||||
(detected_lang or "en").split("_")[0][:2] or "en"
|
||||
).lower()
|
||||
job["source_lang"] = job.get("source_lang_override") or _detected_source_lang(
|
||||
detected_lang
|
||||
)
|
||||
|
||||
scene_cuts = job.get("scene_cuts") or []
|
||||
segments = segment_transcript(result, duration=job.get("duration", 0.0), scene_cuts=scene_cuts)
|
||||
|
||||
@@ -23,6 +23,7 @@ from services.ffmpeg_utils import (
|
||||
find_ffmpeg,
|
||||
run_ffmpeg,
|
||||
)
|
||||
from services.karaoke_ass import build_ass, scale_words
|
||||
from services.video_retime import (
|
||||
DRIFT_TOLERANCE_S,
|
||||
RetimeError,
|
||||
@@ -403,6 +404,27 @@ def _write_burn_srt(job: dict, exports_dir: str, stamp: str, dual: bool,
|
||||
return sub_path
|
||||
|
||||
|
||||
def _write_burn_ass(job: dict, exports_dir: str, stamp: str,
|
||||
fitted_segments: "list[dict] | None" = None,
|
||||
lang: "str | None" = None) -> str | None:
|
||||
"""Karaoke variant of ``_write_burn_srt``: word-timed ASS via ``build_ass``.
|
||||
|
||||
Same text/timing resolution (``_segments_for_lang`` + fitted-cue overlay,
|
||||
which also scales per-word times onto the fitted timeline); the basename
|
||||
is plain ASCII under exports_dir so it is ffmpeg-filter-safe. Returns
|
||||
None if there are no segments to render.
|
||||
"""
|
||||
segments = _segments_for_lang(job, lang)
|
||||
if not segments:
|
||||
return None
|
||||
if fitted_segments:
|
||||
segments = _apply_fitted_times(segments, fitted_segments)
|
||||
sub_path = os.path.join(exports_dir, f"burn_subs_{stamp}.ass")
|
||||
with open(sub_path, "w", encoding="utf-8") as f:
|
||||
f.write(build_ass(segments))
|
||||
return sub_path
|
||||
|
||||
|
||||
def _ffmpeg_filter_escape(path: str) -> str:
|
||||
"""Escape a path for use inside an ffmpeg filter value (subtitles=...).
|
||||
|
||||
@@ -515,6 +537,20 @@ def _apply_fitted_times(segments: list[dict], fitted: list[dict]) -> list[dict]:
|
||||
patched = dict(seg)
|
||||
patched["start"] = float(cue["start"])
|
||||
patched["end"] = float(cue["end"])
|
||||
# Karaoke burn-in: persisted word times live on the original timeline;
|
||||
# scale them linearly onto the fitted cue span so the highlight sweep
|
||||
# follows the retimed audio. Degenerate spans drop the words — export
|
||||
# then falls back to an even split over the fitted span. Inert for
|
||||
# SRT/VTT, which never read ``words``.
|
||||
if isinstance(seg.get("words"), list) and seg.get("words"):
|
||||
scaled = scale_words(
|
||||
seg["words"], seg.get("start", 0.0), seg.get("end", 0.0),
|
||||
patched["start"], patched["end"],
|
||||
)
|
||||
if scaled is not None:
|
||||
patched["words"] = scaled
|
||||
else:
|
||||
patched.pop("words", None)
|
||||
out.append(patched)
|
||||
return out
|
||||
|
||||
@@ -577,6 +613,7 @@ async def dub_download(
|
||||
save_authorization: str = Header("", alias="X-VoiceStudio-Path-Authorization"),
|
||||
burn_subs: bool = Query(False, description="Burn subtitles into the video stream (forces re-encode). Uses dual-subtitle layout when dual=1."),
|
||||
dual: bool = Query(False, description="When burn_subs=1, render translated on top of italicised original."),
|
||||
karaoke: bool = Query(False, description="When burn_subs=1, burn a word-timed karaoke highlight (ASS) instead of line subtitles. Ignored when dual=1 (dual karaoke is unsupported — the line burn renders instead)."),
|
||||
out_format: str = Query("m4a", description="Audio-only jobs (#119): output container — wav, m4a, mp3, or flac. Ignored for video jobs."),
|
||||
):
|
||||
# Strict allowlist on the path param BEFORE it reaches any filesystem
|
||||
@@ -729,7 +766,18 @@ async def dub_download(
|
||||
fitted_segments = _fitted_segments_for(job, default_track) if default_track and default_track != "original" else None
|
||||
# Burn the DEFAULT track's text (P1.2) — it's the audio the viewer hears.
|
||||
_burn_lang = default_track if default_track and default_track != "original" else None
|
||||
sub_path = _write_burn_srt(job, exports_dir, stamp, dual, fitted_segments=fitted_segments, lang=_burn_lang) if burn_subs else None
|
||||
# Karaoke (word-highlight) burn writes an ASS instead of the line SRT.
|
||||
# Dual layout keeps the line burn — dual karaoke is out of scope, matching
|
||||
# the disabled control in the Export drawer. The default (karaoke off)
|
||||
# takes exactly the legacy SRT path.
|
||||
sub_path = None
|
||||
sub_is_ass = False
|
||||
if burn_subs:
|
||||
if karaoke and not dual:
|
||||
sub_path = _write_burn_ass(job, exports_dir, stamp, fitted_segments=fitted_segments, lang=_burn_lang)
|
||||
sub_is_ass = sub_path is not None
|
||||
if sub_path is None:
|
||||
sub_path = _write_burn_srt(job, exports_dir, stamp, dual, fitted_segments=fitted_segments, lang=_burn_lang)
|
||||
|
||||
# ── Smart Fit video retime (two-tier) ─────────────────────────────────
|
||||
# Tier 1 (≤48 chunks): single filter_complex graph inlined into the mux
|
||||
@@ -829,14 +877,16 @@ async def dub_download(
|
||||
esc = _ffmpeg_filter_escape(sub_path)
|
||||
# Burn AFTER any retime so cues (already on the fitted timeline for
|
||||
# Smart Fit) land on the retimed video. Without retime this reduces
|
||||
# to the legacy `[0:v]subtitles=…[vsub]` graph.
|
||||
# to the legacy `[0:v]subtitles=…[vsub]` graph. Karaoke burns the
|
||||
# word-timed ASS through the ass filter at the same graph position.
|
||||
if video_map.startswith("["):
|
||||
sub_src = video_map
|
||||
elif retimed_idx is not None:
|
||||
sub_src = f"[{retimed_idx}:v]"
|
||||
else:
|
||||
sub_src = "[0:v]"
|
||||
filter_parts.append(f"{sub_src}subtitles='{esc}'[vsub]")
|
||||
_sub_filter = "ass" if sub_is_ass else "subtitles"
|
||||
filter_parts.append(f"{sub_src}{_sub_filter}='{esc}'[vsub]")
|
||||
video_map = "[vsub]"
|
||||
if stretch_entry:
|
||||
orig_dur = float(stretch_entry.get("orig_duration") or job.get("duration") or 0.0)
|
||||
@@ -1744,6 +1794,50 @@ async def dub_export_vtt(
|
||||
)
|
||||
|
||||
|
||||
@router.get("/dub/ass/{job_id}")
|
||||
@router.get("/dub/ass/{job_id}/{filename}")
|
||||
async def dub_export_ass(
|
||||
job_id: str,
|
||||
lang: str = Query(None, description="Track language code. Same text/timing resolution as /dub/srt, rendered as a karaoke (word-highlight) ASS sidecar."),
|
||||
):
|
||||
"""Karaoke ASS sidecar — the same script the karaoke burn-in renders.
|
||||
|
||||
Raw text body like /dub/srt and /dub/vtt (the Tauri side writes the file
|
||||
itself; no ?save_path= variant — see the comment above /dub/srt).
|
||||
"""
|
||||
_job_dir_or_400(job_id)
|
||||
lang = _safe_lang_or_400(lang)
|
||||
job = _get_job(job_id)
|
||||
if not job:
|
||||
raise HTTPException(status_code=404, detail="Job not found")
|
||||
|
||||
segments = _segments_for_lang(job, lang)
|
||||
if not segments:
|
||||
raise HTTPException(status_code=400, detail="No transcript segments available")
|
||||
|
||||
# Same strategy-aware cue timing as /dub/srt. The fitted overlay also
|
||||
# scales word times; the stretch_video cue path has no per-word record,
|
||||
# so words are dropped and build_ass even-splits over the new spans.
|
||||
fitted = _fitted_segments_for(job, lang)
|
||||
if fitted:
|
||||
segments = _apply_fitted_times(segments, fitted)
|
||||
else:
|
||||
cues = _fitted_cue_times(job, lang)
|
||||
if cues:
|
||||
segments = [
|
||||
{**{k: v for k, v in seg.items() if k != "words"}, "start": s, "end": e}
|
||||
for seg, (s, e) in zip(segments, cues)
|
||||
]
|
||||
|
||||
base_name = os.path.splitext(job.get('filename', 'video'))[0]
|
||||
dl_name = f"subtitles_{base_name}_karaoke.ass"
|
||||
return Response(
|
||||
content=build_ass(segments),
|
||||
media_type="text/plain",
|
||||
headers={"Content-Disposition": content_disposition(dl_name)},
|
||||
)
|
||||
|
||||
|
||||
@router.get("/dub/export-segments/{job_id}")
|
||||
async def dub_export_segments_zip(job_id: str, lang: str = Query(None)):
|
||||
import zipfile
|
||||
|
||||
@@ -42,10 +42,26 @@ _FAMILIES = {
|
||||
}
|
||||
|
||||
|
||||
def _catalogue_active_id(family: str, module) -> str:
|
||||
"""Return the active id represented by the public engine catalogue."""
|
||||
active = module.active_backend_id()
|
||||
if family != "tts" or active != "omnivoice-subprocess":
|
||||
return active
|
||||
|
||||
from core.device_caps import detect_host_caps
|
||||
|
||||
try:
|
||||
return "omnivoice" if detect_host_caps().family == "mps" else active
|
||||
except Exception:
|
||||
return active
|
||||
|
||||
|
||||
def _family_payload(family: str, module):
|
||||
"""Public inventory plus whether an environment pin owns this family."""
|
||||
return {
|
||||
"active": module.active_backend_id(),
|
||||
# MPS hides the explicit compatibility row, so legacy configs report
|
||||
# the visible canonical equivalent as active to picker consumers.
|
||||
"active": _catalogue_active_id(family, module),
|
||||
"env_override": bool(os.environ.get(f"OMNIVOICE_{family.upper()}_BACKEND")),
|
||||
"backends": public_backends(module.list_backends()),
|
||||
}
|
||||
@@ -188,6 +204,11 @@ async def uninstall_translation_engine(engine_id: str):
|
||||
pkg = entry.get("pip_package")
|
||||
if not pkg:
|
||||
return {"status": "no_op", "engine": engine_id}
|
||||
# The builtin flag is a promise someone has to remember to make; this
|
||||
# check does not depend on it (#2019).
|
||||
blocked = translation_engines.uninstall_blocker(engine_id)
|
||||
if blocked:
|
||||
raise HTTPException(status_code=blocked[0], detail=blocked[1])
|
||||
rc, out = await translation_engines.run_pip(["uninstall", "-y", pkg])
|
||||
if rc != 0:
|
||||
raise HTTPException(status_code=500, detail=f"pip uninstall {pkg} failed ({rc}): {out[-1000:]}")
|
||||
@@ -230,6 +251,11 @@ def install_sidecar_engine(engine_id: str):
|
||||
from services import sidecar_install
|
||||
try:
|
||||
return sidecar_install.start_install(engine_id)
|
||||
except sidecar_install.HostUnsupported as exc:
|
||||
# The engine has an installer, but not one that can work on this
|
||||
# machine. 409, not 404: the route is right, the host is the problem,
|
||||
# and the message (a VoiceStudio-owned sentence) says what to do.
|
||||
raise HTTPException(status_code=409, detail=str(exc))
|
||||
except KeyError:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
@@ -354,6 +380,9 @@ def engine_health(engine_id: str):
|
||||
)
|
||||
|
||||
t0 = perf_counter()
|
||||
# Stable exception class when the probe itself raised, None when it merely
|
||||
# returned not-available. Never the exception text — see the log line below.
|
||||
raised_class: str | None = None
|
||||
if hasattr(cls, "health_check"):
|
||||
# SubprocessBackend path — spawn sidecar (if not running) and ping.
|
||||
# ``health_check`` already swallows its own exceptions per Plan
|
||||
@@ -364,6 +393,7 @@ def engine_health(engine_id: str):
|
||||
ok, msg = instance.health_check()
|
||||
except Exception as exc:
|
||||
ok, msg = False, f"{type(exc).__name__}: {exc}"
|
||||
raised_class = type(exc).__name__
|
||||
else:
|
||||
# In-process backend — `is_available()` is the classmethod-level
|
||||
# liveness check. Cheap and side-effect-free for every shipping
|
||||
@@ -372,6 +402,7 @@ def engine_health(engine_id: str):
|
||||
ok, msg = cls.is_available()
|
||||
except Exception as exc:
|
||||
ok, msg = False, f"{type(exc).__name__}: {exc}"
|
||||
raised_class = type(exc).__name__
|
||||
|
||||
# Engine-owned output can contain much more than shaped HF tokens: local
|
||||
# paths, arbitrary credentials, source lines, or a nested traceback.
|
||||
@@ -379,7 +410,38 @@ def engine_health(engine_id: str):
|
||||
|
||||
latency_ms = (perf_counter() - t0) * 1000.0
|
||||
if not ok:
|
||||
logger.warning("Engine health check failed; details withheld")
|
||||
# The response tells the user to "check the backend log for details",
|
||||
# and docs/engines/*.md asks a user diagnosing an unavailable engine to
|
||||
# copy that engine's log lines. The old line named neither the engine
|
||||
# nor anything about the probe, so neither instruction could be
|
||||
# followed (#1866).
|
||||
#
|
||||
# `probe=` reports what the PROBE DID, not what went wrong. It cannot
|
||||
# classify the cause: SubprocessBackend.health_check() swallows its own
|
||||
# exceptions per Plan 02-01's contract, so a dead sidecar and a package
|
||||
# that was never installed both arrive here as `returned-unavailable`.
|
||||
# Separating those needs structured failure metadata from the probes
|
||||
# themselves, which is a wider change than this one.
|
||||
#
|
||||
# Still no diagnostic text and still not the caller-supplied id: the
|
||||
# engine id comes off the resolved registry class and a raised probe
|
||||
# contributes only its exception class, the same shape
|
||||
# core.public_errors.public_failure() logs as `class=`.
|
||||
# tests/test_response_safety.py pins that boundary and passes
|
||||
# unchanged.
|
||||
#
|
||||
# The id is a class attribute off the registry rather than caller
|
||||
# input, but this line is a log-injection surface either way, so it is
|
||||
# flattened to a single token before it goes in.
|
||||
engine_label = str(getattr(cls, "id", None) or cls.__name__)
|
||||
engine_label = "".join(
|
||||
c if (c.isalnum() or c in "-_.") else "-" for c in engine_label
|
||||
)[:64]
|
||||
logger.warning(
|
||||
"Engine health check failed; engine=%s probe=%s, details withheld",
|
||||
engine_label or "unknown",
|
||||
f"raised:{raised_class}" if raised_class else "returned-unavailable",
|
||||
)
|
||||
return {
|
||||
"id": engine_id,
|
||||
"ok": bool(ok),
|
||||
@@ -589,7 +651,15 @@ def select_engine(req: SelectEngineRequest):
|
||||
if not family:
|
||||
raise HTTPException(400, f"Unknown family: {req.family}. Expected one of tts/asr/llm.")
|
||||
module, pref_key = family
|
||||
available = {b["id"]: b for b in module.list_backends()}
|
||||
# MPS intentionally hides the redundant explicit OmniVoice sidecar from
|
||||
# the picker, but existing scripts and saved preferences may still submit
|
||||
# that supported compatibility id directly.
|
||||
rows = (
|
||||
module.list_backends(include_hidden=True)
|
||||
if req.family == "tts"
|
||||
else module.list_backends()
|
||||
)
|
||||
available = {b["id"]: b for b in rows}
|
||||
if req.backend_id not in available:
|
||||
raise HTTPException(400, f"Unknown {req.family} backend: {req.backend_id!r}")
|
||||
entry = available[req.backend_id]
|
||||
|
||||
@@ -125,6 +125,96 @@ def _profile_instruct(row):
|
||||
return heal_design_instruct(row["instruct"], vd)
|
||||
|
||||
|
||||
def _resolve_profile_conditioning(row, *, ref_text=None, instruct=None,
|
||||
seed=None, language=None):
|
||||
"""Resolve a ``voice_profiles`` row into generation conditioning.
|
||||
|
||||
Extracted verbatim from /generate's inline profile-resolution block so
|
||||
other synthesis routes (POST /convert) share the exact same semantics —
|
||||
lock wins, ``kind`` is authoritative (0005), legacy pre-0004 rows fall
|
||||
back to the is_locked/instruct inference, and #533's language fill.
|
||||
|
||||
Request-supplied values (``ref_text``/``instruct``/``seed``/``language``)
|
||||
always win over the stored row; only gaps are filled. Returns a dict with
|
||||
``ref_audio_path`` / ``ref_text`` / ``instruct`` / ``seed`` / ``language``
|
||||
/ ``kind`` plus ``persist_ref_text`` — True when the caller should cache
|
||||
an auto-transcribed reference transcript back onto the row (#1032).
|
||||
"""
|
||||
out = {
|
||||
"ref_audio_path": None, "ref_text": ref_text, "instruct": instruct,
|
||||
"seed": seed, "language": language, "kind": None,
|
||||
"persist_ref_text": False,
|
||||
}
|
||||
# `kind` is authoritative (0005): 'design' profiles condition on their
|
||||
# deterministic rendered sample + instruct; 'clone' on the user's
|
||||
# reference. Lock always wins (it pins a specific take). Rows from
|
||||
# pre-0004 DBs mid-upgrade may lack the column → fall back to the legacy
|
||||
# is_locked/instruct inference.
|
||||
try:
|
||||
profile_kind = row["kind"] or "clone"
|
||||
except (KeyError, IndexError):
|
||||
profile_kind = "design" if (
|
||||
row["instruct"] and not row["is_locked"] and not row["ref_audio_path"]
|
||||
) else "clone"
|
||||
out["kind"] = profile_kind
|
||||
if row["is_locked"] and row["locked_audio_path"]:
|
||||
out["ref_audio_path"] = os.path.join(VOICES_DIR, row["locked_audio_path"])
|
||||
if not out["ref_text"]:
|
||||
out["ref_text"] = row["ref_text"]
|
||||
if not out["instruct"]:
|
||||
out["instruct"] = _profile_instruct(row)
|
||||
if out["seed"] is None and row["seed"] is not None:
|
||||
out["seed"] = row["seed"]
|
||||
elif profile_kind == "design":
|
||||
# Rendered sample (if present) carries the voice identity; instruct
|
||||
# alone is the fallback for legacy archetype rows.
|
||||
out["ref_audio_path"] = (
|
||||
os.path.join(VOICES_DIR, row["ref_audio_path"]) if row["ref_audio_path"] else None
|
||||
)
|
||||
if out["ref_audio_path"] and not out["ref_text"] and row["ref_text"]:
|
||||
out["ref_text"] = row["ref_text"]
|
||||
if not out["instruct"]:
|
||||
out["instruct"] = _profile_instruct(row)
|
||||
if out["seed"] is None and row["seed"] is not None:
|
||||
out["seed"] = row["seed"]
|
||||
elif row["instruct"] and not row["is_locked"] and not row["ref_audio_path"]:
|
||||
# Legacy design-shaped row (pre-0004 archetype materialization failure
|
||||
# path): instruct-only conditioning.
|
||||
if not out["instruct"]:
|
||||
out["instruct"] = _profile_instruct(row)
|
||||
if out["seed"] is None and row["seed"] is not None:
|
||||
out["seed"] = row["seed"]
|
||||
else:
|
||||
out["ref_audio_path"] = (
|
||||
os.path.join(VOICES_DIR, row["ref_audio_path"]) if row["ref_audio_path"] else None
|
||||
)
|
||||
if not out["ref_text"] and row["ref_text"]:
|
||||
out["ref_text"] = row["ref_text"]
|
||||
elif out["ref_audio_path"] and not out["ref_text"]:
|
||||
# Empty stored transcript → the caller's auto-transcribe will run;
|
||||
# cache its result onto the profile so it runs ONCE, not on every
|
||||
# generate (#1032 perf regression).
|
||||
out["persist_ref_text"] = True
|
||||
if not out["instruct"] and row["instruct"]:
|
||||
out["instruct"] = row["instruct"]
|
||||
if out["seed"] is None and row["seed"] is not None:
|
||||
out["seed"] = row["seed"]
|
||||
if out["language"] == "Auto":
|
||||
out["language"] = None
|
||||
# #533: a profile's stored language must drive generation when the request
|
||||
# didn't pin one. An EXPLICIT non-Auto request language still wins; we
|
||||
# only fill the gap. `row` is a sqlite3.Row, so guard the column lookup
|
||||
# for pre-language DBs mid-upgrade.
|
||||
if out["language"] is None:
|
||||
try:
|
||||
prof_lang = row["language"]
|
||||
except (KeyError, IndexError):
|
||||
prof_lang = None
|
||||
if prof_lang and prof_lang != "Auto":
|
||||
out["language"] = prof_lang
|
||||
return out
|
||||
|
||||
|
||||
def _note_generate_progress() -> None:
|
||||
"""Tell the pool guard this render just finished a unit of work (#1391).
|
||||
|
||||
@@ -714,16 +804,34 @@ def _oom_friendly_reraise(e):
|
||||
) from e
|
||||
|
||||
|
||||
def _generate_timeout_s(text: str, *, execution_device=None) -> float:
|
||||
def _generate_timeout_s(
|
||||
text: str,
|
||||
*,
|
||||
execution_device=None,
|
||||
min_vram_gb=0.0,
|
||||
hardware_family=None,
|
||||
vram_gb=None,
|
||||
) -> float:
|
||||
"""Wall-clock budget for one generate, scaled to the request.
|
||||
|
||||
Thin alias for the canonical helper, which moved to
|
||||
``services.model_manager.generate_timeout_s`` (#1190) so /v1/audio/speech,
|
||||
batch, dub and archetype previews share it instead of each re-deriving (or,
|
||||
as they did, silently keeping the flat 300s).
|
||||
|
||||
``min_vram_gb`` is the engine's declared VRAM floor. A GPU below it pages to
|
||||
system RAM and renders slower than this machine's CPU, so it must not be
|
||||
budgeted as fast hardware (#1804) — the same figure the dispatch already
|
||||
hands the guard so a timeout message can name the card (#1226/#1222).
|
||||
"""
|
||||
from services.model_manager import generate_timeout_s
|
||||
return generate_timeout_s(text, execution_device=execution_device)
|
||||
return generate_timeout_s(
|
||||
text,
|
||||
execution_device=execution_device,
|
||||
min_vram_gb=min_vram_gb,
|
||||
hardware_family=hardware_family,
|
||||
vram_gb=vram_gb,
|
||||
)
|
||||
|
||||
|
||||
def _run_inference(
|
||||
@@ -1344,6 +1452,8 @@ async def generate_speech(
|
||||
# local fallback call's timeout device-neutral so the closure is valid
|
||||
# without pretending the control plane describes the remote worker.
|
||||
_routing = {"effective_device": None}
|
||||
_routing_hardware_family = None
|
||||
_routing_vram_gb = None
|
||||
|
||||
if not _remote:
|
||||
# Single-active-engine memory discipline: hand back any OTHER resident
|
||||
@@ -1400,11 +1510,16 @@ async def generate_speech(
|
||||
# 4090 from a Mac control plane would be refused by a gate describing
|
||||
# a machine that is about to do nothing.
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_routing import resolve_routing, routing_notice
|
||||
_routing = resolve_routing(
|
||||
getattr(backend_cls, "gpu_compat", ("cpu",)), detect_host_caps(),
|
||||
_engine_min_vram_gb,
|
||||
from services.engine_routing import (
|
||||
routing_notice,
|
||||
runtime_compute_profile_async,
|
||||
)
|
||||
_routing = await runtime_compute_profile_async(
|
||||
backend_cls, detect_host_caps()
|
||||
)
|
||||
_engine_min_vram_gb = _routing["min_vram_gb"]
|
||||
_routing_hardware_family = _routing.get("runtime_hardware_family")
|
||||
_routing_vram_gb = _routing.get("runtime_vram_gb")
|
||||
if _routing["routing_status"] == "unavailable":
|
||||
# The engine needs an accelerator this host lacks and has no CPU path.
|
||||
raise HTTPException(status_code=400, detail=_routing["routing_reason"])
|
||||
@@ -1461,70 +1576,21 @@ async def generate_speech(
|
||||
row = conn.execute("SELECT * FROM voice_profiles WHERE id=?", (profile_id,)).fetchone()
|
||||
if row:
|
||||
resolved_profile_id = profile_id
|
||||
# `kind` is authoritative (0005): 'design' profiles condition on
|
||||
# their deterministic rendered sample + instruct; 'clone' on the
|
||||
# user's reference. Lock always wins (it pins a specific take).
|
||||
# Rows from pre-0004 DBs mid-upgrade may lack the column → fall
|
||||
# back to the legacy is_locked/instruct inference.
|
||||
try:
|
||||
profile_kind = row["kind"] or "clone"
|
||||
except (KeyError, IndexError):
|
||||
profile_kind = "design" if (row["instruct"] and not row["is_locked"] and not row["ref_audio_path"]) else "clone"
|
||||
history_mode = profile_kind
|
||||
if row["is_locked"] and row["locked_audio_path"]:
|
||||
ref_audio_path = os.path.join(VOICES_DIR, row["locked_audio_path"])
|
||||
if not ref_text:
|
||||
ref_text = row["ref_text"]
|
||||
if not instruct:
|
||||
instruct = _profile_instruct(row)
|
||||
if used_seed is None and row["seed"] is not None:
|
||||
used_seed = row["seed"]
|
||||
elif profile_kind == "design":
|
||||
# Rendered sample (if present) carries the voice identity;
|
||||
# instruct alone is the fallback for legacy archetype rows.
|
||||
ref_audio_path = os.path.join(VOICES_DIR, row["ref_audio_path"]) if row["ref_audio_path"] else None
|
||||
if ref_audio_path and not ref_text and row["ref_text"]:
|
||||
ref_text = row["ref_text"]
|
||||
if not instruct:
|
||||
instruct = _profile_instruct(row)
|
||||
if used_seed is None and row["seed"] is not None:
|
||||
used_seed = row["seed"]
|
||||
elif row["instruct"] and not row["is_locked"] and not row["ref_audio_path"]:
|
||||
# Legacy design-shaped row (pre-0004 archetype materialization
|
||||
# failure path): instruct-only conditioning.
|
||||
if not instruct:
|
||||
instruct = _profile_instruct(row)
|
||||
if used_seed is None and row["seed"] is not None:
|
||||
used_seed = row["seed"]
|
||||
else:
|
||||
ref_audio_path = os.path.join(VOICES_DIR, row["ref_audio_path"]) if row["ref_audio_path"] else None
|
||||
if not ref_text and row["ref_text"]:
|
||||
ref_text = row["ref_text"]
|
||||
elif ref_audio_path and not ref_text:
|
||||
# Empty stored transcript → the auto-transcribe below will
|
||||
# run; cache its result onto the profile so it runs ONCE,
|
||||
# not on every generate (#1032 perf regression).
|
||||
persist_ref_text_profile_id = profile_id
|
||||
if not instruct and row["instruct"]:
|
||||
instruct = row["instruct"]
|
||||
if used_seed is None and row["seed"] is not None:
|
||||
used_seed = row["seed"]
|
||||
if language == "Auto":
|
||||
language = None
|
||||
# #533: a profile's stored language must drive generation when the
|
||||
# request didn't pin one. Without this the German (etc.) archetype
|
||||
# generates with language=None and the model drifts to English —
|
||||
# even though the archetype PREVIEW renders correctly (archetypes.py
|
||||
# passes the language). An EXPLICIT non-Auto request language still
|
||||
# wins; we only fill the gap. `row` is a sqlite3.Row, so guard the
|
||||
# column lookup for pre-language DBs mid-upgrade.
|
||||
if language is None:
|
||||
try:
|
||||
prof_lang = row["language"]
|
||||
except (KeyError, IndexError):
|
||||
prof_lang = None
|
||||
if prof_lang and prof_lang != "Auto":
|
||||
language = prof_lang
|
||||
# Shared with POST /convert — see _resolve_profile_conditioning
|
||||
# for the resolution rules (kind-authoritative, lock wins, #533
|
||||
# language fill, #1032 transcript-cache signal).
|
||||
_cond = _resolve_profile_conditioning(
|
||||
row, ref_text=ref_text, instruct=instruct, seed=used_seed,
|
||||
language=language,
|
||||
)
|
||||
history_mode = _cond["kind"]
|
||||
ref_audio_path = _cond["ref_audio_path"]
|
||||
ref_text = _cond["ref_text"]
|
||||
instruct = _cond["instruct"]
|
||||
used_seed = _cond["seed"]
|
||||
language = _cond["language"]
|
||||
if _cond["persist_ref_text"]:
|
||||
persist_ref_text_profile_id = profile_id
|
||||
elif ref_audio is not None:
|
||||
try:
|
||||
with tempfile.NamedTemporaryFile(delete=False, suffix=".wav") as f:
|
||||
@@ -1677,7 +1743,13 @@ async def generate_speech(
|
||||
local=gpu_gateway.LocalCall(
|
||||
_remote_only_local_call(_target_label),
|
||||
what="TTS generate",
|
||||
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
|
||||
timeout=_generate_timeout_s(
|
||||
text,
|
||||
execution_device=_routing["effective_device"],
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
hardware_family=_routing_hardware_family,
|
||||
vram_gb=_routing_vram_gb,
|
||||
),
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
),
|
||||
remote=_remote_call,
|
||||
@@ -1971,7 +2043,13 @@ async def generate_speech(
|
||||
),
|
||||
what="TTS generate",
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
|
||||
timeout=_generate_timeout_s(
|
||||
text,
|
||||
execution_device=_routing["effective_device"],
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
hardware_family=_routing_hardware_family,
|
||||
vram_gb=_routing_vram_gb,
|
||||
),
|
||||
on_abandon=release,
|
||||
)
|
||||
)
|
||||
@@ -1991,7 +2069,13 @@ async def generate_speech(
|
||||
),
|
||||
what="TTS generate",
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
|
||||
timeout=_generate_timeout_s(
|
||||
text,
|
||||
execution_device=_routing["effective_device"],
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
hardware_family=_routing_hardware_family,
|
||||
vram_gb=_routing_vram_gb,
|
||||
),
|
||||
on_abandon=release,
|
||||
)
|
||||
)
|
||||
@@ -2031,7 +2115,13 @@ async def generate_speech(
|
||||
# Budget scaled to THIS chunk (#1190) — the flat
|
||||
# 300s here is what made long streamed renders fail
|
||||
# even after the v0.3.22 scaled budget shipped.
|
||||
timeout=_generate_timeout_s(chunk_text, execution_device=_routing["effective_device"]),
|
||||
timeout=_generate_timeout_s(
|
||||
chunk_text,
|
||||
execution_device=_routing["effective_device"],
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
hardware_family=_routing_hardware_family,
|
||||
vram_gb=_routing_vram_gb,
|
||||
),
|
||||
on_abandon=release,
|
||||
)
|
||||
)
|
||||
@@ -2191,7 +2281,13 @@ async def generate_speech(
|
||||
_REMOTE_OP,
|
||||
local=gpu_gateway.LocalCall(
|
||||
_local_render, what="TTS generate",
|
||||
timeout=_generate_timeout_s(text, execution_device=_routing["effective_device"]),
|
||||
timeout=_generate_timeout_s(
|
||||
text,
|
||||
execution_device=_routing["effective_device"],
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
hardware_family=_routing_hardware_family,
|
||||
vram_gb=_routing_vram_gb,
|
||||
),
|
||||
min_vram_gb=_engine_min_vram_gb,
|
||||
on_abandon=release,
|
||||
),
|
||||
@@ -2316,7 +2412,16 @@ async def generate_speech(
|
||||
raise HTTPException(status_code=503, detail=str(e)) from e
|
||||
except ValueError as e:
|
||||
logger.error("Validation failed: %s", e)
|
||||
raise HTTPException(status_code=400, detail=str(e)) from e
|
||||
# Most ValueErrors here are VoiceStudio's own validation messages and
|
||||
# are exactly what the user should read. A few are raw library text
|
||||
# naming parameters and files the user cannot act on — those get the
|
||||
# owned remedy for their class instead (#1879). Unclassified ones keep
|
||||
# passing through, so this cannot swallow a good message.
|
||||
from core.failure import classify, public_hint_for_topic
|
||||
|
||||
_topic = classify(str(e))
|
||||
_owned = public_hint_for_topic(_topic) if _topic else ""
|
||||
raise HTTPException(status_code=400, detail=_owned or str(e)) from e
|
||||
except Exception as e:
|
||||
tb = traceback.format_exc()
|
||||
logger.error("Inference failed: %s\n%s", e, tb)
|
||||
|
||||
@@ -325,10 +325,9 @@ async def create_speech(req: SpeechRequest):
|
||||
|
||||
# Routing gate (#21 — no silent CPU fallback), identical to REST /generate.
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_routing import resolve_routing, routing_notice
|
||||
_routing = resolve_routing(
|
||||
getattr(backend, "gpu_compat", ("cpu",)), detect_host_caps(),
|
||||
getattr(backend, "min_vram_gb", 0.0),
|
||||
from services.engine_routing import routing_notice, runtime_compute_profile_async
|
||||
_routing = await runtime_compute_profile_async(
|
||||
backend, detect_host_caps()
|
||||
)
|
||||
if _routing["routing_status"] == "unavailable":
|
||||
raise HTTPException(status_code=400, detail=_routing["routing_reason"])
|
||||
|
||||
@@ -32,7 +32,11 @@ from pydantic import BaseModel
|
||||
|
||||
from api.dependencies import require_admin
|
||||
from core.db import db_conn
|
||||
from services.pronunciation import apply_pronunciation, entries_for_language
|
||||
from services.pronunciation import (
|
||||
apply_pronunciation,
|
||||
entries_for_language,
|
||||
inert_entries_for_language,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("omnivoice.pronunciation")
|
||||
router = APIRouter(dependencies=[Depends(require_admin)])
|
||||
@@ -250,11 +254,18 @@ def test_substitution(req: PronTestRequest):
|
||||
).fetchall()
|
||||
substituted = apply_pronunciation(req.text, rows, req.language)
|
||||
applied = entries_for_language(rows, req.language)
|
||||
# IPA/CMU rows are validated and stored but not applied yet, so a term that
|
||||
# DOES match can still change nothing. Reporting them separately keeps the
|
||||
# dry run honest — otherwise it says "no entries match", which is wrong and
|
||||
# sends the user to re-type an entry that was already correct (#1949).
|
||||
inert = inert_entries_for_language(rows, req.language)
|
||||
return {
|
||||
"input": req.text,
|
||||
"substituted": substituted,
|
||||
"changed": substituted != req.text,
|
||||
"applied_terms": sorted(applied.keys(), key=len, reverse=True),
|
||||
# Present but not honoured: [{term, type}, …]. Empty on the happy path.
|
||||
"inert_entries": inert,
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -35,11 +35,11 @@ class _HFTokenBody(BaseModel):
|
||||
token: str = Field(..., min_length=1, description="HuggingFace access token")
|
||||
|
||||
|
||||
def _state_response() -> dict:
|
||||
def _state_response(*, validate: bool = False) -> dict:
|
||||
"""Return the same shape the React panel renders. Never includes raw token."""
|
||||
from services import token_resolver
|
||||
|
||||
s = token_resolver.state()
|
||||
s = token_resolver.state(validate=validate)
|
||||
return {
|
||||
"active": s["active"],
|
||||
"sources": [asdict(row) for row in s["sources"]],
|
||||
@@ -65,8 +65,7 @@ def save_hf_token(body: _HFTokenBody):
|
||||
|
||||
@router.delete("/hf-token")
|
||||
def clear_hf_token(also_clear_hf_cli: bool = Query(False)):
|
||||
"""Clear the App-source token. Optionally also call huggingface_hub.logout
|
||||
to clear the canonical HF file. Returns the updated cascade state."""
|
||||
"""Clear the App token and optionally recognized local Hub token files."""
|
||||
from services import token_resolver
|
||||
try:
|
||||
token_resolver.clear_app_token(also_clear_hf_cli=also_clear_hf_cli)
|
||||
@@ -82,13 +81,12 @@ def get_hf_token_state(fresh: bool = Query(False)):
|
||||
|
||||
``fresh=1`` drops the resolver's whoami validation cache first so the
|
||||
response re-runs whoami for every source — this is what the panel's
|
||||
"Test now" button sends. Plain GETs (panel mounts) keep the 300s cache
|
||||
so repeat Settings visits don't hammer the HF API.
|
||||
"Test now" button sends. Plain GETs only inspect local token presence.
|
||||
"""
|
||||
from services import token_resolver
|
||||
if fresh:
|
||||
token_resolver.invalidate_cache()
|
||||
return _state_response()
|
||||
return _state_response(validate=fresh)
|
||||
|
||||
|
||||
# ── Performance settings (INST-12) ────────────────────────────────────────
|
||||
@@ -150,7 +148,7 @@ def _compute_device_state() -> dict:
|
||||
caps = device_caps.detect_host_caps()
|
||||
env_pin = (os.environ.get("OMNIVOICE_DEVICE") or "").strip().lower()
|
||||
auto_family = next(
|
||||
(f for f in ("cuda", "rocm", "xpu", "mps") if f in caps.available_families),
|
||||
(f for f in device_caps.ACCELERATOR_PRIORITY if f in caps.available_families),
|
||||
"cpu",
|
||||
)
|
||||
value = device_caps.requested_device_override()
|
||||
@@ -1035,9 +1033,10 @@ def set_asr_openai_compat(body: _ASROpenAICompatBody):
|
||||
from services import asr_backend, settings_store
|
||||
|
||||
if body.base_url is not None:
|
||||
url = body.base_url.strip().rstrip("/")
|
||||
if url and not url.startswith(("http://", "https://")):
|
||||
raise HTTPException(status_code=400, detail="Base URL must start with http(s)://")
|
||||
try:
|
||||
url = asr_backend.normalize_openai_compat_asr_base_url(body.base_url)
|
||||
except ValueError as exc:
|
||||
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||
settings_store.set_text(asr_backend._ASR_OPENAI_COMPAT_BASE_URL_KEY, url)
|
||||
if body.model is not None:
|
||||
settings_store.set_text(
|
||||
|
||||
@@ -381,11 +381,60 @@ def _is_retryable_download_error(exc: BaseException) -> bool:
|
||||
return is_hf_connectivity_error(str(exc))
|
||||
|
||||
|
||||
def _segmented_retry_plan(
|
||||
exc: BaseException, attempt: int, max_attempts: int
|
||||
) -> tuple[bool, bool]:
|
||||
"""What to do after the segmented accelerator failed on ``attempt``.
|
||||
|
||||
Returns ``(disable_accelerator, reraise)``.
|
||||
|
||||
A dropped connection is not the accelerator's fault, so the error is
|
||||
re-raised for the outer retry: the next attempt re-enters
|
||||
:func:`_segmented_snapshot`, which resumes from the ``.part`` manifest.
|
||||
Falling straight through to ``snapshot_download`` instead would finish the
|
||||
install from a separate ``.incomplete`` file and strand that manifest — the
|
||||
restart-from-zero this exists to prevent.
|
||||
|
||||
The final attempt is always reserved for the plain path, so the accelerator
|
||||
can never be the reason an install fails outright. The two flags are
|
||||
decoupled for that handover: the attempt that exhausts the accelerator still
|
||||
re-raises, so the plain path starts on the LAST attempt rather than the
|
||||
second-to-last. Disabling and falling through in the same attempt would
|
||||
abandon the resumable manifest one attempt early and restart through a
|
||||
separate file — which is the failure this whole helper exists to avoid.
|
||||
"""
|
||||
if not _is_retryable_download_error(exc):
|
||||
return True, False # the accelerator cannot work here at all
|
||||
if attempt >= max_attempts:
|
||||
# Nothing left to hand over to: take the plain path now rather than
|
||||
# re-raising out of the loop with no fallback ever tried.
|
||||
return True, False
|
||||
return attempt >= max_attempts - 1, True
|
||||
|
||||
|
||||
def _segmented_retry_note(disable: bool, reraise: bool) -> str:
|
||||
"""How to describe the outcome of :func:`_segmented_retry_plan` in the log.
|
||||
|
||||
Three distinct states, and reading only ``disable`` conflates two of them:
|
||||
the attempt that exhausts the accelerator is disabled AND re-raises, so the
|
||||
fallback starts on the NEXT attempt, not this one.
|
||||
"""
|
||||
if not disable:
|
||||
return "kept for the next attempt (resumes from its manifest)"
|
||||
if reraise:
|
||||
return "exhausted — retrying once more, then snapshot_download takes over"
|
||||
return "disabled for this install — falling back to snapshot_download now"
|
||||
|
||||
|
||||
@router.post("/models/install")
|
||||
async def install_model(req: InstallModelRequest):
|
||||
"""Download one HF repo snapshot; progress goes through the shared
|
||||
``/setup/download-stream`` SSE feed."""
|
||||
if req.repo_id not in [m["repo_id"] for m in KNOWN_MODELS]:
|
||||
model_spec = next(
|
||||
(model for model in KNOWN_MODELS if model["repo_id"] == req.repo_id),
|
||||
None,
|
||||
)
|
||||
if model_spec is None:
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail=(
|
||||
@@ -393,6 +442,7 @@ async def install_model(req: InstallModelRequest):
|
||||
+ ", ".join(m["repo_id"] for m in KNOWN_MODELS)
|
||||
),
|
||||
)
|
||||
allow_patterns = list(model_spec.get("allow_patterns") or []) or None
|
||||
target = (req.target or "").strip()
|
||||
if target != "local":
|
||||
from services import gpu_gateway # noqa: PLC0415
|
||||
@@ -450,6 +500,8 @@ async def install_model(req: InstallModelRequest):
|
||||
"revision": revision_for(req.repo_id),
|
||||
"max_workers": _download_max_workers(),
|
||||
}
|
||||
if allow_patterns:
|
||||
dl_kwargs["allow_patterns"] = allow_patterns
|
||||
_tqdm_cls = hf_progress.tracked_tqdm_class()
|
||||
if _tqdm_cls is not None:
|
||||
dl_kwargs["tqdm_class"] = _tqdm_cls
|
||||
@@ -493,6 +545,8 @@ async def install_model(req: InstallModelRequest):
|
||||
"revision": dl_kwargs["revision"],
|
||||
"dry_run": True,
|
||||
}
|
||||
if allow_patterns:
|
||||
_preflight_kwargs["allow_patterns"] = allow_patterns
|
||||
if _endpoint:
|
||||
_preflight_kwargs["endpoint"] = _endpoint
|
||||
try:
|
||||
@@ -548,6 +602,11 @@ async def install_model(req: InstallModelRequest):
|
||||
|
||||
_max_attempts = 5
|
||||
_attempt = 0
|
||||
# The accelerator is retried across attempts so its manifest-based
|
||||
# resume actually gets used; it is disabled for the rest of the
|
||||
# install only when it fails for a reason that is NOT transient
|
||||
# network trouble (i.e. the accelerator itself is unusable here).
|
||||
_segmented_off = False
|
||||
while True:
|
||||
if req.repo_id in _cancelled:
|
||||
raise _InstallCancelled()
|
||||
@@ -555,11 +614,17 @@ async def install_model(req: InstallModelRequest):
|
||||
try:
|
||||
# Segmented accelerator (FDL-09, default ON): parallel
|
||||
# byte-range fetch with real live progress, for the
|
||||
# legacy-LFS path. Any failure falls through to
|
||||
# snapshot_download — the accelerator can never compromise a
|
||||
# correct install.
|
||||
# legacy-LFS path. A failure that is not transient network
|
||||
# trouble falls through to snapshot_download, and so does the
|
||||
# install's last attempt — the accelerator can never
|
||||
# compromise a correct install (see _segmented_retry_plan).
|
||||
_snapshot_path = None
|
||||
if _attempt == 1 and _segmented_enabled() and not _xet_active():
|
||||
if (
|
||||
not _segmented_off
|
||||
and not allow_patterns
|
||||
and _segmented_enabled()
|
||||
and not _xet_active()
|
||||
):
|
||||
try:
|
||||
_snapshot_path = _segmented_snapshot(
|
||||
req.repo_id,
|
||||
@@ -569,10 +634,16 @@ async def install_model(req: InstallModelRequest):
|
||||
except _InstallCancelled:
|
||||
raise
|
||||
except Exception as _seg_err:
|
||||
logger.info(
|
||||
"segmented download for %s failed (%s); falling back to snapshot_download",
|
||||
req.repo_id, _seg_err,
|
||||
_segmented_off, _seg_reraise = _segmented_retry_plan(
|
||||
_seg_err, _attempt, _max_attempts
|
||||
)
|
||||
logger.info(
|
||||
"segmented download for %s failed (%s); accelerator %s",
|
||||
req.repo_id, _seg_err,
|
||||
_segmented_retry_note(_segmented_off, _seg_reraise),
|
||||
)
|
||||
if _seg_reraise:
|
||||
raise
|
||||
_snapshot_path = None
|
||||
if _snapshot_path is None:
|
||||
_snapshot_path = snapshot_download(**dl_kwargs) # nosec B615 -- immutable revision_for pin
|
||||
|
||||
@@ -19,6 +19,7 @@ import sys
|
||||
from fastapi import APIRouter
|
||||
|
||||
from api.schemas import SetupStatusResponse, PreflightResponse
|
||||
from core.device_caps import KERNEL_RISK_MARKER
|
||||
# MIN_FREE_GB + disk_free_bytes are single-sourced in ``.models`` (the lowest
|
||||
# module in the setup import graph) so the wizard gate, the /models header, and
|
||||
# the per-install disk guard can't drift apart.
|
||||
@@ -174,8 +175,8 @@ def _detect_gpu() -> dict:
|
||||
return info
|
||||
|
||||
|
||||
def _probe_network(host: str = "huggingface.co", port: int = 443, timeout: float = 2.0) -> bool:
|
||||
"""Tiny TCP connect test."""
|
||||
def _probe_network(host: str = "huggingface.co", port: int = 443, timeout: float = 8.0) -> bool:
|
||||
"""Tiny TCP connect test. 8s default — high-latency / China paths often exceed 2–3s."""
|
||||
import socket
|
||||
try:
|
||||
with socket.create_connection((host, port), timeout=timeout):
|
||||
@@ -498,10 +499,14 @@ def preflight():
|
||||
_why = gpu_routing.get("routing_reason")
|
||||
if _rs == "accelerated" and not _why:
|
||||
r_status, r_detail, r_fix = "pass", f"{_eng} → {_dev} (accelerated)", None
|
||||
elif _rs == "accelerated": # driver/arch caveat
|
||||
elif _rs == "accelerated" and KERNEL_RISK_MARKER in (_why or ""):
|
||||
r_status, r_detail, r_fix = "warn", f"{_eng} → {_dev}: {_why}", (
|
||||
"GPU selected but may fail at kernel launch — update drivers / "
|
||||
"reinstall torch for this GPU architecture.")
|
||||
elif _rs == "accelerated": # low-VRAM caveat — not a driver/arch issue
|
||||
r_status, r_detail, r_fix = "warn", f"{_eng} → {_dev}: {_why}", (
|
||||
"Unload other models before generating, keep the text short, "
|
||||
"or pick a lighter engine.")
|
||||
elif _rs == "cpu_fallback":
|
||||
r_status, r_detail, r_fix = "warn", (
|
||||
f"{_eng} runs on CPU here: {_why or 'no GPU path for this host'}"), (
|
||||
|
||||
+252
-46
@@ -203,8 +203,22 @@ def system_info():
|
||||
"""
|
||||
try:
|
||||
_ffmpeg = find_ffmpeg()
|
||||
from services import model_manager as _mm
|
||||
from core import prefs as _prefs_mod
|
||||
return {
|
||||
"app_version": APP_VERSION,
|
||||
"generate_timeout_s": _mm.GPU_JOB_TIMEOUT_S,
|
||||
"cpu_generate_timeout_s": _mm.CPU_JOB_TIMEOUT_S,
|
||||
# #1787 review fix: a saved prefs.json value for either key can be
|
||||
# silently shadowed by an external env var (os.environ.setdefault
|
||||
# in core.prefs.restore_env is a no-op when one is already
|
||||
# present) — the Settings panel must say so rather than promise a
|
||||
# restart will apply a value that never will.
|
||||
"generate_timeout_shadowed": _prefs_mod.is_env_shadowed(
|
||||
"OMNIVOICE_GENERATE_TIMEOUT_S"),
|
||||
"cpu_generate_timeout_shadowed": _prefs_mod.is_env_shadowed(
|
||||
"OMNIVOICE_CPU_GENERATE_TIMEOUT_S"),
|
||||
"code_fingerprint": os.environ.get("OMNIVOICE_BUILD_FINGERPRINT", ""),
|
||||
"data_dir": DATA_DIR,
|
||||
"outputs_dir": OUTPUTS_DIR,
|
||||
"crash_log_path": CRASH_LOG_PATH,
|
||||
@@ -240,6 +254,11 @@ def system_info():
|
||||
logger.exception("system_info failed — returning safe defaults")
|
||||
return {
|
||||
"app_version": APP_VERSION,
|
||||
"generate_timeout_s": 300.0,
|
||||
"cpu_generate_timeout_s": 600.0,
|
||||
"generate_timeout_shadowed": False,
|
||||
"cpu_generate_timeout_shadowed": False,
|
||||
"code_fingerprint": os.environ.get("OMNIVOICE_BUILD_FINGERPRINT", ""),
|
||||
"data_dir": DATA_DIR,
|
||||
"outputs_dir": OUTPUTS_DIR,
|
||||
"crash_log_path": str(CRASH_LOG_PATH),
|
||||
@@ -278,6 +297,142 @@ def _tail_file(path: str, tail: int):
|
||||
return all_lines[-tail:], len(all_lines)
|
||||
|
||||
|
||||
# Must track main.py's _WindowsSafeRotatingFileHandler(backupCount=3). The
|
||||
# handler rolls omnivoice.log at 2 MB into .1/.2/.3, so up to 6 MB of history
|
||||
# lives in files this module used to ignore entirely.
|
||||
_LOG_BACKUP_COUNT = 3
|
||||
|
||||
|
||||
def _rotated_log_paths(base: str) -> list[str]:
|
||||
"""Existing `<base>.1 … .N`, newest first."""
|
||||
return [p for p in (f"{base}.{i}" for i in range(1, _LOG_BACKUP_COUNT + 1)) if os.path.exists(p)]
|
||||
|
||||
|
||||
def _tail_rolling(base: str, tail: int):
|
||||
"""Tail `base`, reaching into its rotated siblings when it runs short.
|
||||
|
||||
A rollover leaves omnivoice.log nearly empty, and the Backend tab then
|
||||
showed a handful of lines — or none — while the failure the user was asked
|
||||
to copy sat in omnivoice.log.1. Reading the current file first keeps the
|
||||
common case at one file read; the backups are only touched when they are
|
||||
the only place the requested lines can come from.
|
||||
|
||||
Returns (lines oldest-first, total lines across the files read, paths read
|
||||
oldest-first). The total counts only the files it had to open — it stops as
|
||||
soon as `tail` is satisfied, so it is "how much is behind these lines",
|
||||
not the size of the whole rotation set.
|
||||
"""
|
||||
chunks: list[list[str]] = []
|
||||
paths: list[str] = []
|
||||
total = 0
|
||||
remaining = tail
|
||||
candidates = [p for p in [base, *_rotated_log_paths(base)] if os.path.exists(p)]
|
||||
for path in candidates:
|
||||
if remaining <= 0:
|
||||
break
|
||||
try:
|
||||
lines, count = _tail_file(path, remaining)
|
||||
except FileNotFoundError:
|
||||
# A rollover can rename a candidate between the existence check
|
||||
# above and this open, and the handler holds no lock we can take
|
||||
# from a route. Skip the vanished file rather than 500 the whole
|
||||
# panel over one member of the set — the previous single-file
|
||||
# version failed the request outright in the same situation.
|
||||
#
|
||||
# A roll landing mid-walk can also shift which chunk a file holds,
|
||||
# so a tail taken at that instant may repeat or miss a block. The
|
||||
# panel re-polls every 5s and the next read is clean; buying strict
|
||||
# consistency here would mean reaching into logging's internals.
|
||||
continue
|
||||
except PermissionError as exc:
|
||||
# Windows only, and only the sharing violation: the handler still
|
||||
# holds the file it is rolling. Any other permission failure is a
|
||||
# real misconfiguration and must not be hidden.
|
||||
if os.name == "nt" and getattr(exc, "winerror", None) == 32:
|
||||
continue
|
||||
raise
|
||||
if count == 0:
|
||||
continue
|
||||
chunks.append(lines)
|
||||
paths.append(path)
|
||||
total += count
|
||||
remaining -= len(lines)
|
||||
# Files were visited newest-first; the reader wants oldest-first.
|
||||
out: list[str] = []
|
||||
for chunk in reversed(chunks):
|
||||
out.extend(chunk)
|
||||
return out, total, list(reversed(paths))
|
||||
|
||||
def _tauri_plugin_log_candidates():
|
||||
"""The `tauri-plugin-log` files — the shell's own log, and the only thing
|
||||
the Tauri tab actually displays.
|
||||
|
||||
Split out from :func:`_tauri_log_candidates` so Clear can touch these and
|
||||
leave the backend stdout/stderr redirect alone. See
|
||||
:func:`clear_tauri_logs`.
|
||||
"""
|
||||
home = os.path.expanduser("~")
|
||||
bid = "com.debpalash.omnivoice-studio"
|
||||
if sys.platform == "darwin":
|
||||
return [
|
||||
os.path.join(home, "Library/Logs", bid, "tauri.log"),
|
||||
os.path.join(home, "Library/Logs", bid, "VoiceStudio.log"),
|
||||
]
|
||||
if sys.platform.startswith("linux"):
|
||||
data_dir = os.environ.get("XDG_DATA_HOME") or os.path.join(home, ".local/share")
|
||||
return [
|
||||
os.path.join(data_dir, bid, "logs", "tauri.log"),
|
||||
os.path.join(home, ".config", bid, "logs", "tauri.log"),
|
||||
]
|
||||
if sys.platform.startswith("win"):
|
||||
appdata = os.environ.get("APPDATA", home)
|
||||
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
|
||||
return [
|
||||
os.path.join(localappdata, bid, "logs", "tauri.log"),
|
||||
os.path.join(appdata, bid, "logs", "tauri.log"),
|
||||
]
|
||||
return []
|
||||
|
||||
|
||||
def _backend_redirect_log_candidates():
|
||||
"""`backend.log` / `backend_err.log` — the spawned backend's stdout and
|
||||
stderr, written by `src-tauri/src/backend.rs::backend_log_path()`.
|
||||
|
||||
Deliberately NOT cleared by the Tauri tab's Clear button.
|
||||
`open_err_log_for_run()` opens `backend_err.log` **append-only** so "a
|
||||
respawn must not destroy the previous run's evidence" (#1510), rotates it
|
||||
to `.1` rather than truncating, and its spawn diagnostics are described
|
||||
there as "retained in backend_err.log across runs and lands verbatim in bug
|
||||
reports". A native death (a Windows access violation, a SIGSEGV) writes
|
||||
nothing to the Python log by construction, so this file is the only record
|
||||
of it.
|
||||
|
||||
`OMNIVOICE_LOG_DIR` is honoured first, in the same precedence
|
||||
`backend_log_path()` uses. The backend is a child of the shell, so an
|
||||
ambient override reaches both — and a resolver that ignored it would look
|
||||
in the per-OS default while the writer wrote somewhere else, which is the
|
||||
divergence class this file already has one of (see #1782).
|
||||
"""
|
||||
override = (os.environ.get("OMNIVOICE_LOG_DIR") or "").strip()
|
||||
if override:
|
||||
return [
|
||||
os.path.join(override, "backend.log"),
|
||||
os.path.join(override, "backend_err.log"),
|
||||
]
|
||||
home = os.path.expanduser("~")
|
||||
if sys.platform == "darwin":
|
||||
base = os.path.join(home, "Library/Logs/OmniVoice")
|
||||
elif sys.platform.startswith("linux"):
|
||||
state_dir = os.environ.get("XDG_STATE_HOME") or os.path.join(home, ".local/state")
|
||||
base = os.path.join(state_dir, "OmniVoice")
|
||||
elif sys.platform.startswith("win"):
|
||||
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
|
||||
base = os.path.join(localappdata, "OmniVoice", "Logs")
|
||||
else:
|
||||
return []
|
||||
return [os.path.join(base, "backend.log"), os.path.join(base, "backend_err.log")]
|
||||
|
||||
|
||||
def _tauri_log_candidates():
|
||||
"""Likely paths for Tauri-side logs, most useful first.
|
||||
|
||||
@@ -289,40 +444,15 @@ def _tauri_log_candidates():
|
||||
`com.debpalash.omnivoice-studio` (frontend/src-tauri/tauri.conf.json).
|
||||
- backend.rs::backend_log_path() redirects the spawned backend's
|
||||
stdout/stderr to `backend.log` / `backend_err.log` under
|
||||
`~/Library/Logs/OmniVoice` (macOS), `$XDG_STATE_HOME/VoiceStudio` falling
|
||||
`~/Library/Logs/OmniVoice` (macOS), `$XDG_STATE_HOME/OmniVoice` falling
|
||||
back to `~/.local/state/OmniVoice` (Linux), and
|
||||
`%LOCALAPPDATA%\\OmniVoice\\Logs` (Windows). This is where uvicorn
|
||||
startup banners and hard-crash tracebacks land — keep all three OS
|
||||
shapes listed or sidecar crashes become invisible off-macOS.
|
||||
"""
|
||||
home = os.path.expanduser("~")
|
||||
bid = "com.debpalash.omnivoice-studio"
|
||||
if sys.platform == "darwin":
|
||||
return [
|
||||
os.path.join(home, "Library/Logs", bid, "tauri.log"),
|
||||
os.path.join(home, "Library/Logs", bid, "VoiceStudio.log"),
|
||||
os.path.join(home, "Library/Logs/OmniVoice/backend.log"),
|
||||
os.path.join(home, "Library/Logs/OmniVoice/backend_err.log"),
|
||||
]
|
||||
if sys.platform.startswith("linux"):
|
||||
data_dir = os.environ.get("XDG_DATA_HOME") or os.path.join(home, ".local/share")
|
||||
state_dir = os.environ.get("XDG_STATE_HOME") or os.path.join(home, ".local/state")
|
||||
return [
|
||||
os.path.join(data_dir, bid, "logs", "tauri.log"),
|
||||
os.path.join(home, ".config", bid, "logs", "tauri.log"),
|
||||
os.path.join(state_dir, "OmniVoice", "backend.log"),
|
||||
os.path.join(state_dir, "OmniVoice", "backend_err.log"),
|
||||
]
|
||||
if sys.platform.startswith("win"):
|
||||
appdata = os.environ.get("APPDATA", home)
|
||||
localappdata = os.environ.get("LOCALAPPDATA") or os.path.join(home, "AppData", "Local")
|
||||
return [
|
||||
os.path.join(localappdata, bid, "logs", "tauri.log"),
|
||||
os.path.join(appdata, bid, "logs", "tauri.log"),
|
||||
os.path.join(localappdata, "OmniVoice", "Logs", "backend.log"),
|
||||
os.path.join(localappdata, "OmniVoice", "Logs", "backend_err.log"),
|
||||
]
|
||||
return []
|
||||
# Composed from the two halves so the read path keeps seeing every file
|
||||
# while Clear can be narrowed to the shell's own log.
|
||||
return _tauri_plugin_log_candidates() + _backend_redirect_log_candidates()
|
||||
|
||||
|
||||
@router.get("/system/logs")
|
||||
@@ -337,12 +467,24 @@ async def system_logs(tail: int = 200):
|
||||
except Exception:
|
||||
tail = 200
|
||||
|
||||
path = LOG_PATH if os.path.exists(LOG_PATH) else CRASH_LOG_PATH
|
||||
if not os.path.exists(path):
|
||||
if os.path.exists(LOG_PATH) or _rotated_log_paths(LOG_PATH):
|
||||
base = LOG_PATH
|
||||
else:
|
||||
base = CRASH_LOG_PATH
|
||||
if not os.path.exists(base) and not _rotated_log_paths(base):
|
||||
return {"lines": [], "path": LOG_PATH, "exists": False}
|
||||
path = base
|
||||
try:
|
||||
lines, total = await asyncio.to_thread(_tail_file, path, tail)
|
||||
return {"lines": lines, "path": path, "exists": True, "total_lines": total}
|
||||
lines, total, paths = await asyncio.to_thread(_tail_rolling, base, tail)
|
||||
return {
|
||||
"lines": lines,
|
||||
"path": path,
|
||||
"exists": True,
|
||||
"total_lines": total,
|
||||
# Which files the tail actually came from, oldest first. A bug
|
||||
# report can then say whether it crossed a rollover.
|
||||
"paths": paths,
|
||||
}
|
||||
except Exception as e:
|
||||
raise HTTPException(
|
||||
status_code=500,
|
||||
@@ -448,9 +590,23 @@ def _read_from_pos(path: str, pos: int) -> list[str]:
|
||||
|
||||
@router.post("/system/logs/clear")
|
||||
async def clear_system_logs():
|
||||
"""Truncate the rolling runtime log and the crash log (what the Backend tab reads)."""
|
||||
"""Truncate the rolling runtime log and the crash log (what the Backend tab reads).
|
||||
|
||||
Includes the rotated siblings. Truncating only omnivoice.log left up to
|
||||
6 MB in .1/.2/.3, so Clear freed almost nothing and — now that the tail
|
||||
reaches into those files — would have looked like it did nothing at all.
|
||||
"""
|
||||
cleared_any = False
|
||||
for p in (LOG_PATH, CRASH_LOG_PATH):
|
||||
# The full fixed name set rather than a snapshot of what exists: enumerating
|
||||
# first leaves a window where a rollover creates a backup after the scan and
|
||||
# its history survives a Clear that reported success. Names the handler can
|
||||
# ever write are known up front, so there is nothing to enumerate.
|
||||
targets = [
|
||||
LOG_PATH,
|
||||
*(f"{LOG_PATH}.{i}" for i in range(1, _LOG_BACKUP_COUNT + 1)),
|
||||
CRASH_LOG_PATH,
|
||||
]
|
||||
for p in targets:
|
||||
if os.path.exists(p):
|
||||
try:
|
||||
await asyncio.to_thread(_truncate_file, p)
|
||||
@@ -483,10 +639,20 @@ def _truncate_file(path: str):
|
||||
|
||||
@router.post("/system/logs/tauri/clear")
|
||||
async def clear_tauri_logs():
|
||||
"""Truncate whichever Tauri-side log files we know about. OS-level rotation may recreate them."""
|
||||
"""Truncate the shell's own log files. OS-level rotation may recreate them.
|
||||
|
||||
The backend stdout/stderr redirect is deliberately excluded. This button
|
||||
lives on a tab that shows `tauri.log`, and truncating `backend_err.log`
|
||||
from it destroyed evidence the user was never shown — the one record of a
|
||||
native death, which writes nothing to the Python log. `backend.rs`'s
|
||||
`open_err_log_for_run()` opens that file append-only precisely so "a
|
||||
respawn must not destroy the previous run's evidence" (#1510) and rotates
|
||||
it to `.1` instead of truncating, so it manages its own size and does not
|
||||
need clearing from here.
|
||||
"""
|
||||
cleared = []
|
||||
failed = 0
|
||||
for p in _tauri_log_candidates():
|
||||
for p in _tauri_plugin_log_candidates():
|
||||
if os.path.exists(p):
|
||||
try:
|
||||
await asyncio.to_thread(_truncate_file, p)
|
||||
@@ -848,6 +1014,14 @@ PERSISTENT_KEYS = {
|
||||
# the Rust sidecar reads OMNIVOICE_PORT at startup and the backend derives
|
||||
# the LAN-share/UI ports from the others.
|
||||
"OMNIVOICE_PORT", "OMNIVOICE_SHARE_PORT", "OMNIVOICE_UI_PORT",
|
||||
# Per-job compute-time budgets (#1787). Both are captured at import time
|
||||
# by services/model_manager.py (GPU_JOB_TIMEOUT_S / CPU_JOB_TIMEOUT_S), so
|
||||
# a value saved here takes effect on the NEXT backend restart — same
|
||||
# contract as OMNIVOICE_PORT above. Restored into os.environ during the
|
||||
# "env_prefs" startup step (main.py), which runs before model_manager is
|
||||
# first imported ("ml_imports"), so the restored value is what the module
|
||||
# captures. The Settings UI must say so (RestartBadge).
|
||||
"OMNIVOICE_GENERATE_TIMEOUT_S", "OMNIVOICE_CPU_GENERATE_TIMEOUT_S",
|
||||
}
|
||||
|
||||
# Sidecar-engine install dirs (OMNIVOICE_INDEXTTS_DIR, …). The one-click
|
||||
@@ -865,6 +1039,16 @@ except Exception: # pragma: no cover — defensive: env panel > installer wirin
|
||||
# being set so a bad value never reaches uvicorn / the share listener.
|
||||
_PORT_KEYS = {"OMNIVOICE_PORT", "OMNIVOICE_SHARE_PORT", "OMNIVOICE_UI_PORT"}
|
||||
|
||||
# Keys whose value is a wall-clock compute-time budget in seconds (#1787).
|
||||
# Validated the same way as _PORT_KEYS: reject anything that isn't a
|
||||
# positive number before it reaches services/model_manager.py. Upper bound is
|
||||
# generous — long enough that a legitimate multi-hour, audiobook-length CPU
|
||||
# render is never blocked — but still bounded, so a fat-fingered extra digit
|
||||
# (300 -> 3000000) can't turn a wedged job into one that silently occupies a
|
||||
# worker for days before the guard ever fires.
|
||||
_TIMEOUT_KEYS = {"OMNIVOICE_GENERATE_TIMEOUT_S", "OMNIVOICE_CPU_GENERATE_TIMEOUT_S"}
|
||||
_MAX_GENERATE_TIMEOUT_S = 21600.0 # 6 hours
|
||||
|
||||
|
||||
@router.post("/system/set-env")
|
||||
async def set_env_var(body: dict):
|
||||
@@ -873,7 +1057,7 @@ async def set_env_var(body: dict):
|
||||
Persistent keys (proxy, FFMPEG_PATH, translation provider keys, …) are
|
||||
saved to ``prefs.json`` so they survive backend restarts (restored at
|
||||
startup in ``main.py``). HF_TOKEN is persisted via
|
||||
``huggingface_hub.login()`` (and cleared via ``logout()``). Other keys
|
||||
``huggingface_hub.login()`` (and cleared with the shared token-file helper). Other keys
|
||||
are set on ``os.environ`` for the running process.
|
||||
|
||||
The loopback-origin gate that previously lived inline here is now applied
|
||||
@@ -908,6 +1092,22 @@ async def set_env_var(body: dict):
|
||||
status_code=400,
|
||||
detail=f"Invalid port for {key}: must be between 1024 and 65535.",
|
||||
)
|
||||
if key in _TIMEOUT_KEYS:
|
||||
try:
|
||||
timeout_n = float(value)
|
||||
except (TypeError, ValueError):
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail=f"Invalid timeout for {key}: '{value}' is not a number.",
|
||||
)
|
||||
if not (0 < timeout_n <= _MAX_GENERATE_TIMEOUT_S):
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail=(
|
||||
f"Invalid timeout for {key}: must be greater than 0 "
|
||||
f"and at most {_MAX_GENERATE_TIMEOUT_S:.0f} seconds."
|
||||
),
|
||||
)
|
||||
os.environ[key] = value
|
||||
logger.info("Environment variable set (length=%d)", len(value))
|
||||
|
||||
@@ -932,14 +1132,14 @@ async def set_env_var(body: dict):
|
||||
# Mirror the persistence on clear — wipe the saved token file too.
|
||||
if key == "HF_TOKEN":
|
||||
try:
|
||||
from huggingface_hub import logout as _hf_logout
|
||||
_hf_logout()
|
||||
logger.info("HF token cleared from $HF_HOME/token via logout()")
|
||||
except Exception as e:
|
||||
logger.warning("Could not clear HF token file: %s", e)
|
||||
from services.token_resolver import clear_hf_cli_tokens
|
||||
clear_hf_cli_tokens()
|
||||
logger.info("Local Hugging Face token files cleared")
|
||||
except Exception:
|
||||
raise HTTPException(status_code=500, detail="Could not clear local Hugging Face token files") from None
|
||||
|
||||
# HF_TOKEN persistence is handled above via huggingface_hub.login()/
|
||||
# logout() — it never touches prefs.json. Everything else in
|
||||
# clear_hf_cli_tokens() — it never touches prefs.json. Everything else in
|
||||
# PERSISTENT_KEYS (proxy, FFMPEG_PATH, translation provider keys, …) is
|
||||
# saved to prefs.json so it survives backend restarts (restored at
|
||||
# startup in main.py). Non-persistent keys stay process-local.
|
||||
@@ -950,7 +1150,13 @@ async def set_env_var(body: dict):
|
||||
else:
|
||||
prefs_delete(prefs_key)
|
||||
|
||||
return {"key": key, "set": bool(value)}
|
||||
# #1787 review fix: tell the caller up front when the value just saved is
|
||||
# being shadowed by an external env var — set at THIS process's startup,
|
||||
# before our own prefs restore ran, so it predicts the next restart too.
|
||||
# A response that just said {"set": True} let the Settings panel promise
|
||||
# a restart would apply a value that never will.
|
||||
from core.prefs import is_env_shadowed
|
||||
return {"key": key, "set": bool(value), "shadowed": is_env_shadowed(key)}
|
||||
|
||||
|
||||
@router.post("/clean-audio")
|
||||
@@ -1041,7 +1247,7 @@ def asr_backends():
|
||||
def hf_token_state():
|
||||
"""Return the 3-source HF token cascade state for the Settings UI
|
||||
(Wave 2 React panel consumes this). Never returns the raw token —
|
||||
only a masked preview, whoami username, and per-source validity.
|
||||
only a masked preview and local presence; no outbound validation.
|
||||
"""
|
||||
from dataclasses import asdict
|
||||
from services import token_resolver
|
||||
|
||||
@@ -174,11 +174,14 @@ async def ws_tts(websocket: WebSocket):
|
||||
# close on `unavailable`, a one-time `routing` frame on
|
||||
# cpu_fallback / accelerated-with-caveat (before any audio).
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_routing import resolve_routing, routing_notice
|
||||
from services.engine_routing import (
|
||||
routing_notice,
|
||||
runtime_compute_profile_async,
|
||||
)
|
||||
from core.scrub import scrub_text
|
||||
_routing = resolve_routing(
|
||||
getattr(backend, "gpu_compat", ("cpu",)), detect_host_caps(),
|
||||
getattr(backend, "min_vram_gb", 0.0))
|
||||
_routing = await runtime_compute_profile_async(
|
||||
backend, detect_host_caps()
|
||||
)
|
||||
if _routing["routing_status"] == "unavailable":
|
||||
await websocket.send_json({
|
||||
"type": "error",
|
||||
|
||||
@@ -0,0 +1,394 @@
|
||||
"""Speech-to-speech voice changer — Studio's Convert method (POST /convert).
|
||||
|
||||
The user drops (or records) a source clip, picks an existing voice profile,
|
||||
and gets the same words back in that profile's voice: the active ASR backend
|
||||
transcribes the clip (no word timestamps — the text is all we need), the
|
||||
active TTS engine re-synthesizes it conditioned on the profile's reference
|
||||
audio, and — by default — the take is pitch-preservingly time-stretched
|
||||
(ffmpeg atempo, clamped to one well-behaved 0.5–2.0 stage) so it lands near
|
||||
the source clip's duration.
|
||||
|
||||
Deliberately reuses the /generate choke points instead of re-deriving them:
|
||||
|
||||
* profile row → conditioning via ``generation._resolve_profile_conditioning``
|
||||
(lock wins, ``kind`` authoritative, #533 language fill),
|
||||
* engine resolution via ``services.tts_backend.resolve_generation_backend``
|
||||
(never a silent OmniVoice fallback; ``require_cloning=True`` refuses
|
||||
clone-less engines with the actionable switch-engine message),
|
||||
* synthesis via ``generation._run_backend_inference`` on the guarded GPU
|
||||
pool (#730 bound + reset; busy/timeout → retryable 503),
|
||||
* provenance + persistence via ``services.watermark.mark_synthetic_async``
|
||||
and ``generation._finalize_generation`` (watermark → WAV in OUTPUTS_DIR →
|
||||
history row → retention prune), marked AFTER the stretch so the take users
|
||||
keep carries exactly one whole-take mark.
|
||||
|
||||
Local-first: no network calls; ASR-model-less installs get the same typed
|
||||
409 download CTA as /transcribe; a backend mid-shutdown surfaces the global
|
||||
503 ``[shutting_down]`` (ModelLoadInterruptedByShutdown → main.py handler).
|
||||
Reachability matches /generate: loopback bind by default, with the shared
|
||||
network-share PIN / API-key middleware gating any non-loopback exposure.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import functools
|
||||
import logging
|
||||
import os
|
||||
import tempfile
|
||||
import time
|
||||
|
||||
from fastapi import APIRouter, File, Form, HTTPException, UploadFile
|
||||
|
||||
router = APIRouter()
|
||||
logger = logging.getLogger("omnivoice.convert")
|
||||
|
||||
#: ffmpeg's atempo filter is well-behaved in [0.5, 2.0] per stage. Convert
|
||||
#: clamps to ONE stage by design: needing more than 2× either way means the
|
||||
#: synthesized speech differs so much from the source that "matching" it
|
||||
#: would produce chipmunk/slow-motion artifacts worse than the mismatch.
|
||||
ATEMPO_MIN = 0.5
|
||||
ATEMPO_MAX = 2.0
|
||||
|
||||
#: Within this relative tolerance the durations already match — stretching
|
||||
#: would resample the whole take for an inaudible gain.
|
||||
_MATCH_TOLERANCE = 0.02
|
||||
|
||||
#: Convert clips are short conversational inputs, not long-form media. Stream
|
||||
#: them to disk in bounded chunks so a network-share client cannot make the
|
||||
#: backend materialize an arbitrarily large multipart upload in memory.
|
||||
_MAX_SOURCE_AUDIO_BYTES = 64 * 1024 * 1024
|
||||
_UPLOAD_CHUNK_BYTES = 1024 * 1024
|
||||
|
||||
|
||||
async def _copy_source_upload(audio: UploadFile, destination) -> int:
|
||||
"""Stream ``audio`` into ``destination`` with the Convert upload cap."""
|
||||
total = 0
|
||||
while True:
|
||||
chunk = await audio.read(_UPLOAD_CHUNK_BYTES)
|
||||
if not chunk:
|
||||
return total
|
||||
total += len(chunk)
|
||||
if total > _MAX_SOURCE_AUDIO_BYTES:
|
||||
raise HTTPException(
|
||||
status_code=413,
|
||||
detail="Source audio is too large (maximum 64 MB).",
|
||||
)
|
||||
destination.write(chunk)
|
||||
|
||||
|
||||
def _clamped_tempo_ratio(tts_duration_s: float, source_duration_s: float) -> "float | None":
|
||||
"""The atempo ratio that fits the take into the source duration, or None.
|
||||
|
||||
ratio > 1 speeds the take up (it came out longer than the source),
|
||||
ratio < 1 slows it down. Clamped to a single atempo stage's [0.5, 2.0];
|
||||
None when either duration is unusable or they already match.
|
||||
"""
|
||||
if not source_duration_s or source_duration_s <= 0:
|
||||
return None
|
||||
if not tts_duration_s or tts_duration_s <= 0:
|
||||
return None
|
||||
ratio = tts_duration_s / source_duration_s
|
||||
if abs(ratio - 1.0) <= _MATCH_TOLERANCE:
|
||||
return None
|
||||
return min(ATEMPO_MAX, max(ATEMPO_MIN, ratio))
|
||||
|
||||
|
||||
async def _match_source_duration(audio_tensor, sample_rate: int, source_duration_s: float):
|
||||
"""Best-effort pitch-preserving stretch of the take toward the source
|
||||
clip's duration. Returns the input unchanged when no stretch is needed
|
||||
or ffmpeg fails — a duration mismatch is better than a failed convert."""
|
||||
n_samples = int(audio_tensor.shape[-1])
|
||||
ratio = _clamped_tempo_ratio(n_samples / sample_rate, source_duration_s)
|
||||
if ratio is None:
|
||||
return audio_tensor
|
||||
target_samples = max(1, int(round(n_samples / ratio)))
|
||||
from services.ffmpeg_utils import _pitch_preserving_stretch
|
||||
try:
|
||||
return await _pitch_preserving_stretch(audio_tensor, target_samples, sample_rate)
|
||||
except Exception as e: # noqa: BLE001 — stretch is opt-in polish, never fatal
|
||||
logger.warning("duration match skipped — atempo stretch failed: %s", e)
|
||||
return audio_tensor
|
||||
|
||||
|
||||
async def _transcribe_source(tmp_path: str, *, source_lease=None) -> dict:
|
||||
"""Active-ASR transcription of the uploaded clip (no word timestamps).
|
||||
|
||||
Mirrors POST /transcribe: typed 409 + download CTA before any backend
|
||||
is constructed (never a silent multi-GB auto-download), the guarded GPU
|
||||
pool dispatch (#730), 504 on timeout, and the same 409 when the loader
|
||||
degrades onto an engine with no weights on disk (#1185).
|
||||
"""
|
||||
from services.asr_backend import (
|
||||
ASRModelMissingError,
|
||||
ASRTimeoutError,
|
||||
asr_model_missing_detail,
|
||||
asr_model_missing_error,
|
||||
run_transcribe_guarded,
|
||||
)
|
||||
|
||||
missing = await asyncio.to_thread(asr_model_missing_error, purpose="transcribe")
|
||||
if missing is not None:
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail={**missing, "message": asr_model_missing_detail(missing)},
|
||||
)
|
||||
|
||||
def _run():
|
||||
# `load_*`, not `get_*`: the loader runs ensure_loaded() and degrades
|
||||
# past an engine whose deep import chain is broken (#1185).
|
||||
from services.asr_backend import load_active_asr_backend
|
||||
backend = load_active_asr_backend()
|
||||
return backend.transcribe(tmp_path, word_timestamps=False)
|
||||
|
||||
from services.model_manager import _gpu_pool
|
||||
release = source_lease.acquire() if source_lease is not None else None
|
||||
abandoned = False
|
||||
try:
|
||||
return await run_transcribe_guarded(
|
||||
_gpu_pool,
|
||||
_run,
|
||||
what="Voice convert",
|
||||
on_abandon=release,
|
||||
)
|
||||
except asyncio.CancelledError:
|
||||
# The guard now owns the lease token until the native worker drains.
|
||||
abandoned = True
|
||||
raise
|
||||
except ASRTimeoutError as e:
|
||||
abandoned = True
|
||||
logger.warning("Convert transcription timed out: %s", e)
|
||||
raise HTTPException(status_code=504, detail=str(e))
|
||||
except ASRModelMissingError as e:
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail={**e.payload, "message": asr_model_missing_detail(e.payload)},
|
||||
)
|
||||
finally:
|
||||
if release is not None and not abandoned:
|
||||
release()
|
||||
|
||||
|
||||
@router.post("/convert")
|
||||
async def convert_speech(
|
||||
audio: UploadFile = File(...),
|
||||
profile_id: str = Form(...),
|
||||
match_duration: bool = Form(True),
|
||||
):
|
||||
"""Convert a spoken clip into an existing voice profile's voice.
|
||||
|
||||
Multipart form: ``audio`` (the source clip), ``profile_id`` (an existing
|
||||
voice profile), optional ``match_duration`` (default on — atempo the take
|
||||
toward the source clip's length, clamped to 0.5–2.0×).
|
||||
|
||||
Returns JSON ``{audio_url, text, duration_s, id}`` — the take is saved to
|
||||
OUTPUTS_DIR and served from the ``/audio`` mount like every other take.
|
||||
"""
|
||||
from core.db import db_conn
|
||||
from api.routers.generation import _resolve_profile_conditioning, _TempReferenceLease
|
||||
|
||||
# ── Profile first: strict 404, unlike /generate's silent skip — Convert
|
||||
# has no meaning without a target voice.
|
||||
with db_conn() as conn:
|
||||
row = conn.execute(
|
||||
"SELECT * FROM voice_profiles WHERE id=?", (profile_id,)
|
||||
).fetchone()
|
||||
if not row:
|
||||
raise HTTPException(
|
||||
status_code=404,
|
||||
detail="That voice profile doesn't exist. It may have been deleted from another tab.",
|
||||
)
|
||||
cond = _resolve_profile_conditioning(row)
|
||||
|
||||
# ── Save the upload before loading an engine. Every ASR backend (and
|
||||
# ffprobe) needs a file path; the bounded streaming copy rejects oversized
|
||||
# network-share requests without materializing them in process memory or
|
||||
# starting heavyweight model work.
|
||||
ext = os.path.splitext(audio.filename or "audio.wav")[1] or ".wav"
|
||||
tmp = tempfile.NamedTemporaryFile(delete=False, suffix=ext)
|
||||
source_lease = None
|
||||
try:
|
||||
try:
|
||||
await _copy_source_upload(audio, tmp)
|
||||
finally:
|
||||
tmp.close()
|
||||
source_lease = _TempReferenceLease(tmp.name)
|
||||
|
||||
# ── Engine gate before ASR/TTS work: the shared resolver refuses a
|
||||
# clone-less engine with the actionable switch-engine message (→ 400),
|
||||
# and a backend mid-shutdown raises ModelLoadInterruptedByShutdown out
|
||||
# of the model load → the global 503 [shutting_down] handler.
|
||||
from services.tts_backend import resolve_generation_backend
|
||||
try:
|
||||
backend = await resolve_generation_backend(
|
||||
require_cloning=True, cloning_purpose="voice conversion",
|
||||
)
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e))
|
||||
|
||||
result = await _transcribe_source(tmp.name, source_lease=source_lease)
|
||||
|
||||
segments = result.get("segments", [])
|
||||
text = result.get("text", "")
|
||||
if not text and segments:
|
||||
text = " ".join(s.get("text", "") for s in segments).strip()
|
||||
# Same final-text hygiene as /transcribe: strip Whisper hallucination
|
||||
# loops, then deterministic polish (leading capital + terminal
|
||||
# punctuation) so the TTS input reads as typed text.
|
||||
from services.refinement import collapse_repetitive_artifacts
|
||||
from services.text_polish import polish_text
|
||||
text = polish_text(collapse_repetitive_artifacts(text))
|
||||
if not text or not text.strip():
|
||||
raise HTTPException(
|
||||
status_code=422,
|
||||
detail=(
|
||||
"No speech was recognized in the source clip, so there is "
|
||||
"nothing to convert. Record or drop a clip with clear, "
|
||||
"audible speech and try again."
|
||||
),
|
||||
)
|
||||
|
||||
# #308/#1032 parity with /generate: a clone profile saved without a
|
||||
# transcript conditions better when its reference clip is transcribed,
|
||||
# and that transcript is cached onto the row so it happens ONCE, not
|
||||
# per convert. Best-effort exactly like /generate — a timeout/failure
|
||||
# degrades to ref_text=None and the engine's own fallback. The ASR
|
||||
# model is already warm here (the source transcribe above just used it).
|
||||
if cond["ref_audio_path"] and not cond["ref_text"]:
|
||||
from api.routers.generation import (
|
||||
_generate_timeout_s,
|
||||
_persist_profile_ref_text,
|
||||
)
|
||||
from services.asr_backend import transcribe_reference
|
||||
from services.model_manager import run_on_gpu_pool_guarded
|
||||
try:
|
||||
cond["ref_text"] = await run_on_gpu_pool_guarded(
|
||||
functools.partial(transcribe_reference, cond["ref_audio_path"]),
|
||||
what="Reference transcribe",
|
||||
timeout=_generate_timeout_s(""),
|
||||
)
|
||||
except TimeoutError as e:
|
||||
logger.warning(
|
||||
"reference transcribe hung (%s); using engine ASR fallback", e,
|
||||
)
|
||||
cond["ref_text"] = None
|
||||
if cond["ref_text"] and cond["persist_ref_text"]:
|
||||
_persist_profile_ref_text(profile_id, cond["ref_text"])
|
||||
|
||||
# Source duration for the optional match: the container's own length
|
||||
# (ffprobe), falling back to the last ASR segment end. Best-effort —
|
||||
# None just skips the stretch.
|
||||
source_duration_s = None
|
||||
if match_duration:
|
||||
from services.ffmpeg_utils import probe_duration
|
||||
source_duration_s = await probe_duration(
|
||||
tmp.name, allowed_root=os.path.dirname(tmp.name),
|
||||
)
|
||||
if not source_duration_s and segments:
|
||||
source_duration_s = max((s.get("end", 0) or 0) for s in segments) or None
|
||||
|
||||
# ── Same text choke point as /generate: engine-agnostic normalization
|
||||
# (numbers→words, junk strip) on the fully resolved language.
|
||||
from services.text_normalization import normalize_for_tts
|
||||
language = cond["language"]
|
||||
text = normalize_for_tts(text, language)
|
||||
|
||||
used_seed = cond["seed"]
|
||||
if used_seed is None:
|
||||
import random
|
||||
used_seed = random.randint(0, 2**31 - 1)
|
||||
|
||||
from api.routers.generation import (
|
||||
_finalize_generation,
|
||||
_generate_timeout_s,
|
||||
_run_backend_inference,
|
||||
)
|
||||
from services.model_manager import (
|
||||
GpuJobTimeoutError,
|
||||
GpuPoolBusyError,
|
||||
run_on_gpu_pool_guarded,
|
||||
)
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_routing import runtime_compute_profile_async
|
||||
compute_profile = await runtime_compute_profile_async(
|
||||
backend, detect_host_caps()
|
||||
)
|
||||
if compute_profile["routing_status"] == "unavailable":
|
||||
raise HTTPException(
|
||||
status_code=400,
|
||||
detail=compute_profile["routing_reason"],
|
||||
)
|
||||
|
||||
start_time = time.time()
|
||||
_render = functools.partial(
|
||||
_run_backend_inference,
|
||||
backend, text, language, cond["ref_audio_path"], cond["ref_text"],
|
||||
cond["instruct"],
|
||||
None, # duration — the model picks; match_duration owns pacing
|
||||
16, 2.0, # num_step / guidance_scale (the /generate defaults)
|
||||
1.0, # speed
|
||||
True, True, # denoise / postprocess_output
|
||||
used_seed,
|
||||
)
|
||||
try:
|
||||
audio_tensor = await run_on_gpu_pool_guarded(
|
||||
_render,
|
||||
what="Voice convert",
|
||||
timeout=_generate_timeout_s(
|
||||
text,
|
||||
execution_device=compute_profile["effective_device"],
|
||||
min_vram_gb=compute_profile["min_vram_gb"],
|
||||
hardware_family=compute_profile.get("runtime_hardware_family"),
|
||||
vram_gb=compute_profile.get("runtime_vram_gb"),
|
||||
),
|
||||
min_vram_gb=compute_profile["min_vram_gb"],
|
||||
)
|
||||
except GpuPoolBusyError as e:
|
||||
raise HTTPException(
|
||||
status_code=503, detail=str(e),
|
||||
headers={"Retry-After": str(e.retry_after),
|
||||
"X-OmniVoice-Retryable": "true"},
|
||||
) from e
|
||||
except GpuJobTimeoutError as e:
|
||||
raise HTTPException(
|
||||
status_code=503, detail=str(e),
|
||||
headers={"Retry-After": "30", "X-OmniVoice-Retryable": "true"},
|
||||
) from e
|
||||
except ValueError as e:
|
||||
raise HTTPException(status_code=400, detail=str(e)) from e
|
||||
sample_rate = backend.sample_rate
|
||||
|
||||
if match_duration and source_duration_s:
|
||||
audio_tensor = await _match_source_duration(
|
||||
audio_tensor, sample_rate, source_duration_s,
|
||||
)
|
||||
|
||||
# Provenance mark AFTER the stretch (one whole-take mark on the audio
|
||||
# the user actually keeps), then the shared finalize tail — WAV in
|
||||
# OUTPUTS_DIR, self-healing history row, retention prune, event emit.
|
||||
from services.watermark import mark_synthetic_async
|
||||
audio_tensor = await mark_synthetic_async(
|
||||
audio_tensor, sample_rate, context="convert.finalize",
|
||||
)
|
||||
_, meta = await _finalize_generation(
|
||||
audio_tensor, sample_rate, text=text, history_mode="convert",
|
||||
ref_audio_path=cond["ref_audio_path"], language=language,
|
||||
instruct=cond["instruct"], resolved_profile_id=profile_id,
|
||||
used_seed=used_seed, start_time=start_time,
|
||||
already_marked=True,
|
||||
)
|
||||
|
||||
return {
|
||||
"id": meta["id"],
|
||||
"audio_url": f"/audio/{meta['filename']}",
|
||||
"text": text,
|
||||
"duration_s": meta["duration"],
|
||||
"gen_time_s": meta["gen_time"],
|
||||
}
|
||||
finally:
|
||||
if source_lease is not None:
|
||||
source_lease.finish_request()
|
||||
else:
|
||||
try:
|
||||
os.unlink(tmp.name)
|
||||
except OSError:
|
||||
pass
|
||||
@@ -26,6 +26,21 @@ class SystemInfoResponse(BaseModel):
|
||||
model_config = ConfigDict(extra="allow")
|
||||
|
||||
app_version: str = ""
|
||||
# Effective compute-time budgets (seconds) for one synthesis job — the
|
||||
# values services/model_manager.py's GPU_JOB_TIMEOUT_S / CPU_JOB_TIMEOUT_S
|
||||
# captured at backend import time (#1787). A value just saved via
|
||||
# /system/set-env is NOT reflected here until the next restart.
|
||||
generate_timeout_s: float = 300.0
|
||||
cpu_generate_timeout_s: float = 600.0
|
||||
# True when an external env var (shell, `.env`, Docker, …) is currently
|
||||
# shadowing a prefs.json save for this key — see core.prefs.is_env_shadowed.
|
||||
generate_timeout_shadowed: bool = False
|
||||
cpu_generate_timeout_shadowed: bool = False
|
||||
# #1770: the desktop attach handshake's code fingerprint — whatever
|
||||
# Tauri set OMNIVOICE_BUILD_FINGERPRINT to when it spawned this process,
|
||||
# echoed back verbatim. Blank when unset (dev mode, a manually started
|
||||
# backend). See frontend/src-tauri/src/backend.rs::code_fingerprint_is_current.
|
||||
code_fingerprint: str = ""
|
||||
data_dir: str
|
||||
outputs_dir: str
|
||||
crash_log_path: str
|
||||
|
||||
@@ -28,6 +28,9 @@
|
||||
# their own (weights live in referenced sub-repos). Such
|
||||
# a cache is legitimately tiny, so the truncated-download
|
||||
# (weights-missing) detector must NOT flag it incomplete.
|
||||
# allow_patterns (optional) — restrict installation to these repository paths.
|
||||
# Use for multi-package repos so an explicit install
|
||||
# never downloads unrelated model variants.
|
||||
# ─────────────────────────────────────────────────────────────────────────
|
||||
|
||||
models:
|
||||
@@ -40,6 +43,14 @@ models:
|
||||
required: true
|
||||
curated_on: [all]
|
||||
|
||||
- repo_id: "audio-cpp/audio.cpp-gguf"
|
||||
label: "Breeze-TTS-2 Q8_0 for audio.cpp (English + Chinese, clone + design)"
|
||||
role: TTS
|
||||
size_gb: 4.73
|
||||
allow_patterns:
|
||||
- "Breeze-TTS-2-GGUF/breeze-tts-2-q8_0.gguf"
|
||||
note: "Optional audio.cpp model. Research/non-commercial weights and self-hosted outputs; install only after reviewing the license."
|
||||
|
||||
# ── ASR (optional — curated per platform) ─────────────────────────────
|
||||
# No ASR model is required to boot: TTS-only installs work. Dubbing,
|
||||
# dictation, and clone-reference transcription prompt for the curated
|
||||
|
||||
@@ -15,6 +15,18 @@ def get_app_data_dir():
|
||||
return os.path.expanduser("~/.omnivoice")
|
||||
|
||||
|
||||
def _configured_hf_token_path():
|
||||
"""Match Hub's token location without importing or refreshing credentials."""
|
||||
default_cache = os.path.join(os.path.expanduser("~"), ".cache")
|
||||
hf_home = os.environ.get("HF_HOME", os.path.join(os.environ.get("XDG_CACHE_HOME", default_cache), "huggingface"))
|
||||
return os.path.expandvars(os.path.expanduser(os.environ.get("HF_TOKEN_PATH", os.path.join(hf_home, "token"))))
|
||||
|
||||
|
||||
# Snapshot recognized locations before automatic model-cache redirection.
|
||||
# Explicit cache/token overrides restrict clearing to their selected location.
|
||||
HF_CLI_TOKEN_PATHS = (_configured_hf_token_path(),)
|
||||
|
||||
|
||||
def _ensure_short_hf_cache_on_windows():
|
||||
"""Redirect HuggingFace cache to a short path on Windows.
|
||||
|
||||
@@ -38,6 +50,14 @@ def _ensure_short_hf_cache_on_windows():
|
||||
return
|
||||
short_cache = os.path.join(local_app, "OmniVoice", "hf_cache")
|
||||
os.makedirs(short_cache, exist_ok=True)
|
||||
if "HF_TOKEN_PATH" not in os.environ:
|
||||
global HF_CLI_TOKEN_PATHS
|
||||
canonical = HF_CLI_TOKEN_PATHS[0]
|
||||
legacy = os.path.join(short_cache, "token")
|
||||
HF_CLI_TOKEN_PATHS = tuple(dict.fromkeys((canonical, legacy)))
|
||||
# Keep existing app-written logins usable without copying credentials.
|
||||
selected = canonical if os.path.exists(canonical) or not os.path.exists(legacy) else legacy
|
||||
os.environ.setdefault("HF_TOKEN_PATH", selected)
|
||||
os.environ["HF_HOME"] = short_cache
|
||||
os.environ["HF_HUB_CACHE"] = short_cache
|
||||
|
||||
|
||||
@@ -3,13 +3,14 @@
|
||||
The desktop owns the backend with an OS process group/Job. Engine and
|
||||
installer operations also need an independently terminable subtree: killing
|
||||
only their direct child on a timeout leaves uv/git/model workers holding pipes
|
||||
and mutating files. A small direct-child supervisor bridges both lifetimes.
|
||||
and mutating files.
|
||||
|
||||
On POSIX the supervisor is the unreaped leader of a nested process group. A
|
||||
control-pipe EOF (including kernel EOF when the backend dies) kills that group;
|
||||
the parent also drains the group before reaping its stable leader. On Windows
|
||||
the supervisor assigns the operation, while suspended, to a nested
|
||||
kill-on-close Job. The outer desktop Job still contains both levels.
|
||||
On POSIX a small supervisor is the unreaped leader of a nested process group.
|
||||
A control-pipe EOF (including kernel EOF when the backend dies) kills that
|
||||
group; the parent also drains the group before reaping its stable leader. On
|
||||
Windows the backend retains a nested kill-on-close Job directly and assigns
|
||||
the suspended operation before resuming it. The outer desktop Job remains the
|
||||
terminal fallback.
|
||||
|
||||
Standalone/server launches use the same nested owner, preserving their
|
||||
independently terminable subtree without relying on ``taskkill`` or discovery.
|
||||
@@ -263,42 +264,148 @@ class OwnedPopen:
|
||||
pass
|
||||
|
||||
|
||||
def spawn_owned(argv: list[str], **kwargs: Any) -> "subprocess.Popen | OwnedPopen":
|
||||
class WindowsJobPopen:
|
||||
"""Popen-compatible handle whose child tree lives in a retained Job.
|
||||
|
||||
Windows Job handles already provide the stable ownership that POSIX needs
|
||||
a supervisor process group for. Keeping the handle in the backend means an
|
||||
abrupt backend exit closes it in the kernel and kills the whole operation
|
||||
tree, without inserting a second Python process in the sidecar loader path
|
||||
(#1734).
|
||||
"""
|
||||
|
||||
def __init__(self, proc: subprocess.Popen, job: Any, kernel32: Any) -> None:
|
||||
self._proc = proc
|
||||
self._job = job
|
||||
self._kernel32 = kernel32
|
||||
self._lock = threading.RLock()
|
||||
self.stdin = proc.stdin
|
||||
self.stdout = proc.stdout
|
||||
self.stderr = proc.stderr
|
||||
|
||||
@property
|
||||
def pid(self) -> int:
|
||||
return self._proc.pid
|
||||
|
||||
@property
|
||||
def args(self) -> Any:
|
||||
return self._proc.args
|
||||
|
||||
@property
|
||||
def returncode(self) -> Optional[int]:
|
||||
return self._proc.returncode
|
||||
|
||||
def _close_job(self, *, terminate: bool) -> None:
|
||||
job, self._job = self._job, None
|
||||
if job is None:
|
||||
return
|
||||
try:
|
||||
if terminate:
|
||||
self._kernel32.TerminateJobObject(job, 1)
|
||||
finally:
|
||||
self._kernel32.CloseHandle(job)
|
||||
|
||||
def poll(self) -> Optional[int]:
|
||||
with self._lock:
|
||||
rc = self._proc.poll()
|
||||
if rc is None:
|
||||
return None
|
||||
# A successful direct child may leave helpers behind. Match the
|
||||
# supervisor contract by draining the retained Job before return.
|
||||
self._close_job(terminate=True)
|
||||
return rc
|
||||
|
||||
def wait(self, timeout: Optional[float] = None) -> int:
|
||||
try:
|
||||
rc = self._proc.wait(timeout=timeout)
|
||||
except subprocess.TimeoutExpired:
|
||||
raise
|
||||
with self._lock:
|
||||
self._close_job(terminate=True)
|
||||
return rc
|
||||
|
||||
def terminate(self) -> None:
|
||||
with self._lock:
|
||||
self._close_job(terminate=True)
|
||||
|
||||
def kill(self) -> None:
|
||||
self.terminate()
|
||||
|
||||
def __getattr__(self, name: str) -> Any:
|
||||
return getattr(self._proc, name)
|
||||
|
||||
def __del__(self) -> None:
|
||||
try:
|
||||
self._close_job(terminate=True)
|
||||
except Exception:
|
||||
pass # interpreter shutdown; closing the OS handle is best-effort
|
||||
|
||||
|
||||
def _spawn_windows_owned(argv: list[str], kwargs: dict[str, Any]) -> WindowsJobPopen:
|
||||
"""Start *argv* suspended, assign its tree to a Job, then resume it."""
|
||||
import ctypes
|
||||
|
||||
job, kernel32, wintypes = _windows_job()
|
||||
child: Optional[subprocess.Popen] = None
|
||||
popen_kwargs = dict(kwargs)
|
||||
supplied_env = popen_kwargs.get("env")
|
||||
operation_env = dict(os.environ if supplied_env is None else supplied_env)
|
||||
operation_env.pop(_DRAIN_FD_ENV, None)
|
||||
operation_env.pop(_DESKTOP_MARKER, None)
|
||||
popen_kwargs["env"] = operation_env
|
||||
supplied_flags = int(popen_kwargs.pop("creationflags", 0))
|
||||
popen_kwargs["creationflags"] = supplied_flags | 0x08000000 | 0x00000004
|
||||
try:
|
||||
child = subprocess.Popen(argv, **popen_kwargs)
|
||||
assign = kernel32.AssignProcessToJobObject
|
||||
assign.argtypes = (wintypes.HANDLE, wintypes.HANDLE)
|
||||
assign.restype = wintypes.BOOL
|
||||
if not assign(job, wintypes.HANDLE(child._handle)):
|
||||
raise OSError(ctypes.get_last_error(), "AssignProcessToJobObject")
|
||||
_resume_windows_process(kernel32, wintypes, child.pid)
|
||||
return WindowsJobPopen(child, job, kernel32)
|
||||
except BaseException:
|
||||
kernel32.TerminateJobObject(job, 1)
|
||||
if child is not None:
|
||||
try:
|
||||
child.kill()
|
||||
except OSError:
|
||||
pass # the suspended child may already have exited
|
||||
try:
|
||||
child.wait(timeout=5)
|
||||
except (OSError, subprocess.TimeoutExpired):
|
||||
pass # Job termination remains the authoritative cleanup
|
||||
kernel32.CloseHandle(job)
|
||||
raise
|
||||
|
||||
|
||||
def spawn_owned(
|
||||
argv: list[str], **kwargs: Any
|
||||
) -> "subprocess.Popen | OwnedPopen | WindowsJobPopen":
|
||||
"""Spawn an operation with a stable, independently terminable owner."""
|
||||
|
||||
drain_fd = backend_drain_fd(required=True) if os.name == "posix" else None
|
||||
if os.name == "nt":
|
||||
return _spawn_windows_owned(argv, kwargs)
|
||||
|
||||
drain_fd = backend_drain_fd(required=True)
|
||||
control_read, control_write = os.pipe()
|
||||
result_read, result_write = os.pipe()
|
||||
control_token = control_read
|
||||
result_token = result_write
|
||||
if os.name == "nt":
|
||||
import msvcrt
|
||||
|
||||
control_token = msvcrt.get_osfhandle(control_read)
|
||||
result_token = msvcrt.get_osfhandle(result_write)
|
||||
wrapper_argv = _supervisor_argv(
|
||||
control_token,
|
||||
result_token,
|
||||
control_read,
|
||||
result_write,
|
||||
argv,
|
||||
)
|
||||
wrapper_kwargs = dict(kwargs)
|
||||
if os.name == "posix":
|
||||
wrapper_kwargs["start_new_session"] = True
|
||||
pass_fds = [control_read, result_write]
|
||||
if drain_fd is not None:
|
||||
pass_fds.append(drain_fd)
|
||||
if wrapper_kwargs.get("env") is not None:
|
||||
wrapper_env = dict(wrapper_kwargs["env"])
|
||||
wrapper_env[_DESKTOP_MARKER] = "1"
|
||||
wrapper_env[_DRAIN_FD_ENV] = str(drain_fd)
|
||||
wrapper_kwargs["env"] = wrapper_env
|
||||
wrapper_kwargs["pass_fds"] = tuple(pass_fds)
|
||||
else:
|
||||
# Python's Windows fd inheritance requires inheritable CRT handles.
|
||||
# All unrelated descriptors are non-inheritable by default (PEP 446).
|
||||
os.set_handle_inheritable(control_token, True)
|
||||
os.set_handle_inheritable(result_token, True)
|
||||
wrapper_kwargs["close_fds"] = False
|
||||
wrapper_kwargs["start_new_session"] = True
|
||||
pass_fds = [control_read, result_write]
|
||||
if drain_fd is not None:
|
||||
pass_fds.append(drain_fd)
|
||||
if wrapper_kwargs.get("env") is not None:
|
||||
wrapper_env = dict(wrapper_kwargs["env"])
|
||||
wrapper_env[_DESKTOP_MARKER] = "1"
|
||||
wrapper_env[_DRAIN_FD_ENV] = str(drain_fd)
|
||||
wrapper_kwargs["env"] = wrapper_env
|
||||
wrapper_kwargs["pass_fds"] = tuple(pass_fds)
|
||||
try:
|
||||
proc = subprocess.Popen(wrapper_argv, **wrapper_kwargs)
|
||||
except BaseException:
|
||||
|
||||
+2
-1
@@ -274,7 +274,8 @@ _BASE_SCHEMA = """
|
||||
started_at REAL,
|
||||
finished_at REAL,
|
||||
lease_expires_at REAL,
|
||||
grace_expires_at REAL
|
||||
grace_expires_at REAL,
|
||||
deadlines_json TEXT
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_remote_attempts_task ON remote_task_attempts(task_id);
|
||||
CREATE INDEX IF NOT EXISTS idx_remote_attempts_worker ON remote_task_attempts(worker_id, state);
|
||||
|
||||
@@ -35,7 +35,8 @@ import sys
|
||||
from dataclasses import dataclass
|
||||
from typing import Literal
|
||||
|
||||
DeviceFamily = Literal["cuda", "rocm", "mps", "xpu", "cpu"]
|
||||
DeviceFamily = Literal["cuda", "rocm", "mps", "xpu", "npu", "cpu"]
|
||||
ACCELERATOR_PRIORITY = ("cuda", "rocm", "xpu", "npu", "mps")
|
||||
|
||||
# Stable substring stamped onto notes that represent a real kernel-launch risk
|
||||
# (arch/driver mismatch) — as opposed to advisory notes (multi-GPU, VRAM query
|
||||
@@ -533,9 +534,13 @@ def _probe() -> HostCaps:
|
||||
# is the whole truth in that case (CodeRabbit, #1425).
|
||||
notes.extend(why_no_gpu(torch))
|
||||
|
||||
# ── Intel XPU via IPEX ───────────────────────────────────────────────
|
||||
# Older builds register XPU through IPEX; modern torch exposes it directly.
|
||||
try:
|
||||
import intel_extension_for_pytorch # noqa: F401
|
||||
except Exception:
|
||||
# Optional IPEX may be absent or incompatible; still probe native torch XPU.
|
||||
pass
|
||||
try:
|
||||
if hasattr(torch, "xpu") and torch.xpu.is_available():
|
||||
detected.append("xpu")
|
||||
if not device_name:
|
||||
@@ -546,7 +551,23 @@ def _probe() -> HostCaps:
|
||||
pass
|
||||
notes.append("XPU VRAM not queried (unreliable across IPEX versions)")
|
||||
except Exception:
|
||||
# IPEX absent or XPU probe failed — no XPU on this host.
|
||||
# XPU probe failed — no usable XPU on this host.
|
||||
pass
|
||||
|
||||
# Vendor extensions may register an NPU with torch. Probe only an already
|
||||
# registered backend; never install or import an optional vendor package.
|
||||
try:
|
||||
if hasattr(torch, "npu") and torch.npu.is_available():
|
||||
detected.append("npu")
|
||||
if not device_name:
|
||||
try:
|
||||
device_name = torch.npu.get_device_name(0)
|
||||
except Exception:
|
||||
# An unavailable display name does not invalidate a usable NPU.
|
||||
pass
|
||||
notes.append("NPU VRAM not queried")
|
||||
except Exception:
|
||||
# Missing or broken vendor backends mean no usable NPU; continue probing.
|
||||
pass
|
||||
|
||||
# ── Apple Silicon MPS ────────────────────────────────────────────────
|
||||
@@ -579,7 +600,7 @@ def _probe() -> HostCaps:
|
||||
|
||||
# Preferred family by priority; cpu when nothing accelerated was detected.
|
||||
family: DeviceFamily = "cpu"
|
||||
for pref in ("cuda", "rocm", "xpu", "mps"):
|
||||
for pref in ACCELERATOR_PRIORITY:
|
||||
if pref in detected:
|
||||
family = pref # type: ignore[assignment]
|
||||
break
|
||||
|
||||
@@ -28,6 +28,7 @@ import shutil
|
||||
import sys
|
||||
|
||||
from core.config import DATA_DIR
|
||||
from core.device_caps import KERNEL_RISK_MARKER
|
||||
from core.scrub import scrub_text
|
||||
from core.version import APP_VERSION
|
||||
|
||||
@@ -189,7 +190,7 @@ def _check_ram() -> dict:
|
||||
def _check_engines() -> dict:
|
||||
try:
|
||||
from services.tts_backend import list_backends, active_backend_id
|
||||
backends = list_backends()
|
||||
backends = list_backends(include_hidden=True)
|
||||
active = active_backend_id()
|
||||
except Exception as e:
|
||||
return _check("engines", "TTS engines", WARN, f"could not enumerate: {e}")
|
||||
@@ -230,11 +231,16 @@ def _check_gpu_routing() -> dict:
|
||||
host = v.get("host_family", "cpu")
|
||||
|
||||
if status == "accelerated":
|
||||
if reason: # driver/arch caveat — accelerated but at risk
|
||||
if reason and KERNEL_RISK_MARKER in reason: # driver/arch caveat — at risk
|
||||
return _check("gpu_routing", "GPU routing", WARN,
|
||||
f"{engine} -> {dev}: {reason}",
|
||||
"The GPU is selected but may fail at kernel launch — "
|
||||
"update drivers / reinstall torch for this GPU arch.")
|
||||
if reason: # low-VRAM caveat — not a driver/arch issue
|
||||
return _check("gpu_routing", "GPU routing", WARN,
|
||||
f"{engine} -> {dev}: {reason}",
|
||||
"Unload other models before generating, keep the text "
|
||||
"short, or pick a lighter engine.")
|
||||
return _check("gpu_routing", "GPU routing", OK, f"{engine} -> {dev} (accelerated)")
|
||||
if status == "cpu_fallback":
|
||||
return _check("gpu_routing", "GPU routing", WARN,
|
||||
@@ -374,7 +380,12 @@ def run_diagnostics(include_network: bool = True, deep: bool = False) -> dict:
|
||||
try:
|
||||
module = importlib.import_module(f"services.{family}_backend")
|
||||
active = module.active_backend_id()
|
||||
row = next((item for item in module.list_backends() if item.get("id") == active), None)
|
||||
rows = (
|
||||
module.list_backends(include_hidden=True)
|
||||
if family == "tts"
|
||||
else module.list_backends()
|
||||
)
|
||||
row = next((item for item in rows if item.get("id") == active), None)
|
||||
if row is not None:
|
||||
engine_execution.append({
|
||||
"family": family,
|
||||
|
||||
@@ -97,6 +97,9 @@ _HINTS: dict[str, str] = {
|
||||
"TRANSFORMERS_IMPORT": "Your transformers install is incomplete, or a package it loads models through (torchaudio, torchvision) is missing or mismatched with your torch — a torch/torchvision version mismatch fails with exactly this wording. Reinstall them together at the pinned versions (`uv pip install --python .venv --reinstall torch==2.8.0 torchaudio==2.8.0 torchvision==0.23.0 transformers` in the project folder), then restart the backend. If only transcription is affected, switching ASR to faster-whisper (Model Catalogue → Models) also works around it.",
|
||||
"WINDOWS_APP_CONTROL_BLOCKED": "Windows refused to load a file VoiceStudio needs — an Application Control policy (Smart App Control, WDAC, or AppLocker) blocked it. On a personal PC: Windows Security → App & browser control → Smart App Control → Off (Windows only lets you turn it off once — re-enabling requires a Windows reset), then restart VoiceStudio. On a managed/work PC, ask IT to allow the VoiceStudio install folder.",
|
||||
"WINDOWS_PAGING_FILE_TOO_SMALL": "Windows ran out of virtual memory while mapping the model into memory — its paging file is smaller than the model needs. This is not the same as your RAM being full, and closing other apps usually won't fix it: Windows has to be allowed to back the mapping. Set a bigger paging file — Settings → System → About → Advanced system settings → Performance → Settings → Advanced → Virtual memory → Change: untick \"Automatically manage\", pick your system drive, choose \"Custom size\" and set both Initial and Maximum to at least 32768 MB (more than the model's size), then OK and restart Windows. A smaller/quantized engine (OmniVoice GGUF, Supertonic-3) also avoids the large mapping entirely.",
|
||||
"WINDOWS_UNTRUSTED_MOUNT": "Windows refused to walk a folder on the way to this file because the path crosses a mount point it does not trust (WinError 448). That is a Windows rule about the VOLUME, not about VoiceStudio or the file itself — it turns up on Dev Drives, on mounted VHD/ReFS volumes, and on junctions pointing into another user profile, so retrying the same link cannot help. Point VoiceStudio at a folder on an ordinary local drive instead: Settings → Storage → data directory, or the download/output folder named in the message. If that folder has to stay where it is, trust the volume with `fsutil devdrv trust <drive>:` from an elevated prompt and restart.",
|
||||
"INPUT_TOO_SHORT": "The input was too short for this engine to process — its first convolution needs more frames than the text (or the reference clip) produced. This is a hard limit of the model, not a transient failure, so retrying the same input will fail the same way. Give it a few more words, or a longer reference clip: a short phrase rather than one or two characters, and about a second of speech rather than a fragment.",
|
||||
"CLONE_REFERENCE_MISSING": "This engine was asked to clone a voice but got no reference audio to clone FROM, and the model folder carries no built-in voice either. Pick a voice profile that has a saved reference clip, or record/upload a few seconds of clean speech as the reference, then generate again. A designed voice with no saved reference cannot be cloned from — synthesize with it directly instead.",
|
||||
"MEDIA_TOOL_MISSING": "VoiceStudio's media engine (ffmpeg/ffprobe) wasn't on the system path when a component went looking for it. Open Settings → Audio tools and use Download/Repair to fetch the bundled copy, then retry — a restart picks it up for everything. If you'd rather use a system install, install ffmpeg (macOS: `brew install ffmpeg`; Windows: `winget install Gyan.FFmpeg`; Linux: your package manager) and restart VoiceStudio, or point FFMPEG_PATH / OMNIVOICE_FFPROBE_PATH at the binaries in Settings.",
|
||||
"AUDIO_IO_FAILED": "An audio file couldn't be read or written at the OS level. Check the drive isn't full, that the output and temp folders exist and are writable, and that antivirus or OneDrive isn't locking them (add a VoiceStudio exclusion if you use one).",
|
||||
"VIDEO_DOWNLOAD_OS_ERROR": "The OS refused a file operation while saving the downloaded video — this is a disk/folder problem, not a network one, so retrying the same link won't help. The download is written to a job folder under your VoiceStudio data directory (Settings → Storage shows the path): check that drive isn't full, that the folder exists and is writable, and that antivirus or a cloud-sync client (OneDrive, Dropbox) isn't locking it — add a VoiceStudio exclusion if you use one. If your data directory sits on a synced or network drive, move it to a local one.",
|
||||
@@ -308,6 +311,23 @@ _CONTEXT_FREE_HINT_CLASSES = frozenset({
|
||||
# a Windows virtual-memory setting rather than a connectivity problem, and
|
||||
# the detailed hint we already had for it never reached them.
|
||||
"WINDOWS_PAGING_FILE_TOO_SMALL",
|
||||
# #1957: triggered by WinError 448 or the literal "untrusted mount
|
||||
# point" — both unmistakable, and it reaches the user as a bare
|
||||
# download failure with only the OS sentence attached.
|
||||
"WINDOWS_UNTRUSTED_MOUNT",
|
||||
# #1826: torch's own conv wording, which nothing else produces, and it
|
||||
# reaches the user through the generic 500.
|
||||
"INPUT_TOO_SHORT",
|
||||
# #1879: matched on wording no other failure produces, and it reaches the
|
||||
# user as a bare 400 carrying only the library sentence.
|
||||
"CLONE_REFERENCE_MISSING",
|
||||
# Its trigger is a VoiceStudio-authored sentence — "the TTS model cache
|
||||
# for … is incomplete" plus "could not be auto-repaired" / "weights
|
||||
# missing" — so it cannot be produced by an unrelated library. The 500
|
||||
# handler is the surface a corrupt cache actually reaches, and dropping
|
||||
# its hint there would leave the user with no way to know a redownload
|
||||
# is the fix.
|
||||
"MODEL_CACHE_CORRUPT",
|
||||
})
|
||||
|
||||
|
||||
@@ -569,6 +589,34 @@ def classify(reason: str) -> str:
|
||||
or "application control policy" in low
|
||||
):
|
||||
return "WINDOWS_APP_CONTROL_BLOCKED"
|
||||
# #1957: the path to a download or output file crosses a mount point
|
||||
# Windows will not traverse (Dev Drive, mounted VHD/ReFS, a junction into
|
||||
# another profile). Matched on the numeric code first because the OS
|
||||
# translates the sentence, with the English phrase as a fallback.
|
||||
if "[winerror 448]" in low or "untrusted mount point" in low:
|
||||
return "WINDOWS_UNTRUSTED_MOUNT"
|
||||
# #1826: a degenerate-length input reaches a conv layer whose kernel is
|
||||
# wider than the tensor, and torch says so in its own terms — "Calculated
|
||||
# padded input size per channel: (1). Kernel size: (2). Kernel size can't
|
||||
# be greater than actual input size". That arrived doubly wrapped in
|
||||
# "Underlying error:" and told the user nothing they could act on, when
|
||||
# the fix is simply "type more than one character".
|
||||
if "kernel size can't be greater than actual input size" in low or (
|
||||
"calculated padded input size per channel" in low
|
||||
):
|
||||
return "INPUT_TOO_SHORT"
|
||||
# #1879: mlx-audio (and the Chatterbox-family models under it) raise a
|
||||
# bare ValueError naming their own parameters — "No conditionals
|
||||
# available. Either provide audio_prompt/audio_prompt_sr ... or ensure
|
||||
# conds.safetensors is in the model directory." The generate route passed
|
||||
# that straight through as the 400 detail, so the user was told to supply
|
||||
# an argument they have no way to name and to check for a file they have
|
||||
# never heard of. What actually happened is "you asked to clone without a
|
||||
# reference clip".
|
||||
if "no conditionals available" in low or (
|
||||
"audio_prompt" in low and "conds.safetensors" in low
|
||||
):
|
||||
return "CLONE_REFERENCE_MISSING"
|
||||
# #1221: libsndfile failed an OS-level audio read/write. Its own wording is
|
||||
# a bare "System error.", so match the library name — audio_io already
|
||||
# prefixes the target path and free space onto the write-path failures.
|
||||
|
||||
@@ -4,7 +4,14 @@ from __future__ import annotations
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
from typing import BinaryIO, Callable
|
||||
import time
|
||||
from typing import Any, BinaryIO, Callable, Optional
|
||||
|
||||
# Poll cadence for the Windows pipe watcher. Exit latency after the desktop
|
||||
# closes its end is bounded by this; the desktop's own kill-on-close Job is the
|
||||
# hard backstop, so a quarter second is plenty and costs nothing measurable.
|
||||
WINDOWS_PIPE_POLL_INTERVAL_S = 0.25
|
||||
_FILE_TYPE_PIPE = 3 # winbase.h FILE_TYPE_PIPE
|
||||
|
||||
|
||||
def _watch_parent_pipe(reader: BinaryIO, exit_process: Callable[[int], None]) -> None:
|
||||
@@ -18,6 +25,64 @@ def _watch_parent_pipe(reader: BinaryIO, exit_process: Callable[[int], None]) ->
|
||||
exit_process(0)
|
||||
|
||||
|
||||
def _watch_parent_pipe_handle(
|
||||
handle: int,
|
||||
exit_process: Callable[[int], None],
|
||||
*,
|
||||
peek: Optional[Callable[[int], Any]] = None,
|
||||
read_file: Optional[Callable[[int, int], Any]] = None,
|
||||
sleep: Callable[[float], None] = time.sleep,
|
||||
interval: float = WINDOWS_PIPE_POLL_INTERVAL_S,
|
||||
) -> None:
|
||||
"""Windows twin of :func:`_watch_parent_pipe` that never leaves a read
|
||||
pending on the pipe.
|
||||
|
||||
A synchronous ``ReadFile`` parked on the stdin pipe — whether issued through
|
||||
the C runtime's ``read()`` or straight to the kernel — deadlocks the
|
||||
OpenBLAS DLL initializer that ``import torch`` reaches (numpy's
|
||||
``_multiarray_umath``) in the startup worker: every desktop-spawned backend
|
||||
on Windows froze at "Loading ML runtime (PyTorch)" while the identical
|
||||
command from a terminal, with no stdin pipe and no watchdog, started in
|
||||
seconds. A thread that merely sleeps does not trigger it; only the pending
|
||||
read on that pipe does. So instead of blocking in a read, poll with
|
||||
``PeekNamedPipe``: it returns immediately, holds no I/O on the file object,
|
||||
drains any keepalive bytes the desktop might write, and fails with
|
||||
``ERROR_BROKEN_PIPE`` the moment the desktop closes its end — which is the
|
||||
same EOF signal the POSIX reader gets.
|
||||
"""
|
||||
if peek is None or read_file is None:
|
||||
import _winapi # Windows-only stdlib module; the caller gates on the platform
|
||||
|
||||
peek = peek or _winapi.PeekNamedPipe
|
||||
read_file = read_file or _winapi.ReadFile
|
||||
try:
|
||||
while True:
|
||||
available, _ = peek(handle)
|
||||
if available:
|
||||
# Bytes are already buffered, so this read cannot block.
|
||||
read_file(handle, available)
|
||||
else:
|
||||
sleep(interval)
|
||||
except OSError:
|
||||
# ERROR_BROKEN_PIPE (109) is how the closed parent end surfaces here.
|
||||
pass
|
||||
exit_process(0)
|
||||
|
||||
|
||||
def _windows_pipe_handle(reader: Any) -> Optional[int]:
|
||||
"""The OS handle behind ``reader`` when it is a pipe, else None."""
|
||||
try:
|
||||
import msvcrt
|
||||
import _winapi
|
||||
|
||||
handle = msvcrt.get_osfhandle(reader.fileno())
|
||||
if _winapi.GetFileType(handle) != _FILE_TYPE_PIPE:
|
||||
return None
|
||||
return handle
|
||||
except (OSError, ValueError, AttributeError, ImportError):
|
||||
return None
|
||||
|
||||
|
||||
def arm_desktop_parent_watchdog() -> bool:
|
||||
"""Use stdin EOF as an unforgeable parent-liveness signal for desktop runs."""
|
||||
if os.environ.get("OMNIVOICE_DESKTOP_CONTAINED") != "1":
|
||||
@@ -25,9 +90,18 @@ def arm_desktop_parent_watchdog() -> bool:
|
||||
reader = getattr(sys.stdin, "buffer", None)
|
||||
if reader is None:
|
||||
return False
|
||||
target: Callable[..., None] = _watch_parent_pipe
|
||||
args: tuple = (reader, os._exit)
|
||||
if os.name == "nt":
|
||||
handle = _windows_pipe_handle(reader)
|
||||
if handle is not None:
|
||||
target = _watch_parent_pipe_handle
|
||||
args = (handle, os._exit)
|
||||
# A non-pipe stdin (file, NUL) cannot have a read pending against a
|
||||
# pipe file object, so the blocking reader stays correct there.
|
||||
threading.Thread(
|
||||
target=_watch_parent_pipe,
|
||||
args=(reader, os._exit),
|
||||
target=target,
|
||||
args=args,
|
||||
name="desktop-parent-watchdog",
|
||||
daemon=True,
|
||||
).start()
|
||||
|
||||
@@ -7,6 +7,7 @@ only the unguessable capability token crosses loopback HTTP.
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import secrets
|
||||
@@ -14,6 +15,8 @@ import stat
|
||||
|
||||
from core.config import DATA_DIR
|
||||
|
||||
logger = logging.getLogger("omnivoice.path_authorization")
|
||||
|
||||
_TOKEN_RE = re.compile(r"[0-9a-f]{64}\Z")
|
||||
_KINDS = {
|
||||
"models_dir",
|
||||
@@ -40,25 +43,49 @@ def consume(token: str, expected_kind: str) -> str:
|
||||
if expected_kind not in _KINDS or not _TOKEN_RE.fullmatch(token or ""):
|
||||
raise PathAuthorizationError("Invalid or expired desktop authorization")
|
||||
root = _AUTH_DIR
|
||||
# Distinguish "the store exists but this token isn't in it" (expired /
|
||||
# already consumed / never issued — normal, no server-side signal) from
|
||||
# "the store doesn't exist at all" (the desktop app and this backend are
|
||||
# very likely pointed at different data directories, e.g. a dev backend
|
||||
# started without OMNIVOICE_DATA_DIR, or a stale custom data folder — see
|
||||
# #1781). The client-facing message is byte-identical either way (never
|
||||
# leak local filesystem paths, or even which case occurred, over HTTP —
|
||||
# CWE-200); the mismatch case additionally gets a server log line so it's
|
||||
# diagnosable instead of a silent 403. That log line is deliberately
|
||||
# path-free too (CWE-532: per-user filesystem paths, e.g. a home
|
||||
# directory username, are sensitive and don't belong in application
|
||||
# logs) — it names the failure mode, not the directory.
|
||||
try:
|
||||
entries = os.scandir(root)
|
||||
except FileNotFoundError as exc:
|
||||
logger.warning(
|
||||
"path authorization store does not exist; the desktop app and "
|
||||
"this backend likely resolved different data directories "
|
||||
"(see #1781)"
|
||||
)
|
||||
raise PathAuthorizationError("Invalid or expired desktop authorization") from exc
|
||||
except OSError as exc:
|
||||
raise PathAuthorizationError("Invalid or expired desktop authorization") from exc
|
||||
candidate = None
|
||||
try:
|
||||
for entry in os.scandir(root):
|
||||
if not _TOKEN_RE.fullmatch(entry.name.removesuffix(".json")):
|
||||
continue
|
||||
if not entry.is_file(follow_symlinks=False):
|
||||
continue
|
||||
try:
|
||||
with open(entry.path, "r", encoding="utf-8") as handle:
|
||||
probe = json.load(handle)
|
||||
except (OSError, UnicodeError, json.JSONDecodeError):
|
||||
continue # Ignore corrupt/stale capabilities; they authorize nothing.
|
||||
if isinstance(probe, dict) and secrets.compare_digest(
|
||||
str(probe.get("token", "")), token
|
||||
):
|
||||
candidate = entry.path
|
||||
break
|
||||
with entries:
|
||||
for entry in entries:
|
||||
if not _TOKEN_RE.fullmatch(entry.name.removesuffix(".json")):
|
||||
continue
|
||||
if not entry.is_file(follow_symlinks=False):
|
||||
continue
|
||||
try:
|
||||
with open(entry.path, "r", encoding="utf-8") as handle:
|
||||
probe = json.load(handle)
|
||||
except (OSError, UnicodeError, json.JSONDecodeError):
|
||||
continue # Ignore corrupt/stale capabilities; they authorize nothing.
|
||||
if isinstance(probe, dict) and secrets.compare_digest(
|
||||
str(probe.get("token", "")), token
|
||||
):
|
||||
candidate = entry.path
|
||||
break
|
||||
if candidate is None:
|
||||
raise OSError("capability not found")
|
||||
raise PathAuthorizationError("Invalid or expired desktop authorization")
|
||||
claimed = os.path.join(root, f".consuming-{os.getpid()}-{secrets.token_hex(16)}")
|
||||
os.replace(candidate, claimed)
|
||||
except OSError as exc:
|
||||
|
||||
@@ -91,3 +91,53 @@ def resolve(key: str, *, env: Optional[str] = None, default: Any = None) -> Any:
|
||||
if v:
|
||||
return v
|
||||
return get(key, default)
|
||||
|
||||
|
||||
# ── external-override detection (#1787 review fix) ──────────────────────────
|
||||
# restore_env() below uses os.environ.setdefault(), so a value already present
|
||||
# in the process's environment (shell profile, `.env`, Docker `-e`, systemd
|
||||
# unit, …) silently wins over anything saved in prefs.json — the setdefault
|
||||
# call is a no-op. That is the right behavior (env stays authoritative,
|
||||
# matching resolve()'s contract above), but a Settings control that persists a
|
||||
# value to prefs.json must not tell the user it "took effect after restart"
|
||||
# when an external source will keep shadowing it on every future restart too.
|
||||
#
|
||||
# _EXTERNALLY_PROVIDED records, once per process start, every bare key that
|
||||
# was ALREADY present in os.environ the moment restore_env() ran — i.e.
|
||||
# before our own setdefault() calls could have put it there, and before any
|
||||
# value our Settings UI ever wrote (Settings only ever writes prefs.json plus
|
||||
# the CURRENT process's os.environ; it never touches a shell profile or `.env`
|
||||
# file). Snapshotting unconditionally — not only for keys prefs.json already
|
||||
# has an entry for — means is_env_shadowed() also answers correctly for a key
|
||||
# a user is about to save for the FIRST time. Membership is stable for the
|
||||
# life of the process (nothing removes an inherited env var), and since a
|
||||
# plain restart re-inherits the same shell / container environment, it is
|
||||
# also a reliable predictor for the NEXT start: if the external source is
|
||||
# still exporting the key, the next restart will be shadowed again the same
|
||||
# way.
|
||||
_EXTERNALLY_PROVIDED: frozenset[str] = frozenset()
|
||||
|
||||
|
||||
def restore_env(data: dict) -> None:
|
||||
"""Restore ``env.*`` prefs into ``os.environ`` (startup only).
|
||||
|
||||
Called once from main.py's ``env_prefs`` step, before any user code reads
|
||||
``os.environ``. Snapshots which keys were already externally provided —
|
||||
see :func:`is_env_shadowed` — then applies every saved ``env.*`` pref via
|
||||
``setdefault`` (never overriding an explicitly-set env var).
|
||||
"""
|
||||
global _EXTERNALLY_PROVIDED
|
||||
_EXTERNALLY_PROVIDED = frozenset(os.environ.keys())
|
||||
for k, v in data.items():
|
||||
if not k.startswith("env.") or not v:
|
||||
continue
|
||||
os.environ.setdefault(k[len("env."):], str(v))
|
||||
|
||||
|
||||
def is_env_shadowed(key: str) -> bool:
|
||||
"""Whether *key* was already present in the environment from a source
|
||||
other than our own prefs restore, as of the last time :func:`restore_env`
|
||||
ran. If prefs.json holds (or will hold) a saved value for *key*, that
|
||||
value is being silently ignored — and will be again on the next restart —
|
||||
unless the external source is removed."""
|
||||
return key in _EXTERNALLY_PROVIDED
|
||||
|
||||
@@ -32,8 +32,8 @@ def stream_failure(code: str) -> dict[str, object]:
|
||||
"code": "generation_timeout",
|
||||
"detail": (
|
||||
"Generation exceeded the compute-time limit. The backend is "
|
||||
"still running; try a shorter passage or raise the generation "
|
||||
"timeout."
|
||||
"still running; try a shorter passage, or raise the "
|
||||
"compute-time budget in Settings → Performance & Device."
|
||||
),
|
||||
"retryable": True,
|
||||
},
|
||||
@@ -94,6 +94,16 @@ def stream_generation_failure(error: BaseException | object) -> dict[str, object
|
||||
replace the failure being diagnosed.
|
||||
"""
|
||||
payload = stream_failure("generation_failed")
|
||||
if isinstance(error, BaseException):
|
||||
# The exception's TYPE NAME, never its message. Two failures that both
|
||||
# render the floor message "Generation failed. Check the selected
|
||||
# engine and try again." are indistinguishable in an auto-filed report,
|
||||
# so every unclassified streaming failure arrives as the same issue and
|
||||
# none of them can be triaged (#1800). A class name is VoiceStudio-safe
|
||||
# by the same reasoning that already puts it on the wire as
|
||||
# `error_class` in the dub routes and on the analytics allowlist: it is
|
||||
# a Python type, not user text, and no substring of `error` is copied.
|
||||
payload["error_class"] = type(error).__name__
|
||||
try:
|
||||
enriched = public_exception_response(error, fallback=str(payload["detail"]))
|
||||
except Exception:
|
||||
@@ -149,12 +159,31 @@ def public_exception_response(error: BaseException, *, fallback: str) -> dict[st
|
||||
Classification may inspect the private diagnostic locally, but response
|
||||
values come exclusively from VoiceStudio-owned constants. No substring of
|
||||
``error`` is copied into the payload.
|
||||
|
||||
Every caller is a CONTEXT-FREE surface — the global 500 handler, the
|
||||
streaming generate error frame, the dub GPU-OOM 503 — so the topic is
|
||||
filtered through ``failure._CONTEXT_FREE_HINT_CLASSES`` before its hint is
|
||||
attached. Without that filter a topic whose trigger is a generic phrase
|
||||
stamps a confidently wrong remediation on an unrelated failure: #1943 is a
|
||||
macOS mlx-audio TTS 500 that came back advising the user that "the
|
||||
connection to the video server dropped mid-download", because
|
||||
VIDEO_DOWNLOAD_NETWORK triggers on a bare "timed out" / "connection
|
||||
reset". The allowlist already existed and already named that class as the
|
||||
example of what must not appear here; only :func:`failure.append_hint`
|
||||
honoured it, and this helper replaced ``append_hint`` on the 500 path
|
||||
without carrying the rule across.
|
||||
|
||||
HF_MIRROR_UNREACHABLE is allowed alongside it: its hint is dynamic (it
|
||||
names the configured mirror) and its trigger requires that a mirror is
|
||||
configured at all, so it cannot fire on an unrelated failure (#874).
|
||||
"""
|
||||
from core.failure import classify, public_hint_for_topic
|
||||
from core.failure import _CONTEXT_FREE_HINT_CLASSES, classify, public_hint_for_topic
|
||||
|
||||
try:
|
||||
topic = classify(str(error))
|
||||
hint = public_hint_for_topic(topic)
|
||||
if topic and topic not in _CONTEXT_FREE_HINT_CLASSES and topic != "HF_MIRROR_UNREACHABLE":
|
||||
topic = ""
|
||||
hint = public_hint_for_topic(topic) if topic else ""
|
||||
except Exception:
|
||||
topic = ""
|
||||
hint = ""
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
"""The PyTorch wheel index VoiceStudio installs CUDA builds from.
|
||||
|
||||
A local-version pin such as ``torch==2.9.1+cu128`` exists only on PyTorch's
|
||||
own index, never on PyPI. The app's own ``pyproject.toml`` routes torch there
|
||||
through ``[tool.uv.sources]``, but a sidecar engine is installed with
|
||||
``uv pip install`` into its own venv, which knows nothing about that config —
|
||||
so every CUDA-pinned sidecar install has to name the index itself.
|
||||
|
||||
MOSS-TTS-v1.5's install did not, and its ``[torch-runtime]`` extra
|
||||
(``torch==2.9.1+cu128``) could never resolve: ``uv pip compile`` reports it
|
||||
unsatisfiable without this index and resolves it with it. One definition here,
|
||||
imported by the one-click installer and by the engine's own bootstrap, so the
|
||||
two cannot drift apart again. ``tests/test_sidecar_install.py`` pins the URL
|
||||
to the ``pytorch-cuda`` index declared in the app's ``pyproject.toml``.
|
||||
"""
|
||||
|
||||
PYTORCH_CU128_INDEX_URL = "https://download.pytorch.org/whl/cu128"
|
||||
|
||||
# `unsafe-best-match`: the PyTorch index also mirrors common dependencies
|
||||
# (numpy, pillow, sympy, …) at a narrower range of versions than PyPI. uv's
|
||||
# default first-index strategy would stop at whichever index lists a name first
|
||||
# and could pin an old mirror copy or fail outright. The index is PyTorch's
|
||||
# official one, so the dependency-confusion risk the name warns about does not
|
||||
# apply to it.
|
||||
UV_PIP_CU128_ARGS: tuple[str, ...] = (
|
||||
"--extra-index-url",
|
||||
PYTORCH_CU128_INDEX_URL,
|
||||
"--index-strategy",
|
||||
"unsafe-best-match",
|
||||
)
|
||||
|
||||
PYTORCH_CPU_INDEX_URL = "https://download.pytorch.org/whl/cpu"
|
||||
|
||||
# For an engine that runs torch only on the CPU (PocketTTS). On Linux, PyPI's
|
||||
# torch is the CUDA build and pulls ~15 NVIDIA packages the engine never uses;
|
||||
# this index serves `+cpu` builds for Linux and Windows and the regular build
|
||||
# for macOS.
|
||||
UV_PIP_CPU_ARGS: tuple[str, ...] = (
|
||||
"--extra-index-url",
|
||||
PYTORCH_CPU_INDEX_URL,
|
||||
"--index-strategy",
|
||||
"unsafe-best-match",
|
||||
)
|
||||
@@ -24,7 +24,7 @@ from pathlib import Path
|
||||
# tests/test_app_version.py::test_all_version_files_in_lockstep and bumped by
|
||||
# release.yml's version-bump job, so it stays equal to
|
||||
# pyproject/tauri.conf/Cargo/package.json.
|
||||
_FALLBACK_VERSION = "0.5.1"
|
||||
_FALLBACK_VERSION = "0.5.2"
|
||||
|
||||
|
||||
def _fallback_version() -> str:
|
||||
|
||||
@@ -0,0 +1,590 @@
|
||||
"""audio.cpp TTS backend — Breeze-TTS-2 via a managed native server.
|
||||
|
||||
audio.cpp (0xShug0/audio.cpp) is a pure-C++ ggml runtime: prebuilt
|
||||
``audiocpp_server`` binaries for Windows/macOS/Linux, no Python venv, no
|
||||
``transformers`` pin — so this engine needs neither the venv-isolation
|
||||
(``engines.dots_tts``) nor the per-generate CLI-spawn (``engines
|
||||
.omnivoice_gguf``) patterns. The parent instead:
|
||||
|
||||
1. resolves the binary + GGUF model (``bootstrap.py``),
|
||||
2. spawns ONE long-lived ``audiocpp_server`` on 127.0.0.1 (lazy model load,
|
||||
so model memory is only held after the first generate), and
|
||||
3. speaks its OpenAI-style ``POST /v1/audio/speech`` per generate.
|
||||
|
||||
v1 serves the ``breeze_tts`` family only (Breeze-TTS-2, en+zh, voice clone
|
||||
+ voice design + voice direction). The server is task-agnostic on the
|
||||
speech route — reference-audio presence selects clone/direction vs design —
|
||||
so a single ``task: tts`` model entry covers all three modes.
|
||||
|
||||
License honesty: Breeze-TTS-2 weights (``BreezeBlue/Breeze-TTS-2`` and the
|
||||
audio.cpp GGUF repack) are RESEARCH AND NON-COMMERCIAL ONLY
|
||||
(``BreezeBlue Research and Non-Commercial License``); only the audio.cpp
|
||||
code is Apache-2.0. There is no in-tree acceptance dialog for this engine
|
||||
yet (settings ``/license`` allow-list), so the restriction is surfaced in
|
||||
the display name, the install hint, and ``docs/engines/audio-cpp.md`` —
|
||||
not silently.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import atexit
|
||||
import base64
|
||||
import io
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import secrets
|
||||
|
||||
# Used only for stream constants; spawn_owned performs the process launch.
|
||||
import subprocess # nosec B404
|
||||
import threading
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
from core.contained_subprocess import spawn_owned
|
||||
from services.tts_backend import TTSBackend, TTSInputError
|
||||
|
||||
if TYPE_CHECKING:
|
||||
import torch
|
||||
|
||||
logger = logging.getLogger("omnivoice.audiocpp")
|
||||
|
||||
#: Engine id in the TTS registry.
|
||||
ENGINE_ID = "audiocpp"
|
||||
|
||||
#: How long to wait for ``/health`` after spawning the server (first spawn
|
||||
#: extracts nothing heavy — the model loads lazily on first generate).
|
||||
_HEALTH_TIMEOUT_S = 120.0
|
||||
|
||||
#: Finish the inner HTTP request before the canonical generation guard can
|
||||
#: abandon its worker thread. This leaves enough time to terminate the owned
|
||||
#: native process and release its model memory synchronously.
|
||||
_TERMINATE_GRACE_S = 5.0
|
||||
_TERMINATE_KILL_S = 5.0
|
||||
_GENERATE_TIMEOUT_MARGIN_S = (
|
||||
_TERMINATE_GRACE_S + _TERMINATE_KILL_S + 5.0
|
||||
)
|
||||
|
||||
|
||||
# ── pure request/config builders (unit-tested, no I/O) ──────────────────────
|
||||
|
||||
|
||||
def _cpu_thread_count() -> int:
|
||||
"""Use up to 16 physical cores, with a stdlib fallback."""
|
||||
try:
|
||||
import psutil
|
||||
|
||||
cores = psutil.cpu_count(logical=False)
|
||||
except (ImportError, OSError):
|
||||
cores = None
|
||||
return min(16, max(1, cores or os.cpu_count() or 1))
|
||||
|
||||
|
||||
def _device_min_vram_gb(device) -> float:
|
||||
"""Dedicated-memory comfort floor for one discovered native device."""
|
||||
return 6.0 if (
|
||||
device
|
||||
and device.kind == "GPU"
|
||||
and (
|
||||
device.backend == "vulkan"
|
||||
or device.hardware_family in {"cuda", "rocm"}
|
||||
)
|
||||
) else 0.0
|
||||
|
||||
|
||||
def build_server_config(
|
||||
*, model_id: str, family: str, model_path: str, port: int,
|
||||
backend: str = "cpu", device: int = 0,
|
||||
execution_target: str | None = None,
|
||||
) -> dict:
|
||||
"""``server.json`` dict for the managed ``audiocpp_server``.
|
||||
|
||||
``lazy_load`` defers the ~4.73 GiB GGUF load to the first generate;
|
||||
``max_loaded_models: 1`` bounds residency to the one model we serve.
|
||||
"""
|
||||
return {
|
||||
"host": "127.0.0.1",
|
||||
"port": port,
|
||||
"backend": backend,
|
||||
"device": device,
|
||||
# The pinned CPU runtime scales strongly through 16 workers while
|
||||
# producing byte-identical audio.
|
||||
"threads": _cpu_thread_count()
|
||||
if (execution_target or backend) == "cpu" else 1,
|
||||
"lazy_load": True,
|
||||
"max_loaded_models": 1,
|
||||
"models": [
|
||||
{
|
||||
"id": model_id,
|
||||
"family": family,
|
||||
"path": model_path,
|
||||
"task": "tts",
|
||||
"mode": "offline",
|
||||
}
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def build_speech_payload(
|
||||
*, model_id: str, text: str, ref_audio: str | None = None,
|
||||
ref_text: str | None = None, instructions: str | None = None,
|
||||
guidance_scale: float | None = None, seed: int | None = None,
|
||||
) -> dict:
|
||||
"""``POST /v1/audio/speech`` JSON body.
|
||||
|
||||
Field spellings verified against ``app/server/runtime.cpp``
|
||||
(``build_speech_request``): ``instructions`` (plural, OpenAI spelling)
|
||||
feeds the ``instruction`` request option; ``reference_text`` and
|
||||
``guidance_scale``/``seed`` pass through top-level; ``voice_ref`` takes
|
||||
a ``{"type": "path", ...}`` object so the reference stays on disk
|
||||
(the 5 MiB base64 cap never bites). ``response_format: json`` returns
|
||||
the WAV base64-in-JSON — one round trip, no binary framing.
|
||||
"""
|
||||
payload: dict[str, Any] = {
|
||||
"model": model_id,
|
||||
"input": text,
|
||||
"response_format": "json",
|
||||
}
|
||||
if instructions:
|
||||
payload["instructions"] = instructions
|
||||
if ref_audio:
|
||||
payload["voice_ref"] = {"type": "path", "path": str(ref_audio)}
|
||||
if ref_text:
|
||||
payload["reference_text"] = ref_text
|
||||
if guidance_scale is not None:
|
||||
payload["guidance_scale"] = float(guidance_scale)
|
||||
if seed is not None:
|
||||
payload["seed"] = int(seed)
|
||||
return payload
|
||||
|
||||
|
||||
def decode_speech_json(obj: dict) -> tuple[int, object]:
|
||||
"""``(sample_rate, mono float32 numpy)`` from a ``response_format=json``
|
||||
speech body. Raises ``ValueError`` on a server error payload."""
|
||||
if not isinstance(obj, dict):
|
||||
raise TypeError(f"audio.cpp speech reply is not JSON: {obj!r:.120}")
|
||||
if "audio" not in obj:
|
||||
raise ValueError(f"audio.cpp speech failed: {obj.get('error', obj)!r:.300}")
|
||||
import numpy as np
|
||||
import soundfile as sf
|
||||
|
||||
wav_bytes = base64.b64decode(obj["audio"])
|
||||
wav, sr = sf.read(io.BytesIO(wav_bytes), dtype="float32", always_2d=False)
|
||||
wav = np.asarray(wav, dtype=np.float32)
|
||||
if wav.ndim > 1:
|
||||
wav = wav.mean(axis=-1)
|
||||
return int(sr), wav
|
||||
|
||||
|
||||
# ── backend ─────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
class AudioCPPBackend(TTSBackend):
|
||||
"""Breeze-TTS-2 through a parent-managed ``audiocpp_server``."""
|
||||
|
||||
id = ENGINE_ID
|
||||
display_name = (
|
||||
"audio.cpp · Breeze-TTS-2 (native GGUF, en+zh, clone+design; "
|
||||
"weights research/non-commercial)"
|
||||
)
|
||||
supports_voice_design = True
|
||||
applies_own_mastering = True # model-decoded 24 kHz studio output
|
||||
gpu_compat = ("cpu",)
|
||||
runs_out_of_process = True
|
||||
# Same marker SubprocessBackend sets: this engine lives in another OS
|
||||
# process. Consumers only branch the matrix label and the self-test
|
||||
# route (spawn-and-ping instead of in-process synth) — both correct
|
||||
# here; nothing assumes the stdio protocol from it.
|
||||
_is_subprocess_isolated = True
|
||||
_DEFAULT_SAMPLE_RATE = 24000 # Breeze-TTS-2 native rate
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._proc: Any | None = None
|
||||
self._port: int | None = None
|
||||
self._server_model_id: str | None = None
|
||||
self._sr = self._DEFAULT_SAMPLE_RATE
|
||||
self._lock = threading.RLock()
|
||||
self._server_json: Path | None = None
|
||||
self._selection = None
|
||||
self._device = None
|
||||
self._provider = None
|
||||
|
||||
# ── availability ────────────────────────────────────────────────────
|
||||
|
||||
@classmethod
|
||||
def is_available(cls) -> tuple[bool, str]:
|
||||
from engines.audiocpp import bootstrap
|
||||
|
||||
try:
|
||||
bootstrap.resolve_server_binary()
|
||||
bootstrap.resolve_model_file()
|
||||
except RuntimeError as exc:
|
||||
return False, str(exc)
|
||||
return True, "ready"
|
||||
|
||||
@classmethod
|
||||
def runtime_compute_profile(cls, caps) -> dict:
|
||||
from dataclasses import replace
|
||||
|
||||
from engines.audiocpp import bootstrap
|
||||
from services.engine_routing import low_vram_caveat
|
||||
|
||||
try:
|
||||
selection = bootstrap.resolve_compute_selection(caps)
|
||||
targets = bootstrap.runtime_targets()
|
||||
except RuntimeError as exc:
|
||||
return {
|
||||
"gpu_compat": cls.gpu_compat,
|
||||
"min_vram_gb": 0.0,
|
||||
"effective_device": "cpu",
|
||||
"routing_status": "unavailable",
|
||||
"routing_reason": str(exc),
|
||||
"runtime_backend": None,
|
||||
"runtime_device_index": None,
|
||||
"runtime_device_name": None,
|
||||
"runtime_hardware_family": None,
|
||||
"runtime_vram_gb": None,
|
||||
"runtime_device_verified": False,
|
||||
}
|
||||
selected = selection.device
|
||||
accelerated = selected.target != "cpu"
|
||||
min_vram_gb = _device_min_vram_gb(selected)
|
||||
dedicated = min_vram_gb > 0
|
||||
reason = selection.fallback_reason
|
||||
if accelerated and dedicated and reason is None:
|
||||
selected_caps = replace(
|
||||
caps,
|
||||
device_name=selected.name,
|
||||
vram_gb=selection.verified_vram_gb,
|
||||
)
|
||||
reason = low_vram_caveat(
|
||||
selected_caps,
|
||||
min_vram_gb,
|
||||
family=selected.hardware_family,
|
||||
vram_gb=selection.verified_vram_gb,
|
||||
)
|
||||
status = "accelerated" if accelerated else (
|
||||
"cpu_fallback" if selection.fallback_reason else "cpu_only"
|
||||
)
|
||||
return {
|
||||
"gpu_compat": targets,
|
||||
"min_vram_gb": min_vram_gb,
|
||||
"effective_device": selected.target,
|
||||
"routing_status": status,
|
||||
"routing_reason": reason,
|
||||
"runtime_backend": selected.backend,
|
||||
"runtime_device_index": selected.index,
|
||||
"runtime_device_name": selected.name,
|
||||
"runtime_hardware_family": selected.hardware_family,
|
||||
"runtime_vram_gb": selection.verified_vram_gb,
|
||||
"runtime_device_verified": selection.verified_vram_gb > 0,
|
||||
}
|
||||
|
||||
# ── TTSBackend protocol ─────────────────────────────────────────────
|
||||
|
||||
@property
|
||||
def sample_rate(self) -> int:
|
||||
return self._sr
|
||||
|
||||
@property
|
||||
def supported_languages(self) -> list[str]:
|
||||
return ["en", "zh"]
|
||||
|
||||
def model_identity(self) -> str | None:
|
||||
from engines.audiocpp import bootstrap
|
||||
|
||||
return f"{bootstrap.FAMILY}/{bootstrap.package_filename()}"
|
||||
|
||||
# ── server lifecycle ────────────────────────────────────────────────
|
||||
|
||||
def _base_url(self) -> str:
|
||||
return f"http://127.0.0.1:{self._port}"
|
||||
|
||||
def _ensure_loaded(self) -> None:
|
||||
"""Spawn the server (once) and wait for ``/health``. Idempotent."""
|
||||
with self._lock:
|
||||
if self._proc is not None and self._proc.poll() is None:
|
||||
return
|
||||
self._proc = None # stale handle — respawn below
|
||||
from engines.audiocpp import bootstrap
|
||||
|
||||
binary = bootstrap.resolve_server_binary()
|
||||
selection = bootstrap.resolve_compute_selection()
|
||||
model_file = bootstrap.resolve_model_file()
|
||||
self._port = bootstrap.server_port()
|
||||
# The random model id is a per-launch challenge. Before sending
|
||||
# speech text or a reference path, _verify_server_identity asks
|
||||
# /v1/models to prove this is the child configured by this process,
|
||||
# not an unrelated listener that pre-bound the loopback port.
|
||||
self._server_model_id = f"{bootstrap.MODEL_ID}-{secrets.token_hex(16)}"
|
||||
config = build_server_config(
|
||||
model_id=self._server_model_id,
|
||||
family=bootstrap.FAMILY,
|
||||
model_path=str(model_file),
|
||||
port=self._port,
|
||||
backend=selection.device.backend,
|
||||
device=selection.device.index,
|
||||
execution_target=selection.device.target,
|
||||
)
|
||||
self._selection = selection
|
||||
self._device = selection.device.target
|
||||
self._provider = selection.device.backend
|
||||
from core.config import DATA_DIR
|
||||
|
||||
workdir = Path(str(DATA_DIR)) / "audiocpp"
|
||||
workdir.mkdir(parents=True, exist_ok=True)
|
||||
self._server_json = workdir / "server.json"
|
||||
flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC
|
||||
config_fd = os.open(self._server_json, flags, 0o600)
|
||||
try:
|
||||
if os.name != "nt":
|
||||
os.fchmod(config_fd, 0o600)
|
||||
with os.fdopen(config_fd, "w", encoding="utf-8") as config_fh:
|
||||
config_fd = -1
|
||||
json.dump(config, config_fh, indent=2)
|
||||
finally:
|
||||
if config_fd >= 0:
|
||||
os.close(config_fd)
|
||||
log_path = workdir / "server.log"
|
||||
logger.info(
|
||||
"audio.cpp: starting %s (backend=%s, device=%d, port=%d, model=%s)",
|
||||
binary.name, selection.device.backend, selection.device.index,
|
||||
self._port, model_file.name,
|
||||
)
|
||||
with open(log_path, "ab") as log_fh:
|
||||
self._proc = spawn_owned(
|
||||
[str(binary), "--config", str(self._server_json)],
|
||||
stdout=log_fh,
|
||||
stderr=subprocess.STDOUT,
|
||||
stdin=subprocess.DEVNULL,
|
||||
)
|
||||
atexit.register(self._terminate_server)
|
||||
self._wait_for_health()
|
||||
|
||||
def _wait_for_health(self) -> None:
|
||||
if self._proc is None or self._port is None:
|
||||
raise RuntimeError("managed audio.cpp server was not started")
|
||||
deadline = time.monotonic() + _HEALTH_TIMEOUT_S
|
||||
last_err = "unknown"
|
||||
url = self._base_url() + "/health"
|
||||
while time.monotonic() < deadline:
|
||||
if self._proc.poll() is not None:
|
||||
raise RuntimeError(
|
||||
"audiocpp_server exited during startup "
|
||||
f"(code {self._proc.returncode}). See the server log next "
|
||||
"to server.json under the app data audiocpp/ directory — "
|
||||
"the managed port may already be in use."
|
||||
)
|
||||
try:
|
||||
# ``url`` is always the hard-coded loopback host plus a
|
||||
# validated integer port; arbitrary schemes are impossible.
|
||||
with urllib.request.urlopen(url, timeout=5) as resp: # nosec B310
|
||||
if resp.status == 200:
|
||||
self._verify_server_identity()
|
||||
if self._proc.poll() is None:
|
||||
logger.info(
|
||||
"audio.cpp: managed server is healthy on loopback"
|
||||
)
|
||||
return
|
||||
last_err = f"HTTP {resp.status}"
|
||||
except Exception as exc: # noqa: BLE001 — still starting; retry
|
||||
last_err = f"{type(exc).__name__}: {exc}"
|
||||
time.sleep(1.0)
|
||||
self._terminate_server()
|
||||
raise RuntimeError(
|
||||
f"audiocpp_server did not become healthy within "
|
||||
f"{_HEALTH_TIMEOUT_S:.0f}s (last: {last_err})."
|
||||
)
|
||||
|
||||
def _get_json(self, path: str, timeout: float = 5.0) -> dict:
|
||||
"""GET one loopback JSON endpoint without sending request content."""
|
||||
if self._port is None:
|
||||
raise RuntimeError("managed audio.cpp server port is missing")
|
||||
req = urllib.request.Request(self._base_url() + path, method="GET")
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp: # nosec B310
|
||||
obj = json.loads(resp.read().decode("utf-8"))
|
||||
if not isinstance(obj, dict):
|
||||
raise TypeError("audio.cpp returned an invalid JSON response")
|
||||
return obj
|
||||
|
||||
def _verify_server_identity(self) -> None:
|
||||
"""Prove the loopback listener owns this launch's random model id."""
|
||||
if self._proc is None or self._proc.poll() is not None:
|
||||
raise RuntimeError("managed audio.cpp server is not running")
|
||||
expected = self._server_model_id
|
||||
if not expected:
|
||||
raise RuntimeError("managed audio.cpp server identity is missing")
|
||||
obj = self._get_json("/v1/models")
|
||||
data = obj.get("data", [])
|
||||
if not isinstance(data, list):
|
||||
raise TypeError("managed audio.cpp server identity is invalid")
|
||||
model_ids = {
|
||||
item.get("id") for item in data
|
||||
if isinstance(item, dict)
|
||||
}
|
||||
if expected not in model_ids or self._proc.poll() is not None:
|
||||
raise RuntimeError(
|
||||
"loopback listener did not prove managed audio.cpp ownership"
|
||||
)
|
||||
|
||||
def _post_json(self, path: str, payload: dict, timeout: float) -> dict:
|
||||
"""Verify child ownership, then POST JSON to the managed server."""
|
||||
if self._port is None:
|
||||
raise RuntimeError("managed audio.cpp server port is missing")
|
||||
self._verify_server_identity()
|
||||
body = json.dumps(payload).encode("utf-8")
|
||||
req = urllib.request.Request(
|
||||
self._base_url() + path,
|
||||
data=body,
|
||||
headers={"Content-Type": "application/json"},
|
||||
method="POST",
|
||||
)
|
||||
try:
|
||||
# ``req`` targets only ``_base_url()`` (127.0.0.1 + validated
|
||||
# integer port), never a caller-provided URL.
|
||||
with urllib.request.urlopen(req, timeout=timeout) as resp: # nosec B310
|
||||
return json.loads(resp.read().decode("utf-8"))
|
||||
except urllib.error.HTTPError as exc:
|
||||
detail = exc.read().decode("utf-8", errors="replace")[:500]
|
||||
raise RuntimeError(
|
||||
f"audio.cpp {path} failed (HTTP {exc.code}): {detail}"
|
||||
) from exc
|
||||
except urllib.error.URLError as exc:
|
||||
if isinstance(exc.reason, TimeoutError):
|
||||
raise TimeoutError("audio.cpp request timed out") from exc
|
||||
raise
|
||||
|
||||
def _terminate_server(self) -> None:
|
||||
proc, self._proc = self._proc, None
|
||||
self._server_model_id = None
|
||||
if proc is None:
|
||||
return
|
||||
try:
|
||||
proc.terminate()
|
||||
proc.wait(timeout=_TERMINATE_GRACE_S)
|
||||
except Exception: # noqa: BLE001 — kill as last resort, never raise
|
||||
try:
|
||||
proc.kill()
|
||||
proc.wait(timeout=_TERMINATE_KILL_S)
|
||||
except Exception as exc: # noqa: BLE001 — process is already failing
|
||||
logger.debug("audio.cpp: final server kill failed: %s", exc)
|
||||
|
||||
# ── generate ────────────────────────────────────────────────────────
|
||||
|
||||
def generate(self, text: str, **kw) -> torch.Tensor:
|
||||
import torch
|
||||
from services.model_manager import (
|
||||
GENERATE_PROGRESS_GRACE_S,
|
||||
generate_timeout_s,
|
||||
report_generate_progress,
|
||||
)
|
||||
|
||||
if not text or not text.strip():
|
||||
raise TTSInputError(
|
||||
"audio.cpp: the input contains no speakable text — "
|
||||
"send at least one word."
|
||||
)
|
||||
ref_audio = kw.get("ref_audio")
|
||||
ref_text = kw.get("ref_text")
|
||||
if ref_text and not ref_audio:
|
||||
logger.info(
|
||||
"audio.cpp: ref_text supplied without ref_audio; ignoring."
|
||||
)
|
||||
ref_text = None
|
||||
|
||||
# Voice design: our `description=` (no ref) and voice direction
|
||||
# (`instruct=` + ref) both ride the server's `instructions` field —
|
||||
# verified spelling against app/server/runtime.cpp.
|
||||
instruct = kw.get("instruct") or kw.get("description") or None
|
||||
|
||||
language = kw.get("language")
|
||||
if language and str(language).strip().lower() not in {
|
||||
"auto", "en", "english", "zh", "chinese",
|
||||
}:
|
||||
logger.info(
|
||||
"audio.cpp (Breeze-TTS-2) is en+zh only; ignoring "
|
||||
"language=%r.", language,
|
||||
)
|
||||
if kw.get("speed", 1.0) != 1.0:
|
||||
logger.info("audio.cpp: speed is not supported; ignoring.")
|
||||
|
||||
request_started = time.monotonic()
|
||||
with self._lock:
|
||||
self._ensure_loaded()
|
||||
selected = self._selection.device if self._selection else None
|
||||
min_vram_gb = _device_min_vram_gb(selected)
|
||||
request_budget = generate_timeout_s(
|
||||
text,
|
||||
execution_device=selected.target if selected else "cpu",
|
||||
min_vram_gb=min_vram_gb,
|
||||
hardware_family=selected.hardware_family if selected else None,
|
||||
vram_gb=self._selection.verified_vram_gb
|
||||
if self._selection else 0.0,
|
||||
)
|
||||
if not self._server_model_id:
|
||||
raise RuntimeError("managed audio.cpp server identity is missing")
|
||||
payload = build_speech_payload(
|
||||
model_id=self._server_model_id,
|
||||
text=text,
|
||||
ref_audio=str(ref_audio) if ref_audio else None,
|
||||
ref_text=ref_text,
|
||||
instructions=instruct,
|
||||
guidance_scale=kw.get("guidance_scale", 1.0),
|
||||
seed=kw.get("seed"),
|
||||
)
|
||||
# Device discovery and server startup can consume part of the soft
|
||||
# budget. This fresh synthesis lease gives the lazy model load and
|
||||
# request a bounded window. The inner request always expires early
|
||||
# enough to reap the owned server before the outer guard abandons us.
|
||||
report_generate_progress()
|
||||
soft_remaining = request_budget - (time.monotonic() - request_started)
|
||||
timeout = (
|
||||
max(soft_remaining, GENERATE_PROGRESS_GRACE_S)
|
||||
- _GENERATE_TIMEOUT_MARGIN_S
|
||||
)
|
||||
if timeout <= 0:
|
||||
self._terminate_server()
|
||||
raise TimeoutError(
|
||||
"audio.cpp startup exhausted the generation time budget"
|
||||
)
|
||||
try:
|
||||
obj = self._post_json(
|
||||
"/v1/audio/speech", payload, timeout=timeout,
|
||||
)
|
||||
except TimeoutError:
|
||||
self._terminate_server()
|
||||
raise RuntimeError(
|
||||
"audio.cpp generation timed out; its managed server was reset"
|
||||
) from None
|
||||
sr, wav_np = decode_speech_json(obj)
|
||||
self._sr = sr
|
||||
wav = torch.from_numpy(wav_np).float()
|
||||
if wav.ndim == 0:
|
||||
raise RuntimeError("audio.cpp produced empty audio")
|
||||
return wav.unsqueeze(0)
|
||||
|
||||
# ── lifecycle ───────────────────────────────────────────────────────
|
||||
|
||||
def unload(self) -> None:
|
||||
"""Free the model server-side, then stop it. Idempotent."""
|
||||
with self._lock:
|
||||
if self._port is not None and self._proc is not None \
|
||||
and self._proc.poll() is None:
|
||||
try:
|
||||
self._post_json("/v1/tasks/unload_all_models", {}, timeout=30)
|
||||
except Exception as exc: # noqa: BLE001 — best effort
|
||||
logger.warning("audio.cpp: server unload failed: %s", exc)
|
||||
self._port = None
|
||||
self._terminate_server()
|
||||
super().unload()
|
||||
|
||||
|
||||
__all__ = [
|
||||
"ENGINE_ID",
|
||||
"AudioCPPBackend",
|
||||
"build_server_config",
|
||||
"build_speech_payload",
|
||||
"decode_speech_json",
|
||||
]
|
||||
@@ -0,0 +1,738 @@
|
||||
"""audio.cpp binary probe + model resolution.
|
||||
|
||||
audio.cpp (0xShug0/audio.cpp) is a pure-C++ ggml inference engine with
|
||||
prebuilt release binaries — no Python venv, no ``transformers`` pin, so
|
||||
none of the dependency-isolation machinery in ``engines._venv_probe`` or
|
||||
``services.subprocess_backend`` applies. The parent instead:
|
||||
|
||||
1. locates a user-installed ``audiocpp_server`` (env var, user dir, or this
|
||||
package's ``bin/``), and
|
||||
2. resolves an explicitly installed GGUF model file from a direct path or
|
||||
the shared Hugging Face cache.
|
||||
|
||||
Probe order for the server binary (existing installs win, zero migration):
|
||||
|
||||
1. ``${OMNIVOICE_AUDIOCPP_BIN}`` — absolute path to the binary itself.
|
||||
2. ``${OMNIVOICE_AUDIOCPP_DIR}/audiocpp_server[.exe]`` — a user-managed
|
||||
install dir (e.g. an extracted release zip, or a self-built tree).
|
||||
3. ``backend/engines/audiocpp/bin/audiocpp_server[.exe]`` — an explicitly
|
||||
installed local copy.
|
||||
|
||||
``is_installed()`` is a cheap file-existence check — no spawn, no network.
|
||||
VoiceStudio never downloads executable code for this engine.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import errno
|
||||
import functools
|
||||
import logging
|
||||
import os
|
||||
import platform
|
||||
import subprocess # nosec B404 -- fixed argv probes a user-selected executable
|
||||
import sys
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
logger = logging.getLogger("omnivoice.audiocpp.bootstrap")
|
||||
|
||||
#: Pinned audio.cpp release. BreezeTTS-2 support landed in 0.7.2 — older
|
||||
#: binaries have no ``breeze_tts`` family, so the floor is also the pin.
|
||||
VERSION = "v0.7.2"
|
||||
|
||||
#: GitHub repo serving the prebuilt binaries.
|
||||
GH_REPO = "0xShug0/audio.cpp"
|
||||
|
||||
#: HuggingFace repo serving the GGUF model packages (not gated).
|
||||
HF_MODEL_REPO = "audio-cpp/audio.cpp-gguf"
|
||||
|
||||
# Immutable repository revision used for the v0.7.2 Breeze-TTS-2 package.
|
||||
# Pinning prevents a later upstream file replacement from silently changing
|
||||
# the model exercised by this backend.
|
||||
HF_MODEL_REVISION = "dc6fecccc2b0c6bdda0a8b2f38fa61394fee0b9c"
|
||||
|
||||
#: Model id used in the generated ``server.json`` and in speech requests.
|
||||
MODEL_ID = "breeze-tts-2"
|
||||
|
||||
#: audio.cpp family name for BreezeTTS 2 (``--family`` / server ``family``).
|
||||
FAMILY = "breeze_tts"
|
||||
|
||||
#: GGUF package directory inside :data:`HF_MODEL_REPO`.
|
||||
PACKAGE_DIR = "Breeze-TTS-2-GGUF"
|
||||
|
||||
#: Default package (Q8_0, the upstream-recommended GGUF). ``bf16`` is
|
||||
#: available via ``OMNIVOICE_AUDIOCPP_PACKAGE``.
|
||||
DEFAULT_PACKAGE = "breeze-tts-2-q8_0.gguf"
|
||||
|
||||
#: Env var pointing directly at the ``audiocpp_server`` binary.
|
||||
BIN_ENV = "OMNIVOICE_AUDIOCPP_BIN"
|
||||
|
||||
#: Env var pointing at a directory containing ``audiocpp_server``.
|
||||
DIR_ENV = "OMNIVOICE_AUDIOCPP_DIR"
|
||||
|
||||
#: Env var overriding the GGUF package filename (e.g. the bf16 package).
|
||||
PACKAGE_ENV = "OMNIVOICE_AUDIOCPP_PACKAGE"
|
||||
|
||||
#: Optional advanced overrides for a binary that exposes several runtimes or
|
||||
#: devices. Device indices are local to the selected backend registry.
|
||||
BACKEND_ENV = "OMNIVOICE_AUDIOCPP_BACKEND"
|
||||
DEVICE_ENV = "OMNIVOICE_AUDIOCPP_DEVICE"
|
||||
|
||||
#: Env var overriding the loopback port the managed server binds.
|
||||
PORT_ENV = "OMNIVOICE_AUDIOCPP_PORT"
|
||||
|
||||
#: Default loopback port. High and engine-specific to avoid clashing with
|
||||
#: the app itself or a user-run ``audiocpp_server`` (default 8080).
|
||||
DEFAULT_PORT = 17860
|
||||
|
||||
#: This package's owned binary dir (probe 3).
|
||||
_PKG_BIN_DIR: Path = Path(__file__).parent / "bin"
|
||||
|
||||
# Recommended (asset filename, sha256) per platform slug, from the v0.7.2
|
||||
# release. Windows and Linux use the vendor-neutral Vulkan build, which also
|
||||
# exposes the native CPU backend. Upstream publishes the macOS builds under
|
||||
# the Metal package name. No linux-aarch64 prebuilt exists in v0.7.2.
|
||||
_ASSETS: dict[str, tuple[str, str]] = {
|
||||
"windows-x64": (
|
||||
"audio-v0.7.2-bin-windows-x64-vulkan.zip",
|
||||
"15b8232eae740e21e507d87f827a89966de9451b085a45932d9e214e032962c1",
|
||||
),
|
||||
"linux-x64": (
|
||||
"audio-v0.7.2-bin-ubuntu-x64-vulkan.tar.gz",
|
||||
"fee1f978cee76453cf17f00196554bc2ee294645739538af0726a143b6a69a23",
|
||||
),
|
||||
"darwin-arm64": (
|
||||
"audio-v0.7.2-bin-macos-arm64-metal.tar.gz",
|
||||
"c01e4f82971bedbe341697e63a9cebd5a5d1f72d5a9bcb51a3191f95ddab7a95",
|
||||
),
|
||||
"darwin-x64": (
|
||||
"audio-v0.7.2-bin-macos-x64-metal.tar.gz",
|
||||
"3862270f33439077225324169313f727064f727305b54d8ce920244d75ddcc24",
|
||||
),
|
||||
}
|
||||
|
||||
#: Binary filename per platform.
|
||||
_BINARY_NAMES = {"windows-x64": "audiocpp_server.exe"}
|
||||
|
||||
_REGISTRY_BACKENDS = {
|
||||
"CPU": "cpu",
|
||||
"CUDA": "cuda",
|
||||
"MUSA": "cuda",
|
||||
"HIP": "hip",
|
||||
"ROCm": "hip",
|
||||
"Vulkan": "vulkan",
|
||||
"Metal": "metal",
|
||||
"MTL": "metal",
|
||||
}
|
||||
_BACKEND_ALIASES = {
|
||||
"cpu": "cpu",
|
||||
"cuda": "cuda",
|
||||
"hip": "hip",
|
||||
"rocm": "hip",
|
||||
"vulkan": "vulkan",
|
||||
"metal": "metal",
|
||||
}
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class AudioCPPDevice:
|
||||
"""One immutable device from audio.cpp's backend-local registry."""
|
||||
|
||||
registry: str
|
||||
backend: str
|
||||
index: int
|
||||
name: str
|
||||
kind: str
|
||||
target: str
|
||||
hardware_family: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class AudioCPPSelection:
|
||||
"""The runtime/device chosen for the next managed server."""
|
||||
|
||||
device: AudioCPPDevice
|
||||
fallback_reason: str | None = None
|
||||
verified_vram_gb: float = 0.0
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class _ProbeOutcome:
|
||||
devices: tuple[AudioCPPDevice, ...] = ()
|
||||
error: str | None = None
|
||||
|
||||
|
||||
def _cpu_probe_fallback(error: RuntimeError) -> AudioCPPSelection:
|
||||
"""A usable automatic fallback when native device discovery fails."""
|
||||
return AudioCPPSelection(
|
||||
AudioCPPDevice(
|
||||
registry="CPU",
|
||||
backend="cpu",
|
||||
index=0,
|
||||
name="Host CPU",
|
||||
kind="CPU",
|
||||
target="cpu",
|
||||
hardware_family="cpu",
|
||||
),
|
||||
f"{error}; running on CPU",
|
||||
)
|
||||
|
||||
|
||||
def _vulkan_hardware_family(name: str) -> str:
|
||||
low = name.casefold()
|
||||
if any(token in low for token in ("nvidia", "geforce", "quadro", "tesla")):
|
||||
return "cuda"
|
||||
if any(token in low for token in ("amd", "radeon")):
|
||||
return "rocm"
|
||||
if any(token in low for token in ("intel", "arc ")):
|
||||
return "xpu"
|
||||
return "vulkan"
|
||||
|
||||
|
||||
def _device_families(registry: str, name: str, kind: str) -> tuple[str, str]:
|
||||
# Software adapters such as Vulkan llvmpipe may be listed by a GPU
|
||||
# registry but still execute on the CPU. Keep their runtime backend for
|
||||
# explicit overrides while reporting and routing them as CPU work.
|
||||
if kind == "CPU":
|
||||
return "cpu", "cpu"
|
||||
if registry in {"CUDA", "MUSA"}:
|
||||
return "cuda", "cuda"
|
||||
if registry in {"HIP", "ROCm"}:
|
||||
return "rocm", "rocm"
|
||||
if registry in {"Metal", "MTL"}:
|
||||
return "mps", "mps"
|
||||
if registry == "Vulkan":
|
||||
return "vulkan", _vulkan_hardware_family(name)
|
||||
return "cpu", "cpu"
|
||||
|
||||
|
||||
def parse_device_list(output: str) -> tuple[AudioCPPDevice, ...]:
|
||||
"""Parse the stable stdout contract of ``--list-devices``.
|
||||
|
||||
Backend diagnostics are emitted on stderr and deliberately never enter
|
||||
this parser. Unknown future registries are ignored; malformed entries for
|
||||
a registry we understand fail closed instead of selecting the wrong GPU.
|
||||
"""
|
||||
devices: list[AudioCPPDevice] = []
|
||||
seen: set[tuple[str, int]] = set()
|
||||
for raw in str(output or "").splitlines():
|
||||
line = raw.strip()
|
||||
registry, colon, detail = line.partition(":")
|
||||
if not colon or registry not in _REGISTRY_BACKENDS:
|
||||
continue
|
||||
index_text, space, remainder = detail.strip().partition(" ")
|
||||
if not space or not index_text.isascii() or not index_text.isdecimal():
|
||||
raise RuntimeError(
|
||||
f"malformed audio.cpp {registry} device entry"
|
||||
)
|
||||
index = int(index_text)
|
||||
remainder = remainder.strip()
|
||||
kind_start = remainder.rfind("[")
|
||||
if kind_start < 0 or not remainder.endswith("]"):
|
||||
raise RuntimeError(
|
||||
f"malformed audio.cpp {registry} device entry"
|
||||
)
|
||||
name_field = remainder[:kind_start].strip()
|
||||
if name_field:
|
||||
if len(name_field) < 2 or name_field[0] != '"' or name_field[-1] != '"':
|
||||
raise RuntimeError(
|
||||
f"malformed audio.cpp {registry} device entry"
|
||||
)
|
||||
name = name_field[1:-1]
|
||||
else:
|
||||
name = ""
|
||||
kind = remainder[kind_start + 1:-1].strip().upper()
|
||||
if kind not in {"CPU", "GPU", "IGPU", "ACCEL", "META"}:
|
||||
raise RuntimeError("unknown audio.cpp device kind")
|
||||
# Registry aliases such as HIP/ROCm share one backend-local index
|
||||
# namespace and therefore cannot safely describe different devices.
|
||||
key = (_REGISTRY_BACKENDS[registry], index)
|
||||
if key in seen:
|
||||
raise RuntimeError(
|
||||
f"duplicate audio.cpp device entry: {registry}:{index}"
|
||||
)
|
||||
seen.add(key)
|
||||
target, hardware_family = _device_families(registry, name, kind)
|
||||
devices.append(AudioCPPDevice(
|
||||
registry=registry,
|
||||
backend=_REGISTRY_BACKENDS[registry],
|
||||
index=index,
|
||||
name=name,
|
||||
kind=kind,
|
||||
target=target,
|
||||
hardware_family=hardware_family,
|
||||
))
|
||||
if not devices:
|
||||
raise RuntimeError("audio.cpp reported no recognized compute devices")
|
||||
return tuple(devices)
|
||||
|
||||
|
||||
def _platform_slug() -> str:
|
||||
system = sys.platform
|
||||
machine = platform.machine().lower()
|
||||
if system == "win32":
|
||||
return "windows-x64"
|
||||
if system == "darwin":
|
||||
return "darwin-arm64" if machine in ("arm64", "aarch64") else "darwin-x64"
|
||||
if machine in ("x86_64", "amd64"):
|
||||
return "linux-x64"
|
||||
return f"linux-{machine}"
|
||||
|
||||
|
||||
def binary_name(slug: str | None = None) -> str:
|
||||
"""``audiocpp_server`` filename for ``slug`` (``.exe`` on Windows)."""
|
||||
return _BINARY_NAMES.get(slug or _platform_slug(), "audiocpp_server")
|
||||
|
||||
|
||||
def _probe_paths() -> list[Path]:
|
||||
out: list[Path] = []
|
||||
direct = os.environ.get(BIN_ENV, "").strip()
|
||||
if direct:
|
||||
out.append(Path(direct))
|
||||
user_dir = os.environ.get(DIR_ENV, "").strip()
|
||||
if user_dir:
|
||||
out.append(Path(user_dir) / binary_name())
|
||||
out.append(_PKG_BIN_DIR / binary_name())
|
||||
return out
|
||||
|
||||
|
||||
def is_installed() -> bool:
|
||||
"""Cheap precedence-aware check for a usable server binary."""
|
||||
try:
|
||||
resolve_server_binary()
|
||||
except RuntimeError:
|
||||
return False
|
||||
return True
|
||||
|
||||
|
||||
def resolve_server_binary() -> Path:
|
||||
"""Resolve the ``audiocpp_server`` binary. Raises ``RuntimeError`` with
|
||||
install instructions when none is found."""
|
||||
for cand in _probe_paths():
|
||||
if cand.is_file():
|
||||
if os.name == "nt" or os.access(cand, os.X_OK):
|
||||
return cand
|
||||
raise RuntimeError(
|
||||
"audiocpp_server is not executable. Run `chmod +x "
|
||||
"audiocpp_server` on the configured binary, then restart "
|
||||
"VoiceStudio. See docs/engines/audio-cpp.md."
|
||||
)
|
||||
slug = _platform_slug()
|
||||
asset = _ASSETS.get(slug)
|
||||
if asset is None:
|
||||
raise RuntimeError(
|
||||
f"audio.cpp ships no prebuilt binary for this platform ({slug}). "
|
||||
"Build from https://github.com/0xShug0/audio.cpp and set "
|
||||
f"{BIN_ENV} to your audiocpp_server binary. See "
|
||||
"docs/engines/audio-cpp.md."
|
||||
)
|
||||
raise RuntimeError(
|
||||
"audiocpp_server not found. Download "
|
||||
f"https://github.com/{GH_REPO}/releases/download/{VERSION}/{asset[0]} "
|
||||
f"(SHA-256 {asset[1]}), verify and extract it, and set {BIN_ENV} to the "
|
||||
"audiocpp_server binary (or "
|
||||
f"{DIR_ENV} to its directory). See docs/engines/audio-cpp.md."
|
||||
)
|
||||
|
||||
|
||||
@functools.lru_cache(maxsize=4)
|
||||
def _probe_device_outcome(binary: str) -> _ProbeOutcome:
|
||||
try:
|
||||
proc = subprocess.run( # nosec B603 -- executable is the resolved engine binary
|
||||
[binary, "--list-devices"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
timeout=10,
|
||||
check=False,
|
||||
)
|
||||
except subprocess.TimeoutExpired:
|
||||
return _ProbeOutcome(
|
||||
error="audiocpp_server device discovery timed out after 10 seconds"
|
||||
)
|
||||
except OSError as exc:
|
||||
return _ProbeOutcome(
|
||||
error=(
|
||||
"audiocpp_server device discovery could not start: "
|
||||
f"{type(exc).__name__}"
|
||||
)
|
||||
)
|
||||
if proc.returncode != 0:
|
||||
return _ProbeOutcome(
|
||||
error=(
|
||||
"audiocpp_server device discovery failed "
|
||||
f"(code {proc.returncode}). Check the audio.cpp server log "
|
||||
"for details."
|
||||
)
|
||||
)
|
||||
try:
|
||||
return _ProbeOutcome(devices=parse_device_list(proc.stdout))
|
||||
except RuntimeError as exc:
|
||||
return _ProbeOutcome(error=str(exc))
|
||||
|
||||
|
||||
def _probe_devices(binary: str) -> tuple[AudioCPPDevice, ...]:
|
||||
outcome = _probe_device_outcome(binary)
|
||||
if outcome.error:
|
||||
raise RuntimeError(outcome.error)
|
||||
return outcome.devices
|
||||
|
||||
|
||||
def probe_devices() -> tuple[AudioCPPDevice, ...]:
|
||||
"""Return the installed binary's devices without loading a model."""
|
||||
return _probe_devices(str(resolve_server_binary()))
|
||||
|
||||
|
||||
def _priority(device: AudioCPPDevice) -> tuple[int, int]:
|
||||
if device.kind == "META":
|
||||
# Tensor-parallel meta devices are valid explicit targets, but their
|
||||
# resource footprint is not safe to choose implicitly over CPU.
|
||||
rank = 8
|
||||
elif device.backend != "cpu" and device.kind == "CPU":
|
||||
# Native CPU is the predictable fallback. Software adapters remain
|
||||
# available to an explicit backend override but never win auto mode.
|
||||
rank = 7
|
||||
elif device.backend == "cuda":
|
||||
rank = 0
|
||||
elif device.backend == "hip":
|
||||
rank = 1
|
||||
elif device.backend == "metal":
|
||||
rank = 2
|
||||
elif device.backend == "vulkan" and device.kind == "GPU":
|
||||
rank = 3
|
||||
elif device.backend == "vulkan" and device.kind in {"IGPU", "ACCEL"}:
|
||||
rank = 4
|
||||
elif device.backend == "cpu":
|
||||
rank = 6
|
||||
else:
|
||||
rank = 5
|
||||
return rank, device.index
|
||||
|
||||
|
||||
def select_device(
|
||||
devices: tuple[AudioCPPDevice, ...],
|
||||
*,
|
||||
requested_family: str = "auto",
|
||||
backend_override: str | None = None,
|
||||
device_override: int | None = None,
|
||||
preferred_name: str = "",
|
||||
) -> AudioCPPSelection:
|
||||
"""Resolve one device with explicit overrides and discrete-GPU priority."""
|
||||
if backend_override:
|
||||
normalized = _BACKEND_ALIASES.get(backend_override.strip().lower())
|
||||
if normalized is None:
|
||||
valid = ", ".join(_BACKEND_ALIASES)
|
||||
raise RuntimeError(
|
||||
f"unknown audio.cpp backend '{backend_override}' (valid: {valid})"
|
||||
)
|
||||
candidates = [device for device in devices if device.backend == normalized]
|
||||
if device_override is not None:
|
||||
candidates = [
|
||||
device for device in candidates if device.index == device_override
|
||||
]
|
||||
if not candidates:
|
||||
suffix = "" if device_override is None else f" device {device_override}"
|
||||
available = ", ".join(
|
||||
f"{device.backend}:{device.index}" for device in devices
|
||||
)
|
||||
raise RuntimeError(
|
||||
f"audio.cpp backend '{backend_override}'{suffix} is unavailable "
|
||||
f"(available: {available})"
|
||||
)
|
||||
# An explicit runtime request should still prefer a compute device to
|
||||
# a software adapter when no backend-local index was supplied. META is
|
||||
# valid here because the user explicitly chose this registry.
|
||||
return AudioCPPSelection(min(
|
||||
candidates,
|
||||
key=lambda device: (device.kind == "CPU", _priority(device)),
|
||||
))
|
||||
|
||||
if device_override is not None:
|
||||
raise RuntimeError(
|
||||
f"{DEVICE_ENV} requires {BACKEND_ENV} because device indices are "
|
||||
"backend-local"
|
||||
)
|
||||
|
||||
family = (requested_family or "auto").strip().lower()
|
||||
if family != "auto":
|
||||
candidates = [
|
||||
device for device in devices if device.hardware_family == family
|
||||
]
|
||||
if candidates:
|
||||
preferred = preferred_name.casefold().strip()
|
||||
if preferred:
|
||||
named = [
|
||||
device for device in candidates
|
||||
if device.name
|
||||
and (
|
||||
preferred in device.name.casefold()
|
||||
or device.name.casefold() in preferred
|
||||
)
|
||||
]
|
||||
if named:
|
||||
candidates = named
|
||||
return AudioCPPSelection(min(candidates, key=_priority))
|
||||
cpu = [device for device in devices if device.backend == "cpu"]
|
||||
if cpu:
|
||||
return AudioCPPSelection(
|
||||
min(cpu, key=_priority),
|
||||
f"requested {family.upper()} device is not exposed by the "
|
||||
"installed audio.cpp binary; running on CPU",
|
||||
)
|
||||
raise RuntimeError(
|
||||
f"requested {family.upper()} device is not exposed by the "
|
||||
"installed audio.cpp binary"
|
||||
)
|
||||
|
||||
return AudioCPPSelection(min(devices, key=_priority))
|
||||
|
||||
|
||||
def resolve_compute_selection(caps=None) -> AudioCPPSelection:
|
||||
"""Select the runtime from engine env overrides, Settings, then auto."""
|
||||
backend_override = os.environ.get(BACKEND_ENV, "").strip() or None
|
||||
raw_device = os.environ.get(DEVICE_ENV, "").strip()
|
||||
device_override: int | None = None
|
||||
if raw_device:
|
||||
try:
|
||||
device_override = int(raw_device)
|
||||
except ValueError as exc:
|
||||
raise RuntimeError(
|
||||
f"{DEVICE_ENV} must be a non-negative integer"
|
||||
) from exc
|
||||
if device_override < 0:
|
||||
raise RuntimeError(f"{DEVICE_ENV} must be a non-negative integer")
|
||||
|
||||
if caps is None:
|
||||
from core.device_caps import detect_host_caps
|
||||
|
||||
caps = detect_host_caps()
|
||||
requested = getattr(caps, "requested_family", "auto") or "auto"
|
||||
try:
|
||||
devices = probe_devices()
|
||||
except RuntimeError as exc:
|
||||
if backend_override or raw_device or requested != "auto":
|
||||
raise
|
||||
return _cpu_probe_fallback(exc)
|
||||
selection = select_device(
|
||||
devices,
|
||||
requested_family=requested,
|
||||
backend_override=backend_override,
|
||||
device_override=device_override,
|
||||
preferred_name=getattr(caps, "device_name", "") or "",
|
||||
)
|
||||
# HostCaps measures the preferred accelerator's device 0. Reuse that VRAM
|
||||
# only when the selected native registry has exactly one device with the
|
||||
# same normalized name. Multi-GPU peers with identical names stay unknown.
|
||||
selected_name = " ".join(selection.device.name.casefold().split())
|
||||
host_name = " ".join(
|
||||
str(getattr(caps, "device_name", "") or "").casefold().split()
|
||||
)
|
||||
peers = [
|
||||
device for device in devices
|
||||
if device.backend == selection.device.backend
|
||||
and " ".join(device.name.casefold().split()) == host_name
|
||||
]
|
||||
if (
|
||||
selected_name
|
||||
and selected_name == host_name
|
||||
and len(peers) == 1
|
||||
and float(getattr(caps, "vram_gb", 0.0) or 0.0) > 0
|
||||
):
|
||||
return AudioCPPSelection(
|
||||
selection.device,
|
||||
selection.fallback_reason,
|
||||
float(caps.vram_gb),
|
||||
)
|
||||
return selection
|
||||
|
||||
|
||||
def runtime_targets(devices: tuple[AudioCPPDevice, ...] | None = None) -> tuple[str, ...]:
|
||||
"""Actual compute backends compiled into the selected binary."""
|
||||
if devices is not None:
|
||||
found = devices
|
||||
else:
|
||||
try:
|
||||
found = probe_devices()
|
||||
except RuntimeError:
|
||||
if (
|
||||
os.environ.get(BACKEND_ENV, "").strip()
|
||||
or os.environ.get(DEVICE_ENV, "").strip()
|
||||
):
|
||||
raise
|
||||
return ("cpu",)
|
||||
ordered: list[str] = []
|
||||
for device in sorted(found, key=_priority):
|
||||
if device.target not in ordered:
|
||||
ordered.append(device.target)
|
||||
return tuple(ordered)
|
||||
|
||||
|
||||
def invalidate() -> None:
|
||||
"""Forget cached binary capability discovery after an install change."""
|
||||
_probe_device_outcome.cache_clear()
|
||||
|
||||
|
||||
def default_asset() -> tuple[str, str] | None:
|
||||
"""``(filename, sha256)`` of the release asset for this host, or None
|
||||
when upstream ships no prebuilt for it."""
|
||||
return _ASSETS.get(_platform_slug())
|
||||
|
||||
|
||||
def server_port() -> int:
|
||||
"""Loopback port for the managed server (env override or default)."""
|
||||
raw = os.environ.get(PORT_ENV, "").strip()
|
||||
if raw:
|
||||
try:
|
||||
port = int(raw)
|
||||
if 1 <= port <= 65535:
|
||||
return port
|
||||
logger.warning("Ignoring %s=%r: out of range.", PORT_ENV, raw)
|
||||
except ValueError:
|
||||
logger.warning("Ignoring %s=%r: not a number.", PORT_ENV, raw)
|
||||
return DEFAULT_PORT
|
||||
|
||||
|
||||
def package_filename() -> str:
|
||||
"""GGUF package filename (env override or the Q8_0 default)."""
|
||||
return os.environ.get(PACKAGE_ENV, "").strip() or DEFAULT_PACKAGE
|
||||
|
||||
|
||||
def _materialize_gguf_cache_path(model_file: Path) -> Path:
|
||||
"""Return a real ``.gguf`` path when the HF snapshot is a symlink.
|
||||
|
||||
audio.cpp canonicalizes model paths before inspecting the suffix. The
|
||||
Hugging Face cache points the friendly ``.gguf`` snapshot name at an
|
||||
extensionless content-addressed blob, so passing that symlink makes the
|
||||
server reject a valid model. A hard link beside the snapshot keeps the
|
||||
required suffix without copying a multi-gigabyte model or escaping the
|
||||
snapshot's cleanup lifecycle.
|
||||
"""
|
||||
resolved = model_file.resolve()
|
||||
if resolved.suffix.lower() == ".gguf":
|
||||
return model_file
|
||||
if model_file.suffix.lower() != ".gguf":
|
||||
raise RuntimeError(f"audio.cpp model must be a .gguf file: {model_file}")
|
||||
|
||||
def _link(alias: Path) -> Path:
|
||||
for attempt in range(2):
|
||||
try:
|
||||
os.link(resolved, alias)
|
||||
except FileExistsError:
|
||||
if (
|
||||
not alias.is_symlink()
|
||||
and alias.is_file()
|
||||
and os.path.samefile(resolved, alias)
|
||||
):
|
||||
return alias
|
||||
if attempt == 0 and alias.is_symlink():
|
||||
alias.unlink()
|
||||
continue
|
||||
raise RuntimeError(
|
||||
f"audio.cpp model alias points at a different file: {alias}"
|
||||
) from None
|
||||
return alias
|
||||
raise RuntimeError(f"audio.cpp model alias could not be created: {alias}")
|
||||
|
||||
alias = model_file.with_name(
|
||||
f".{model_file.stem}-{HF_MODEL_REVISION[:12]}.audiocpp.gguf"
|
||||
)
|
||||
try:
|
||||
return _link(alias)
|
||||
except OSError as exc:
|
||||
if exc.errno == errno.EXDEV:
|
||||
# An explicit symlink may live on a different filesystem from its
|
||||
# target. Put the suffix-preserving hard link beside the resolved
|
||||
# file so no multi-gigabyte copy is needed.
|
||||
target_alias = resolved.with_name(
|
||||
f".{resolved.name}-{HF_MODEL_REVISION[:12]}.audiocpp.gguf"
|
||||
)
|
||||
try:
|
||||
return _link(target_alias)
|
||||
except OSError as target_exc:
|
||||
exc = target_exc
|
||||
raise RuntimeError(
|
||||
"audio.cpp cannot materialize the Hugging Face cache symlink as "
|
||||
f"a .gguf hard link: {exc}"
|
||||
) from exc
|
||||
|
||||
|
||||
def resolve_model_file() -> Path:
|
||||
"""Resolve an explicitly installed Breeze-TTS-2 GGUF file.
|
||||
|
||||
An explicit ``OMNIVOICE_AUDIOCPP_MODEL`` path wins (file or directory
|
||||
containing the package file). Otherwise only the local Hugging Face cache
|
||||
is inspected. Downloads must be started explicitly from Model Catalogue →
|
||||
Models, so generation can never silently transfer the 4.73 GiB package.
|
||||
"""
|
||||
override = os.environ.get("OMNIVOICE_AUDIOCPP_MODEL", "").strip()
|
||||
if override:
|
||||
cand = Path(override)
|
||||
if cand.is_file():
|
||||
return _materialize_gguf_cache_path(cand)
|
||||
if cand.is_dir():
|
||||
inner = cand / package_filename()
|
||||
if inner.is_file():
|
||||
return _materialize_gguf_cache_path(inner)
|
||||
raise RuntimeError(
|
||||
f"OMNIVOICE_AUDIOCPP_MODEL={override} is not a GGUF file or a "
|
||||
"directory containing one."
|
||||
)
|
||||
|
||||
from huggingface_hub import snapshot_download
|
||||
from huggingface_hub.utils import LocalEntryNotFoundError
|
||||
|
||||
try:
|
||||
cached = Path(
|
||||
snapshot_download(
|
||||
repo_id=HF_MODEL_REPO,
|
||||
# Full immutable commit SHA declared above; Bandit cannot follow
|
||||
# the module constant through this call.
|
||||
revision=HF_MODEL_REVISION, # nosec B615
|
||||
allow_patterns=[f"{PACKAGE_DIR}/{package_filename()}"],
|
||||
local_files_only=True,
|
||||
)
|
||||
)
|
||||
except (LocalEntryNotFoundError, OSError) as exc:
|
||||
raise RuntimeError(
|
||||
"Breeze-TTS-2 is not installed. Install the audio.cpp Breeze-TTS-2 "
|
||||
"model from Model Catalogue → Models, or set "
|
||||
"OMNIVOICE_AUDIOCPP_MODEL to an existing GGUF file."
|
||||
) from exc
|
||||
model_file = cached / PACKAGE_DIR / package_filename()
|
||||
if not model_file.is_file():
|
||||
raise RuntimeError(
|
||||
f"Breeze-TTS-2 package {package_filename()} is not completely "
|
||||
"installed. Reinstall it from Model Catalogue → Models."
|
||||
)
|
||||
return _materialize_gguf_cache_path(model_file)
|
||||
|
||||
|
||||
__all__ = [
|
||||
"AudioCPPDevice",
|
||||
"AudioCPPSelection",
|
||||
"BACKEND_ENV",
|
||||
"BIN_ENV",
|
||||
"DEFAULT_PACKAGE",
|
||||
"DEFAULT_PORT",
|
||||
"DEVICE_ENV",
|
||||
"DIR_ENV",
|
||||
"FAMILY",
|
||||
"HF_MODEL_REPO",
|
||||
"HF_MODEL_REVISION",
|
||||
"MODEL_ID",
|
||||
"PACKAGE_DIR",
|
||||
"PACKAGE_ENV",
|
||||
"PORT_ENV",
|
||||
"VERSION",
|
||||
"_materialize_gguf_cache_path",
|
||||
"binary_name",
|
||||
"default_asset",
|
||||
"invalidate",
|
||||
"is_installed",
|
||||
"package_filename",
|
||||
"parse_device_list",
|
||||
"probe_devices",
|
||||
"resolve_compute_selection",
|
||||
"resolve_model_file",
|
||||
"resolve_server_binary",
|
||||
"runtime_targets",
|
||||
"server_port",
|
||||
]
|
||||
@@ -56,15 +56,15 @@ class Confucius4Backend(SubprocessBackend):
|
||||
|
||||
id = "confucius4-tts"
|
||||
display_name = (
|
||||
"Confucius4-TTS (LLM, 14 langs, cross-lingual zero-shot clone, CUDA/CPU, Apache-2.0)"
|
||||
"Confucius4-TTS (LLM, 14 langs, cross-lingual zero-shot clone, Apache-2.0)"
|
||||
)
|
||||
supports_voice_design = False # timbre comes from a reference clip
|
||||
# Upstream vocoder rate (config target_sample_rate) — confirmed 22 050 Hz by
|
||||
# a live run (2026-07-02); still re-read from the sidecar's ready/audio frames.
|
||||
_DEFAULT_SAMPLE_RATE = 22050
|
||||
# CUDA fast path + CPU fallback, both exercised (CPU end-to-end validated).
|
||||
# No MPS claim — upstream has no Metal path.
|
||||
gpu_compat = ("cuda", "cpu")
|
||||
# Match device propagation into upstream .to(device). XPU/NPU routing is
|
||||
# contract-tested, not a claim of physical-hardware synthesis validation.
|
||||
gpu_compat = ("cuda", "rocm", "xpu", "npu", "cpu")
|
||||
|
||||
@classmethod
|
||||
def is_available(cls) -> tuple[bool, str]:
|
||||
|
||||
@@ -104,7 +104,7 @@ def _ensure_clone_on_sys_path() -> None:
|
||||
|
||||
|
||||
def _load_model(stdout):
|
||||
"""Cold-construct the Confucius4 model (CUDA, else CPU — both validated)."""
|
||||
"""Cold-construct using an available torch accelerator, with CPU fallback."""
|
||||
global _model
|
||||
if _model is not None:
|
||||
return _model
|
||||
@@ -115,7 +115,18 @@ def _load_model(stdout):
|
||||
import torch
|
||||
from confuciustts.cli.inference import ConfuciusTTS # type: ignore[import-not-found]
|
||||
|
||||
device = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
try:
|
||||
# Existing manually provisioned venvs may predate torch.accelerator.
|
||||
current_accelerator = getattr(getattr(torch, "accelerator", None), "current_accelerator", None)
|
||||
if current_accelerator is None:
|
||||
device = torch.device("cuda") if torch.cuda.is_available() else None
|
||||
else:
|
||||
device = current_accelerator(check_available=True)
|
||||
device = device.type if device is not None else "cpu" # 'cuda', 'npu', 'mps', 'xpu', 'cpu'
|
||||
except Exception:
|
||||
device = "cpu" # Broken accelerator drivers must not block CPU loading.
|
||||
if device == "mps":
|
||||
device = "cpu" # MPS was slower than CPU in the existing validation run
|
||||
_send(stdout, {"op": "progress", "stage": "loading_model", "percent": 50})
|
||||
|
||||
_model = ConfuciusTTS(config_path=_config_path(), device=device)
|
||||
|
||||
@@ -115,7 +115,12 @@ def _load_runtime(stdout):
|
||||
from dots_tts.runtime import DotsTtsRuntime # type: ignore[import-not-found]
|
||||
|
||||
repo = os.environ.get("OMNIVOICE_DOTS_TTS_MODEL", _DEFAULT_REPO)
|
||||
default_precision = "bfloat16" if torch.cuda.is_available() else "float32"
|
||||
# Match DotsTtsRuntime's own CUDA/CPU selection. Its _check_torch_env
|
||||
# rejects half precision without CUDA, even when an XPU/NPU is available.
|
||||
try:
|
||||
default_precision = "bfloat16" if torch.cuda.is_available() else "float32"
|
||||
except Exception:
|
||||
default_precision = "float32" # Probe failure must not force half precision.
|
||||
precision = os.environ.get("OMNIVOICE_DOTS_TTS_PRECISION", default_precision)
|
||||
optimize = os.environ.get("OMNIVOICE_DOTS_TTS_OPTIMIZE", "0") == "1"
|
||||
|
||||
|
||||
@@ -29,15 +29,11 @@ Do NOT import ``main.py`` from the parent process — it runs under a
|
||||
different venv (``transformers==5.0.0``) and importing it in-process would
|
||||
re-introduce the exact conflict this isolation exists to avoid.
|
||||
|
||||
Hardware honesty (cross-platform rule): MOSS-TTS-v1.5's upstream documents
|
||||
only CUDA and CPU. There is **no documented or tested MPS path** — the
|
||||
custom ``trust_remote_code`` modelling code and the separate audio
|
||||
tokenizer are unverified on Apple Silicon. We therefore advertise
|
||||
``gpu_compat = ("cuda", "cpu")`` and the sidecar selects ``cuda`` when
|
||||
present else ``cpu`` — it never silently routes to MPS where it might
|
||||
crash. On Apple Silicon the engine honestly resolves to CPU (slow but
|
||||
correct), and the engine is opt-in regardless, so it never becomes a
|
||||
broken default on any platform.
|
||||
Hardware routing follows the sidecar's runtime-available PyTorch accelerator:
|
||||
CUDA/ROCm, XPU, or a registered NPU. MPS remains excluded; CPU is the fallback.
|
||||
XPU/NPU routing is covered with mocked device contracts, not physical-hardware
|
||||
synthesis certification; users need a compatible torch/vendor runtime in the
|
||||
isolated engine venv.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -85,13 +81,13 @@ class MossTTSV15Backend(SubprocessBackend):
|
||||
|
||||
id = "moss-tts-v15"
|
||||
display_name = (
|
||||
"MOSS-TTS-v1.5 (8B, 31 langs, zero-shot clone, CUDA/CPU, Apache-2.0)"
|
||||
"MOSS-TTS-v1.5 (8B, 31 langs, zero-shot clone, Apache-2.0)"
|
||||
)
|
||||
supports_voice_design = False # requires ref audio for timbre cloning
|
||||
_DEFAULT_SAMPLE_RATE = 24000
|
||||
# Honest hardware surface: upstream documents CUDA + CPU only. MPS is
|
||||
# undocumented / untested, so we do NOT claim it (cross-platform rule).
|
||||
gpu_compat = ("cuda", "cpu")
|
||||
# Accelerator routing requires its matching runtime in the isolated venv.
|
||||
# MPS remains untested and is deliberately excluded.
|
||||
gpu_compat = ("cuda", "rocm", "xpu", "npu", "cpu")
|
||||
|
||||
# ── availability ───────────────────────────────────────────────────────
|
||||
|
||||
@@ -111,7 +107,7 @@ class MossTTSV15Backend(SubprocessBackend):
|
||||
return False, (
|
||||
"MOSS-TTS-v1.5 venv not found. Set OMNIVOICE_MOSS_TTS_V15_DIR "
|
||||
"to your MOSS-TTS clone (the directory containing pyproject.toml) "
|
||||
"and restart VoiceStudio. CUDA or CPU only (no MPS). See "
|
||||
"and restart VoiceStudio. Install the matching PyTorch runtime. See "
|
||||
"docs/engines/moss-tts-v15.md for the full install walk-through."
|
||||
)
|
||||
if not MOSS_TTS_V15_SIDECAR_SCRIPT.exists():
|
||||
@@ -119,7 +115,7 @@ class MossTTSV15Backend(SubprocessBackend):
|
||||
"MOSS-TTS-v1.5 sidecar script missing at "
|
||||
f"{MOSS_TTS_V15_SIDECAR_SCRIPT} — reinstall VoiceStudio."
|
||||
)
|
||||
return True, "ok (CUDA when present, else CPU)"
|
||||
return True, "ok (runtime-available accelerator or CPU; no MPS)"
|
||||
|
||||
@classmethod
|
||||
def venv_python(cls):
|
||||
|
||||
@@ -222,8 +222,7 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
|
||||
Runs ``uv venv <engines_venv>`` then ``uv pip install --python
|
||||
<engines_venv>/bin/python -e "<clone>[torch-runtime]"``. Verifies the
|
||||
result by re-probing the import — a successful uv invocation that still
|
||||
can't import the stack indicates a deeper environment problem (e.g. the
|
||||
``+cu128`` torch-runtime extra can't resolve on a non-CUDA host) and we
|
||||
can't import the stack indicates a deeper environment problem, and we
|
||||
raise with whatever stderr we captured plus a docs pointer.
|
||||
"""
|
||||
uv = _locate_uv()
|
||||
@@ -254,6 +253,8 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
|
||||
f"{exc.stderr.decode('utf-8', errors='replace') if exc.stderr else exc}"
|
||||
) from exc
|
||||
|
||||
from core.torch_indexes import UV_PIP_CU128_ARGS
|
||||
|
||||
python_path = _venv_python_path(_ENGINES_VENV_DIR)
|
||||
try:
|
||||
subprocess.run(
|
||||
@@ -261,6 +262,10 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
|
||||
uv, "pip", "install",
|
||||
"--python", str(python_path),
|
||||
"-e", f"{clone_dir}[torch-runtime]",
|
||||
# The extra pins torch==2.9.1+cu128, which exists only on
|
||||
# PyTorch's index — without it this could never resolve, on
|
||||
# any host (core.torch_indexes).
|
||||
*UV_PIP_CU128_ARGS,
|
||||
],
|
||||
check=True,
|
||||
timeout=_UV_PIP_INSTALL_TIMEOUT_S,
|
||||
@@ -270,9 +275,10 @@ def _bootstrap_engines_venv(clone_dir: Path) -> Path:
|
||||
except subprocess.CalledProcessError as exc:
|
||||
raise RuntimeError(
|
||||
"uv pip install -e failed during MOSS-TTS-v1.5 bootstrap "
|
||||
f"({clone_dir}). On a non-CUDA host the upstream '[torch-runtime]' "
|
||||
"extra (cu128) cannot resolve — set up the venv manually per "
|
||||
"docs/engines/moss-tts-v15.md. Error: "
|
||||
# uv's own error names what failed; the PyTorch index is always
|
||||
# supplied now, so a guess about the host would only mislead.
|
||||
f"({clone_dir}). See docs/engines/moss-tts-v15.md for the manual "
|
||||
"install. Error: "
|
||||
f"{exc.stderr.decode('utf-8', errors='replace') if exc.stderr else exc}"
|
||||
) from exc
|
||||
|
||||
|
||||
@@ -138,11 +138,12 @@ _state = None
|
||||
def _load_model(stdout):
|
||||
"""Cold-construct the MOSS-TTS-v1.5 processor + model.
|
||||
|
||||
Device selection is CUDA-or-CPU only — MOSS's upstream documents no MPS
|
||||
path and the custom ``trust_remote_code`` modelling code is untested on
|
||||
Apple Silicon, so we never route to MPS where it might crash. dtype is
|
||||
bf16 on CUDA, fp32 on CPU (bf16 CPU ops are spotty). Emits progress
|
||||
frames so the parent can surface the multi-GB cold-load latency.
|
||||
Device selection uses the torch.accelerator API to support any backend
|
||||
(CUDA, NPU, XPU, etc.) automatically. MPS is excluded — MOSS's upstream
|
||||
``trust_remote_code`` modelling code is untested on Apple Silicon. dtype is
|
||||
bf16 on GPU-class accelerators, fp32 on CPU (bf16 CPU ops are spotty).
|
||||
Emits progress frames so the parent can surface the multi-GB cold-load
|
||||
latency.
|
||||
"""
|
||||
global _state
|
||||
if _state is not None:
|
||||
@@ -154,8 +155,22 @@ def _load_model(stdout):
|
||||
from transformers import AutoModel, AutoProcessor
|
||||
|
||||
repo, revision = _model_source()
|
||||
device = "cuda" if torch.cuda.is_available() else "cpu"
|
||||
dtype = torch.bfloat16 if device == "cuda" else torch.float32
|
||||
# current_accelerator() returns None on CPU-only builds (no accelerator
|
||||
# compiled in) or when no accelerator is available; fall back to "cpu".
|
||||
# Existing manually provisioned venvs may predate torch.accelerator.
|
||||
current_accelerator = getattr(getattr(torch, "accelerator", None), "current_accelerator", None)
|
||||
try:
|
||||
if current_accelerator is None:
|
||||
accel = torch.device("cuda") if torch.cuda.is_available() else None
|
||||
else:
|
||||
accel = current_accelerator(check_available=True)
|
||||
except Exception:
|
||||
# Optional drivers can fail during probing; CPU loading remains usable.
|
||||
accel = None
|
||||
device = accel.type if accel is not None else "cpu" # 'cuda', 'npu', 'mps', 'xpu', 'cpu'
|
||||
if device == "mps":
|
||||
device = "cpu" # MOSS is untested on MPS; fall back to CPU for safety
|
||||
dtype = torch.bfloat16 if device != "cpu" else torch.float32
|
||||
# "sdpa" works on CUDA + CPU and needs no extra dep. flash_attention_2
|
||||
# (Ampere+ CUDA, optional flash-attn) is opt-in via env.
|
||||
attn = os.environ.get("OMNIVOICE_MOSS_TTS_V15_ATTN", "sdpa")
|
||||
|
||||
@@ -48,6 +48,15 @@ from services.subprocess_backend import SubprocessBackend
|
||||
|
||||
logger = logging.getLogger("omnivoice.engines.pockettts")
|
||||
|
||||
_VENV_ENV_VAR = "OMNIVOICE_POCKETTTS_DIR"
|
||||
|
||||
|
||||
def _own_venv_python() -> "Path | None":
|
||||
"""The venv the one-click installer made for this engine, if any."""
|
||||
from services.sidecar_install import engine_venv_python
|
||||
|
||||
return engine_venv_python(_VENV_ENV_VAR)
|
||||
|
||||
if TYPE_CHECKING:
|
||||
import torch # noqa: F401
|
||||
|
||||
@@ -121,16 +130,17 @@ class PocketTTSBackend(SubprocessBackend):
|
||||
def is_available(cls) -> tuple[bool, str]:
|
||||
if platform_error := cls._platform_error():
|
||||
return False, platform_error
|
||||
# Optional-dep gate: the pocket-tts wheel is installed only when the user
|
||||
# opted in. The interpreter is the parent's own (sys.executable), so
|
||||
# there is no separate venv to validate.
|
||||
try:
|
||||
import pocket_tts # type: ignore[import-not-found] # noqa: F401
|
||||
except Exception as e:
|
||||
return False, (
|
||||
f"pocket_tts package not installed or failed to import ({e}). "
|
||||
f"Enable in Settings -> Engines (uv sync --extra pockettts)."
|
||||
)
|
||||
# Installed either into its own venv by the one-click installer, which
|
||||
# verified `import pocket_tts` there before saving the path, or into the
|
||||
# app's environment by `uv sync --extra pockettts`.
|
||||
if _own_venv_python() is None:
|
||||
try:
|
||||
import pocket_tts # type: ignore[import-not-found] # noqa: F401
|
||||
except Exception as e:
|
||||
return False, (
|
||||
f"pocket_tts package not installed or failed to import ({e}). "
|
||||
"Install it from Model Catalogue → Engines."
|
||||
)
|
||||
|
||||
# The model repository has an additional gated-access agreement and
|
||||
# prohibited-use conditions beyond its CC-BY-4.0 license. Keep first
|
||||
@@ -145,10 +155,10 @@ class PocketTTSBackend(SubprocessBackend):
|
||||
|
||||
@classmethod
|
||||
def venv_python(cls) -> Path:
|
||||
# Parent interpreter: pocket-tts deps (torch>=2.5, scipy, beartype) sit
|
||||
# happily at the parent's pins, so this isolates for crash recovery, not
|
||||
# dependency pins (same rationale as omnivoice-subprocess).
|
||||
return Path(sys.executable)
|
||||
# Its own venv when the one-click installer made one. Otherwise the
|
||||
# parent interpreter, where `uv sync --extra pockettts` installs it
|
||||
# (its deps sit happily at the parent's pins).
|
||||
return _own_venv_python() or Path(sys.executable)
|
||||
|
||||
@classmethod
|
||||
def sidecar_script(cls) -> Path:
|
||||
|
||||
@@ -48,6 +48,15 @@ if TYPE_CHECKING:
|
||||
|
||||
logger = logging.getLogger("omnivoice.supertonic3")
|
||||
|
||||
_VENV_ENV_VAR = "OMNIVOICE_SUPERTONIC3_DIR"
|
||||
|
||||
|
||||
def _own_venv_python() -> "Path | None":
|
||||
"""The venv the one-click installer made for this engine, if any."""
|
||||
from services.sidecar_install import engine_venv_python
|
||||
|
||||
return engine_venv_python(_VENV_ENV_VAR)
|
||||
|
||||
|
||||
# Absolute path to the sidecar script ‑‑ same pattern as IndexTTS's
|
||||
# ``INDEXTTS_SIDECAR_SCRIPT``. SubprocessBackend spawns it with the
|
||||
@@ -80,11 +89,11 @@ class Supertonic3Backend(SubprocessBackend):
|
||||
|
||||
@classmethod
|
||||
def venv_python(cls) -> Path:
|
||||
"""Supertonic-3 lives in the main OmniVoice venv ‑‑ no dedicated
|
||||
venv. ``sys.executable`` is the parent interpreter, which is the
|
||||
same Python that ``uv sync --extra supertonic`` populated.
|
||||
"""Its own venv when the one-click installer made one. Otherwise the
|
||||
parent interpreter, the same Python ``uv sync --extra supertonic``
|
||||
populated.
|
||||
"""
|
||||
return Path(sys.executable)
|
||||
return _own_venv_python() or Path(sys.executable)
|
||||
|
||||
@classmethod
|
||||
def sidecar_script(cls) -> Path:
|
||||
@@ -96,14 +105,16 @@ class Supertonic3Backend(SubprocessBackend):
|
||||
def is_available(cls) -> tuple[bool, str]:
|
||||
# 1. Optional-dep gate (TTS-02). The ``supertonic`` wheel is only
|
||||
# installed when the user opted in via ``--extra supertonic``.
|
||||
try:
|
||||
import supertonic # type: ignore[import-not-found] # noqa: F401
|
||||
except ImportError:
|
||||
return False, (
|
||||
"supertonic package not installed. Enable in "
|
||||
"Model Catalogue → Engines (installs `supertonic` via `uv add --optional "
|
||||
"supertonic supertonic==1.3.1`)."
|
||||
)
|
||||
# Its own venv (made by the one-click installer, which verified the
|
||||
# import there) or the app's environment (`uv sync --extra`).
|
||||
if _own_venv_python() is None:
|
||||
try:
|
||||
import supertonic # type: ignore[import-not-found] # noqa: F401
|
||||
except ImportError:
|
||||
return False, (
|
||||
"supertonic package not installed. Install it from "
|
||||
"Model Catalogue → Engines."
|
||||
)
|
||||
|
||||
# 2. License acceptance gate (TTS-05). Defence in depth: the
|
||||
# settings_store helper handles the read; we just refuse
|
||||
|
||||
@@ -137,9 +137,17 @@ def _resolve_pinned_sha() -> str:
|
||||
# Final fallback ‑‑ relative import for when the file is invoked
|
||||
# via ``python backend/engines/supertonic3/sidecar.py`` rather
|
||||
# than via ``python -m backend.engines.supertonic3.sidecar``.
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
|
||||
from engines.supertonic3.constants import PINNED_REVISION_SHA # type: ignore[import-not-found]
|
||||
return PINNED_REVISION_SHA
|
||||
# Load constants.py by path. Importing it as `engines.supertonic3…`
|
||||
# runs the package __init__, which imports the app's backend, and that
|
||||
# is absent from the engine's own venv (one-click install).
|
||||
import importlib.util
|
||||
|
||||
spec = importlib.util.spec_from_file_location(
|
||||
"_supertonic3_constants", Path(__file__).resolve().with_name("constants.py"),
|
||||
)
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module) # type: ignore[union-attr]
|
||||
return module.PINNED_REVISION_SHA
|
||||
|
||||
|
||||
# ── model loading (lazy, on first synthesize) ─────────────────────────────
|
||||
|
||||
+52
-15
@@ -585,13 +585,14 @@ def _phase_a_build_inner() -> None:
|
||||
pass # never block startup on the migration; it retries next launch
|
||||
# Restore persisted env vars from prefs.json (Settings UI writes them
|
||||
# there so they survive backend restarts) — before any user code reads
|
||||
# os.environ, and never overriding an explicitly-set env var.
|
||||
# os.environ, and never overriding an explicitly-set env var. Also
|
||||
# snapshots which keys an external source (shell, `.env`, Docker, …)
|
||||
# already provided, so a Settings control can tell the user their saved
|
||||
# value is being shadowed instead of silently promising it will apply
|
||||
# (core.prefs.is_env_shadowed — #1787 review fix).
|
||||
try:
|
||||
from core.prefs import _load as _load_all_prefs
|
||||
_prefs = _load_all_prefs()
|
||||
for _k, _v in _prefs.items():
|
||||
if _k.startswith("env.") and _v:
|
||||
os.environ.setdefault(_k[len("env."):], str(_v))
|
||||
from core.prefs import _load as _load_all_prefs, restore_env
|
||||
restore_env(_load_all_prefs())
|
||||
except Exception:
|
||||
pass # prefs.json missing or broken — fine on first run
|
||||
# yt-dlp user-update overlay: must run before anything imports yt_dlp so
|
||||
@@ -623,7 +624,19 @@ def _phase_a_build_inner() -> None:
|
||||
_startup_progress.begin_step("ml_imports")
|
||||
import torchaudio
|
||||
warnings.filterwarnings("ignore", category=UserWarning)
|
||||
torchaudio.set_audio_backend("soundfile")
|
||||
# torchaudio 2.9 REMOVED set_audio_backend(); soundfile has been the only
|
||||
# backend since 2.0, so the call was already a no-op there and is simply
|
||||
# absent now. Unguarded it raises AttributeError inside `ml_imports`, and a
|
||||
# failure in that phase takes the whole backend down — the desktop app sits
|
||||
# on "starting backend" forever and /health stays 503.
|
||||
#
|
||||
# That is not a hypothetical version: RTX 50-series (Blackwell, sm_120)
|
||||
# users have no choice but to move off the pinned torch 2.8.0, which has no
|
||||
# sm_120 kernels, and the torch 2.9.x they land on brings torchaudio 2.9
|
||||
# with it. So the one group forced to upgrade hit a hard startup crash for
|
||||
# a line that does nothing (#1931).
|
||||
if hasattr(torchaudio, "set_audio_backend"):
|
||||
torchaudio.set_audio_backend("soundfile")
|
||||
from utils import hf_progress
|
||||
# HF tqdm patch before any library import that can trigger
|
||||
# hf_hub_download (transformers, mlx_whisper, …).
|
||||
@@ -685,11 +698,12 @@ def _phase_a_build_inner() -> None:
|
||||
settings as settings_router, # Phase 1 AUTH-03: HF token save/clear/state
|
||||
media_tools as media_tools_router, # Audio tools: ffmpeg/ffprobe/yt-dlp
|
||||
auth as auth_router,
|
||||
voice_convert, # Studio Convert: speech-to-speech via ASR → TTS
|
||||
)
|
||||
from api.routers import mcp_bindings as _mcp_bindings_router # noqa: E402
|
||||
from api.routers import workers as workers_router # noqa: E402
|
||||
_router_modules.extend([
|
||||
system, profiles, exports, generation, dub_core, dub_generate,
|
||||
system, profiles, exports, generation, voice_convert, dub_core, dub_generate,
|
||||
dub_export, dub_translate, projects, glossary, engines, tools,
|
||||
stories, setup, gallery, archetypes, describe_voice, community,
|
||||
batch, watermark, events, capture, capture_ws, speech_platform, dictation,
|
||||
@@ -1078,6 +1092,33 @@ async def lifespan(app: FastAPI):
|
||||
app.state.startup_task = asyncio.create_task(_deferred_startup(app))
|
||||
yield
|
||||
# ── Graceful shutdown (SIGTERM from Tauri, Ctrl+C, etc.) ────────────
|
||||
# Retire the run sentinel FIRST, before any bounded wait below (#1895):
|
||||
# once uvicorn has begun graceful shutdown the exit is deliberate by
|
||||
# definition, so the sentinel has already done its job. This is one
|
||||
# os.remove, against a ~50s worst-case tail of bounded waits plus model
|
||||
# unload / free_vram() / gc.collect() below. Measured on macOS: a normal
|
||||
# shutdown takes 5.25s end to end, while the desktop shell allows 2s
|
||||
# (bootstrap.rs terminate_process_tree) before SIGKILL — so the old
|
||||
# placement at the very end was killed every time on any run that had
|
||||
# reached a working state. Doing the deadline-sensitive step first makes
|
||||
# correctness independent of how much of that tail runs, instead of
|
||||
# depending on the shell-side deadline being long enough to cover it.
|
||||
#
|
||||
# SCOPE, explicitly: this only helps platforms where lifespan teardown
|
||||
# actually BEGINS. On Windows it does not — tools.rs terminates the job
|
||||
# object with no graceful phase at all, so this line is never reached and
|
||||
# a deliberate quit is still misreported as a crash there. That needs the
|
||||
# shell to signal deliberate intent before the hard kill, which is a
|
||||
# separate Rust-side change and is tracked separately; nothing here
|
||||
# should be read as fixing Windows.
|
||||
#
|
||||
# sentinel_cleared feeds the truthful "Shutdown: done."/degraded log at
|
||||
# the end of this function; nothing below re-clears the sentinel, so a
|
||||
# later failure can't mask this result.
|
||||
try:
|
||||
sentinel_cleared = run_sentinel.clear_sentinel()
|
||||
except Exception:
|
||||
sentinel_cleared = False
|
||||
# May run after a startup that never finished (SIGTERM mid-Phase-A/B), so
|
||||
# every handle is read from app.state with a None default and every
|
||||
# deferred-phase name is guarded.
|
||||
@@ -1204,13 +1245,9 @@ async def lifespan(app: FastAPI):
|
||||
await close_http_client()
|
||||
except Exception:
|
||||
pass
|
||||
# Last thing on a clean shutdown: retire the run sentinel so the next
|
||||
# startup doesn't misread this exit as a crash (#1164). If clearing fails,
|
||||
# retain the sentinel and report a degraded shutdown truthfully.
|
||||
try:
|
||||
sentinel_cleared = run_sentinel.clear_sentinel()
|
||||
except Exception:
|
||||
sentinel_cleared = False
|
||||
# Sentinel was already retired at the TOP of this block (#1895) — report
|
||||
# truthfully using that result rather than clearing (or re-checking) it
|
||||
# again here, so a failure in the steps above can't mask it as "done."
|
||||
if sentinel_cleared:
|
||||
logger.info("Shutdown: done.")
|
||||
else:
|
||||
|
||||
+297
-37
@@ -7,8 +7,8 @@ Run standalone:
|
||||
|
||||
Tools exposed:
|
||||
generate_speech — text → WAV audio (voice clone or design)
|
||||
clone_voice — base64 reference audio → new voice profile
|
||||
transcribe — base64 audio → text
|
||||
clone_voice — reference audio (base64, or a file path) → new voice profile
|
||||
transcribe — audio (base64, or a file path) → text
|
||||
list_voices — enumerate saved voice profiles
|
||||
list_languages — available TTS languages
|
||||
list_personalities — voice personality presets
|
||||
@@ -17,6 +17,18 @@ Tools exposed:
|
||||
Resources exposed:
|
||||
voice://{profile_id} — voice profile metadata
|
||||
history://recent — last 20 generated audio items
|
||||
|
||||
Output mode (OMNIVOICE_MCP_OUTPUT_MODE):
|
||||
resources — generate_speech returns the WAV as base64 inline (the original
|
||||
contract; default)
|
||||
files — it returns a URL to the render (and, with a base path, a WAV
|
||||
written there); nothing large ever enters the agent's context
|
||||
both — both of the above
|
||||
|
||||
File inputs (OMNIVOICE_MCP_BASE_PATH):
|
||||
One directory that agents may read audio from (transcribe / clone_voice
|
||||
`*_path` arguments) and receive files in (files mode). It is the security
|
||||
boundary: with no base path configured, path-shaped inputs are refused.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -25,6 +37,8 @@ import base64
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
import stat
|
||||
import sys
|
||||
|
||||
logger = logging.getLogger("omnivoice.mcp")
|
||||
@@ -69,6 +83,244 @@ def _sniff_audio_ext(raw: bytes) -> str:
|
||||
return ".wav"
|
||||
|
||||
|
||||
# ── Output mode + the base path boundary ─────────────────────────────────
|
||||
# An LLM agent that receives a WAV as base64 pays for every byte in context:
|
||||
# a 1.4 s clip already brushes per-result token caps, and a paragraph of
|
||||
# narration blows them outright. The ElevenLabs MCP settled this with an
|
||||
# OUTPUT_MODE (files / resources / both) and a BASE_PATH that doubles as the
|
||||
# security boundary for file-shaped inputs; the same two knobs here, named in
|
||||
# the OMNIVOICE_* family the rest of the server reads.
|
||||
|
||||
_OUTPUT_MODES = ("resources", "files", "both")
|
||||
_MAX_INPUT_BYTES = 200 * 1024 * 1024
|
||||
_SAFE_AUDIO_ID = re.compile(r"^[A-Za-z0-9_-]{1,64}$")
|
||||
|
||||
|
||||
def _output_mode() -> str:
|
||||
"""How generate_speech hands audio back (OMNIVOICE_MCP_OUTPUT_MODE).
|
||||
|
||||
'resources' is the original base64-inline contract and stays the default
|
||||
so existing integrations see no change; 'files' returns a URL to the
|
||||
render (plus a WAV under the base path when one is configured); 'both'
|
||||
returns everything. Anything unrecognized falls back to 'resources' with
|
||||
a warning rather than failing the tool."""
|
||||
mode = os.environ.get("OMNIVOICE_MCP_OUTPUT_MODE", "resources").strip().lower()
|
||||
if mode not in _OUTPUT_MODES:
|
||||
logger.warning(
|
||||
"OMNIVOICE_MCP_OUTPUT_MODE=%r is not one of %s; using 'resources'",
|
||||
mode, _OUTPUT_MODES,
|
||||
)
|
||||
return "resources"
|
||||
return mode
|
||||
|
||||
|
||||
def _base_path() -> "str | None":
|
||||
"""The one directory agents may read audio from and receive files in
|
||||
(OMNIVOICE_MCP_BASE_PATH), realpath'd; None when unset."""
|
||||
raw = os.environ.get("OMNIVOICE_MCP_BASE_PATH", "").strip()
|
||||
if not raw:
|
||||
return None
|
||||
return os.path.realpath(os.path.expanduser(raw))
|
||||
|
||||
|
||||
def _resolve_under_base(path: str) -> str:
|
||||
"""Absolute realpath of ``path`` when it lies inside the base path.
|
||||
|
||||
Relative paths resolve against the base; absolute paths must already be
|
||||
inside it. Both sides are realpath'd, so a symlink pointing outward cannot
|
||||
smuggle a read in. Raises ValueError with an agent-legible reason when no
|
||||
base path is configured or the path escapes it."""
|
||||
base = _base_path()
|
||||
if base is None:
|
||||
raise ValueError(
|
||||
"OMNIVOICE_MCP_BASE_PATH is not set; file paths are refused until it "
|
||||
"names a directory"
|
||||
)
|
||||
candidate = os.path.realpath(os.path.join(base, os.path.expanduser(path)))
|
||||
if not _path_is_under_base(base, candidate):
|
||||
raise ValueError(f"{path!r} resolves outside OMNIVOICE_MCP_BASE_PATH")
|
||||
return candidate
|
||||
|
||||
|
||||
def _opened_file_is_confined(fd: int, resolved: str, base: str) -> bool:
|
||||
"""Verify that an opened descriptor still names a file under ``base``."""
|
||||
proc_fd = f"/proc/self/fd/{fd}"
|
||||
if os.path.exists(proc_fd):
|
||||
return _path_is_under_base(base, os.path.realpath(proc_fd))
|
||||
try:
|
||||
current = os.path.realpath(resolved)
|
||||
return _path_is_under_base(base, current) and os.path.samestat(
|
||||
os.fstat(fd), os.stat(current, follow_symlinks=False)
|
||||
)
|
||||
except OSError:
|
||||
return False
|
||||
|
||||
|
||||
def _path_is_under_base(base: str, candidate: str) -> bool:
|
||||
try:
|
||||
common = os.path.commonpath([base, candidate])
|
||||
except ValueError: # different drives on Windows
|
||||
return False
|
||||
return os.path.normcase(common) == os.path.normcase(base)
|
||||
|
||||
|
||||
def _open_under_base(path: str, flags: int, *, mode: int = 0o600) -> tuple[int, str]:
|
||||
"""Open ``path`` without following a component replaced after validation."""
|
||||
base = _base_path()
|
||||
if base is None:
|
||||
raise ValueError(
|
||||
"OMNIVOICE_MCP_BASE_PATH is not set; file paths are refused until it "
|
||||
"names a directory"
|
||||
)
|
||||
resolved = _resolve_under_base(path)
|
||||
relative = os.path.relpath(resolved, base)
|
||||
parts = [part for part in relative.split(os.sep) if part not in ("", ".")]
|
||||
if not parts or parts[0] == os.pardir:
|
||||
raise ValueError(f"{path!r} resolves outside OMNIVOICE_MCP_BASE_PATH")
|
||||
|
||||
no_follow = getattr(os, "O_NOFOLLOW", 0)
|
||||
close_on_exec = getattr(os, "O_CLOEXEC", 0)
|
||||
binary = getattr(os, "O_BINARY", 0)
|
||||
file_flags = flags | no_follow | close_on_exec | binary
|
||||
supports_dir_fd = os.open in getattr(os, "supports_dir_fd", ())
|
||||
directory_flag = getattr(os, "O_DIRECTORY", 0)
|
||||
|
||||
if supports_dir_fd and directory_flag:
|
||||
directory_flags = os.O_RDONLY | directory_flag | no_follow | close_on_exec
|
||||
directory_fd = os.open(base, directory_flags)
|
||||
try:
|
||||
for component in parts[:-1]:
|
||||
next_fd = os.open(component, directory_flags, dir_fd=directory_fd)
|
||||
os.close(directory_fd)
|
||||
directory_fd = next_fd
|
||||
fd = os.open(parts[-1], file_flags, mode, dir_fd=directory_fd)
|
||||
finally:
|
||||
os.close(directory_fd)
|
||||
else:
|
||||
fd = os.open(resolved, file_flags, mode)
|
||||
|
||||
if not _opened_file_is_confined(fd, resolved, base):
|
||||
os.close(fd)
|
||||
raise ValueError(f"{path!r} resolves outside OMNIVOICE_MCP_BASE_PATH")
|
||||
return fd, resolved
|
||||
|
||||
|
||||
def _read_input_audio(
|
||||
audio_base64: "str | None",
|
||||
audio_path: "str | None",
|
||||
*,
|
||||
label: str = "audio_base64",
|
||||
too_big: str = "audio exceeds 200 MB limit",
|
||||
) -> "tuple[bytes | None, str | None]":
|
||||
"""Audio bytes from exactly one of the two input lanes, or (None, error).
|
||||
|
||||
The base64 lane keeps its data-URI tolerance and 200 MB cap; the path lane
|
||||
is honored only inside the base path (the security boundary) and applies
|
||||
the same cap to the file's size before reading it."""
|
||||
if bool(audio_base64) == bool(audio_path):
|
||||
return None, f"pass exactly one of {label} or the matching *_path argument"
|
||||
if audio_path:
|
||||
try:
|
||||
fd, _resolved = _open_under_base(audio_path, os.O_RDONLY)
|
||||
except ValueError as e:
|
||||
return None, str(e)
|
||||
except FileNotFoundError:
|
||||
return None, f"no such file under OMNIVOICE_MCP_BASE_PATH: {audio_path!r}"
|
||||
except OSError as e:
|
||||
return None, f"could not safely read {audio_path!r}: {e}"
|
||||
with os.fdopen(fd, "rb") as handle:
|
||||
info = os.fstat(handle.fileno())
|
||||
if not stat.S_ISREG(info.st_mode):
|
||||
return None, f"{audio_path!r} is not a regular file"
|
||||
if info.st_size > _MAX_INPUT_BYTES:
|
||||
return None, too_big
|
||||
raw = handle.read(_MAX_INPUT_BYTES + 1)
|
||||
if len(raw) > _MAX_INPUT_BYTES:
|
||||
return None, too_big
|
||||
if not raw:
|
||||
return None, f"{label} is empty"
|
||||
return raw, None
|
||||
encoded = (
|
||||
audio_base64.split(",", 1)[-1]
|
||||
if audio_base64.startswith("data:")
|
||||
else audio_base64
|
||||
)
|
||||
max_encoded_bytes = 4 * ((_MAX_INPUT_BYTES + 2) // 3)
|
||||
if len(encoded) > max_encoded_bytes:
|
||||
return None, too_big
|
||||
raw = _decode_ref_audio(audio_base64)
|
||||
if raw is None:
|
||||
return None, f"{label} is not valid base64"
|
||||
if not raw:
|
||||
return None, f"{label} is empty"
|
||||
if len(raw) > _MAX_INPUT_BYTES:
|
||||
return None, too_big
|
||||
return raw, None
|
||||
|
||||
|
||||
def _write_output(audio_id: str, raw: bytes) -> str:
|
||||
"""Land a render under the base path as ``<audio_id>.wav``; returns the path."""
|
||||
if not _SAFE_AUDIO_ID.fullmatch(audio_id):
|
||||
raise ValueError("backend returned an invalid X-Audio-Id header")
|
||||
base = _base_path()
|
||||
os.makedirs(base, exist_ok=True)
|
||||
filename = f"{audio_id}.wav"
|
||||
fd, path = _open_under_base(filename, os.O_WRONLY | os.O_CREAT | os.O_EXCL)
|
||||
with os.fdopen(fd, "wb") as handle:
|
||||
handle.write(raw)
|
||||
return path
|
||||
|
||||
|
||||
def _post_timeout_s() -> float:
|
||||
"""Seconds the tools wait on a backend POST (OMNIVOICE_MCP_TIMEOUT_S,
|
||||
default 120). A CPU host renders a paragraph in minutes and serializes
|
||||
generations, so an agent behind another render used to hit the fixed
|
||||
budget with an empty-message timeout; the knob follows the backend's own
|
||||
OMNIVOICE_GENERATE_TIMEOUT_S when a deployment raises that."""
|
||||
raw = os.environ.get("OMNIVOICE_MCP_TIMEOUT_S", "").strip()
|
||||
try:
|
||||
value = float(raw) if raw else 120.0
|
||||
except ValueError:
|
||||
logger.warning("OMNIVOICE_MCP_TIMEOUT_S=%r is not a number; using 120", raw)
|
||||
return 120.0
|
||||
return value if value > 0 else 120.0
|
||||
|
||||
|
||||
def _maybe_number(value):
|
||||
"""A response-header number as a number, or the raw text (e.g. '?')."""
|
||||
try:
|
||||
return float(value)
|
||||
except (TypeError, ValueError):
|
||||
return value
|
||||
|
||||
|
||||
def _speech_result(audio_id: str, gen_time, duration, raw: bytes, api_base: str) -> dict:
|
||||
"""The generate_speech reply shaped by the output mode.
|
||||
|
||||
The backend already keeps every render on disk and serves it at
|
||||
``/audio/<audio_id>.wav``, so files mode costs nothing but a URL - plus one
|
||||
write when a base path invites the WAV into the agent's own directory."""
|
||||
if not _SAFE_AUDIO_ID.fullmatch(audio_id):
|
||||
raise ValueError("backend returned an invalid X-Audio-Id header")
|
||||
mode = _output_mode()
|
||||
out = {
|
||||
"audio_id": audio_id,
|
||||
"generation_time_s": gen_time,
|
||||
"audio_duration_s": duration,
|
||||
"format": "wav",
|
||||
"output_mode": mode,
|
||||
}
|
||||
if mode in ("files", "both"):
|
||||
out["audio_url"] = f"{api_base.rstrip('/')}/audio/{audio_id}.wav"
|
||||
if _base_path() is not None:
|
||||
out["output_path"] = _write_output(audio_id, raw)
|
||||
else:
|
||||
out["note"] = "set OMNIVOICE_MCP_BASE_PATH to also receive the WAV as a file"
|
||||
if mode in ("resources", "both"):
|
||||
out["wav_base64"] = base64.b64encode(raw).decode("ascii")
|
||||
return out
|
||||
|
||||
|
||||
# ── Lazy imports — keeps startup fast when not using MCP ────────────────
|
||||
|
||||
|
||||
@@ -147,7 +399,7 @@ def create_mcp_server():
|
||||
|
||||
async def _api_post_form(path: str, data: dict, files: dict | None = None):
|
||||
import httpx
|
||||
async with httpx.AsyncClient(base_url=_api_base(), timeout=120) as c:
|
||||
async with httpx.AsyncClient(base_url=_api_base(), timeout=_post_timeout_s()) as c:
|
||||
r = await c.post(path, data=data, files=files or {})
|
||||
r.raise_for_status()
|
||||
return r
|
||||
@@ -190,8 +442,12 @@ def create_mcp_server():
|
||||
steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
|
||||
|
||||
Returns:
|
||||
JSON with audio_id, generation_time, audio_duration, and
|
||||
base64-encoded WAV data.
|
||||
JSON with audio_id, generation_time_s, audio_duration_s and the
|
||||
audio itself shaped by OMNIVOICE_MCP_OUTPUT_MODE: base64 WAV data
|
||||
('resources', the default), a URL to the render plus a WAV under
|
||||
OMNIVOICE_MCP_BASE_PATH when one is set ('files'), or all of the
|
||||
above ('both'). Prefer 'files' for LLM agents: nothing large
|
||||
enters the context.
|
||||
"""
|
||||
# Per-agent voice binding (Wave 2.2): explicit arg wins; otherwise
|
||||
# resolve this client's bound profile, then the global default.
|
||||
@@ -218,18 +474,10 @@ def create_mcp_server():
|
||||
r = await _api_post_form("/generate", data=form)
|
||||
|
||||
audio_id = r.headers.get("X-Audio-Id", "unknown")
|
||||
gen_time = r.headers.get("X-Gen-Time", "?")
|
||||
duration = r.headers.get("X-Audio-Duration", "?")
|
||||
gen_time = _maybe_number(r.headers.get("X-Gen-Time", "?"))
|
||||
duration = _maybe_number(r.headers.get("X-Audio-Duration", "?"))
|
||||
|
||||
wav_b64 = base64.b64encode(r.content).decode("ascii")
|
||||
|
||||
return (
|
||||
f'{{"audio_id":"{audio_id}",'
|
||||
f'"generation_time_s":{gen_time},'
|
||||
f'"audio_duration_s":{duration},'
|
||||
f'"format":"wav",'
|
||||
f'"wav_base64":"{wav_b64}"}}'
|
||||
)
|
||||
return json.dumps(_speech_result(audio_id, gen_time, duration, r.content, _api_base()))
|
||||
|
||||
@mcp.tool()
|
||||
async def list_voices() -> str:
|
||||
@@ -266,30 +514,39 @@ def create_mcp_server():
|
||||
)
|
||||
|
||||
@mcp.tool()
|
||||
async def transcribe(audio_base64: str, language: str | None = None) -> str:
|
||||
async def transcribe(
|
||||
audio_base64: str | None = None,
|
||||
audio_path: str | None = None,
|
||||
language: str | None = None,
|
||||
) -> str:
|
||||
"""Transcribe spoken audio to text.
|
||||
|
||||
Pass exactly one of audio_base64 or audio_path.
|
||||
|
||||
Args:
|
||||
audio_base64: Base64-encoded audio bytes (wav/mp3/webm/m4a).
|
||||
audio_path: Path to an audio file under OMNIVOICE_MCP_BASE_PATH
|
||||
(relative to it, or absolute inside it). The base path is the
|
||||
security boundary: with none configured, paths are refused.
|
||||
Prefer this lane for LLM agents - the audio never enters the
|
||||
agent's context.
|
||||
language: Optional language hint; omit for auto-detect.
|
||||
|
||||
Returns:
|
||||
JSON with the recognized text, language, and duration.
|
||||
"""
|
||||
try:
|
||||
raw = base64.b64decode(audio_base64, validate=True)
|
||||
except Exception:
|
||||
return '{"error":"audio_base64 is not valid base64"}'
|
||||
# 200 MB cap — same spirit as voicebox's transcribe gate. Keeps a
|
||||
# buggy/hostile agent from posting an unbounded blob.
|
||||
if len(raw) > 200 * 1024 * 1024:
|
||||
return '{"error":"audio exceeds 200 MB limit"}'
|
||||
# 200 MB cap on both lanes — same spirit as voicebox's transcribe
|
||||
# gate. Keeps a buggy/hostile agent from posting an unbounded blob.
|
||||
raw, err = _read_input_audio(audio_base64, audio_path)
|
||||
if err:
|
||||
return json.dumps({"error": err})
|
||||
data = {}
|
||||
if language:
|
||||
data["language"] = language
|
||||
r = await _api_post_form(
|
||||
"/transcribe", data=data,
|
||||
files={"audio": ("audio.wav", raw, "application/octet-stream")},
|
||||
files={"audio": (f"audio{_sniff_audio_ext(raw)}", raw,
|
||||
"application/octet-stream")},
|
||||
)
|
||||
return str(r.json())
|
||||
|
||||
@@ -319,15 +576,17 @@ def create_mcp_server():
|
||||
@mcp.tool()
|
||||
async def clone_voice(
|
||||
name: str,
|
||||
ref_audio_base64: str,
|
||||
ref_audio_base64: str | None = None,
|
||||
ref_text: str = "",
|
||||
instruct: str = "",
|
||||
language: str = "Auto",
|
||||
ref_audio_path: str | None = None,
|
||||
) -> str:
|
||||
"""Clone a new voice profile from a reference audio sample.
|
||||
|
||||
The new voice is immediately available for use with generate_speech
|
||||
(pass the returned profile_id as the profile_id argument).
|
||||
(pass the returned profile_id as the profile_id argument). Pass
|
||||
exactly one of ref_audio_base64 or ref_audio_path.
|
||||
|
||||
Args:
|
||||
name: A human-friendly name for the cloned voice.
|
||||
@@ -338,19 +597,20 @@ def create_mcp_server():
|
||||
quality for some engines).
|
||||
instruct: Optional style instruction (e.g. 'whisper', 'excited').
|
||||
language: Language of the reference audio (ISO code or 'Auto').
|
||||
ref_audio_path: Path to the reference audio under
|
||||
OMNIVOICE_MCP_BASE_PATH (relative to it, or absolute inside
|
||||
it); refused when no base path is configured. Prefer this
|
||||
lane for LLM agents - the clip never enters the context.
|
||||
|
||||
Returns:
|
||||
JSON with the new profile's id, name, and kind.
|
||||
"""
|
||||
# Reject oversized inputs before decoding (base64 is always larger
|
||||
# than raw, so this is a safe lower bound on the decoded size).
|
||||
if len(ref_audio_base64) > 200 * 1024 * 1024:
|
||||
return '{"error":"reference audio exceeds 200 MB limit"}'
|
||||
raw = _decode_ref_audio(ref_audio_base64)
|
||||
if raw is None:
|
||||
return '{"error":"ref_audio_base64 is not valid base64"}'
|
||||
if not raw:
|
||||
return '{"error":"ref_audio_base64 is empty"}'
|
||||
raw, err = _read_input_audio(
|
||||
ref_audio_base64, ref_audio_path,
|
||||
label="ref_audio_base64", too_big="reference audio exceeds 200 MB limit",
|
||||
)
|
||||
if err:
|
||||
return json.dumps({"error": err})
|
||||
import httpx
|
||||
try:
|
||||
r = await _api_post_form(
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
"""Retain the dispatch-time deadline policy across worker/control-plane loss."""
|
||||
from alembic import op
|
||||
import sqlalchemy as sa
|
||||
|
||||
revision = "0011_remote_attempt_deadlines"
|
||||
down_revision = "0010_remote_worker_schema"
|
||||
branch_labels = None
|
||||
depends_on = None
|
||||
|
||||
|
||||
def upgrade() -> None:
|
||||
columns = op.get_bind().execute(sa.text("PRAGMA table_info(remote_task_attempts)"))
|
||||
if not any(row[1] == "deadlines_json" for row in columns):
|
||||
op.add_column("remote_task_attempts", sa.Column("deadlines_json", sa.Text(), nullable=True))
|
||||
|
||||
|
||||
def downgrade() -> None:
|
||||
op.drop_column("remote_task_attempts", "deadlines_json")
|
||||
@@ -33,8 +33,11 @@ WS_TICKET_PREFIX = "ovs_ws_ticket_"
|
||||
_TOKEN_BYTES = 32
|
||||
_ENCODED_TOKEN_LENGTH = 43
|
||||
_TOKEN_BODY_RE = re.compile(rf"^[A-Za-z0-9_-]{{{_ENCODED_TOKEN_LENGTH}}}$")
|
||||
# Every ticketed WebSocket route. The first-party mirror is ``ALLOWED_WS_PATHS``
|
||||
# in frontend/src/api/authSession.ts — a route missing here mints a 422 and the
|
||||
# UI consumer fails silently (#1769 added /ws/tts for the live dub preview).
|
||||
_ALLOWED_WS_PATHS = frozenset(
|
||||
{"/ws/events", "/ws/transcribe", "/v1/audio/transcriptions/stream"}
|
||||
{"/ws/events", "/ws/transcribe", "/ws/tts", "/v1/audio/transcriptions/stream"}
|
||||
)
|
||||
_ADMIN_CAPABILITIES = frozenset({"consume", "admin"})
|
||||
_KEY_GENERATION_INFO = b"omnivoice-admin-key-generation-v1"
|
||||
|
||||
+145
-19
@@ -24,6 +24,7 @@ faster-whisper because it's available on every platform we ship to).
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import ipaddress
|
||||
import logging
|
||||
import os
|
||||
import re
|
||||
@@ -31,6 +32,7 @@ import contextlib
|
||||
import threading
|
||||
import time
|
||||
import weakref
|
||||
from urllib.parse import urlsplit
|
||||
from utils.containment import contain_system_exit
|
||||
|
||||
from abc import ABC, abstractmethod
|
||||
@@ -147,25 +149,78 @@ def _isolated_engine_hint(streak: int) -> str:
|
||||
async def run_transcribe_guarded(executor, fn, *, what: str = "ASR",
|
||||
timeout: float = ASR_TRANSCRIBE_TIMEOUT_S,
|
||||
timeout_env: str = "OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S",
|
||||
reset_on_timeout: bool = False):
|
||||
reset_on_timeout: bool = False,
|
||||
on_abandon=None):
|
||||
"""Run a blocking transcribe ``fn`` in ``executor`` with a hard wall-clock
|
||||
bound. On timeout, raise :class:`ASRTimeoutError` with guidance instead of
|
||||
letting the request hang forever.
|
||||
|
||||
``run_in_executor`` cannot cancel the underlying thread, so a timed-out
|
||||
A future cannot cancel the underlying thread, so a timed-out
|
||||
in-process CTranslate2/whisperx call still owns its model and device. The
|
||||
default deliberately leaves that worker accounted for: swapping in a fresh
|
||||
pool and immediately retrying the same backend overlaps two native calls,
|
||||
which produced the Windows access violation in #1669. A caller backed by a
|
||||
genuinely killable process may opt into ``reset_on_timeout``.
|
||||
|
||||
``on_abandon`` is called once after a timed-out or cancelled worker can no
|
||||
longer access its inputs. Queued work cancelled before it starts calls it
|
||||
immediately; running work calls it from the worker finalizer. Normal
|
||||
completion leaves cleanup with the caller.
|
||||
"""
|
||||
loop = asyncio.get_running_loop()
|
||||
# Same SystemExit containment as the TTS pool (#1133 class): an ASR
|
||||
# dependency written as a CLI must not be able to shut the backend down.
|
||||
fut = loop.run_in_executor(executor, contain_system_exit(fn, what))
|
||||
inner = contain_system_exit(fn, what)
|
||||
abandon_lock = threading.Lock()
|
||||
abandon_state = {
|
||||
"requested": False,
|
||||
"finished": False,
|
||||
"callback_called": False,
|
||||
}
|
||||
|
||||
def _fire_abandon_callback() -> None:
|
||||
if on_abandon is None:
|
||||
return
|
||||
with abandon_lock:
|
||||
if abandon_state["callback_called"]:
|
||||
return
|
||||
abandon_state["callback_called"] = True
|
||||
try:
|
||||
on_abandon()
|
||||
except Exception: # noqa: BLE001 — cleanup cannot hide the ASR result
|
||||
logger.exception("%s abandon cleanup failed", what)
|
||||
|
||||
def _job():
|
||||
try:
|
||||
return inner()
|
||||
finally:
|
||||
with abandon_lock:
|
||||
abandon_state["finished"] = True
|
||||
abandoned = abandon_state["requested"]
|
||||
if abandoned:
|
||||
_fire_abandon_callback()
|
||||
|
||||
concurrent_fut = executor.submit(_job)
|
||||
fut = asyncio.wrap_future(concurrent_fut, loop=loop)
|
||||
|
||||
def _abandon() -> None:
|
||||
cancelled_before_start = concurrent_fut.cancel()
|
||||
with abandon_lock:
|
||||
abandon_state["requested"] = True
|
||||
finished = abandon_state["finished"]
|
||||
fut.cancel()
|
||||
if cancelled_before_start or finished:
|
||||
_fire_abandon_callback()
|
||||
|
||||
try:
|
||||
result = await asyncio.wait_for(fut, timeout=timeout)
|
||||
# Shield the wrapper so timeout does not discard our ability to tell a
|
||||
# queued cancellation from a native thread that is still running.
|
||||
result = await asyncio.wait_for(asyncio.shield(fut), timeout=timeout)
|
||||
except asyncio.CancelledError:
|
||||
_abandon()
|
||||
raise
|
||||
except asyncio.TimeoutError:
|
||||
_abandon()
|
||||
if reset_on_timeout:
|
||||
reset_pool_after_wedge(executor, what=what)
|
||||
streak = _note_transcribe_timeout()
|
||||
@@ -1746,9 +1801,7 @@ class SherpaDictationBackend(ASRBackend):
|
||||
|
||||
def __init__(self, model_id: str | None = None):
|
||||
from services import sherpa_dictation as _sd
|
||||
mid = model_id or os.environ.get(
|
||||
"OMNIVOICE_SHERPA_ASR_MODEL", _sd.DEFAULT_MODEL_ID
|
||||
)
|
||||
mid = model_id or sherpa_engine_model_id()
|
||||
spec = _sd.get_spec(mid)
|
||||
if spec is None:
|
||||
raise ValueError(
|
||||
@@ -2000,6 +2053,42 @@ _ASR_OPENAI_COMPAT_MODEL_KEY = "asr.openai_compat.model"
|
||||
_ASR_OPENAI_COMPAT_SECRET_NAME = "asr_openai_compat_key"
|
||||
|
||||
|
||||
def normalize_openai_compat_asr_base_url(value: str) -> str:
|
||||
"""Normalize a safe ASR endpoint, allowing plain HTTP only on loopback."""
|
||||
base = (value or "").strip().rstrip("/")
|
||||
if not base:
|
||||
return ""
|
||||
try:
|
||||
parsed = urlsplit(base)
|
||||
_ = parsed.port
|
||||
except (TypeError, ValueError) as exc:
|
||||
raise ValueError("Invalid OpenAI-compatible ASR base URL") from exc
|
||||
scheme = parsed.scheme.lower()
|
||||
if (
|
||||
scheme not in {"http", "https"}
|
||||
or not parsed.hostname
|
||||
or parsed.username is not None
|
||||
or parsed.password is not None
|
||||
or parsed.query
|
||||
or parsed.fragment
|
||||
):
|
||||
raise ValueError(
|
||||
"OpenAI-compatible ASR base URL must be a credential-free HTTP(S) URL"
|
||||
)
|
||||
host = parsed.hostname.lower()
|
||||
loopback = host == "localhost"
|
||||
if not loopback:
|
||||
try:
|
||||
address = ipaddress.ip_address(host)
|
||||
address = getattr(address, "ipv4_mapped", None) or address
|
||||
loopback = address.is_loopback
|
||||
except ValueError:
|
||||
loopback = False
|
||||
if scheme == "http" and not loopback:
|
||||
raise ValueError("Non-loopback OpenAI-compatible ASR endpoints require HTTPS")
|
||||
return base
|
||||
|
||||
|
||||
def resolve_openai_compat_asr_base_url() -> str:
|
||||
from services import settings_store
|
||||
return (
|
||||
@@ -2063,7 +2152,7 @@ def probe_openai_compat_server(
|
||||
maps to a translated message:
|
||||
|
||||
not_configured no base URL anywhere
|
||||
invalid_url base URL without an http(s):// scheme
|
||||
invalid_url malformed URL or non-loopback HTTP endpoint
|
||||
ok 2xx — ``model_found`` says whether the configured
|
||||
model appears in the server's list (None = unknown)
|
||||
ok_no_models 404/405/501 — reachable, but no /models endpoint
|
||||
@@ -2078,7 +2167,7 @@ def probe_openai_compat_server(
|
||||
|
||||
from core.scrub import scrub_text
|
||||
|
||||
base = (base_url if base_url is not None else resolve_openai_compat_asr_base_url()).strip().rstrip("/")
|
||||
configured_base = base_url if base_url is not None else resolve_openai_compat_asr_base_url()
|
||||
mdl = (model if model is not None else resolve_openai_compat_asr_model()).strip()
|
||||
if api_key is None:
|
||||
key = resolve_openai_compat_asr_api_key()
|
||||
@@ -2094,9 +2183,11 @@ def probe_openai_compat_server(
|
||||
"model_found": None,
|
||||
"detail": None,
|
||||
}
|
||||
if not base:
|
||||
if not configured_base.strip():
|
||||
return out
|
||||
if not base.startswith(("http://", "https://")):
|
||||
try:
|
||||
base = normalize_openai_compat_asr_base_url(configured_base)
|
||||
except ValueError:
|
||||
out["status"] = "invalid_url"
|
||||
return out
|
||||
|
||||
@@ -2107,7 +2198,7 @@ def probe_openai_compat_server(
|
||||
try:
|
||||
with httpx.Client(
|
||||
timeout=httpx.Timeout(timeout_s, connect=min(5.0, timeout_s)),
|
||||
follow_redirects=True,
|
||||
follow_redirects=False,
|
||||
) as client:
|
||||
resp = client.get(f"{base}/models", headers=headers)
|
||||
except httpx.TimeoutException as exc:
|
||||
@@ -2170,13 +2261,20 @@ class OpenAICompatASRBackend(ASRBackend):
|
||||
gpu_compat = ("cpu",) # network client only — no local compute
|
||||
|
||||
def __init__(self):
|
||||
self._base_url = resolve_openai_compat_asr_base_url()
|
||||
self._base_url = normalize_openai_compat_asr_base_url(
|
||||
resolve_openai_compat_asr_base_url()
|
||||
)
|
||||
self._model = resolve_openai_compat_asr_model()
|
||||
|
||||
@classmethod
|
||||
def is_available(cls) -> tuple[bool, str]:
|
||||
if not resolve_openai_compat_asr_base_url():
|
||||
base_url = resolve_openai_compat_asr_base_url()
|
||||
if not base_url:
|
||||
return False, "Configure a server endpoint in Model Catalogue → Engines"
|
||||
try:
|
||||
normalize_openai_compat_asr_base_url(base_url)
|
||||
except ValueError as exc:
|
||||
return False, str(exc)
|
||||
try:
|
||||
import openai # noqa: F401
|
||||
except ImportError:
|
||||
@@ -2184,13 +2282,18 @@ class OpenAICompatASRBackend(ASRBackend):
|
||||
return True, "ready"
|
||||
|
||||
def _client(self):
|
||||
from openai import OpenAI
|
||||
from openai import DefaultHttpxClient, OpenAI
|
||||
api_key = resolve_openai_compat_asr_api_key() or "not-needed"
|
||||
# max_retries=0: mirrors llm_skills.resolve_skill_client — a
|
||||
# rate-limited/slow server retrying inside the SDK would blow past
|
||||
# whatever bounded timeout the caller (dub transcribe, dictation)
|
||||
# expects from a single call.
|
||||
return OpenAI(base_url=self._base_url, api_key=api_key, max_retries=0)
|
||||
return OpenAI(
|
||||
base_url=self._base_url,
|
||||
api_key=api_key,
|
||||
max_retries=0,
|
||||
http_client=DefaultHttpxClient(follow_redirects=False),
|
||||
)
|
||||
|
||||
def transcribe(self, audio_path: str, *, word_timestamps: bool = True) -> dict:
|
||||
logger.info(
|
||||
@@ -2942,6 +3045,31 @@ def get_sherpa_dictation_backend(model_id: str) -> "SherpaDictationBackend":
|
||||
return backend
|
||||
|
||||
|
||||
def sherpa_engine_model_id() -> str:
|
||||
"""The sherpa model the ``sherpa-onnx-asr`` engine loads when nothing pins
|
||||
one explicitly: env var (power-user pin) → the dictation model the user
|
||||
picked in Settings / the Engines menu → the catalogue default.
|
||||
|
||||
Unlike :func:`dictation_model_id` this ignores ``dictation.enabled`` — a
|
||||
user who turned the hotkey off but chose the Sherpa engine for dub/batch
|
||||
transcription still means *this* model — and never returns None: the
|
||||
engine needs *some* model to construct. A demoted model (decoded nothing
|
||||
on this host) falls through to the default rather than being re-picked.
|
||||
"""
|
||||
from services import sherpa_dictation as _sd
|
||||
explicit = os.environ.get("OMNIVOICE_SHERPA_ASR_MODEL")
|
||||
if explicit:
|
||||
return explicit
|
||||
try:
|
||||
from core import prefs
|
||||
mid = prefs.get("dictation.model_id")
|
||||
except Exception: # noqa: BLE001 — prefs store unavailable → default
|
||||
return _sd.DEFAULT_MODEL_ID
|
||||
if _sd.is_sherpa_model(mid) and not _sd.is_demoted(mid):
|
||||
return _sd.get_spec(mid).id
|
||||
return _sd.DEFAULT_MODEL_ID
|
||||
|
||||
|
||||
def dictation_model_id() -> str | None:
|
||||
"""The selected sherpa dictation model id, or None when dictation is off /
|
||||
no sherpa model is chosen. Env var wins (power-user pin), then prefs."""
|
||||
@@ -3216,9 +3344,7 @@ def _offline_asr_repo(backend_id: str | None = None) -> str | None:
|
||||
# Unknown/none → fail open.
|
||||
try:
|
||||
from services import sherpa_dictation as _sd
|
||||
spec = _sd.get_spec(
|
||||
os.environ.get("OMNIVOICE_SHERPA_ASR_MODEL", _sd.DEFAULT_MODEL_ID)
|
||||
)
|
||||
spec = _sd.get_spec(sherpa_engine_model_id())
|
||||
return spec.repo_id if spec is not None else None
|
||||
except Exception: # noqa: BLE001 — preflight must stay best-effort
|
||||
return None
|
||||
|
||||
@@ -65,7 +65,7 @@ _MODE_PREF = "hf_endpoint_mode" # "auto" | "manual"; absent → default
|
||||
_DECISION_PREF = "hf_endpoint_auto" # cached decision dict (see race())
|
||||
|
||||
DECISION_MAX_AGE_S = 7 * 24 * 3600.0 # re-race a decision older than 7 days
|
||||
PROBE_TIMEOUT_S = 3.0 # short: a probe is not a download
|
||||
PROBE_TIMEOUT_S = 8.0 # high-latency / China paths often need >3s
|
||||
MIRROR_SPEEDUP_FACTOR = 3.0 # mirror must be ≥3× faster to win
|
||||
|
||||
# Small, stable, long-lived public file for the optional ranged-GET
|
||||
|
||||
@@ -33,10 +33,20 @@ _ESTIMATES: dict[str, dict] = {
|
||||
"destination": "hf_model_cache",
|
||||
"deduplication": None,
|
||||
},
|
||||
"audiocpp": {
|
||||
"package_download_bytes": None,
|
||||
"unique_installed_bytes": None,
|
||||
"potentially_shared_bytes": None,
|
||||
"temporary_free_bytes": None,
|
||||
"confidence": "estimated",
|
||||
"destination": "hf_model_cache",
|
||||
"deduplication": None,
|
||||
},
|
||||
}
|
||||
_MODEL_REPOS = {
|
||||
"omnivoice": "k2-fsa/OmniVoice",
|
||||
"kittentts": "KittenML/kitten-tts-mini-0.8",
|
||||
"audiocpp": "audio-cpp/audio.cpp-gguf",
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -88,14 +88,34 @@ def snapshot(
|
||||
evidence_state = "loaded"
|
||||
if isolated and provider is None and actual_device is None:
|
||||
evidence_state = "subprocess_loaded_provider_unreported"
|
||||
from core.scrub import scrub_text
|
||||
|
||||
runtime_device_name = routing.get("runtime_device_name")
|
||||
device_name = (
|
||||
getattr(caps, "device_name", "")
|
||||
if runtime_device_name is None
|
||||
else runtime_device_name
|
||||
)
|
||||
public_device_name = scrub_text(device_name)[:256]
|
||||
return {
|
||||
"implementation_variant": f"{engine_cls.__module__}.{engine_cls.__name__}",
|
||||
"declared_device_families": list(getattr(engine_cls, "gpu_compat", ("cpu",))),
|
||||
"declared_device_families": list(
|
||||
routing.get("gpu_compat", getattr(engine_cls, "gpu_compat", ("cpu",)))
|
||||
),
|
||||
"evidence_state": evidence_state,
|
||||
"actual_execution_provider": provider,
|
||||
"actual_execution_device": actual_device,
|
||||
"gpu_name": getattr(caps, "device_name", "") or None,
|
||||
"gpu_architecture": _gpu_architecture(getattr(caps, "family", "cpu")),
|
||||
"gpu_name": public_device_name or None,
|
||||
"gpu_architecture": None
|
||||
if (
|
||||
routing.get("runtime_hardware_family")
|
||||
and not routing.get("runtime_device_verified")
|
||||
)
|
||||
else _gpu_architecture(
|
||||
routing.get("runtime_hardware_family")
|
||||
or getattr(caps, "family", "cpu")
|
||||
),
|
||||
"runtime_vram_gb": routing.get("runtime_vram_gb"),
|
||||
"precision_or_quantization": precision,
|
||||
"cpu_fallback_reason": runtime_fallback_reason or (routing.get("routing_reason") if fallback else None),
|
||||
"cpu_fallback_stage": runtime_fallback_stage or ("routing_preflight" if fallback else None),
|
||||
|
||||
@@ -14,6 +14,7 @@ carry a home path.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
from typing import Literal, TypedDict
|
||||
|
||||
from core.device_caps import (
|
||||
@@ -31,6 +32,93 @@ class RoutingResult(TypedDict):
|
||||
routing_reason: str | None # raw, pre-scrub
|
||||
|
||||
|
||||
def runtime_compute_profile(engine_or_cls, caps: HostCaps) -> dict:
|
||||
"""Return one engine's runtime-aware compute contract.
|
||||
|
||||
Native executables may discover providers independently of PyTorch. They
|
||||
override ``runtime_compute_profile``; all existing engines retain the
|
||||
exact static routing contract.
|
||||
"""
|
||||
hook = getattr(engine_or_cls, "runtime_compute_profile", None)
|
||||
if callable(hook):
|
||||
return hook(caps)
|
||||
cls = engine_or_cls if isinstance(engine_or_cls, type) else type(engine_or_cls)
|
||||
compat = tuple(getattr(cls, "gpu_compat", ("cpu",)))
|
||||
floor = float(getattr(cls, "min_vram_gb", 0.0) or 0.0)
|
||||
return {
|
||||
"gpu_compat": compat,
|
||||
"min_vram_gb": floor,
|
||||
**resolve_routing(compat, caps, floor),
|
||||
"runtime_backend": None,
|
||||
"runtime_device_index": None,
|
||||
"runtime_device_name": None,
|
||||
"runtime_hardware_family": None,
|
||||
"runtime_vram_gb": None,
|
||||
"runtime_device_verified": None,
|
||||
}
|
||||
|
||||
|
||||
async def runtime_compute_profile_async(engine_or_cls, caps: HostCaps) -> dict:
|
||||
"""Resolve runtime compute metadata without blocking the event loop."""
|
||||
return await asyncio.to_thread(runtime_compute_profile, engine_or_cls, caps)
|
||||
|
||||
|
||||
def under_provisioned_vram(
|
||||
caps: HostCaps,
|
||||
min_vram_gb: float = 0.0,
|
||||
*,
|
||||
family: str | None = None,
|
||||
vram_gb: float | None = None,
|
||||
) -> bool:
|
||||
"""Is this host's DEDICATED VRAM below the engine's declared floor?
|
||||
|
||||
The one definition of "under-provisioned", shared by everything that acts
|
||||
on the verdict: the routing caveat below, the timeout guidance, and — since
|
||||
#1804 — the compute-time budget itself (``model_manager
|
||||
.generate_timeout_s``). It was written out inline in each of them, which is
|
||||
how the budget came to disagree with the warning printed next to it.
|
||||
|
||||
Dedicated-VRAM families ONLY. CUDA, ROCm, XPU, and a native Vulkan device
|
||||
report dedicated memory. On MPS, ``HostCaps.vram_gb`` is a heuristic
|
||||
(system RAM / 2, see device_caps) for a UNIFIED memory pool; comparing it
|
||||
against a floor measured on discrete CUDA hardware would tell every 8 GB Mac
|
||||
its 4 GB "VRAM" is too small for an engine that runs fine there. A VRAM
|
||||
figure of 0 means the probe failed — don't guess from it. A floor of 0 means
|
||||
the engine declares none, and inventing one is worse than staying quiet.
|
||||
"""
|
||||
if not min_vram_gb or min_vram_gb <= 0:
|
||||
return False
|
||||
if (family or getattr(caps, "family", None)) not in (
|
||||
"cuda", "rocm", "xpu", "vulkan",
|
||||
):
|
||||
return False
|
||||
raw_vram_gb = getattr(caps, "vram_gb", 0.0) if vram_gb is None else vram_gb
|
||||
available_vram_gb = float(raw_vram_gb or 0.0)
|
||||
return 0 < available_vram_gb < float(min_vram_gb)
|
||||
|
||||
|
||||
def low_vram_caveat(
|
||||
caps: HostCaps,
|
||||
min_vram_gb: float = 0.0,
|
||||
*,
|
||||
family: str | None = None,
|
||||
vram_gb: float | None = None,
|
||||
) -> str | None:
|
||||
"""User-facing advisory for a known under-provisioned dedicated GPU."""
|
||||
if not under_provisioned_vram(
|
||||
caps, min_vram_gb, family=family, vram_gb=vram_gb,
|
||||
):
|
||||
return None
|
||||
device = caps.device_name or (family or caps.family).upper()
|
||||
available_vram_gb = caps.vram_gb if vram_gb is None else vram_gb
|
||||
return (
|
||||
f"{device} has {available_vram_gb:.1f} GB VRAM; this engine wants about "
|
||||
f"{min_vram_gb:.0f} GB. It will run, but expect slow generations "
|
||||
f"that may time out. Unload other models before generating, keep "
|
||||
f"the text short, or pick a lighter engine."
|
||||
)
|
||||
|
||||
|
||||
def _caveat(caps: HostCaps, min_vram_gb: float = 0.0) -> str | None:
|
||||
"""A caveat string for an otherwise-accelerated host, or None.
|
||||
|
||||
@@ -51,24 +139,7 @@ def _caveat(caps: HostCaps, min_vram_gb: float = 0.0) -> str | None:
|
||||
for note in caps.notes:
|
||||
if KERNEL_RISK_MARKER in note:
|
||||
return f"{caps.family.upper()} selected, but: {note}"
|
||||
# Dedicated-VRAM families ONLY. On MPS, HostCaps.vram_gb is a heuristic
|
||||
# (system RAM / 2, see device_caps) for a UNIFIED memory pool — comparing
|
||||
# it against a floor measured on discrete CUDA hardware would tell every
|
||||
# 8 GB Mac its 4 GB "VRAM" is too small for an engine that runs fine there.
|
||||
# Different memory model, different (unmeasured) floor; don't guess.
|
||||
if (
|
||||
caps.family in ("cuda", "rocm")
|
||||
and min_vram_gb > 0
|
||||
and 0 < caps.vram_gb < min_vram_gb
|
||||
):
|
||||
device = caps.device_name or caps.family.upper()
|
||||
return (
|
||||
f"{device} has {caps.vram_gb:.1f} GB VRAM; this engine wants about "
|
||||
f"{min_vram_gb:.0f} GB. It will run, but expect slow generations "
|
||||
f"that may time out. Unload other models before generating, keep "
|
||||
f"the text short, or pick a lighter engine."
|
||||
)
|
||||
return None
|
||||
return low_vram_caveat(caps, min_vram_gb)
|
||||
|
||||
|
||||
def resolve_routing(
|
||||
@@ -208,5 +279,7 @@ def routing_fields(
|
||||
|
||||
__all__ = [
|
||||
"RoutingStatus", "RoutingResult", "resolve_routing", "routing_fields",
|
||||
"routing_notice", "header_safe_reason",
|
||||
"routing_notice", "header_safe_reason", "low_vram_caveat",
|
||||
"runtime_compute_profile", "runtime_compute_profile_async",
|
||||
"under_provisioned_vram",
|
||||
]
|
||||
|
||||
@@ -15,6 +15,7 @@ _SHA = re.compile(r"[0-9a-f]{40}\Z")
|
||||
CURATED_REVISIONS: dict[str, str] = {
|
||||
"facebook/nllb-200-distilled-600M": "f8d333a098d19b4fd9a8b18f94170487ad3f821d",
|
||||
"k2-fsa/OmniVoice": "c5fdb5ccb189668d56333f77ba2629f4cd7535f4",
|
||||
"audio-cpp/audio.cpp-gguf": "dc6fecccc2b0c6bdda0a8b2f38fa61394fee0b9c",
|
||||
"Systran/faster-whisper-large-v3": "edaa852ec7e145841d8ffdb056a99866b5f0a478",
|
||||
"mlx-community/whisper-large-v3-mlx": "49e6aa286ad60c14352c404340ded53710378a11",
|
||||
"mlx-community/whisper-large-v3-turbo": "a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb",
|
||||
|
||||
@@ -0,0 +1,220 @@
|
||||
"""Karaoke (word-highlight) ASS builder for dub hardsub export.
|
||||
|
||||
Pure text-in/text-out: no ffmpeg, no models, no filesystem. ``build_ass``
|
||||
turns subtitle cues into an ASS script whose lines carry ``\\k``/``\\kf``
|
||||
karaoke tags, so ffmpeg's ``ass=`` filter burns a word-by-word highlight
|
||||
sweep instead of the static line the SRT path renders.
|
||||
|
||||
Word timing sources, in order:
|
||||
|
||||
1. ``cue["words"]`` — per-word ``{text, start, end}`` persisted at
|
||||
transcribe time (services.segmentation). Used only when the words still
|
||||
spell the cue's display text: after translation the persisted ASR words
|
||||
are source-language tokens, so re-using their timing would burn the
|
||||
wrong language. The display text is always authoritative.
|
||||
2. Even split — the cue text's whitespace tokens spread uniformly across
|
||||
``[start, end]``. This is the compatibility path for jobs transcribed
|
||||
before word persistence and for translated tracks.
|
||||
|
||||
Dual-layout karaoke is intentionally unsupported (out of scope): callers
|
||||
must fall back to the line (SRT) burn when the dual layout is requested.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
from typing import Optional, Sequence
|
||||
|
||||
_WS = re.compile(r"\s+")
|
||||
|
||||
#: Default ASS canvas. libass scales the script to the real video size, so
|
||||
#: one reference resolution keeps font/margin proportions stable everywhere.
|
||||
DEFAULT_PLAY_RES = (1920, 1080)
|
||||
|
||||
_HEADER_TEMPLATE = """[Script Info]
|
||||
; Generated by VoiceStudio karaoke burn-in
|
||||
ScriptType: v4.00+
|
||||
PlayResX: {res_x}
|
||||
PlayResY: {res_y}
|
||||
WrapStyle: 0
|
||||
ScaledBorderAndShadow: yes
|
||||
|
||||
[V4+ Styles]
|
||||
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
|
||||
Style: Default,Arial,64,&H0000E7FF,&H00FFFFFF,&H00101010,&H7F000000,0,0,0,0,100,100,0,0,1,3,1,2,96,96,48,1
|
||||
|
||||
[Events]
|
||||
Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
|
||||
"""
|
||||
|
||||
|
||||
def _norm(text: object) -> str:
|
||||
return _WS.sub(" ", str(text or "").strip())
|
||||
|
||||
|
||||
def _ass_time(seconds: float) -> str:
|
||||
"""``H:MM:SS.CC`` (centiseconds) — the ASS event timestamp format."""
|
||||
cs = max(0, int(round(float(seconds) * 100)))
|
||||
h, rem = divmod(cs, 360000)
|
||||
m, rem = divmod(rem, 6000)
|
||||
s, c = divmod(rem, 100)
|
||||
return f"{h}:{m:02d}:{s:02d}.{c:02d}"
|
||||
|
||||
|
||||
def _ass_escape(text: str) -> str:
|
||||
"""Escape a display token for an ASS Dialogue text field.
|
||||
|
||||
Braces would open an override block (user text like ``{\\b1}`` must render
|
||||
literally, never execute); newlines become ASS hard line breaks.
|
||||
"""
|
||||
return (
|
||||
str(text)
|
||||
.replace("{", "\\{")
|
||||
.replace("}", "\\}")
|
||||
.replace("\r\n", "\\N")
|
||||
.replace("\n", "\\N")
|
||||
.replace("\r", "\\N")
|
||||
)
|
||||
|
||||
|
||||
def _cs(seconds: float) -> int:
|
||||
"""Karaoke tag duration in centiseconds; ≥1 so a tag never renders as 0."""
|
||||
return max(1, int(round(float(seconds) * 100)))
|
||||
|
||||
|
||||
def even_split_words(text: str, start: float, end: float) -> list[dict]:
|
||||
"""Uniformly distribute the cue text's whitespace tokens over [start, end].
|
||||
|
||||
The export fallback for jobs transcribed before per-word persistence and
|
||||
for translated tracks (whose persisted words are source-language tokens).
|
||||
"""
|
||||
tokens = [tok for tok in _WS.split(str(text or "").strip()) if tok]
|
||||
if not tokens:
|
||||
return []
|
||||
start = float(start)
|
||||
dur = max(0.0, float(end) - start) / len(tokens)
|
||||
return [
|
||||
{"text": tok, "start": start + i * dur, "end": start + (i + 1) * dur}
|
||||
for i, tok in enumerate(tokens)
|
||||
]
|
||||
|
||||
|
||||
def scale_words(
|
||||
words: Sequence[dict],
|
||||
orig_start: float,
|
||||
orig_end: float,
|
||||
new_start: float,
|
||||
new_end: float,
|
||||
) -> Optional[list[dict]]:
|
||||
"""Map word times linearly from [orig_start, orig_end] → [new_start, new_end].
|
||||
|
||||
Used when Smart Fit moves a cue onto the fitted timeline: the persisted
|
||||
word times live on the original timeline and must ride along. Returns
|
||||
``None`` when either span is degenerate (caller should drop the words so
|
||||
export falls back to an even split over the new span).
|
||||
"""
|
||||
orig_span = float(orig_end) - float(orig_start)
|
||||
new_span = float(new_end) - float(new_start)
|
||||
if orig_span <= 0 or new_span <= 0:
|
||||
return None
|
||||
ratio = new_span / orig_span
|
||||
out: list[dict] = []
|
||||
for w in words:
|
||||
try:
|
||||
ws = float(w["start"])
|
||||
we = float(w["end"])
|
||||
except (KeyError, TypeError, ValueError):
|
||||
return None
|
||||
out.append({
|
||||
**w,
|
||||
"start": round(float(new_start) + (ws - float(orig_start)) * ratio, 3),
|
||||
"end": round(float(new_start) + (we - float(orig_start)) * ratio, 3),
|
||||
})
|
||||
return out
|
||||
|
||||
|
||||
def _usable_words(cue: dict, text: str) -> Optional[list[tuple[str, float, float]]]:
|
||||
"""Persisted words, iff well-formed AND they spell the cue's display text."""
|
||||
words = cue.get("words")
|
||||
if not isinstance(words, list) or not words:
|
||||
return None
|
||||
clean: list[tuple[str, float, float]] = []
|
||||
for w in words:
|
||||
if not isinstance(w, dict):
|
||||
return None
|
||||
wtext = _norm(w.get("text"))
|
||||
try:
|
||||
ws = float(w["start"])
|
||||
we = float(w["end"])
|
||||
except (KeyError, TypeError, ValueError):
|
||||
return None
|
||||
if wtext:
|
||||
clean.append((wtext, ws, we))
|
||||
if not clean:
|
||||
return None
|
||||
if _norm(" ".join(t for t, _, _ in clean)) != text:
|
||||
return None
|
||||
return clean
|
||||
|
||||
|
||||
def _karaoke_text(cue: dict, text: str, start: float, end: float) -> str:
|
||||
"""One Dialogue text field: ``{\\k…}`` lead-in + per-word ``{\\kf…}`` tags.
|
||||
|
||||
Each word's sweep runs until the next word starts (the classic karaoke
|
||||
layout — inter-word gaps finish the previous word's fill), and the last
|
||||
word sweeps out to the cue end.
|
||||
"""
|
||||
words = _usable_words(cue, text) or [
|
||||
(w["text"], w["start"], w["end"]) for w in even_split_words(text, start, end)
|
||||
]
|
||||
# Clamp into the cue span and enforce monotonic starts so malformed
|
||||
# persisted data can only mistime the sweep, never corrupt the script.
|
||||
clamped: list[tuple[str, float]] = []
|
||||
prev = start
|
||||
for wtext, ws, _ in words:
|
||||
ws = min(max(ws, prev), end)
|
||||
clamped.append((wtext, ws))
|
||||
prev = ws
|
||||
parts: list[str] = []
|
||||
lead = clamped[0][1] - start
|
||||
if lead > 0.005:
|
||||
parts.append(f"{{\\k{_cs(lead)}}}")
|
||||
for i, (wtext, ws) in enumerate(clamped):
|
||||
nxt = clamped[i + 1][1] if i + 1 < len(clamped) else end
|
||||
sep = " " if i + 1 < len(clamped) else ""
|
||||
parts.append(f"{{\\kf{_cs(max(nxt, ws) - ws)}}}{_ass_escape(wtext)}{sep}")
|
||||
return "".join(parts)
|
||||
|
||||
|
||||
def build_ass(
|
||||
cues: Sequence[dict],
|
||||
*,
|
||||
dual: bool = False,
|
||||
play_res: tuple[int, int] = DEFAULT_PLAY_RES,
|
||||
) -> str:
|
||||
"""Build a karaoke ASS script from subtitle cues ({text, start, end, words?}).
|
||||
|
||||
One ``Default`` style; one Dialogue event per cue. ``dual`` exists for
|
||||
signature parity with the line burn but dual-layout karaoke is out of
|
||||
scope — callers must keep the SRT line burn for dual, so requesting it
|
||||
here is a contract violation, not a rendering mode.
|
||||
"""
|
||||
if dual:
|
||||
raise ValueError(
|
||||
"dual-layout karaoke is not supported; use the line (SRT) burn for dual subtitles"
|
||||
)
|
||||
res_x, res_y = play_res
|
||||
lines = [_HEADER_TEMPLATE.format(res_x=int(res_x), res_y=int(res_y))]
|
||||
for cue in cues or []:
|
||||
text = _norm(cue.get("text"))
|
||||
if not text:
|
||||
continue
|
||||
start = float(cue["start"])
|
||||
end = float(cue["end"])
|
||||
if end <= start:
|
||||
end = start + 0.1
|
||||
lines.append(
|
||||
f"Dialogue: 0,{_ass_time(start)},{_ass_time(end)},Default,,0,0,0,,"
|
||||
f"{_karaoke_text(cue, text, start, end)}"
|
||||
)
|
||||
return "\n".join(lines) + "\n"
|
||||
@@ -404,7 +404,23 @@ _CONFIGURED_GPU_JOB_TIMEOUT_S = GPU_JOB_TIMEOUT_S
|
||||
# CPU synthesis is healthy but substantially slower than accelerated inference.
|
||||
# Keep a separate, bounded floor so a short render on CPU is not abandoned at
|
||||
# the GPU-oriented five-minute deadline (#1588).
|
||||
#
|
||||
# #1787 review fix: an explicit OMNIVOICE_CPU_GENERATE_TIMEOUT_S must ALWAYS
|
||||
# govern CPU dispatches, even when OMNIVOICE_GENERATE_TIMEOUT_S is ALSO
|
||||
# explicit. Before this flag existed, `universal_override` below treated any
|
||||
# explicit GENERATE_TIMEOUT_S as authoritative for CPU too, so the Settings
|
||||
# panel's "CPU budget" row could be saved and silently never apply whenever
|
||||
# the "Accelerated" row was also set — the exact defect (a control that looks
|
||||
# like it works and doesn't) issue #1787 exists to remove. Setting ONLY
|
||||
# OMNIVOICE_GENERATE_TIMEOUT_S keeps its historical "universal" behavior
|
||||
# unchanged (test_explicit_universal_generate_timeout_wins_on_cpu) — nobody
|
||||
# who already relies on that single-var override loses it. The only case that
|
||||
# changes is the previously-undocumented, previously-broken combination of
|
||||
# setting BOTH: the more specific (CPU) value now wins for CPU jobs, matching
|
||||
# what a user who filled in both Settings rows was told would happen.
|
||||
_CPU_GENERATE_TIMEOUT_EXPLICIT = "OMNIVOICE_CPU_GENERATE_TIMEOUT_S" in os.environ
|
||||
CPU_JOB_TIMEOUT_S = float(os.environ.get("OMNIVOICE_CPU_GENERATE_TIMEOUT_S", "600.0"))
|
||||
_CONFIGURED_CPU_JOB_TIMEOUT_S = CPU_JOB_TIMEOUT_S
|
||||
|
||||
# Queue-wait budget — a SEPARATE, deliberately generous clock (#1190/#1202).
|
||||
# The execution bound above must never be spent waiting in line: a job queued
|
||||
@@ -509,6 +525,8 @@ class GpuPoolBusyError(TimeoutError):
|
||||
|
||||
def generate_timeout_s(
|
||||
text: "str | None", *, engine: object = None, execution_device: "str | None" = None,
|
||||
min_vram_gb: float = 0.0, hardware_family: "str | None" = None,
|
||||
vram_gb: "float | None" = None,
|
||||
) -> float:
|
||||
"""THE wall-clock execution budget for one synthesis job, scaled to input.
|
||||
|
||||
@@ -520,33 +538,74 @@ def generate_timeout_s(
|
||||
on long inputs. Lives here (not in a router) so every router shares it
|
||||
without importing generation.py.
|
||||
|
||||
Policy: floor at the configured OMNIVOICE_GENERATE_TIMEOUT_S, plus 1s per
|
||||
40 characters past a 1200-character free allowance — generous enough for
|
||||
Policy: floor at the configured OMNIVOICE_GENERATE_TIMEOUT_S (accelerated
|
||||
hosts) or OMNIVOICE_CPU_GENERATE_TIMEOUT_S (CPU hosts — the latter wins
|
||||
for CPU whenever it is itself explicit, even if the former also is; see
|
||||
the #1787 comment on the module-level constants), plus 1s per 40
|
||||
characters past a 1200-character free allowance — generous enough for
|
||||
CPU-class hardware, still bounded (a wedged job is caught in minutes, not
|
||||
hours).
|
||||
|
||||
#1804: "accelerated" is not one performance class. A card with less VRAM
|
||||
than the engine declares it needs pages to system RAM over PCIe and renders
|
||||
SLOWER than the same machine's CPU would — yet, judged by device family
|
||||
alone, it was handed HALF the CPU budget. That inversion is what three 4 GB
|
||||
reporters hit (#1226 GTX 1650 Ti, #1222 Quadro P2000, #1804 GTX 1650), all
|
||||
on the engine that declares a 6 GB floor. Every layer already knew: routing
|
||||
raises a caveat, the preflight toast warns, and the timeout message names
|
||||
the card. Only the budget ignored it. So an under-provisioned accelerator
|
||||
now floors at the CPU budget — the class of hardware it actually performs
|
||||
like. ``min_vram_gb`` is the engine's declared floor; callers that pass
|
||||
``engine`` get it read off the engine automatically. Native runtimes pass
|
||||
an explicit ``vram_gb=0`` when their dedicated-memory probe failed; that
|
||||
unknown capacity gets the same conservative CPU-class budget without
|
||||
claiming the card is under-provisioned in user-facing diagnostics.
|
||||
"""
|
||||
base = GPU_JOB_TIMEOUT_S
|
||||
try:
|
||||
from core.device_caps import detect_host_caps
|
||||
family = execution_device or detect_host_caps().family
|
||||
caps = detect_host_caps()
|
||||
family = execution_device or caps.family
|
||||
if not min_vram_gb and engine is not None:
|
||||
min_vram_gb = float(getattr(engine, "min_vram_gb", 0.0) or 0.0)
|
||||
if execution_device is None and engine is not None:
|
||||
from services.engine_routing import resolve_routing
|
||||
compat = getattr(engine, "gpu_compat", None)
|
||||
if compat is None:
|
||||
compat = getattr(type(engine), "gpu_compat", (family, "cpu"))
|
||||
if tuple(compat) == ("cpu",):
|
||||
family = "cpu"
|
||||
else:
|
||||
family = resolve_routing(
|
||||
compat, detect_host_caps(),
|
||||
float(getattr(engine, "min_vram_gb", 0.0) or 0.0),
|
||||
)["effective_device"]
|
||||
from services.engine_routing import runtime_compute_profile
|
||||
profile = runtime_compute_profile(engine, caps)
|
||||
family = profile["effective_device"]
|
||||
min_vram_gb = profile["min_vram_gb"]
|
||||
hardware_family = profile.get("runtime_hardware_family")
|
||||
vram_gb = profile.get("runtime_vram_gb")
|
||||
universal_override = (
|
||||
_GENERATE_TIMEOUT_EXPLICIT
|
||||
or GPU_JOB_TIMEOUT_S != _CONFIGURED_GPU_JOB_TIMEOUT_S
|
||||
)
|
||||
if family == "cpu" and not universal_override:
|
||||
# An explicit (env-set, or runtime-changed the same way tests do)
|
||||
# CPU budget is more specific than the universal override and always
|
||||
# wins for CPU dispatches — see the #1787 comment above.
|
||||
cpu_explicit = (
|
||||
_CPU_GENERATE_TIMEOUT_EXPLICIT
|
||||
or CPU_JOB_TIMEOUT_S != _CONFIGURED_CPU_JOB_TIMEOUT_S
|
||||
)
|
||||
if family == "cpu" and (cpu_explicit or not universal_override):
|
||||
base = CPU_JOB_TIMEOUT_S
|
||||
elif not universal_override and family in (
|
||||
"cuda", "rocm", "vulkan", "xpu",
|
||||
):
|
||||
from services.engine_routing import under_provisioned_vram
|
||||
|
||||
runtime_family = hardware_family or family
|
||||
unknown_dedicated_vram = (
|
||||
min_vram_gb > 0
|
||||
and runtime_family in ("cuda", "rocm", "xpu", "vulkan")
|
||||
and vram_gb is not None
|
||||
and float(vram_gb or 0.0) <= 0
|
||||
)
|
||||
if unknown_dedicated_vram or under_provisioned_vram(
|
||||
caps, min_vram_gb, family=hardware_family, vram_gb=vram_gb,
|
||||
):
|
||||
# `max`, never a plain assignment: an operator who raised the
|
||||
# accelerated budget above the CPU one must not have it cut.
|
||||
base = max(base, CPU_JOB_TIMEOUT_S)
|
||||
except Exception:
|
||||
# Device probing is advisory here; the configured universal bound is
|
||||
# still safe when a platform probe is unavailable during startup.
|
||||
@@ -1050,6 +1109,7 @@ def _timeout_guidance(
|
||||
"""
|
||||
family = "cuda" # conservative default: GPU wording if the probe fails
|
||||
device_name, vram_gb = "", 0.0
|
||||
_caps = None # a failed probe stays None; under_provisioned_vram() reads it safely
|
||||
try:
|
||||
from core.device_caps import detect_host_caps
|
||||
_caps = detect_host_caps()
|
||||
@@ -1085,7 +1145,7 @@ def _timeout_guidance(
|
||||
"compute-bound. For a durable fix try shorter text or a lighter "
|
||||
"engine (OmniVoice GGUF and Supertonic-3 are CPU-tuned). If you "
|
||||
"expect very long single generations, raise "
|
||||
"OMNIVOICE_GENERATE_TIMEOUT_S."
|
||||
"the compute-time budget in Settings → Performance & Device."
|
||||
)
|
||||
# #1226/#1222: two users on 4 GB cards were told, generically, that the GPU
|
||||
# "is VRAM-starved" — true, but it read as a transient contention problem
|
||||
@@ -1097,11 +1157,9 @@ def _timeout_guidance(
|
||||
# a threshold applied without knowing whose job it is would confidently
|
||||
# misdiagnose most of them. And on MPS `vram_gb` is a unified-memory
|
||||
# heuristic (RAM/2), not a dedicated pool to compare against.
|
||||
if (
|
||||
min_vram_gb > 0
|
||||
and family in ("cuda", "rocm")
|
||||
and 0 < vram_gb < min_vram_gb
|
||||
):
|
||||
from services.engine_routing import under_provisioned_vram
|
||||
|
||||
if under_provisioned_vram(_caps, min_vram_gb):
|
||||
return common + (
|
||||
f"{device_name or 'this GPU'} has {vram_gb:.1f} GB of VRAM and "
|
||||
f"this engine wants about {min_vram_gb:.0f} GB — generations here "
|
||||
@@ -1110,7 +1168,8 @@ def _timeout_guidance(
|
||||
f"Supertonic-3 are tuned for small/no GPU) or shorter text; "
|
||||
f"Flush caches / Unload the resident model (top toolbar or "
|
||||
f"Model Catalogue → Models) frees what little headroom there is. (Raise "
|
||||
f"OMNIVOICE_GENERATE_TIMEOUT_S if you'd rather let long "
|
||||
f"the compute-time budget in Settings → Performance & Device if "
|
||||
f"you'd rather let long "
|
||||
f"generations run.)"
|
||||
)
|
||||
return common + (
|
||||
@@ -1118,7 +1177,8 @@ def _timeout_guidance(
|
||||
"contend for memory). For a durable fix, Flush caches / Unload the "
|
||||
"resident model (top toolbar or Model Catalogue → Models) before retrying, "
|
||||
"try shorter text, a lighter engine, or set the engine to CPU in "
|
||||
"Model Catalogue → Models. (Raise OMNIVOICE_GENERATE_TIMEOUT_S for very "
|
||||
"Model Catalogue → Models. (Raise the compute-time budget in "
|
||||
"Settings → Performance & Device for very "
|
||||
"long single generations.)"
|
||||
)
|
||||
|
||||
@@ -1480,14 +1540,17 @@ def get_best_device():
|
||||
# ── DirectML — universal Windows GPU (probe reports this as "cpu") ─
|
||||
# Reached only when no torch family was detected (family == "cpu"), which is
|
||||
# exactly the DirectML case — the probe classifies DirectML hosts as cpu.
|
||||
try:
|
||||
import torch_directml
|
||||
if torch_directml.device_count() > 0:
|
||||
logger.info("Using DirectML device (GPU %d)", 0)
|
||||
return str(torch_directml.device(0))
|
||||
except ImportError:
|
||||
pass
|
||||
if family == "cpu":
|
||||
try:
|
||||
import torch_directml
|
||||
if torch_directml.device_count() > 0:
|
||||
logger.info("Using DirectML device (GPU %d)", 0)
|
||||
return str(torch_directml.device(0))
|
||||
except ImportError:
|
||||
# DirectML is optional; an absent package leaves CPU available.
|
||||
pass
|
||||
|
||||
# Other families need an explicitly compatible loader (e.g. NPU sidecars).
|
||||
return "cpu"
|
||||
|
||||
_COMPILE_ERR_MODULE_PREFIXES = ("torch._dynamo", "torch._inductor", "torch.fx", "triton")
|
||||
|
||||
@@ -222,6 +222,46 @@ def entries_for_language(entries, language: Optional[str]) -> dict[str, str]:
|
||||
return merged
|
||||
|
||||
|
||||
def inert_entries_for_language(entries, language: str | None) -> list[dict]:
|
||||
"""Enabled entries that MATCH the language but cannot be applied yet.
|
||||
|
||||
Settings offers three notations — Respelling, IPA, CMU — and only
|
||||
respelling substitutes text today. IPA and CMU rows save cleanly, are
|
||||
validated, get a badge and can be toggled on, and are then dropped before
|
||||
term matching. Nothing downstream reads them.
|
||||
|
||||
That is Phase 1 behaving as designed; the gap is that it is INVISIBLE.
|
||||
"Test a sentence" reported "No entries match; spoken as written" for a
|
||||
term that does match, which is not a degraded answer but a wrong one, and
|
||||
it sent the user off to re-type an entry that was already correct (#1949).
|
||||
|
||||
docs/specs/01-expressive-tts.md asked for exactly the opposite — such
|
||||
entries "passed through and flagged 'phoneme not honored on this
|
||||
engine' (parity-rule: visible degradation)". This is that flag: the
|
||||
caller can now say WHY nothing happened instead of implying nothing
|
||||
matched.
|
||||
"""
|
||||
req_prefix = _lang_prefix(language)
|
||||
out: list[dict] = []
|
||||
for e in entries or []:
|
||||
try:
|
||||
if not int(e["enabled"]):
|
||||
continue
|
||||
except (KeyError, IndexError, TypeError, ValueError):
|
||||
continue
|
||||
term = (e["term"] or "").strip()
|
||||
if not term:
|
||||
continue
|
||||
etype = (e["type"] or "respelling").strip().lower()
|
||||
if etype == "respelling":
|
||||
continue
|
||||
scope = (e["language"] or _ALL_LANG).strip() or _ALL_LANG
|
||||
if scope != _ALL_LANG and (req_prefix is None or scope[:2].lower() != req_prefix):
|
||||
continue
|
||||
out.append({"term": term, "type": etype})
|
||||
return out
|
||||
|
||||
|
||||
# ── Inline one-off override: [[term|replacement]] / [[replacement]] ─────────
|
||||
#
|
||||
# Double brackets are unambiguous against the single-bracket grammar
|
||||
|
||||
@@ -100,10 +100,33 @@ class Segment:
|
||||
}
|
||||
|
||||
|
||||
def _serialize_words(words: Sequence[Word]) -> list[dict]:
|
||||
"""Word objects → the ``{text, start, end}`` dicts persisted on segments.
|
||||
|
||||
Per-word timing is kept on each segment (``Segment.extra["words"]``, so
|
||||
``to_dict`` carries it onto the job) to drive the karaoke hardsub export.
|
||||
"""
|
||||
return [
|
||||
{"text": w.text, "start": round(w.start, 3), "end": round(w.end, 3)}
|
||||
for w in words
|
||||
]
|
||||
|
||||
|
||||
def _merge_segment_extra(target: Segment, incoming: Segment, *, prepend: bool) -> None:
|
||||
"""Preserve editor metadata when cleanup folds ``incoming`` into ``target``."""
|
||||
# Word lists must CONCATENATE in text order (the setdefault below would
|
||||
# otherwise adopt the incoming list wholesale when the target has none,
|
||||
# then double it). Capture both sides before setdefault runs.
|
||||
raw_target_words = target.extra.get("words")
|
||||
raw_incoming_words = incoming.extra.get("words")
|
||||
for key, value in incoming.extra.items():
|
||||
target.extra.setdefault(key, value)
|
||||
target_words = raw_target_words if isinstance(raw_target_words, list) else []
|
||||
incoming_words = raw_incoming_words if isinstance(raw_incoming_words, list) else []
|
||||
if target_words or incoming_words:
|
||||
target.extra["words"] = (
|
||||
incoming_words + target_words if prepend else target_words + incoming_words
|
||||
)
|
||||
|
||||
def joined(left: object, right: object) -> str:
|
||||
return _clean(f"{left or ''} {right or ''}")
|
||||
@@ -233,7 +256,10 @@ def _build_segments_from_words(words: Sequence[Word]) -> List[Segment]:
|
||||
if not text:
|
||||
buf = []
|
||||
return
|
||||
segments.append(Segment(start=buf_start, end=buf[-1].end, text=text))
|
||||
segments.append(Segment(
|
||||
start=buf_start, end=buf[-1].end, text=text,
|
||||
extra={"words": _serialize_words(buf)},
|
||||
))
|
||||
buf = []
|
||||
if not force:
|
||||
buf_start = 0.0
|
||||
@@ -291,6 +317,7 @@ def _build_segments_from_words(words: Sequence[Word]) -> List[Segment]:
|
||||
start=buf_start,
|
||||
end=left_buf[-1].end,
|
||||
text=_clean(" ".join(x.text for x in left_buf)),
|
||||
extra={"words": _serialize_words(left_buf)},
|
||||
))
|
||||
buf = list(right_buf)
|
||||
buf_start = right_buf[0].start
|
||||
@@ -479,11 +506,23 @@ def _apply_scene_cuts(segments: List[Segment], scene_cuts: Iterable[float]) -> L
|
||||
or (remaining.end - cut) < MIN_DUR
|
||||
):
|
||||
continue
|
||||
# Segment text is the joined word texts, so a whitespace-boundary
|
||||
# text split maps exactly onto a word-count split of the list.
|
||||
words = remaining.extra.get("words")
|
||||
left_extra: dict = {}
|
||||
right_extra: dict = {}
|
||||
if isinstance(words, list) and words:
|
||||
n_left = len(left_text.split())
|
||||
if n_left and len(words) > n_left:
|
||||
left_extra = {"words": words[:n_left]}
|
||||
right_extra = {"words": words[n_left:]}
|
||||
out.append(Segment(
|
||||
start=remaining.start, end=cut, text=left_text, speaker_id=remaining.speaker_id,
|
||||
extra=left_extra,
|
||||
))
|
||||
remaining = Segment(
|
||||
start=cut, end=remaining.end, text=right_text, speaker_id=remaining.speaker_id,
|
||||
extra=right_extra,
|
||||
)
|
||||
out.append(remaining)
|
||||
return out
|
||||
@@ -715,6 +754,10 @@ def _resplit_core(
|
||||
piece["text"] = text
|
||||
piece["start"] = s0 if k == 0 else ws[0].start
|
||||
piece["end"] = s1 if k == n_runs - 1 else ws[-1].end
|
||||
# dict(seg) copied the WHOLE segment's word list into every piece;
|
||||
# each piece keeps only its own run's words (karaoke burn-in).
|
||||
if "words" in piece:
|
||||
piece["words"] = _serialize_words(ws)
|
||||
if label:
|
||||
piece["speaker_id"] = label
|
||||
if piece_no > 0:
|
||||
|
||||
@@ -29,6 +29,14 @@ import httpx
|
||||
_HF_AUTH_HOSTS = ("huggingface.co", "hf.co")
|
||||
_DEFAULT_CONNECTIONS = 8
|
||||
_MIN_SEGMENT_BYTES = 4 * 1024 * 1024 # don't split below this — overhead > gain
|
||||
# Cap on a single segment. Progress is committed to the manifest only when a
|
||||
# whole segment lands, so the segment size is also the MOST bytes a dropped
|
||||
# connection can throw away. Sizing segments as size/num_connections made that
|
||||
# ~100 MB on an 800 MB blob: on a link that drops every ~50 MB no segment ever
|
||||
# completed, the manifest was never written, and every retry restarted from
|
||||
# zero (#1224 follow-up). Bounded segments turn the same flaky link into steady
|
||||
# forward progress.
|
||||
_MAX_SEGMENT_BYTES = 16 * 1024 * 1024
|
||||
_READ_CHUNK = 1024 * 1024
|
||||
|
||||
|
||||
@@ -67,8 +75,16 @@ async def _resolve(client: httpx.AsyncClient, url: str, token: Optional[str], ma
|
||||
|
||||
|
||||
def _plan_segments(size: int, num_connections: int) -> list[tuple[int, int]]:
|
||||
"""Byte ranges to fetch, each at most ``_MAX_SEGMENT_BYTES``.
|
||||
|
||||
``num_connections`` controls how many run at once (see the semaphore in
|
||||
:func:`segmented_download`), NOT how many segments exist — a large file is
|
||||
split into many bounded segments so each one commits to the manifest
|
||||
quickly and a dropped connection costs at most one segment.
|
||||
"""
|
||||
n = max(1, min(num_connections, max(1, size // _MIN_SEGMENT_BYTES)))
|
||||
step = -(-size // n) # ceil
|
||||
step = max(_MIN_SEGMENT_BYTES, min(step, _MAX_SEGMENT_BYTES))
|
||||
segs = []
|
||||
start = 0
|
||||
while start < size:
|
||||
@@ -143,6 +159,10 @@ async def segmented_download(
|
||||
_preallocate(part, size)
|
||||
segments = [s for s in _plan_segments(size, num_connections) if s not in done]
|
||||
lock = asyncio.Lock()
|
||||
# Segments are bounded, so a big file yields many more of them than
|
||||
# there are connections. The semaphore — not the segment count — is
|
||||
# what keeps concurrency at num_connections.
|
||||
sem = asyncio.Semaphore(max(1, num_connections))
|
||||
|
||||
async def _fetch(seg: tuple[int, int]):
|
||||
start, end = seg
|
||||
@@ -169,8 +189,12 @@ async def segmented_download(
|
||||
done.add(seg)
|
||||
_save_done(part, size, done)
|
||||
|
||||
async def _fetch_limited(seg: tuple[int, int]):
|
||||
async with sem:
|
||||
await _fetch(seg)
|
||||
|
||||
if segments:
|
||||
await asyncio.gather(*(_fetch(s) for s in segments))
|
||||
await asyncio.gather(*(_fetch_limited(s) for s in segments))
|
||||
|
||||
# ── verify ──────────────────────────────────────────────────────
|
||||
actual = os.path.getsize(part)
|
||||
|
||||
@@ -60,7 +60,7 @@ from pathlib import Path
|
||||
from typing import Callable, Optional
|
||||
|
||||
from core.config import DATA_DIR
|
||||
from core.contained_subprocess import OwnedPopen, spawn_owned
|
||||
from core.contained_subprocess import OwnedPopen, WindowsJobPopen, spawn_owned
|
||||
|
||||
logger = logging.getLogger("omnivoice.sidecar_install")
|
||||
|
||||
@@ -126,6 +126,33 @@ class SidecarSpec:
|
||||
invalidate: Callable[[], None] = field(default=lambda: None)
|
||||
# Cheap "is a healthy install already present?" probe (file existence only).
|
||||
installed_probe: Callable[[], bool] = field(default=lambda: False)
|
||||
# Extra `uv venv` arguments — an interpreter pin for an upstream that
|
||||
# declares one, e.g. ("--python", "3.10").
|
||||
venv_args: tuple[str, ...] = ()
|
||||
# `uv pip install` target, "{checkout}" substituted. Each upstream installs
|
||||
# differently (editable, editable with an extra, a requirements file, a
|
||||
# constraints file); the default is the editable install IndexTTS uses.
|
||||
install_args: tuple[str, ...] = ("-e", "{checkout}")
|
||||
# Add PyTorch's CUDA index on a CUDA host. Plain PyPI torch is CPU-only on
|
||||
# Windows, and `+cuNNN` local-version pins exist nowhere else.
|
||||
uses_cuda_index: bool = False
|
||||
# Python that proves the venv works; "{checkout}" / "{checkout_repr}"
|
||||
# substituted. None means `import <probe_module>`.
|
||||
probe_code: Optional[str] = None
|
||||
# The file whose presence proves a fetched checkout is the whole
|
||||
# repository. Most upstreams ship a pyproject.toml; Confucius4 ships
|
||||
# only requirements.txt and setup.py.
|
||||
source_manifest: str = "pyproject.toml"
|
||||
# False for an engine that is a PyPI package, not a repository: nothing
|
||||
# is fetched, and the managed root holds only the engine's own venv.
|
||||
has_source: bool = True
|
||||
# Add PyTorch's CPU index on every host, for an engine that only ever
|
||||
# runs torch on the CPU (see core.torch_indexes).
|
||||
cpu_torch_index: bool = False
|
||||
# Can the one-click install work on THIS machine? (ok, reason). Consulted
|
||||
# before an Install button is offered and again when an install starts, so
|
||||
# a host the upstream does not support never gets a job that can only fail.
|
||||
host_supported: Callable[[], tuple[bool, str]] = field(default=lambda: (True, ""))
|
||||
|
||||
|
||||
def _indextts_invalidate() -> None:
|
||||
@@ -138,6 +165,86 @@ def _indextts_installed() -> bool:
|
||||
return is_indextts_installed()
|
||||
|
||||
|
||||
def _moss_invalidate() -> None:
|
||||
from engines.moss_tts_v15 import bootstrap
|
||||
bootstrap.invalidate()
|
||||
|
||||
|
||||
def _moss_installed() -> bool:
|
||||
from engines.moss_tts_v15.bootstrap import is_moss_tts_v15_installed
|
||||
return is_moss_tts_v15_installed()
|
||||
|
||||
|
||||
def _confucius4_invalidate() -> None:
|
||||
from engines.confucius4 import bootstrap
|
||||
bootstrap.invalidate()
|
||||
|
||||
|
||||
def _confucius4_installed() -> bool:
|
||||
from engines.confucius4.bootstrap import is_confucius4_installed
|
||||
return is_confucius4_installed()
|
||||
|
||||
|
||||
def _dots_invalidate() -> None:
|
||||
from engines.dots_tts import bootstrap
|
||||
bootstrap.invalidate()
|
||||
|
||||
|
||||
def _dots_installed() -> bool:
|
||||
from engines.dots_tts.bootstrap import is_dots_tts_installed
|
||||
return is_dots_tts_installed()
|
||||
|
||||
|
||||
def _host_family() -> str:
|
||||
"""The accelerator family this host runs, or "cpu" when it cannot tell."""
|
||||
try:
|
||||
from core.device_caps import detect_host_caps
|
||||
return str(detect_host_caps().family)
|
||||
except Exception: # noqa: BLE001 — a probe failure must not break installs
|
||||
return "cpu"
|
||||
|
||||
|
||||
def _moss_host() -> tuple[bool, str]:
|
||||
if _host_family() == "cuda":
|
||||
return True, ""
|
||||
return False, (
|
||||
"MOSS-TTS-v1.5's one-click install uses its CUDA build of PyTorch, and "
|
||||
"this machine has no NVIDIA GPU available. Its guide covers a manual "
|
||||
"CPU install."
|
||||
)
|
||||
|
||||
|
||||
def _dots_host() -> tuple[bool, str]:
|
||||
if sys.platform != "win32":
|
||||
return True, ""
|
||||
return False, (
|
||||
"dots.tts publishes no Windows install. Run VoiceStudio on Linux or "
|
||||
"macOS, or under WSL2, to use it."
|
||||
)
|
||||
|
||||
|
||||
def _pockettts_host() -> tuple[bool, str]:
|
||||
import platform
|
||||
if sys.platform == "darwin" and platform.machine().lower() == "x86_64":
|
||||
return False, (
|
||||
"PocketTTS needs a PyTorch version that has no Intel Mac build."
|
||||
)
|
||||
return True, ""
|
||||
|
||||
|
||||
def _in_app_env(module: str) -> Callable[[], bool]:
|
||||
"""An install made with ``uv sync --extra`` lives in the app's own
|
||||
environment. It counts as installed, so the installer never provisions a
|
||||
second copy over one that works."""
|
||||
def probe() -> bool:
|
||||
import importlib.util
|
||||
try:
|
||||
return importlib.util.find_spec(module) is not None
|
||||
except (ImportError, ValueError):
|
||||
return False
|
||||
return probe
|
||||
|
||||
|
||||
SPECS: dict[str, SidecarSpec] = {
|
||||
"indextts2": SidecarSpec(
|
||||
engine_id="indextts2",
|
||||
@@ -170,9 +277,165 @@ SPECS: dict[str, SidecarSpec] = {
|
||||
invalidate=_indextts_invalidate,
|
||||
installed_probe=_indextts_installed,
|
||||
),
|
||||
# Pinned to the upstream commits current on 2026-09-10. Weights are not
|
||||
# fetched here: each engine downloads them into the shared HF cache on its
|
||||
# first synthesis, as its manual install always has.
|
||||
"moss-tts-v15": SidecarSpec(
|
||||
engine_id="moss-tts-v15",
|
||||
display_name="MOSS-TTS-v1.5",
|
||||
repo_url="https://github.com/OpenMOSS/MOSS-TTS.git",
|
||||
tarball_url=(
|
||||
"https://github.com/OpenMOSS/MOSS-TTS/archive/"
|
||||
"934d6826b084c46a0d033402174d5f8ac4ed2519.tar.gz"
|
||||
),
|
||||
checkout_dirname="MOSS-TTS",
|
||||
env_var="OMNIVOICE_MOSS_TTS_V15_DIR",
|
||||
probe_module="transformers",
|
||||
probe_code="import transformers, torch",
|
||||
source_revision="934d6826b084c46a0d033402174d5f8ac4ed2519",
|
||||
source_required_path="pyproject.toml",
|
||||
venv_args=("--python", "3.11"),
|
||||
install_args=("-e", "{checkout}[torch-runtime]"),
|
||||
uses_cuda_index=True,
|
||||
host_supported=_moss_host,
|
||||
docs_path="docs/engines/moss-tts-v15.md",
|
||||
# ~7 GB CUDA torch venv now, ~16 GB of weights on first synthesis.
|
||||
required_bytes=24 * _GIB,
|
||||
dependency_bytes=8 * _GIB,
|
||||
temporary_free_bytes=8 * _GIB,
|
||||
disk_confidence="estimated",
|
||||
invalidate=_moss_invalidate,
|
||||
installed_probe=_moss_installed,
|
||||
),
|
||||
"confucius4-tts": SidecarSpec(
|
||||
engine_id="confucius4-tts",
|
||||
display_name="Confucius4-TTS",
|
||||
repo_url="https://github.com/netease-youdao/Confucius4-TTS.git",
|
||||
tarball_url=(
|
||||
"https://github.com/netease-youdao/Confucius4-TTS/archive/"
|
||||
"4fb32c481302d8858c3aec6a1c2a8b4cea8894c0.tar.gz"
|
||||
),
|
||||
checkout_dirname="Confucius4-TTS",
|
||||
env_var="OMNIVOICE_CONFUCIUS4_TTS_DIR",
|
||||
probe_module="confuciustts",
|
||||
# Upstream is not pip-installable; the package resolves from the
|
||||
# checkout on sys.path, exactly as the engine's sidecar imports it.
|
||||
probe_code="import sys; sys.path.insert(0, {checkout_repr}); import confuciustts",
|
||||
source_revision="4fb32c481302d8858c3aec6a1c2a8b4cea8894c0",
|
||||
# No pyproject.toml upstream: requirements.txt is its manifest.
|
||||
source_manifest="requirements.txt",
|
||||
source_required_path="setup.py",
|
||||
venv_args=("--python", "3.10"),
|
||||
install_args=("-r", "{checkout}/requirements.txt"),
|
||||
# torch==2.7.0: CPU-only from PyPI on Windows; the CUDA index supplies
|
||||
# 2.7.0+cu128, which satisfies the same pin.
|
||||
uses_cuda_index=True,
|
||||
docs_path="docs/engines/confucius4-tts.md",
|
||||
# ~7 GB venv now, ~5 GB of weights on first synthesis.
|
||||
required_bytes=14 * _GIB,
|
||||
dependency_bytes=8 * _GIB,
|
||||
temporary_free_bytes=8 * _GIB,
|
||||
disk_confidence="estimated",
|
||||
invalidate=_confucius4_invalidate,
|
||||
installed_probe=_confucius4_installed,
|
||||
),
|
||||
"dots-tts": SidecarSpec(
|
||||
engine_id="dots-tts",
|
||||
display_name="dots.tts",
|
||||
repo_url="https://github.com/rednote-hilab/dots.tts.git",
|
||||
tarball_url=(
|
||||
"https://github.com/rednote-hilab/dots.tts/archive/"
|
||||
"32407a55228630475c48ecdb2c4e2c0f9c09e030.tar.gz"
|
||||
),
|
||||
checkout_dirname="dots.tts",
|
||||
env_var="OMNIVOICE_DOTS_TTS_DIR",
|
||||
probe_module="dots_tts.runtime",
|
||||
source_revision="32407a55228630475c48ecdb2c4e2c0f9c09e030",
|
||||
source_required_path="constraints/recommended.txt",
|
||||
# Upstream requires-python is >=3.10,<3.13.
|
||||
venv_args=("--python", "3.11"),
|
||||
install_args=("-e", "{checkout}", "-c", "{checkout}/constraints/recommended.txt"),
|
||||
host_supported=_dots_host,
|
||||
docs_path="docs/engines/dots-tts.md",
|
||||
# ~7 GB venv now, ~9 GB checkpoint on first synthesis.
|
||||
required_bytes=18 * _GIB,
|
||||
dependency_bytes=8 * _GIB,
|
||||
temporary_free_bytes=8 * _GIB,
|
||||
disk_confidence="estimated",
|
||||
invalidate=_dots_invalidate,
|
||||
installed_probe=_dots_installed,
|
||||
),
|
||||
# PyPI packages rather than repositories: nothing to clone, and the managed
|
||||
# root holds only the engine's own venv. The pins are the app's own
|
||||
# optional extras (a test ties the two together), so the engine runs the
|
||||
# same wheel whichever way it was installed.
|
||||
"supertonic3": SidecarSpec(
|
||||
engine_id="supertonic3",
|
||||
display_name="Supertonic-3",
|
||||
repo_url="",
|
||||
tarball_url="",
|
||||
checkout_dirname="supertonic3",
|
||||
env_var="OMNIVOICE_SUPERTONIC3_DIR",
|
||||
probe_module="supertonic",
|
||||
has_source=False,
|
||||
venv_args=("--python", "3.11"),
|
||||
install_args=("supertonic==1.3.1",),
|
||||
docs_path="docs/engines/supertonic3.md",
|
||||
# onnxruntime + numpy + huggingface_hub, no torch. The ~400 MB of
|
||||
# weights download on first synthesis into the shared HF cache.
|
||||
required_bytes=1 * _GIB,
|
||||
installed_probe=_in_app_env("supertonic"),
|
||||
),
|
||||
"pockettts": SidecarSpec(
|
||||
engine_id="pockettts",
|
||||
display_name="PocketTTS",
|
||||
repo_url="",
|
||||
tarball_url="",
|
||||
checkout_dirname="pockettts",
|
||||
env_var="OMNIVOICE_POCKETTTS_DIR",
|
||||
probe_module="pocket_tts",
|
||||
has_source=False,
|
||||
venv_args=("--python", "3.11"),
|
||||
install_args=("pocket-tts==2.1.0",),
|
||||
cpu_torch_index=True,
|
||||
docs_path="docs/engines/pockettts.md",
|
||||
# CPU torch + scipy. The gated weights download on first use.
|
||||
required_bytes=3 * _GIB,
|
||||
installed_probe=_in_app_env("pocket_tts"),
|
||||
host_supported=_pockettts_host,
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
class HostUnsupported(RuntimeError):
|
||||
"""The one-click install cannot work on this machine. The message is a
|
||||
VoiceStudio-owned sentence from the spec, safe to show the user."""
|
||||
|
||||
|
||||
def host_support(spec: SidecarSpec) -> tuple[bool, str]:
|
||||
"""Whether *spec*'s install can work here. A probe that raises counts as
|
||||
unsupported: offering a button that fails is worse than not offering it."""
|
||||
try:
|
||||
ok, why = spec.host_supported()
|
||||
except Exception: # noqa: BLE001
|
||||
return False, (
|
||||
f"Could not check whether {spec.display_name} can be installed on "
|
||||
f"this machine. Its guide ({spec.docs_path}) has the manual steps."
|
||||
)
|
||||
return bool(ok), (why or "")
|
||||
|
||||
|
||||
def installable_engine_ids() -> frozenset[str]:
|
||||
"""Engines that get an Install button on THIS host."""
|
||||
return frozenset(eid for eid, spec in SPECS.items() if host_support(spec)[0])
|
||||
|
||||
|
||||
def _expand(value: str, checkout: Path) -> str:
|
||||
return value.replace("{checkout_repr}", repr(str(checkout))).replace(
|
||||
"{checkout}", str(checkout)
|
||||
)
|
||||
|
||||
|
||||
def get_spec(engine_id: str) -> Optional[SidecarSpec]:
|
||||
return SPECS.get(engine_id)
|
||||
|
||||
@@ -196,6 +459,20 @@ def managed_checkout(spec: SidecarSpec) -> Path:
|
||||
return managed_root(spec) / spec.checkout_dirname
|
||||
|
||||
|
||||
def engine_venv_python(env_var: str) -> Optional[Path]:
|
||||
"""The interpreter of the install *env_var* points at, if it has one.
|
||||
|
||||
For engines that can live in the app's environment or in a venv of their
|
||||
own (PocketTTS, Supertonic-3): they prefer their own, and fall back to the
|
||||
app's interpreter for an install made with ``uv sync --extra``.
|
||||
"""
|
||||
env_dir = os.environ.get(env_var)
|
||||
if not env_dir:
|
||||
return None
|
||||
py = _venv_python(Path(env_dir) / ".venv")
|
||||
return py if py.is_file() else None
|
||||
|
||||
|
||||
def _legacy_managed_checkouts(spec: SidecarSpec) -> tuple[Path, ...]:
|
||||
"""App-owned predecessor checkouts retained during in-place upgrades."""
|
||||
if spec.engine_id == "indextts2":
|
||||
@@ -272,13 +549,20 @@ def _default_uv_cache_root() -> Path:
|
||||
return Path(os.environ.get("XDG_CACHE_HOME") or Path.home() / ".cache") / "uv"
|
||||
|
||||
|
||||
def uv_subprocess_env(cache_parent: Path) -> "dict[str, str] | None":
|
||||
def uv_subprocess_env(cache_parent: Path) -> "dict[str, str]":
|
||||
"""Environment for ``uv`` subprocesses that install into *cache_parent*'s volume.
|
||||
|
||||
Returns ``None`` (inherit the parent environment untouched) when uv's
|
||||
default cache already shares a volume with *cache_parent* or the user
|
||||
pinned both variables themselves. Otherwise returns a copy of
|
||||
``os.environ`` with the *unset* one(s) of ``UV_CACHE_DIR`` /
|
||||
Always a copy of ``os.environ`` with ``UV_NO_CONFIG=1``: an engine's
|
||||
install resolves its own requirements, never VoiceStudio's. The backend
|
||||
runs inside the app's tree, so uv would otherwise discover the app's
|
||||
``pyproject.toml`` and apply its ``[tool.uv] constraint-dependencies``
|
||||
(``torch==2.8.0``) to the engine's venv. An engine pinning another torch
|
||||
(MOSS-TTS-v1.5, Confucius4) could then never resolve, and one that pins
|
||||
none got the app's torch instead of its own. Mirrors still apply: they
|
||||
arrive as ``UV_INDEX_URL``, an environment variable, not a config file.
|
||||
|
||||
When uv's default cache is on another volume than *cache_parent*, the
|
||||
copy also places the *unset* one(s) of ``UV_CACHE_DIR`` /
|
||||
``UV_PYTHON_INSTALL_DIR`` placed inside *cache_parent*, so downloads, the
|
||||
unpacked wheel cache, managed Pythons, and the venv all stay on the
|
||||
target volume — and same-volume hardlink installs work again. The two
|
||||
@@ -291,17 +575,15 @@ def uv_subprocess_env(cache_parent: Path) -> "dict[str, str] | None":
|
||||
pass the directory that should hold the shared ``.uv-cache`` — typically
|
||||
the common parent of the engine venvs on that volume.
|
||||
"""
|
||||
if _same_volume(cache_parent, _default_uv_cache_root()):
|
||||
return None
|
||||
env = dict(os.environ)
|
||||
overrode = False
|
||||
env["UV_NO_CONFIG"] = "1"
|
||||
if _same_volume(cache_parent, _default_uv_cache_root()):
|
||||
return env
|
||||
if not env.get("UV_CACHE_DIR"): # explicit user choice always wins
|
||||
env["UV_CACHE_DIR"] = str(Path(cache_parent) / ".uv-cache")
|
||||
overrode = True
|
||||
if not env.get("UV_PYTHON_INSTALL_DIR"):
|
||||
env["UV_PYTHON_INSTALL_DIR"] = str(Path(cache_parent) / ".uv-python")
|
||||
overrode = True
|
||||
return env if overrode else None
|
||||
return env
|
||||
|
||||
|
||||
# ── Disk preflight ─────────────────────────────────────────────────────────
|
||||
@@ -520,9 +802,13 @@ def _healthy(spec: SidecarSpec) -> bool:
|
||||
return False
|
||||
if not _venv_python(checkout / ".venv").is_file():
|
||||
return False
|
||||
if spec.weights_repo_id and not _weights_present(spec):
|
||||
return False
|
||||
return True
|
||||
if spec.weights_repo_id:
|
||||
return _weights_present(spec)
|
||||
# Nothing downloaded after the dependencies proves they finished; only
|
||||
# the marker the import probe writes does. IndexTTS (weights) predates
|
||||
# the marker and keeps its own check, so no existing install is asked
|
||||
# to reinstall.
|
||||
return (checkout / _INSTALL_COMPLETE_MARKER).is_file()
|
||||
|
||||
|
||||
def _persist(spec: SidecarSpec) -> None:
|
||||
@@ -548,6 +834,9 @@ def start_install(engine_id: str) -> dict:
|
||||
spec = get_spec(engine_id)
|
||||
if spec is None:
|
||||
raise KeyError(engine_id)
|
||||
ok, why = host_support(spec)
|
||||
if not ok:
|
||||
raise HostUnsupported(why)
|
||||
with _jobs_lock:
|
||||
existing = _jobs.get(engine_id)
|
||||
if existing and existing["state"] == "running":
|
||||
@@ -679,6 +968,11 @@ def _step_preflight(spec: SidecarSpec, job: dict) -> None:
|
||||
def _step_fetch_source(spec: SidecarSpec, job: dict) -> None:
|
||||
step = _job_step(job, "fetch_source")
|
||||
checkout = managed_checkout(spec)
|
||||
if not spec.has_source:
|
||||
checkout.mkdir(parents=True, exist_ok=True)
|
||||
step["state"] = "done"
|
||||
step["detail"] = "PyPI package, no source to fetch"
|
||||
return
|
||||
if _source_present(spec, checkout):
|
||||
step["state"] = "done"
|
||||
step["detail"] = "source already present"
|
||||
@@ -717,7 +1011,7 @@ def _step_fetch_source(spec: SidecarSpec, job: dict) -> None:
|
||||
_fetch_tarball(spec, job, checkout)
|
||||
if not _source_layout_ok(spec, checkout):
|
||||
raise _StepError(
|
||||
f"Fetched source at {checkout} has no pyproject.toml — the download "
|
||||
f"Fetched source at {checkout} has no {spec.source_manifest} — the download "
|
||||
"appears incomplete or the upstream layout changed.",
|
||||
"Re-run the install; if it keeps failing, clone the repository "
|
||||
f"manually and set {spec.env_var} to the clone (see the engine docs).",
|
||||
@@ -727,10 +1021,14 @@ def _step_fetch_source(spec: SidecarSpec, job: dict) -> None:
|
||||
|
||||
|
||||
_SOURCE_REVISION_MARKER = ".voicestudio_source_revision"
|
||||
# Written once the import probe passes. For an engine with no weights
|
||||
# download, the venv interpreter existing proves nothing: a dependency
|
||||
# install that died halfway leaves one behind.
|
||||
_INSTALL_COMPLETE_MARKER = ".voicestudio_install_complete"
|
||||
|
||||
|
||||
def _source_layout_ok(spec: SidecarSpec, checkout: Path) -> bool:
|
||||
if not (checkout / "pyproject.toml").is_file():
|
||||
if not (checkout / spec.source_manifest).is_file():
|
||||
return False
|
||||
return not spec.source_required_path or (checkout / spec.source_required_path).is_file()
|
||||
|
||||
@@ -743,6 +1041,8 @@ def _write_source_marker(spec: SidecarSpec, checkout: Path) -> None:
|
||||
|
||||
|
||||
def _source_present(spec: SidecarSpec, checkout: Path) -> bool:
|
||||
if not spec.has_source:
|
||||
return checkout.is_dir()
|
||||
if not _source_layout_ok(spec, checkout):
|
||||
return False
|
||||
if not spec.source_revision:
|
||||
@@ -836,8 +1136,8 @@ def _step_create_venv(spec: SidecarSpec, job: dict) -> None:
|
||||
# uv_subprocess_env. The cache parent is the shared engines root, so
|
||||
# every sidecar engine reuses one cache.
|
||||
uv_env = uv_subprocess_env(Path(DATA_DIR) / "engines")
|
||||
rc = _run_logged(job, [uv, "venv", str(venv_dir)], timeout=_UV_VENV_TIMEOUT_S,
|
||||
env=uv_env)
|
||||
rc = _run_logged(job, [uv, "venv", str(venv_dir), *spec.venv_args],
|
||||
timeout=_UV_VENV_TIMEOUT_S, env=uv_env)
|
||||
if rc != 0 or not py.is_file():
|
||||
raise _StepError(
|
||||
f"uv venv failed (exit {rc}) at {venv_dir}.",
|
||||
@@ -857,17 +1157,28 @@ def _step_install_deps(spec: SidecarSpec, job: dict) -> None:
|
||||
"""
|
||||
checkout = managed_checkout(spec)
|
||||
py = _venv_python(checkout / ".venv")
|
||||
# A reinstall that fails must not leave the previous run's marker.
|
||||
(checkout / _INSTALL_COMPLETE_MARKER).unlink(missing_ok=True)
|
||||
uv = _locate_uv()
|
||||
_log(job, f"Installing {spec.display_name} into its venv (this can take several minutes) …")
|
||||
target = [_expand(arg, checkout) for arg in spec.install_args]
|
||||
if spec.cpu_torch_index:
|
||||
from core.torch_indexes import UV_PIP_CPU_ARGS
|
||||
target += list(UV_PIP_CPU_ARGS)
|
||||
elif spec.uses_cuda_index and _host_family() == "cuda":
|
||||
from core.torch_indexes import UV_PIP_CU128_ARGS
|
||||
target += list(UV_PIP_CU128_ARGS)
|
||||
# Always `--python <this engine's venv>`: the install can only ever land in
|
||||
# the venv this engine owns, never the app's interpreter.
|
||||
rc = _run_logged(
|
||||
job,
|
||||
[uv, "pip", "install", "--python", str(py), "-e", str(checkout)],
|
||||
[uv, "pip", "install", "--python", str(py), *target],
|
||||
timeout=_UV_PIP_INSTALL_TIMEOUT_S,
|
||||
env=uv_subprocess_env(Path(DATA_DIR) / "engines"),
|
||||
)
|
||||
if rc != 0:
|
||||
raise _StepError(
|
||||
f"uv pip install -e failed (exit {rc}).",
|
||||
f"uv pip install failed (exit {rc}).",
|
||||
"Usually a network hiccup — re-run the install to resume. Behind a "
|
||||
"proxy, set HTTPS_PROXY in Settings → Environment first.",
|
||||
)
|
||||
@@ -879,8 +1190,13 @@ def _step_verify(spec: SidecarSpec, job: dict) -> None:
|
||||
py = _venv_python(checkout / ".venv")
|
||||
_log(job, f"Verifying `import {spec.probe_module}` inside the venv …")
|
||||
try:
|
||||
probe = (
|
||||
_expand(spec.probe_code, checkout)
|
||||
if spec.probe_code
|
||||
else f"import {spec.probe_module}"
|
||||
)
|
||||
proc = subprocess.run(
|
||||
[str(py), "-c", f"import {spec.probe_module}"],
|
||||
[str(py), "-c", probe],
|
||||
capture_output=True, timeout=_IMPORT_PROBE_TIMEOUT_S,
|
||||
)
|
||||
except (subprocess.TimeoutExpired, OSError) as exc:
|
||||
@@ -898,6 +1214,7 @@ def _step_verify(spec: SidecarSpec, job: dict) -> None:
|
||||
"the engine docs.",
|
||||
)
|
||||
_job_step(job, "verify")["detail"] = f"import {spec.probe_module} OK"
|
||||
(checkout / _INSTALL_COMPLETE_MARKER).write_text(f"{spec.probe_module}\n", encoding="utf-8")
|
||||
_log(job, "Venv verified.")
|
||||
|
||||
|
||||
@@ -1057,7 +1374,8 @@ def _run_logged(job: dict, argv: list[str], *, timeout: float,
|
||||
would hang past the timeout waiting for pipe EOF.
|
||||
"""
|
||||
# ``spawn_owned`` creates the local timeout group/Job before the operation
|
||||
# starts and links it to backend death through its control pipe.
|
||||
# starts. POSIX links it to backend death through a control pipe; Windows
|
||||
# retains a kill-on-close Job handle in this backend process.
|
||||
popen_kwargs = _install_containment_kwargs()
|
||||
try:
|
||||
proc = spawn_owned(
|
||||
@@ -1096,14 +1414,14 @@ def _run_logged(job: dict, argv: list[str], *, timeout: float,
|
||||
|
||||
def _kill_tree(proc: "subprocess.Popen") -> None:
|
||||
"""Kill an operation through its stable nested group/Job owner."""
|
||||
if isinstance(proc, OwnedPopen):
|
||||
if isinstance(proc, (OwnedPopen, WindowsJobPopen)):
|
||||
# The retained supervisor/process-group or nested Job is the stable
|
||||
# per-operation owner. Do not fall back to a direct PID kill.
|
||||
proc.kill()
|
||||
try:
|
||||
proc.wait(timeout=5)
|
||||
except subprocess.TimeoutExpired:
|
||||
pass
|
||||
return
|
||||
return
|
||||
# A test double or a legacy caller without the nested owner can only be
|
||||
# stopped through its stable direct-process handle.
|
||||
|
||||
@@ -32,9 +32,9 @@ Threat-model summary (see Plan 02-01 frontmatter):
|
||||
AUTH-05 installed (``HFTokenRedactor``) on the root logger.
|
||||
T-02-04 — compromised sidecar emitting unexpected ops: parent allowlist
|
||||
``PARENT_INBOUND_OPS`` rejects everything else.
|
||||
T-02-05 — nested containment: a retained supervisor process group/Job owns
|
||||
each engine operation and is linked to backend death by a control
|
||||
pipe, while still permitting independent timeout teardown.
|
||||
T-02-05 — nested containment: a retained POSIX supervisor process group or
|
||||
Windows Job owns each engine operation, while still permitting
|
||||
independent timeout teardown and cleanup on backend death.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -385,6 +385,10 @@ class SubprocessBackend(TTSBackend):
|
||||
|
||||
def __init__(self) -> None:
|
||||
self._proc: Optional[subprocess.Popen] = None
|
||||
# A failed bounded reap must retain ownership and forbid reuse. This
|
||||
# lock is separate from _lock: the receive owner joins its watchdog.
|
||||
self._timeout_quarantine: list[subprocess.Popen] = []
|
||||
self._timeout_quarantine_lock = threading.Lock()
|
||||
# Single lock serialises spawn + every send/recv pair so two threads
|
||||
# can't interleave half-frames on the same pipe.
|
||||
self._lock = threading.Lock()
|
||||
@@ -451,6 +455,10 @@ class SubprocessBackend(TTSBackend):
|
||||
def _spawn(self) -> None:
|
||||
"""Launch the sidecar if not already running. Blocks on the ready
|
||||
handshake. Caller must hold self._lock."""
|
||||
if not self._retry_timeout_cleanup():
|
||||
raise RuntimeError(
|
||||
f"{self.id} sidecar is still stopping after a timeout; retry once it exits"
|
||||
)
|
||||
if self._proc is not None and self._proc.poll() is None:
|
||||
return # already up
|
||||
|
||||
@@ -542,6 +550,7 @@ class SubprocessBackend(TTSBackend):
|
||||
"""Idempotent. Sends {op:shutdown}; falls back to terminate/kill."""
|
||||
proc = self._proc
|
||||
if proc is None:
|
||||
self._retry_timeout_cleanup()
|
||||
return
|
||||
try:
|
||||
try:
|
||||
@@ -577,6 +586,7 @@ class SubprocessBackend(TTSBackend):
|
||||
pass
|
||||
finally:
|
||||
self._proc = None
|
||||
self._retry_timeout_cleanup()
|
||||
|
||||
def _force_kill(self) -> None:
|
||||
"""Internal: kill a sidecar that never reached the ready state."""
|
||||
@@ -777,38 +787,59 @@ class SubprocessBackend(TTSBackend):
|
||||
return msg
|
||||
|
||||
def _recv_with_timeout(self, timeout_s: float) -> Optional[dict]:
|
||||
"""Recv that aborts if the sidecar goes silent.
|
||||
"""Read one frame, finishing timeout cleanup before the caller can retry.
|
||||
|
||||
Implemented by polling the proc for liveness with a deadline. We
|
||||
don't block on a `select` of the pipe because Windows can't select
|
||||
on subprocess pipes — keeping the implementation cross-platform
|
||||
means a simpler polling loop here.
|
||||
A watchdog closes the pipe on timeout; Windows cannot select on pipes.
|
||||
EOF alone does not prove the owned process/supervisor has exited.
|
||||
"""
|
||||
# On Unix we could use selectors; on Windows the pipe is not
|
||||
# selectable. Use a watchdog thread that kills the sidecar on
|
||||
# timeout — that triggers EOF on stdout, so _recv returns None
|
||||
# and the caller raises.
|
||||
watchdog = threading.Timer(timeout_s, self._timeout_kill)
|
||||
proc = self._proc
|
||||
watchdog = threading.Timer(timeout_s, self._timeout_kill, args=(proc,))
|
||||
watchdog.daemon = True
|
||||
watchdog.start()
|
||||
try:
|
||||
return self._recv()
|
||||
finally:
|
||||
watchdog.cancel()
|
||||
# cancel() cannot stop an already-running callback. Finish its
|
||||
# bounded reap before another receive or generation starts.
|
||||
watchdog.join()
|
||||
self._touch() # any reply (or attempt) counts as recent activity
|
||||
|
||||
def _timeout_kill(self) -> None:
|
||||
proc = self._proc
|
||||
def _timeout_kill(self, proc: Optional[subprocess.Popen]) -> None:
|
||||
"""Kill only the child this receive captured, then reap its owner."""
|
||||
if proc is None:
|
||||
return
|
||||
logger.error("[%s] sidecar exceeded recv timeout; killing", self.id)
|
||||
try:
|
||||
logger.error(
|
||||
"[%s] sidecar exceeded recv timeout; killing",
|
||||
self.id,
|
||||
)
|
||||
proc.kill()
|
||||
except Exception:
|
||||
# A raced exit can make kill fail, but its owner still needs reaping.
|
||||
pass
|
||||
try:
|
||||
proc.wait(timeout=2)
|
||||
except Exception:
|
||||
# Do not discard a possibly live owner, or replace a newer _proc.
|
||||
with self._timeout_quarantine_lock:
|
||||
if not any(item is proc for item in self._timeout_quarantine):
|
||||
self._timeout_quarantine.append(proc)
|
||||
else:
|
||||
with self._timeout_quarantine_lock:
|
||||
self._timeout_quarantine = [
|
||||
item for item in self._timeout_quarantine if item is not proc
|
||||
]
|
||||
|
||||
def _retry_timeout_cleanup(self) -> bool:
|
||||
"""Retry bounded cleanup, retaining every owner that could still be live."""
|
||||
with self._timeout_quarantine_lock:
|
||||
pending = tuple(self._timeout_quarantine)
|
||||
for proc in pending:
|
||||
self._timeout_kill(proc)
|
||||
with self._timeout_quarantine_lock:
|
||||
return not self._timeout_quarantine
|
||||
|
||||
# ── stderr drain ───────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -112,6 +112,13 @@ _FULL_NAME_TO_CODE = {
|
||||
"vietnamese": "vi",
|
||||
"kazakh": "kz",
|
||||
"standard arabic": "ar",
|
||||
# Below: inert for num2words (absent from _NUM2WORDS_LANGS, which reads
|
||||
# digits natively for these scripts), present so _plain_lang_code can
|
||||
# resolve them for the digit-range rule.
|
||||
"korean": "ko",
|
||||
"japanese": "ja",
|
||||
"chinese": "zh",
|
||||
"mandarin chinese": "zh",
|
||||
}
|
||||
|
||||
# ISO codes whose num2words locale name differs.
|
||||
@@ -178,6 +185,82 @@ def _num2words_lang(language: Optional[str]) -> Optional[str]:
|
||||
return None
|
||||
|
||||
|
||||
def _plain_lang_code(language: Optional[str]) -> Optional[str]:
|
||||
"""Resolve a request language to a bare ISO code, with no num2words gate.
|
||||
|
||||
:func:`_num2words_lang` answers "may I call num2words for this?" and so
|
||||
returns ``None`` for ko/ja/zh/th/vi. Rules that are not num2words-backed
|
||||
need the code itself, which is what this returns.
|
||||
"""
|
||||
if not language:
|
||||
return None
|
||||
s = str(language).strip().lower()
|
||||
if not s or s == "auto":
|
||||
return None
|
||||
code = _FULL_NAME_TO_CODE.get(s)
|
||||
if code:
|
||||
return code
|
||||
m = _ISO_CODE_RE.match(s)
|
||||
if m:
|
||||
return _ISO_ALIASES.get(m.group(1), m.group(1))
|
||||
return None
|
||||
|
||||
|
||||
# ── Digit ranges ─────────────────────────────────────────────────────────────
|
||||
# "20~30" loses its separator at the engine and reads as ONE number: OmniVoice
|
||||
# says "이십삼" (23) for "20~30초". Speak the separator instead. Verified by
|
||||
# rendering each form and transcribing it back (ko, OmniVoice):
|
||||
# "20~30초" → heard "23초" ✗
|
||||
# "20에서 30초" → heard "20에서 30초" ✓
|
||||
# Only the tilde family is rewritten — those are unambiguously range marks
|
||||
# between digits. An ASCII hyphen is left alone on purpose: it also spells
|
||||
# dates, phone numbers and product codes, where "to" would be wrong.
|
||||
#: Spacing is part of the form, not decoration: a Korean postposition binds to
|
||||
#: the numeral ("20에서 30"), Japanese and Chinese set no spaces at all, and
|
||||
#: English needs them on both sides.
|
||||
_RANGE_FORM = {
|
||||
"ko": "{a}에서 {b}",
|
||||
"ja": "{a}から{b}",
|
||||
"zh": "{a}到{b}",
|
||||
"en": "{a} to {b}",
|
||||
}
|
||||
|
||||
#: ASCII tilde, wave dash, fullwidth tilde — Japanese and Korean IMEs emit the
|
||||
#: latter two, so all three have to match.
|
||||
#:
|
||||
#: Match complete signed/decimal endpoints; reject partial numbers and product
|
||||
#: codes while allowing adjacent CJK units. Guard all tilde forms so malformed
|
||||
#: chains cannot be partially rewritten, including when their separators have
|
||||
#: whitespace around them.
|
||||
_RANGE_MARKS = "~\u301c\uff5e"
|
||||
_RANGE_ENDPOINT = r"[+-]?(?:\d{1,6}(?:\.\d{1,6})?|\.\d{1,6})"
|
||||
_NUM_RANGE_RE = re.compile(
|
||||
rf"(?<![\d.,A-Za-z+{_RANGE_MARKS}-])({_RANGE_ENDPOINT})"
|
||||
rf"\s*[{_RANGE_MARKS}]\s*({_RANGE_ENDPOINT})"
|
||||
rf"(?![\d.,A-Za-z+{_RANGE_MARKS}-])"
|
||||
)
|
||||
|
||||
|
||||
def _speak_number_ranges(text: str, lang: str) -> str:
|
||||
"""Speak complete tilde ranges only for languages with a verified form."""
|
||||
form = _RANGE_FORM.get(lang)
|
||||
if not form:
|
||||
return text
|
||||
|
||||
def replace(match: re.Match) -> str:
|
||||
before, after = match.start() - 1, match.end()
|
||||
while before >= 0 and text[before].isspace():
|
||||
before -= 1
|
||||
while after < len(text) and text[after].isspace():
|
||||
after += 1
|
||||
if ((before >= 0 and text[before] in _RANGE_MARKS)
|
||||
or (after < len(text) and text[after] in _RANGE_MARKS)):
|
||||
return match.group(0)
|
||||
return form.format(a=match.group(1), b=match.group(2))
|
||||
|
||||
return _NUM_RANGE_RE.sub(replace, text)
|
||||
|
||||
|
||||
# ── Universal safety filters (all languages) ─────────────────────────────────
|
||||
|
||||
# Zero-width & bidi controls, C0/C1 controls (except \t \n \r), BOM, U+FFFD.
|
||||
@@ -487,6 +570,11 @@ def normalize_text(text: str, language: Optional[str] = None) -> str:
|
||||
if not text:
|
||||
return text or ""
|
||||
out = _safety_filters(text)
|
||||
# Runs outside the num2words gate below: ko/ja/zh keep their digits (that
|
||||
# gate returns None for them) but still need the range mark spoken.
|
||||
plain = _plain_lang_code(language)
|
||||
if plain:
|
||||
out = _outside_brackets(out, lambda t: _speak_number_ranges(t, plain))
|
||||
lang = _num2words_lang(language)
|
||||
if lang:
|
||||
if lang in _ABBREV_COMPILED:
|
||||
|
||||
@@ -4,7 +4,7 @@ Resolution priority (highest → lowest):
|
||||
|
||||
1. app — `settings_store.get_hf_token()` (encrypted in SQLite)
|
||||
2. env — `HF_TOKEN` or the legacy `HUGGING_FACE_HUB_TOKEN` env var
|
||||
3. hf-cli — `huggingface_hub.get_token()` (canonical ~/.cache/huggingface/token)
|
||||
3. hf-cli — the selected local Hub token file (`HF_TOKEN_PATH`)
|
||||
|
||||
For each candidate, the resolver calls `huggingface_hub.whoami(token=...)`
|
||||
to verify the token is live; any HTTP error (401, 403, network) skips to
|
||||
@@ -20,6 +20,7 @@ from __future__ import annotations
|
||||
import hashlib
|
||||
import logging
|
||||
import os
|
||||
from pathlib import Path
|
||||
import threading
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
@@ -46,7 +47,7 @@ class SourceState:
|
||||
set: bool
|
||||
masked: Optional[str]
|
||||
whoami_user: Optional[str]
|
||||
whoami_ok: bool
|
||||
whoami_ok: Optional[bool]
|
||||
|
||||
|
||||
# ── module-level cache ────────────────────────────────────────────────────
|
||||
@@ -79,19 +80,26 @@ def _read_app() -> Optional[str]:
|
||||
return None
|
||||
|
||||
|
||||
def _clean_token(value: Optional[str]) -> Optional[str]:
|
||||
if not value:
|
||||
return None
|
||||
return value.replace("\r", "").replace("\n", "").strip() or None
|
||||
|
||||
|
||||
def _read_env() -> Optional[str]:
|
||||
# HF docs explicitly accept either name; user may have either exported.
|
||||
val = os.environ.get("HF_TOKEN") or os.environ.get("HUGGING_FACE_HUB_TOKEN")
|
||||
return val or None
|
||||
return _clean_token(val)
|
||||
|
||||
|
||||
def _read_hf_cli() -> Optional[str]:
|
||||
try:
|
||||
import huggingface_hub
|
||||
tok = huggingface_hub.get_token()
|
||||
return tok or None
|
||||
from huggingface_hub import constants
|
||||
return _clean_token(Path(constants.HF_TOKEN_PATH).read_text(encoding="utf-8"))
|
||||
except FileNotFoundError:
|
||||
return None
|
||||
except Exception:
|
||||
logger.exception("huggingface_hub.get_token failed")
|
||||
logger.warning("Could not read the local Hugging Face token file")
|
||||
return None
|
||||
|
||||
|
||||
@@ -184,17 +192,17 @@ def on_401(active_source: Source) -> Optional[ResolvedToken]:
|
||||
return resolve(skip=frozenset({active_source}))
|
||||
|
||||
|
||||
def state() -> dict:
|
||||
def state(*, validate: bool = False) -> dict:
|
||||
"""Return one SourceState per priority position so the Settings UI can
|
||||
render the cascade table. Includes a masked token + whoami result;
|
||||
never includes the raw token."""
|
||||
never includes the raw token. Reads are local unless validation is explicitly requested."""
|
||||
rows: list[SourceState] = []
|
||||
active: Optional[Source] = None
|
||||
for source in _PRIORITY:
|
||||
token = _READERS[source]()
|
||||
if token:
|
||||
username = _validate(source, token)
|
||||
ok = username is not None
|
||||
username = _validate(source, token) if validate else None
|
||||
ok = (username is not None) if validate else None
|
||||
rows.append(SourceState(
|
||||
source=source,
|
||||
set=True,
|
||||
@@ -249,15 +257,30 @@ def save_app_token(token: str) -> None:
|
||||
invalidate_cache()
|
||||
|
||||
|
||||
def clear_hf_cli_tokens() -> None:
|
||||
"""Remove recognized Hub token files without refreshing or revoking tokens."""
|
||||
from core.config import HF_CLI_TOKEN_PATHS
|
||||
from huggingface_hub import constants
|
||||
|
||||
# Hub's active path remains authoritative if imported before app config.
|
||||
paths = set(HF_CLI_TOKEN_PATHS) | {constants.HF_TOKEN_PATH}
|
||||
failed = False
|
||||
for token_path in paths:
|
||||
path = Path(token_path)
|
||||
for target in (path, path.parent / "stored_tokens"):
|
||||
try:
|
||||
target.unlink(missing_ok=True)
|
||||
except OSError:
|
||||
failed = True
|
||||
invalidate_cache()
|
||||
if failed:
|
||||
raise OSError("Could not clear all local Hugging Face token files")
|
||||
|
||||
|
||||
def clear_app_token(also_clear_hf_cli: bool = False) -> None:
|
||||
"""Remove from the encrypted settings store; optionally also call
|
||||
`huggingface_hub.logout()` to clear the canonical HF file."""
|
||||
"""Clear the encrypted app token, optionally recognized local Hub files."""
|
||||
from services import settings_store
|
||||
settings_store.clear_hf_token()
|
||||
if also_clear_hf_cli:
|
||||
try:
|
||||
import huggingface_hub
|
||||
huggingface_hub.logout()
|
||||
except Exception:
|
||||
logger.exception("huggingface_hub.logout failed (non-fatal)")
|
||||
clear_hf_cli_tokens()
|
||||
invalidate_cache()
|
||||
|
||||
@@ -17,6 +17,8 @@ from __future__ import annotations
|
||||
import asyncio
|
||||
import importlib
|
||||
import logging
|
||||
import functools
|
||||
import re
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
@@ -91,6 +93,8 @@ REGISTRY: dict[str, dict] = {
|
||||
"probe_module": "openai",
|
||||
"category": "llm",
|
||||
"needs_key": True,
|
||||
# A core dependency: Settings → LLM Providers uses it too.
|
||||
"builtin": True,
|
||||
"notes": (
|
||||
"Uses the LLM provider you configure in Settings → LLM Providers "
|
||||
"(route it via the 'Dub translation' skill in Settings → LLM Skills): "
|
||||
@@ -181,6 +185,65 @@ def list_engines() -> list[dict]:
|
||||
return out
|
||||
|
||||
|
||||
def _normalize(name: str) -> str:
|
||||
"""A distribution name in PEP 503 form (deep_translator == deep-translator)."""
|
||||
return re.sub(r"[-_.]+", "-", name).lower()
|
||||
|
||||
|
||||
@functools.lru_cache(maxsize=1)
|
||||
def _app_dependency_names() -> frozenset[str]:
|
||||
"""Distribution names VoiceStudio itself requires, normalized.
|
||||
|
||||
Read from the installed package metadata, so it follows the lockfile with
|
||||
no second list to keep in step. Without metadata this guards nothing
|
||||
rather than failing.
|
||||
"""
|
||||
try:
|
||||
from importlib.metadata import requires
|
||||
|
||||
reqs = requires("omnivoice") or []
|
||||
except Exception: # noqa: BLE001
|
||||
return frozenset()
|
||||
names = set()
|
||||
for req in reqs:
|
||||
if "extra ==" in req:
|
||||
continue
|
||||
names.add(_normalize(re.split(r"[\s;<>=!~\[@(]", req, maxsplit=1)[0]))
|
||||
return frozenset(names)
|
||||
|
||||
|
||||
def uninstall_blocker(engine_id: str) -> "tuple[int, str] | None":
|
||||
"""Why removing this engine's package would break something, or None.
|
||||
|
||||
`pip uninstall` acts on the app's own environment. A package VoiceStudio
|
||||
depends on (openai, argostranslate) would break the app, and a package
|
||||
other translation engines share (deep_translator backs four) would break
|
||||
those engines too.
|
||||
"""
|
||||
entry = REGISTRY.get(engine_id)
|
||||
pkg = entry.get("pip_package") if entry else None
|
||||
if not pkg:
|
||||
return None
|
||||
if _normalize(pkg) in _app_dependency_names():
|
||||
return 400, (
|
||||
f"{entry['display_name']} uses {pkg}, which VoiceStudio itself "
|
||||
"depends on. Uninstalling it would break the app."
|
||||
)
|
||||
sharing = [
|
||||
other["display_name"]
|
||||
for other_id, other in REGISTRY.items()
|
||||
if other_id != engine_id
|
||||
and other.get("pip_package")
|
||||
and _normalize(other["pip_package"]) == _normalize(pkg)
|
||||
]
|
||||
if sharing:
|
||||
return 409, (
|
||||
f"{entry['display_name']} shares {pkg} with {', '.join(sharing)}. "
|
||||
"Uninstalling it would stop those working too."
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
def get_engine(engine_id: str) -> dict | None:
|
||||
return REGISTRY.get(engine_id)
|
||||
|
||||
|
||||
+187
-22
@@ -300,6 +300,31 @@ class TTSBackend(ABC):
|
||||
#: 0 means "no meaningful floor" (CPU-class engines) and never warns.
|
||||
min_vram_gb: float = 0.0
|
||||
|
||||
@classmethod
|
||||
def runtime_compute_profile(cls, caps) -> dict:
|
||||
"""Resolved compute metadata for this engine on the current host.
|
||||
|
||||
Most engines have one implementation whose static declarations are
|
||||
sufficient. Native adapters may override this single hook when the
|
||||
installed executable determines both the available runtimes and the
|
||||
device actually selected.
|
||||
"""
|
||||
from services.engine_routing import resolve_routing
|
||||
|
||||
gpu_compat = tuple(getattr(cls, "gpu_compat", ("cpu",)))
|
||||
min_vram_gb = float(getattr(cls, "min_vram_gb", 0.0) or 0.0)
|
||||
return {
|
||||
"gpu_compat": gpu_compat,
|
||||
"min_vram_gb": min_vram_gb,
|
||||
**resolve_routing(gpu_compat, caps, min_vram_gb),
|
||||
"runtime_backend": None,
|
||||
"runtime_device_index": None,
|
||||
"runtime_device_name": None,
|
||||
"runtime_hardware_family": None,
|
||||
"runtime_vram_gb": None,
|
||||
"runtime_device_verified": None,
|
||||
}
|
||||
|
||||
#: True when generation allocates in ANOTHER process — a dedicated-venv
|
||||
#: sidecar (SubprocessBackend) or a spawned binary (omnivoice-gguf).
|
||||
#: Parent-process accelerator counters cannot see those allocations, so
|
||||
@@ -566,7 +591,12 @@ def _get_clone_prompt(
|
||||
):
|
||||
"""Return a cached/precomputed ``VoiceClonePrompt`` for
|
||||
(ref_audio, ref_text, preprocess_prompt), or ``None`` to fall back to the
|
||||
inline ref path. Never raises.
|
||||
inline ref path.
|
||||
|
||||
Raises only on a device OOM that survives a cache-drop retry (#1790): the
|
||||
inline path is the same allocation on the same device, so falling back to
|
||||
it after an OOM cannot succeed and has been observed taking the whole
|
||||
process down instead. Every other failure still falls back silently.
|
||||
|
||||
``store=False`` still *reads* the cache (a hit is free) but never inserts:
|
||||
it exists for single-use references — a dub's per-segment ref clips are each
|
||||
@@ -596,10 +626,43 @@ def _get_clone_prompt(
|
||||
ref_audio, ref_text=ref_text, preprocess_prompt=preprocess_prompt
|
||||
)
|
||||
except Exception as e: # noqa: BLE001 — fall back, never break synthesis
|
||||
logger.warning(
|
||||
"voice-clone prompt precompute failed; using inline ref: %s", e
|
||||
)
|
||||
return None
|
||||
# #1790/#1777: a GPU OOM is the one failure this fallback cannot
|
||||
# absorb. `generate()`'s inline ref path runs the SAME encode on the
|
||||
# SAME device — the docstring above says so, because producing
|
||||
# identical output is the point — so returning None after an OOM
|
||||
# guarantees a second OOM moments later, on a device with even less
|
||||
# headroom than the first attempt found. Both reporters' backends
|
||||
# then died with a Windows access violation (exit code
|
||||
# -1073741819) seconds after this exact log line, mid-generation on
|
||||
# a GPU that had just refused an 86 MiB allocation.
|
||||
#
|
||||
# An OOM here is also the most recoverable kind: the allocator is
|
||||
# typically holding reserved-but-unallocated blocks (#1790's own
|
||||
# log reports 90 MiB reserved against an 86 MiB request). Drop them
|
||||
# and try once more. If it still will not fit, raise — the failure
|
||||
# layer turns a device OOM into the actionable GPU_OOM message
|
||||
# ("close other GPU-heavy apps or unload models…"), which is a far
|
||||
# better answer than walking into a native fault.
|
||||
from core.failure import is_gpu_oom
|
||||
|
||||
if is_gpu_oom(e):
|
||||
logger.warning(
|
||||
"voice-clone prompt precompute hit a device OOM (%s) — "
|
||||
"releasing allocator caches and retrying once", e,
|
||||
)
|
||||
try:
|
||||
from services.model_manager import free_vram
|
||||
free_vram()
|
||||
except Exception: # noqa: BLE001 — reclaim is best-effort
|
||||
logger.debug("VRAM reclaim before OOM retry failed", exc_info=True)
|
||||
prompt = model.create_voice_clone_prompt(
|
||||
ref_audio, ref_text=ref_text, preprocess_prompt=preprocess_prompt
|
||||
)
|
||||
else:
|
||||
logger.warning(
|
||||
"voice-clone prompt precompute failed; using inline ref: %s", e
|
||||
)
|
||||
return None
|
||||
if store:
|
||||
_prompt_disk_save(key, prompt)
|
||||
if not store:
|
||||
@@ -2197,6 +2260,13 @@ _LAZY_REGISTRY: dict[str, tuple[str, str]] = {
|
||||
# 2026-07-02 (CPU, Apple Silicon; 22.05 kHz output). Gated behind
|
||||
# OMNIVOICE_CONFUCIUS4_TTS_DIR so it's inert until enabled.
|
||||
"confucius4-tts": ("engines.confucius4", "Confucius4Backend"),
|
||||
# audio.cpp (0xShug0/audio.cpp) — pure-C++ ggml runtime, no Python venv.
|
||||
# v1 serves Breeze-TTS-2 (en+zh, clone+design) through a parent-managed
|
||||
# audiocpp_server over loopback HTTP. Gated behind a server binary
|
||||
# (OMNIVOICE_AUDIOCPP_BIN) so it's inert until enabled. Lazy for the
|
||||
# same import-cycle reason as the entries above (engines.audiocpp
|
||||
# imports services.tts_backend for TTSBackend).
|
||||
"audiocpp": ("engines.audiocpp", "AudioCPPBackend"),
|
||||
}
|
||||
|
||||
|
||||
@@ -2298,9 +2368,10 @@ _INSTALL_HINTS: dict[str, str] = {
|
||||
"omnivoice-gguf":"Bundled — runs the C++ omnivoice-tts binary in bin/. Quants download lazily from Serveurperso/OmniVoice-GGUF on first generate.",
|
||||
"supertonic3": "uv sync --extra supertonic (CPU-only ONNX, 31 langs, ~400 MB model on first use; OpenRAIL-M model license)",
|
||||
"pockettts": "uv sync --extra pockettts (Kyutai, CPU-only, ~100 MB model on first use; MIT code + CC-BY-4.0 weights; HF-gated, review terms and set HF_TOKEN)",
|
||||
"moss-tts-v15": "git clone OpenMOSS/MOSS-TTS + set OMNIVOICE_MOSS_TTS_V15_DIR (own venv, transformers==5.0; 8B, ~16 GB weights; CUDA/CPU, no MPS; Apache-2.0)",
|
||||
"moss-tts-v15": "git clone OpenMOSS/MOSS-TTS + set OMNIVOICE_MOSS_TTS_V15_DIR (own venv, transformers==5.0; 8B, ~16 GB weights; CUDA/ROCm/XPU/NPU/CPU, no MPS; Apache-2.0)",
|
||||
"dots-tts": "git clone rednote-hilab/dots.tts + set OMNIVOICE_DOTS_TTS_DIR (own venv, transformers==4.57; 2B, ~9 GB weights; CUDA/CPU, Linux/macOS only — no Windows; Apache-2.0)",
|
||||
"confucius4-tts":"git clone netease-youdao/Confucius4-TTS + set OMNIVOICE_CONFUCIUS4_TTS_DIR (own Python 3.10 venv; 14-lang cross-lingual zero-shot clone; ~5 GB weights auto-download; CUDA/CPU, no MPS; Apache-2.0)",
|
||||
"confucius4-tts":"git clone netease-youdao/Confucius4-TTS + set OMNIVOICE_CONFUCIUS4_TTS_DIR (own Python 3.10 venv; 14-lang cross-lingual zero-shot clone; ~5 GB weights auto-download; CUDA/ROCm/XPU/NPU/CPU, no MPS; Apache-2.0)",
|
||||
"audiocpp": "download the matching audio.cpp v0.7.2 prebuilt + set OMNIVOICE_AUDIOCPP_BIN, then explicitly install Breeze-TTS-2 in Model Catalogue → Models (native CPU/Vulkan/CUDA/Metal GGUF server, no Python; en+zh clone+design; ~4.73 GiB; weights research/non-commercial only)",
|
||||
}
|
||||
|
||||
|
||||
@@ -2321,6 +2392,51 @@ _SETUP_SNIPPETS: dict[str, str] = {
|
||||
}
|
||||
|
||||
|
||||
# Per-engine documentation page, as a repo-relative path (#1866). Every one of
|
||||
# these docs already exists and several are CI-guarded against the code they
|
||||
# describe (e.g. tests/test_cosyvoice_install_docs.py), but nothing in the app
|
||||
# linked to them, so the point of failure — an unavailable engine row — was a
|
||||
# dead end. Paths rather than URLs so tests/test_engine_docs.py can assert the
|
||||
# file is really there; the URL is built once, at read time, from core.links.
|
||||
#
|
||||
# Keyed on the engine id, so it stays correct when the doc filename does not
|
||||
# match the id (indextts2 → indextts.md).
|
||||
_ENGINE_DOCS: dict[str, str] = {
|
||||
"omnivoice": "docs/engines/omnivoice.md",
|
||||
"omnivoice-subprocess": "docs/engines/omnivoice-subprocess.md",
|
||||
"omnivoice-gguf": "docs/engines/omnivoice-gguf.md",
|
||||
"cosyvoice": "docs/engines/cosyvoice.md",
|
||||
"kittentts": "docs/engines/kittentts.md",
|
||||
"mlx-audio": "docs/engines/mlx-audio.md",
|
||||
"voxcpm2": "docs/engines/voxcpm2.md",
|
||||
"moss-tts-nano": "docs/engines/moss-tts-nano.md",
|
||||
"moss-tts-v15": "docs/engines/moss-tts-v15.md",
|
||||
"dots-tts": "docs/engines/dots-tts.md",
|
||||
"confucius4-tts": "docs/engines/confucius4-tts.md",
|
||||
"indextts2": "docs/engines/indextts.md",
|
||||
"gpt-sovits": "docs/engines/gpt-sovits.md",
|
||||
"sherpa-onnx": "docs/engines/sherpa-onnx.md",
|
||||
"supertonic3": "docs/engines/supertonic3.md",
|
||||
"pockettts": "docs/engines/pockettts.md",
|
||||
"audiocpp": "docs/engines/audio-cpp.md",
|
||||
}
|
||||
|
||||
|
||||
def _engine_docs_url(bid: str) -> str | None:
|
||||
"""Public URL of this engine's doc page, or None when it has none.
|
||||
|
||||
VoiceStudio-owned constant either way: the path comes from the registry
|
||||
above and the base from :mod:`core.links`, so no part of it is derived
|
||||
from an engine probe. That is what lets it cross the public boundary
|
||||
intact (see api.public_engine_metadata).
|
||||
"""
|
||||
path = _ENGINE_DOCS.get(bid)
|
||||
if not path:
|
||||
return None
|
||||
from core import links
|
||||
return f"{links.PROJECT_REPO_BLOB_MAIN}/{path}"
|
||||
|
||||
|
||||
# Short, readable labels for mlx-audio's curated models (#981) — surfaced in
|
||||
# the Model Catalogue → Engines model picker so users see more than a bare key.
|
||||
# Single-sourced here rather than on MLXAudioBackend.CURATED_MODELS itself so
|
||||
@@ -2347,14 +2463,24 @@ def _sidecar_installable_ids() -> frozenset[str]:
|
||||
button into their matrix rows.
|
||||
"""
|
||||
try:
|
||||
from services.sidecar_install import SPECS
|
||||
return frozenset(SPECS)
|
||||
# Host-aware: an engine whose installer cannot work on THIS machine
|
||||
# (dots.tts on Windows, a CUDA-only install on a CPU host) must not get
|
||||
# an Install button that can only fail.
|
||||
from services.sidecar_install import installable_engine_ids
|
||||
return installable_engine_ids()
|
||||
except Exception: # pragma: no cover — defensive only
|
||||
return frozenset()
|
||||
|
||||
|
||||
def list_backends() -> list[dict]:
|
||||
"""Enumerate every registered backend with its availability state.
|
||||
def list_backends(*, include_hidden: bool = False) -> list[dict]:
|
||||
"""Enumerate the engine catalogue with each backend's availability state.
|
||||
|
||||
On MPS, the canonical ``omnivoice`` id already resolves to the killable
|
||||
OmniVoice sidecar. The explicit ``omnivoice-subprocess`` compatibility id
|
||||
is therefore omitted from the normal catalogue so the picker does not
|
||||
advertise two choices with the same runtime behavior. Internal callers
|
||||
that must validate or preserve a stored compatibility id can pass
|
||||
``include_hidden=True``.
|
||||
|
||||
Per-entry shape (ENGINE-05 + ENGINE-06):
|
||||
|
||||
@@ -2368,10 +2494,11 @@ def list_backends() -> list[dict]:
|
||||
# e.g. VoxCPM2's >=2.0.3 upgrade hint)
|
||||
"install_hint": Optional[str],
|
||||
"setup_snippet": Optional[str], # exact `export VAR=...` for path-gated opt-in engines
|
||||
"docs_url": Optional[str], # this engine's doc page (registry-authored constant)
|
||||
"one_click_install": bool, # services.sidecar_install can provision it in-app
|
||||
"last_error": Optional[str], # cached most-recent failure
|
||||
"isolation_mode": "in-process" | "subprocess",
|
||||
"gpu_compat": list[str], # subset of {cuda, rocm, mps, xpu, cpu}
|
||||
"gpu_compat": list[str], # subset of {cuda, rocm, mps, vulkan, xpu, npu, cpu}
|
||||
"supports_cloning": Optional[bool], # True/False from the class attr; None when
|
||||
# model-dependent (property, e.g. mlx-audio)
|
||||
"effective_device": str, # device this engine uses on THIS host
|
||||
@@ -2401,12 +2528,17 @@ def list_backends() -> list[dict]:
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_disk_usage import disk_summary_for
|
||||
from services.engine_evidence import snapshot as execution_snapshot
|
||||
from services.engine_routing import routing_fields
|
||||
caps = detect_host_caps()
|
||||
installable = _sidecar_installable_ids()
|
||||
|
||||
out: list[dict] = []
|
||||
for bid, cls in _REGISTRY.items():
|
||||
if (
|
||||
not include_hidden
|
||||
and caps.family == "mps"
|
||||
and bid == "omnivoice-subprocess"
|
||||
):
|
||||
continue
|
||||
cls = _effective_backend_class(bid, cls, caps.family)
|
||||
try:
|
||||
ok, msg = cls.is_available()
|
||||
@@ -2427,14 +2559,40 @@ def list_backends() -> list[dict]:
|
||||
isolation = "subprocess"
|
||||
else:
|
||||
isolation = "in-process"
|
||||
gpu_compat = getattr(cls, "gpu_compat", ("cpu",))
|
||||
from services.engine_routing import resolve_routing, runtime_compute_profile
|
||||
try:
|
||||
profile = runtime_compute_profile(cls, caps)
|
||||
except Exception:
|
||||
# Runtime-aware native probes remain optional metadata. A broken
|
||||
# provider probe must not take down the engine picker, especially
|
||||
# when availability already explains a missing binary or model.
|
||||
compat = tuple(getattr(cls, "gpu_compat", ("cpu",)))
|
||||
floor = float(getattr(cls, "min_vram_gb", 0.0) or 0.0)
|
||||
profile = {
|
||||
"gpu_compat": compat,
|
||||
"min_vram_gb": floor,
|
||||
**resolve_routing(compat, caps, floor),
|
||||
"runtime_backend": None,
|
||||
"runtime_device_index": None,
|
||||
"runtime_device_name": None,
|
||||
"runtime_hardware_family": None,
|
||||
"runtime_vram_gb": None,
|
||||
"runtime_device_verified": None,
|
||||
}
|
||||
gpu_compat = profile["gpu_compat"]
|
||||
# Cloning capability: same descriptor guard as
|
||||
# cloning_capable_engine_ids() — a class-level getattr on a *property*
|
||||
# (mlx-audio: capability depends on the picked model) returns the
|
||||
# descriptor, not a bool, so report None (= model-dependent) there
|
||||
# instead of an always-truthy false positive.
|
||||
_clone = getattr(cls, "supports_cloning", True)
|
||||
routing = routing_fields(gpu_compat, caps, getattr(cls, "min_vram_gb", 0.0))
|
||||
from core.scrub import scrub_text
|
||||
routing = {
|
||||
"effective_device": profile["effective_device"],
|
||||
"routing_status": profile["routing_status"],
|
||||
"routing_reason": scrub_text(profile["routing_reason"])
|
||||
if profile["routing_reason"] else None,
|
||||
}
|
||||
loaded_instance = None
|
||||
if _active_instance_id == bid:
|
||||
loaded_instance = _active_instance
|
||||
@@ -2455,6 +2613,10 @@ def list_backends() -> list[dict]:
|
||||
"install_hint": _INSTALL_HINTS.get(bid),
|
||||
# Exact `export VAR=...` line for path-gated opt-in engines, or None.
|
||||
"setup_snippet": _SETUP_SNIPPETS.get(bid),
|
||||
# This engine's doc page (#1866). Registry-authored constant, so it
|
||||
# survives api.public_engine_metadata and gives an unavailable row
|
||||
# somewhere to send the user.
|
||||
"docs_url": _engine_docs_url(bid),
|
||||
# True when services.sidecar_install can provision this engine
|
||||
# in-app (Settings renders an Install button instead of leading
|
||||
# with the manual setup snippet).
|
||||
@@ -2464,7 +2626,7 @@ def list_backends() -> list[dict]:
|
||||
"isolation_mode": isolation,
|
||||
"gpu_compat": list(gpu_compat),
|
||||
# effective_device / routing_status / routing_reason (scrubbed):
|
||||
"min_vram_gb": getattr(cls, "min_vram_gb", 0.0) or None,
|
||||
"min_vram_gb": profile["min_vram_gb"] or None,
|
||||
# effective_device / routing_status / routing_reason (scrubbed);
|
||||
# the reason now also carries the under-provisioned-GPU caveat.
|
||||
**routing,
|
||||
@@ -2472,7 +2634,7 @@ def list_backends() -> list[dict]:
|
||||
engine_id=bid,
|
||||
engine_cls=cls,
|
||||
instance=loaded_instance,
|
||||
routing=routing,
|
||||
routing={**profile, **routing},
|
||||
caps=caps,
|
||||
),
|
||||
})
|
||||
@@ -2554,7 +2716,9 @@ def active_routing() -> dict | None:
|
||||
"""
|
||||
try:
|
||||
active = active_backend_id()
|
||||
for b in list_backends():
|
||||
# The MPS picker intentionally hides the redundant compatibility id,
|
||||
# but routing must still describe a saved or environment-pinned id.
|
||||
for b in list_backends(include_hidden=True):
|
||||
if b.get("id") == active:
|
||||
return {
|
||||
"engine": active,
|
||||
@@ -2894,10 +3058,9 @@ async def resolve_generation_backend(
|
||||
raise ValueError(f"TTS engine '{engine_id}' is not available: {_mask_hf_tokens(msg)}")
|
||||
|
||||
from core.device_caps import detect_host_caps
|
||||
from services.engine_routing import resolve_routing
|
||||
routing = resolve_routing(
|
||||
getattr(backend_cls, "gpu_compat", ("cpu",)), detect_host_caps(),
|
||||
getattr(backend_cls, "min_vram_gb", 0.0),
|
||||
from services.engine_routing import runtime_compute_profile_async
|
||||
routing = await runtime_compute_profile_async(
|
||||
backend_cls, detect_host_caps()
|
||||
)
|
||||
if routing["routing_status"] == "unavailable":
|
||||
raise ValueError(routing["routing_reason"])
|
||||
@@ -2935,4 +3098,6 @@ def __getattr__(name: str): # pragma: no cover - exercised via tests
|
||||
return _REGISTRY[name if name in _REGISTRY else None]
|
||||
if name == "IndexTTS2Backend":
|
||||
return _REGISTRY["indextts2"]
|
||||
if name == "AudioCPPBackend":
|
||||
return _REGISTRY["audiocpp"]
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
|
||||
@@ -7,13 +7,15 @@ old silence-only guard missed it (the buzz is loud, not silent) so the garbage
|
||||
was cached and served.
|
||||
|
||||
These tests cover the fix *without the 5 GB model / a GPU*: they drive the pure
|
||||
``_spectral_flatness`` / ``_is_unusable_audio`` helpers with synthetic signals,
|
||||
and assert the render constants didn't regress. The real end-to-end render is
|
||||
verified manually (spectral flatness back in the speech range + Whisper ASR).
|
||||
``_spectral_flatness`` / ``_is_unusable_audio`` helpers with synthetic tones and
|
||||
tracked speech demo renders, and assert the render constants did not regress.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
from pathlib import Path
|
||||
|
||||
import soundfile as sf
|
||||
|
||||
import pytest
|
||||
|
||||
@@ -39,6 +41,13 @@ def _white_noise() -> "torch.Tensor":
|
||||
return 0.5 * (torch.rand(N, generator=g) * 2 - 1)
|
||||
|
||||
|
||||
def _two_tone_buzz() -> "torch.Tensor":
|
||||
"""Two inharmonic partials — the other shape a collapsed render takes."""
|
||||
t = torch.arange(N, dtype=torch.float32) / SR
|
||||
s = torch.sin(2 * math.pi * 180.0 * t) + 0.6 * torch.sin(2 * math.pi * 361.0 * t)
|
||||
return 0.8 * s / s.abs().max()
|
||||
|
||||
|
||||
def _speech_like() -> "torch.Tensor":
|
||||
"""Broadband + harmonic + amplitude-modulated — a coarse stand-in for voiced
|
||||
speech: several harmonics (formant-ish), additive noise (consonants), and a
|
||||
@@ -85,6 +94,71 @@ def test_speech_like_is_usable():
|
||||
assert arch._is_unusable_audio(_speech_like()) is False
|
||||
|
||||
|
||||
# ── Threshold stays between the two things it has to separate ───────────────
|
||||
# Synthetic broadband speech has much higher flatness than real voiced audio.
|
||||
# Measure both sides of the threshold against actual inputs, including the
|
||||
# existing demo renders that the old thresholds rejected.
|
||||
|
||||
|
||||
def test_tonal_ceiling_is_measured_not_assumed():
|
||||
"""Derive the tonal side of the margin instead of trusting a literal.
|
||||
|
||||
A bare constant would keep passing if `_spectral_flatness` stopped scoring
|
||||
tones near zero, so measure the degenerate signals here and require the
|
||||
threshold to clear the worst of them tenfold.
|
||||
"""
|
||||
tones = [
|
||||
arch._spectral_flatness(_pure_tone(80.0)),
|
||||
arch._spectral_flatness(_pure_tone(220.0)),
|
||||
arch._spectral_flatness(_two_tone_buzz()),
|
||||
]
|
||||
assert all(t is not None for t in tones)
|
||||
assert max(tones) * 10 < arch._DEGENERATE_FLATNESS
|
||||
|
||||
|
||||
_SAMPLES = Path(__file__).resolve().parents[1] / "assets" / "samples"
|
||||
_SPEECH_FIXTURES = [
|
||||
"demo_voice.wav",
|
||||
"demo_clone_output.wav",
|
||||
*[f"voice_design/demo_voice_design_{name}.wav" for name in (
|
||||
"audiobook_uk_narrator", "aussie_podcaster", "bedtime_storyteller",
|
||||
"gravelly_villain", "indian_support_agent", "mandarin_sichuan", "us_news_anchor",
|
||||
)],
|
||||
*[f"dictation/{name}.wav" for name in (
|
||||
"en_conversational", "en_technical", "fr_reservation",
|
||||
)],
|
||||
*[f"demo/dubbing/{name}.src.wav" for name in (
|
||||
"source", "dubbed_es", "dubbed_fr", "dubbed_ja", "dubbed_zh",
|
||||
)],
|
||||
]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fixture", _SPEECH_FIXTURES)
|
||||
def test_real_shipped_speech_clears_quality_floor(fixture):
|
||||
# These are existing tracked demo renders, not synthesized stand-ins or
|
||||
# asserted measurements. Loading PCM needs neither a model nor a network.
|
||||
audio, _sample_rate = sf.read(_SAMPLES / fixture, dtype="float32", always_2d=True)
|
||||
speech = torch.from_numpy(audio.T)
|
||||
flatness = arch._spectral_flatness(speech)
|
||||
assert flatness is not None
|
||||
assert flatness > arch._DEGENERATE_FLATNESS * 10
|
||||
assert arch._is_unusable_audio(speech) is False
|
||||
|
||||
|
||||
def test_flatness_is_not_clip_length_dependent():
|
||||
"""Repeating a signal must not change what it measures.
|
||||
|
||||
The whole-clip FFT this replaced failed exactly here: its frequency
|
||||
resolution grew with duration, so the same audio measured 0.0229 at 3 s and
|
||||
~0 at 12 s (100% drift). Framed, the drift is under 0.1%.
|
||||
"""
|
||||
short = _speech_like()
|
||||
long = torch.cat([short] * 4)
|
||||
a, b = arch._spectral_flatness(short), arch._spectral_flatness(long)
|
||||
assert a is not None and b is not None
|
||||
assert abs(a - b) / a < 0.02
|
||||
|
||||
|
||||
# ── Constants didn't regress ────────────────────────────────────────────────
|
||||
def test_preview_render_constants():
|
||||
# 16 steps under-converged on the social script; the fix bumped it.
|
||||
|
||||
@@ -11,6 +11,7 @@ that the error message tells the user what to do.
|
||||
import asyncio
|
||||
import os
|
||||
import sys
|
||||
import threading
|
||||
import time
|
||||
|
||||
import pytest
|
||||
@@ -75,6 +76,62 @@ def test_fast_transcribe_passes_through():
|
||||
pool.shutdown(wait=True)
|
||||
|
||||
|
||||
def test_timeout_defers_abandon_cleanup_until_running_worker_finishes():
|
||||
"""A timed-out native worker may still be reading request-owned inputs."""
|
||||
pool = ThreadPoolExecutor(max_workers=1)
|
||||
started = threading.Event()
|
||||
finish = threading.Event()
|
||||
cleaned = threading.Event()
|
||||
|
||||
def _slow():
|
||||
started.set()
|
||||
finish.wait(timeout=5)
|
||||
assert not cleaned.is_set()
|
||||
return "done"
|
||||
|
||||
async def _go():
|
||||
with pytest.raises(ASRTimeoutError):
|
||||
await run_transcribe_guarded(
|
||||
pool,
|
||||
_slow,
|
||||
what="Convert",
|
||||
timeout=0.05,
|
||||
on_abandon=cleaned.set,
|
||||
)
|
||||
assert started.is_set()
|
||||
assert not cleaned.is_set()
|
||||
finish.set()
|
||||
await asyncio.to_thread(cleaned.wait, 2)
|
||||
assert cleaned.is_set()
|
||||
|
||||
try:
|
||||
asyncio.run(_go())
|
||||
finally:
|
||||
finish.set()
|
||||
pool.shutdown(wait=True)
|
||||
|
||||
|
||||
def test_normal_completion_keeps_abandon_cleanup_with_caller():
|
||||
pool = ThreadPoolExecutor(max_workers=1)
|
||||
cleaned = threading.Event()
|
||||
|
||||
async def _go():
|
||||
result = await run_transcribe_guarded(
|
||||
pool,
|
||||
lambda: "done",
|
||||
what="Convert",
|
||||
timeout=5,
|
||||
on_abandon=cleaned.set,
|
||||
)
|
||||
assert result == "done"
|
||||
assert not cleaned.is_set()
|
||||
|
||||
try:
|
||||
asyncio.run(_go())
|
||||
finally:
|
||||
pool.shutdown(wait=True)
|
||||
|
||||
|
||||
def test_timeout_error_is_a_timeouterror_subclass():
|
||||
# Routers that catch broad TimeoutError (openai_compat) must also catch ours.
|
||||
assert issubclass(ASRTimeoutError, TimeoutError)
|
||||
|
||||
@@ -112,6 +112,42 @@ class TestEnqueue:
|
||||
job = client.get(f"/batch/jobs/{job_id}").json()
|
||||
assert job["filename"] == "test.mp4"
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_upload_is_persisted_in_bounded_chunks(self, batch, tmp_path):
|
||||
class RecordingUpload:
|
||||
def __init__(self):
|
||||
self.read_sizes = []
|
||||
self.remaining = b"video"
|
||||
|
||||
async def read(self, size):
|
||||
self.read_sizes.append(size)
|
||||
chunk, self.remaining = self.remaining[:size], self.remaining[size:]
|
||||
return chunk
|
||||
|
||||
upload = RecordingUpload()
|
||||
destination = tmp_path / "video.mp4"
|
||||
await batch._save_upload(upload, str(destination))
|
||||
|
||||
assert destination.read_bytes() == b"video"
|
||||
assert upload.read_sizes == [batch._UPLOAD_CHUNK_BYTES, batch._UPLOAD_CHUNK_BYTES]
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_failed_upload_removes_partial_file(self, batch, tmp_path):
|
||||
class FailingUpload:
|
||||
calls = 0
|
||||
|
||||
async def read(self, _size):
|
||||
self.calls += 1
|
||||
if self.calls == 1:
|
||||
return b"partial"
|
||||
raise OSError("upload interrupted")
|
||||
|
||||
destination = tmp_path / "video.mp4"
|
||||
with pytest.raises(OSError, match="upload interrupted"):
|
||||
await batch._save_upload(FailingUpload(), str(destination))
|
||||
|
||||
assert not destination.exists()
|
||||
|
||||
|
||||
class TestListJobs:
|
||||
def test_empty(self, client):
|
||||
|
||||
@@ -120,6 +120,7 @@ def test_drain_fd_is_explicitly_inherited_by_wrapper_but_not_operation(monkeypat
|
||||
proc.wait(timeout=5)
|
||||
|
||||
|
||||
@pytest.mark.skipif(os.name != "posix", reason="Unix drain pipe contract")
|
||||
def test_invalid_or_missing_desktop_drain_fd_fails_safe(monkeypatch):
|
||||
monkeypatch.setenv("OMNIVOICE_DESKTOP_CONTAINED", "1")
|
||||
monkeypatch.setenv("OMNIVOICE_DESKTOP_DRAIN_FD", "not-an-fd")
|
||||
@@ -257,3 +258,90 @@ def test_windows_assignment_failure_kills_suspended_unowned_child(monkeypatch):
|
||||
names = [event[0] for event in events]
|
||||
assert names.index("assign") < names.index("terminate") < names.index("kill")
|
||||
assert names.index("kill") < names.index("wait") < names.index("write")
|
||||
|
||||
|
||||
def test_windows_direct_job_owner_assigns_before_resume(monkeypatch):
|
||||
"""Windows skips the extra Python wrapper but retains pre-start Job ownership."""
|
||||
events = []
|
||||
job = 99
|
||||
|
||||
kernel = type("Kernel", (), {})()
|
||||
kernel.AssignProcessToJobObject = _Call(
|
||||
lambda assigned_job, process: events.append(("assign", assigned_job, process)) or True
|
||||
)
|
||||
kernel.TerminateJobObject = _Call(
|
||||
lambda assigned_job, code: events.append(("terminate", assigned_job, code)) or True
|
||||
)
|
||||
kernel.CloseHandle = _Call(
|
||||
lambda handle: events.append(("close", getattr(handle, "value", handle))) or True
|
||||
)
|
||||
monkeypatch.setattr(owned, "_windows_job", lambda: (job, kernel, wintypes))
|
||||
monkeypatch.setattr(
|
||||
owned,
|
||||
"_resume_windows_process",
|
||||
lambda _kernel, _types, pid: events.append(("resume", pid)),
|
||||
)
|
||||
|
||||
class Child:
|
||||
_handle = 77
|
||||
pid = 123
|
||||
args = ["operation.exe"]
|
||||
stdin = None
|
||||
stdout = object()
|
||||
stderr = object()
|
||||
returncode = None
|
||||
|
||||
def poll(self):
|
||||
return self.returncode
|
||||
|
||||
def wait(self, timeout=None):
|
||||
events.append(("wait", timeout))
|
||||
return self.returncode
|
||||
|
||||
def kill(self):
|
||||
events.append(("kill",))
|
||||
|
||||
child = Child()
|
||||
|
||||
def fake_popen(argv, **kwargs):
|
||||
events.append(("spawn", argv, kwargs))
|
||||
return child
|
||||
|
||||
monkeypatch.setattr(owned.subprocess, "Popen", fake_popen)
|
||||
proc = owned._spawn_windows_owned(
|
||||
["operation.exe"],
|
||||
{
|
||||
"env": {
|
||||
"KEEP": "yes",
|
||||
"OMNIVOICE_DESKTOP_CONTAINED": "1",
|
||||
"OMNIVOICE_DESKTOP_DRAIN_FD": "42",
|
||||
},
|
||||
"creationflags": 0x00000200,
|
||||
},
|
||||
)
|
||||
|
||||
names = [event[0] for event in events]
|
||||
assert names[:3] == ["spawn", "assign", "resume"]
|
||||
spawn_argv, spawn_kwargs = events[0][1:]
|
||||
assert spawn_argv == ["operation.exe"]
|
||||
assert spawn_kwargs["creationflags"] == 0x08000204
|
||||
assert spawn_kwargs["env"] == {"KEEP": "yes"}
|
||||
assert proc.stdout is child.stdout
|
||||
|
||||
child.returncode = 0
|
||||
assert proc.poll() == 0
|
||||
assert [event[0] for event in events][-2:] == ["terminate", "close"]
|
||||
|
||||
|
||||
def test_spawn_owned_selects_direct_windows_job_path(monkeypatch):
|
||||
sentinel = object()
|
||||
calls = []
|
||||
monkeypatch.setattr(owned.os, "name", "nt")
|
||||
monkeypatch.setattr(
|
||||
owned,
|
||||
"_spawn_windows_owned",
|
||||
lambda argv, kwargs: calls.append((argv, kwargs)) or sentinel,
|
||||
)
|
||||
|
||||
assert owned.spawn_owned(["sidecar.exe"], text=True) is sentinel
|
||||
assert calls == [(["sidecar.exe"], {"text": True})]
|
||||
|
||||
@@ -15,6 +15,15 @@ import pytest
|
||||
|
||||
from core import contained_subprocess as owned
|
||||
|
||||
# This module simulates macOS by deleting os.waitid, then drives the fallback
|
||||
# with os.waitpid/os.WNOHANG and start_new_session — POSIX-only APIs that
|
||||
# Windows does not have at all (os.WNOHANG raises AttributeError before the
|
||||
# first assertion). CI runs this suite on Linux, so nothing is lost by
|
||||
# skipping; what is gained is a Windows contributor whose checkout runs green.
|
||||
pytestmark = pytest.mark.skipif(
|
||||
os.name != "posix", reason="simulates a POSIX platform without os.waitid"
|
||||
)
|
||||
|
||||
|
||||
def _make_owned(argv):
|
||||
cr, cw = os.pipe()
|
||||
|
||||
@@ -160,6 +160,96 @@ def test_engine_catalogue_reports_effective_mps_isolation(monkeypatch):
|
||||
assert row["isolation_mode"] == "subprocess"
|
||||
|
||||
|
||||
def test_mps_catalogue_hides_redundant_explicit_omnivoice_sidecar(monkeypatch):
|
||||
"""The picker advertises the canonical id, while legacy callers retain both."""
|
||||
from core.device_caps import HostCaps
|
||||
from services import tts_backend
|
||||
|
||||
monkeypatch.setattr(
|
||||
tts_backend,
|
||||
"_REGISTRY",
|
||||
{
|
||||
"omnivoice": OmniVoiceBackend,
|
||||
"omnivoice-subprocess": OmniVoiceSubprocessBackend,
|
||||
},
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
"core.device_caps.detect_host_caps",
|
||||
lambda: HostCaps(family="mps", available_families=("mps", "cpu")),
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
OmniVoiceSubprocessBackend,
|
||||
"is_available",
|
||||
classmethod(lambda cls: (True, "ready")),
|
||||
)
|
||||
|
||||
picker_ids = {item["id"] for item in list_backends()}
|
||||
assert picker_ids == {"omnivoice"}
|
||||
assert get_backend_class("omnivoice") is OmniVoiceMPSSubprocessBackend
|
||||
|
||||
all_ids = {item["id"] for item in list_backends(include_hidden=True)}
|
||||
assert all_ids == {"omnivoice", "omnivoice-subprocess"}
|
||||
assert get_backend_class("omnivoice-subprocess") is OmniVoiceSubprocessBackend
|
||||
|
||||
|
||||
def test_mps_active_routing_preserves_hidden_compatibility_id(monkeypatch):
|
||||
from core.device_caps import HostCaps
|
||||
from services import tts_backend
|
||||
|
||||
monkeypatch.setattr(
|
||||
tts_backend,
|
||||
"_REGISTRY",
|
||||
{"omnivoice-subprocess": OmniVoiceSubprocessBackend},
|
||||
)
|
||||
monkeypatch.setattr(tts_backend, "active_backend_id", lambda: "omnivoice-subprocess")
|
||||
monkeypatch.setattr(
|
||||
"core.device_caps.detect_host_caps",
|
||||
lambda: HostCaps(family="mps", available_families=("mps", "cpu")),
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
OmniVoiceSubprocessBackend,
|
||||
"is_available",
|
||||
classmethod(lambda cls: (True, "ready")),
|
||||
)
|
||||
|
||||
assert tts_backend.active_routing() == {
|
||||
"engine": "omnivoice-subprocess",
|
||||
"available": True,
|
||||
"effective_device": "mps",
|
||||
"routing_status": "accelerated",
|
||||
"routing_reason": None,
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("family", ("cuda", "cpu"))
|
||||
def test_non_mps_catalogue_keeps_explicit_omnivoice_sidecar(monkeypatch, family):
|
||||
from core.device_caps import HostCaps
|
||||
from services import tts_backend
|
||||
|
||||
monkeypatch.setattr(
|
||||
tts_backend,
|
||||
"_REGISTRY",
|
||||
{
|
||||
"omnivoice": OmniVoiceBackend,
|
||||
"omnivoice-subprocess": OmniVoiceSubprocessBackend,
|
||||
},
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
"core.device_caps.detect_host_caps",
|
||||
lambda: HostCaps(family=family, available_families=(family, "cpu")),
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
OmniVoiceSubprocessBackend,
|
||||
"is_available",
|
||||
classmethod(lambda cls: (True, "ready")),
|
||||
)
|
||||
|
||||
assert {item["id"] for item in list_backends()} == {
|
||||
"omnivoice",
|
||||
"omnivoice-subprocess",
|
||||
}
|
||||
|
||||
|
||||
def test_mps_startup_does_not_preload_native_model(monkeypatch):
|
||||
from core.device_caps import HostCaps
|
||||
from services import model_manager
|
||||
@@ -382,7 +472,17 @@ def test_mps_proxy_survives_fatal_child_exit_and_recovers(stub_sidecar, monkeypa
|
||||
try:
|
||||
with pytest.raises(RuntimeError, match="backend is still running"):
|
||||
b.generate("CRASH")
|
||||
assert b._proc is not None and b._proc.poll() is not None
|
||||
assert b._proc is not None
|
||||
# The child called os._exit; the parent raised the moment its pipe hit
|
||||
# EOF, which is BEFORE the OS has reaped the process. Asserting poll()
|
||||
# on the next line is a race the test happened to win on Linux and lost
|
||||
# every time on Windows. Wait for the death instead of assuming it has
|
||||
# already been observed — the claim is that the child is gone, not that
|
||||
# it is gone within one instruction.
|
||||
deadline = time.monotonic() + 5
|
||||
while b._proc.poll() is None and time.monotonic() < deadline:
|
||||
time.sleep(0.02)
|
||||
assert b._proc.poll() is not None, "the crashed sidecar never died"
|
||||
assert b.generate("ok").shape[1] == 24000
|
||||
finally:
|
||||
b.shutdown()
|
||||
@@ -523,3 +623,163 @@ def test_generation_proxy_forwards_native_controls_and_seed():
|
||||
"class_temperature": 0.8,
|
||||
"seed": 321,
|
||||
})]
|
||||
|
||||
|
||||
def test_timeout_reaps_captured_process_before_recv_returns(monkeypatch):
|
||||
import threading
|
||||
|
||||
class Process:
|
||||
def __init__(self):
|
||||
self.killed = threading.Event()
|
||||
self.reaped = False
|
||||
self.wait_entered = threading.Event()
|
||||
self.release_wait = threading.Event()
|
||||
|
||||
def kill(self):
|
||||
self.killed.set() # EOF may arrive before the process is reaped.
|
||||
|
||||
def wait(self, timeout):
|
||||
assert timeout is not None
|
||||
self.wait_entered.set()
|
||||
assert self.release_wait.wait(2)
|
||||
self.reaped = True
|
||||
return -9
|
||||
|
||||
proc = Process()
|
||||
backend = OmniVoiceSubprocessBackend()
|
||||
backend._proc = proc
|
||||
|
||||
def recv():
|
||||
assert proc.killed.wait(2)
|
||||
return None
|
||||
|
||||
monkeypatch.setattr(backend, '_recv', recv)
|
||||
returned = threading.Event()
|
||||
results = []
|
||||
|
||||
def receive():
|
||||
results.append(backend._recv_with_timeout(0.01))
|
||||
returned.set()
|
||||
|
||||
reader = threading.Thread(target=receive)
|
||||
reader.start()
|
||||
try:
|
||||
assert proc.wait_entered.wait(2)
|
||||
assert not returned.wait(0.05), "EOF must not release the caller before process cleanup"
|
||||
finally:
|
||||
proc.release_wait.set()
|
||||
reader.join(2)
|
||||
backend._proc = None
|
||||
assert not reader.is_alive()
|
||||
assert returned.is_set()
|
||||
assert results == [None]
|
||||
assert proc.reaped
|
||||
|
||||
|
||||
def test_timeout_never_kills_a_replacement_process(monkeypatch):
|
||||
from unittest.mock import Mock
|
||||
import services.subprocess_backend as module
|
||||
|
||||
class ManualTimer:
|
||||
def __init__(self, _timeout, callback, args=()):
|
||||
self.callback = lambda: callback(*args)
|
||||
self.daemon = False
|
||||
|
||||
def start(self):
|
||||
pass
|
||||
|
||||
def cancel(self):
|
||||
pass
|
||||
|
||||
def join(self):
|
||||
pass
|
||||
|
||||
timers = []
|
||||
def timer(*args, **kwargs):
|
||||
result = ManualTimer(*args, **kwargs)
|
||||
timers.append(result)
|
||||
return result
|
||||
|
||||
monkeypatch.setattr(module.threading, 'Timer', timer)
|
||||
backend = OmniVoiceSubprocessBackend()
|
||||
original, replacement = Mock(), Mock()
|
||||
backend._proc = original
|
||||
|
||||
def recv():
|
||||
backend._proc = replacement
|
||||
timers[0].callback()
|
||||
return None
|
||||
|
||||
monkeypatch.setattr(backend, '_recv', recv)
|
||||
try:
|
||||
backend._recv_with_timeout(1)
|
||||
original.kill.assert_called_once()
|
||||
replacement.kill.assert_not_called()
|
||||
finally:
|
||||
backend._proc = None
|
||||
|
||||
|
||||
@pytest.mark.parametrize("failure", ["wait", "kill"])
|
||||
def test_timeout_quarantine_blocks_reuse_and_retains_cleanup_handle(failure):
|
||||
class StuckProcess:
|
||||
stdin = None
|
||||
def __init__(self):
|
||||
self.exited = False
|
||||
self.kill_calls = 0
|
||||
def poll(self):
|
||||
return 0 if self.exited else None
|
||||
def kill(self):
|
||||
self.kill_calls += 1
|
||||
if failure == "kill" and not self.exited:
|
||||
raise PermissionError("kill failed")
|
||||
def terminate(self):
|
||||
pass
|
||||
def wait(self, timeout):
|
||||
if not self.exited:
|
||||
raise subprocess.TimeoutExpired("stuck-sidecar", timeout)
|
||||
return 0
|
||||
|
||||
backend = OmniVoiceSubprocessBackend()
|
||||
proc = StuckProcess()
|
||||
backend._proc = proc
|
||||
try:
|
||||
backend._timeout_kill(proc)
|
||||
with pytest.raises(RuntimeError, match="still stopping"):
|
||||
backend._spawn()
|
||||
backend.shutdown()
|
||||
# Even after shutdown clears the current slot, ownership survives;
|
||||
# retry must not silently start a second process next to this one.
|
||||
before = proc.kill_calls
|
||||
with pytest.raises(RuntimeError, match="still stopping"):
|
||||
backend._spawn()
|
||||
assert proc.kill_calls > before
|
||||
finally:
|
||||
proc.exited = True
|
||||
backend.shutdown()
|
||||
|
||||
|
||||
def test_timeout_quarantine_does_not_clear_or_kill_replacement():
|
||||
from unittest.mock import Mock
|
||||
backend = OmniVoiceSubprocessBackend()
|
||||
original = Mock()
|
||||
original.wait.side_effect = subprocess.TimeoutExpired("old-sidecar", 2)
|
||||
replacement = Mock()
|
||||
replacement.poll.return_value = None
|
||||
backend._proc = replacement
|
||||
try:
|
||||
backend._timeout_kill(original)
|
||||
with pytest.raises(RuntimeError, match="still stopping"):
|
||||
backend._spawn()
|
||||
assert backend._proc is replacement
|
||||
replacement.kill.assert_not_called()
|
||||
# Once the captured owner is reaped, reuse of the healthy replacement
|
||||
# is allowed without starting or terminating another process.
|
||||
original.wait.side_effect = None
|
||||
original.wait.return_value = 0
|
||||
backend._spawn()
|
||||
assert backend._proc is replacement
|
||||
replacement.kill.assert_not_called()
|
||||
finally:
|
||||
original.wait.side_effect = None
|
||||
backend._proc = None
|
||||
backend.shutdown()
|
||||
|
||||
@@ -113,6 +113,93 @@ def test_unclean_shutdown_yields_crash_record(sentinel_env, monkeypatch):
|
||||
assert acked is False, "a fresh crash record must be unacknowledged"
|
||||
|
||||
|
||||
def test_lifespan_clears_sentinel_even_if_later_shutdown_raises(monkeypatch, tmp_path):
|
||||
"""THE #1895 regression: before this fix, ``clear_sentinel()`` was the
|
||||
LAST statement of ``main.py``'s lifespan shutdown, behind ~50s of bounded
|
||||
waits plus model unload / ``free_vram()`` / ``gc.collect()`` / httpx
|
||||
close. The desktop shell's quit path grants only a 2s grace before
|
||||
SIGKILL (``frontend/src-tauri/src/bootstrap.rs``
|
||||
``terminate_process_tree``), and Windows grants no graceful phase at all
|
||||
(``tools.rs``) — nowhere near enough, so a deliberate, clean quit
|
||||
routinely got killed before reaching that last line, leaving the
|
||||
sentinel behind for the NEXT startup to misreport as "did not shut down
|
||||
cleanly... likely crashed".
|
||||
|
||||
Simulates that class of interruption without an actual SIGKILL: a later
|
||||
shutdown step (``model_loads_begin_shutdown()``, called unguarded well
|
||||
after the sentinel clear) raises, so nothing past it in the shutdown
|
||||
body ever runs — for this purpose, the same effect as being killed
|
||||
mid-teardown.
|
||||
|
||||
Fail-before/pass-after: with ``clear_sentinel()`` moved to the TOP of
|
||||
the shutdown block (immediately after ``yield``), the sentinel is
|
||||
already gone by the time this raise happens, so the next startup must
|
||||
not fabricate a crash record.
|
||||
"""
|
||||
import asyncio
|
||||
from fastapi import FastAPI
|
||||
|
||||
# Fresh `main`/`core`/`api`/`services` import, mirroring
|
||||
# tests/test_model_load_shutdown.py's `_reimported_backend_modules`: a
|
||||
# sibling suite may have purged these names from sys.modules, leaving a
|
||||
# collection-time alias stale. Purging and re-importing here makes this
|
||||
# test self-consistent in isolation, not dependent on suite order.
|
||||
purge_names = ("main", "core", "api", "services")
|
||||
purge_prefixes = ("core.", "api.", "services.")
|
||||
saved = {
|
||||
name: mod for name, mod in sys.modules.items()
|
||||
if name in purge_names or name.startswith(purge_prefixes)
|
||||
}
|
||||
|
||||
def _purge():
|
||||
for name in [
|
||||
n for n in sys.modules
|
||||
if n in purge_names or n.startswith(purge_prefixes)
|
||||
]:
|
||||
sys.modules.pop(name, None)
|
||||
|
||||
_purge()
|
||||
try:
|
||||
import main as main_mod
|
||||
from core import run_sentinel as fresh_run_sentinel
|
||||
|
||||
monkeypatch.setattr(
|
||||
fresh_run_sentinel, "SENTINEL_PATH", str(tmp_path / "run_sentinel.json")
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
fresh_run_sentinel, "CRASH_RECORD_PATH", str(tmp_path / "last_run_crash.json")
|
||||
)
|
||||
monkeypatch.setattr(
|
||||
fresh_run_sentinel, "LOG_PATH", str(tmp_path / "omnivoice.log")
|
||||
)
|
||||
fresh_run_sentinel._reset_for_tests()
|
||||
|
||||
def _boom():
|
||||
raise RuntimeError("simulated kill: interrupted after the early clear")
|
||||
|
||||
monkeypatch.setattr(main_mod, "model_loads_begin_shutdown", _boom)
|
||||
|
||||
async def scenario():
|
||||
app = FastAPI()
|
||||
async with main_mod.lifespan(app):
|
||||
pass
|
||||
|
||||
with pytest.raises(RuntimeError, match="simulated kill"):
|
||||
asyncio.run(scenario())
|
||||
|
||||
assert not os.path.exists(fresh_run_sentinel.SENTINEL_PATH), (
|
||||
"sentinel must already be cleared even though a later shutdown "
|
||||
"step raised before ever reaching the old clear-sentinel line"
|
||||
)
|
||||
assert fresh_run_sentinel.detect_unclean_shutdown() is None, (
|
||||
"a deliberate quit interrupted after the early clear must never "
|
||||
"be reported as a crash on the next startup"
|
||||
)
|
||||
finally:
|
||||
_purge()
|
||||
sys.modules.update(saved)
|
||||
|
||||
|
||||
def test_live_pid_means_second_instance_not_a_crash(sentinel_env):
|
||||
"""A sentinel owned by a LIVE process is a concurrent second instance
|
||||
sharing DATA_DIR — never a crash, and we must not take over or delete
|
||||
|
||||
@@ -23,6 +23,8 @@ from __future__ import annotations
|
||||
import logging
|
||||
from typing import Optional
|
||||
|
||||
from worker.capacity import derive_concurrency
|
||||
|
||||
logger = logging.getLogger("omnivoice.worker")
|
||||
|
||||
# gpu_compat families that mean "this would run on the CPU here", which is
|
||||
@@ -74,6 +76,12 @@ def discover(*, include_unavailable: bool = False) -> list[dict]:
|
||||
gpu_compat = set(entry.get("gpu_compat") or [])
|
||||
repo_ids = repo_ids_for(entry)
|
||||
downloaded = _downloaded(repo_ids)
|
||||
runtime_vram_gb = (entry.get("execution_evidence") or {}).get(
|
||||
"runtime_vram_gb"
|
||||
)
|
||||
engine_free_bytes = free_bytes if runtime_vram_gb is None else int(
|
||||
float(runtime_vram_gb or 0.0) * 1024**3
|
||||
)
|
||||
discovered.append(
|
||||
{
|
||||
"engine": engine_id,
|
||||
@@ -97,7 +105,17 @@ def discover(*, include_unavailable: bool = False) -> list[dict]:
|
||||
"min_memory_bytes": int(float(entry.get("min_vram_gb") or 0) * 1024**3),
|
||||
"precision": "",
|
||||
"backend": entry.get("effective_device") or family,
|
||||
"free_memory_bytes": free_bytes,
|
||||
"free_memory_bytes": engine_free_bytes,
|
||||
# A native provider that works independently of torch may not
|
||||
# expose memory telemetry. Unknown capacity still gets one
|
||||
# serial slot; zero must not turn a working Vulkan engine into
|
||||
# an unschedulable capability.
|
||||
"derived_concurrency": 1
|
||||
if (
|
||||
runtime_vram_gb is not None
|
||||
and float(runtime_vram_gb or 0.0) <= 0
|
||||
and routing == "accelerated"
|
||||
) else 0,
|
||||
# Capability is not acceleration: an engine present but routed
|
||||
# to the CPU here should not be preferred for GPU work.
|
||||
"cpu_fallback": routing in ("cpu_fallback", "cpu_only")
|
||||
@@ -291,9 +309,18 @@ def max_concurrent_tasks(capabilities: Optional[list[dict]] = None) -> int:
|
||||
caps = capabilities if capabilities is not None else discover()
|
||||
if not caps:
|
||||
return 1
|
||||
derived = [int(c.get("derived_concurrency") or 0) for c in caps]
|
||||
positive = [d for d in derived if d > 0]
|
||||
return min(positive) if positive else 1
|
||||
derived: list[int] = []
|
||||
for cap in caps:
|
||||
concurrency = int(cap.get("derived_concurrency") or 0)
|
||||
if concurrency <= 0:
|
||||
concurrency = derive_concurrency(
|
||||
backend=str(cap.get("backend") or ""),
|
||||
free_memory_bytes=int(cap.get("free_memory_bytes") or 0),
|
||||
min_model_bytes=int(cap.get("min_memory_bytes") or 0),
|
||||
compiled=bool(cap.get("compiled")),
|
||||
)
|
||||
derived.append(concurrency)
|
||||
return min(derived)
|
||||
|
||||
|
||||
__all__ = [
|
||||
|
||||
@@ -97,18 +97,16 @@ def derive_concurrency(
|
||||
) -> int:
|
||||
"""How many jobs of this model may run at once on this worker.
|
||||
|
||||
Returns 0 when the model cannot run here at all — a capability mismatch,
|
||||
which the scheduler must treat as "send it elsewhere", never as a worker
|
||||
fault.
|
||||
Under-provisioned accelerators remain usable with one serial slot and the
|
||||
longer CPU-class deadline. Zero memory means telemetry is unknown, not
|
||||
that the engine cannot run.
|
||||
"""
|
||||
family = (backend or "").strip().lower()
|
||||
if min_model_bytes and free_memory_bytes < min_model_bytes:
|
||||
return 0
|
||||
if compiled:
|
||||
# Thread-affinity pinning (#315). One job, always.
|
||||
return 1
|
||||
if family in _ALWAYS_SERIAL:
|
||||
return 1 if (not min_model_bytes or free_memory_bytes >= min_model_bytes) else 0
|
||||
return 1
|
||||
budget = max(min_model_bytes, _VRAM_PER_JOB_BYTES)
|
||||
if budget <= 0:
|
||||
return 1
|
||||
|
||||
@@ -127,16 +127,31 @@ class Deadlines:
|
||||
|
||||
|
||||
def _base_execution_seconds(
|
||||
text: Optional[str], *, execution_device: Optional[str] = None
|
||||
text: Optional[str], *, execution_device: Optional[str] = None,
|
||||
under_provisioned: bool = False,
|
||||
) -> float:
|
||||
"""Delegate to model_manager's budget; fall back to its formula.
|
||||
|
||||
The lazy import keeps this module usable in a process that has no torch —
|
||||
the control plane schedules work it never executes.
|
||||
|
||||
``under_provisioned`` is the worker's own verdict that its card sits below
|
||||
the engine's declared VRAM floor (``ConnectedWorker.under_provisioned``).
|
||||
It floors the budget at what the same job would get on a CPU, because that
|
||||
is what a card paging to system RAM performs like (#1804). Derived by asking
|
||||
for the CPU budget rather than by probing VRAM here: this process is the
|
||||
control plane, and its hardware is not the worker's.
|
||||
"""
|
||||
target_device = str(execution_device or "cpu").lower()
|
||||
if target_device not in {"cpu", "cuda", "mps", "mlx", "directml", "rocm", "xpu"}:
|
||||
if target_device not in {
|
||||
"cpu", "cuda", "mps", "mlx", "directml", "rocm", "vulkan", "xpu",
|
||||
}:
|
||||
target_device = "cpu"
|
||||
if under_provisioned and target_device != "cpu":
|
||||
return max(
|
||||
_base_execution_seconds(text, execution_device=target_device),
|
||||
_base_execution_seconds(text, execution_device="cpu"),
|
||||
)
|
||||
try:
|
||||
from services import model_manager # noqa: PLC0415 — intentionally lazy
|
||||
|
||||
@@ -171,6 +186,7 @@ def for_task(
|
||||
model_downloaded: bool = True,
|
||||
input_seconds: float = 0.0,
|
||||
execution_device: Optional[str] = None,
|
||||
under_provisioned: bool = False,
|
||||
) -> Deadlines:
|
||||
"""Compute the deadlines for one attempt.
|
||||
|
||||
@@ -183,7 +199,8 @@ def for_task(
|
||||
multiplier, grace = _PROFILE[op]
|
||||
|
||||
execution = _base_execution_seconds(
|
||||
text, execution_device=execution_device
|
||||
text, execution_device=execution_device,
|
||||
under_provisioned=under_provisioned,
|
||||
) * multiplier
|
||||
# Media-length operations scale on duration, not characters.
|
||||
if input_seconds > 0:
|
||||
|
||||
@@ -32,6 +32,7 @@ import uuid
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Iterable, Optional
|
||||
|
||||
from worker.deadlines import Deadlines
|
||||
from worker.clock import resolve
|
||||
from worker.errors import ErrorClass, WorkerError
|
||||
|
||||
@@ -224,6 +225,9 @@ class Attempt:
|
||||
stage: str = ""
|
||||
error: Optional[WorkerError] = None
|
||||
|
||||
# Snapshot the lease policy granted at dispatch, including after restart.
|
||||
deadlines: Optional[Deadlines] = None
|
||||
|
||||
def matches(self, *, session_epoch: Optional[int] = None) -> bool:
|
||||
"""Fence check: reject messages from a superseded session."""
|
||||
if session_epoch is None:
|
||||
|
||||
+53
-20
@@ -34,7 +34,7 @@ _HEARTBEAT_MISS_SECONDS = 90.0
|
||||
# enough that one slow answer cannot move it.
|
||||
_LATENCY_WINDOW = 5
|
||||
_KNOWN_EXECUTION_DEVICES = frozenset(
|
||||
{"cpu", "cuda", "mps", "mlx", "directml", "rocm", "xpu"}
|
||||
{"cpu", "cuda", "mps", "mlx", "directml", "rocm", "vulkan", "xpu"}
|
||||
)
|
||||
|
||||
|
||||
@@ -93,12 +93,11 @@ class ConnectedWorker:
|
||||
return "busy"
|
||||
return "ready"
|
||||
|
||||
def supports(self, engine: str, model_id: str, operation: str) -> bool:
|
||||
"""Can this worker run this work at all?
|
||||
def _capability_for(self, engine: str, model_id: str, operation: str):
|
||||
"""The advertised capability this task would actually be run by.
|
||||
|
||||
``supported`` alone is not enough — an engine whose weights are not on
|
||||
disk cannot start without a download, and one that is not installed
|
||||
cannot start at all. Both are capability mismatches, not failures.
|
||||
One selection rule, so the answers below cannot describe different
|
||||
capabilities of the same worker.
|
||||
"""
|
||||
for cap in self.record.capabilities:
|
||||
if cap.get("engine") != engine:
|
||||
@@ -107,23 +106,57 @@ class ConnectedWorker:
|
||||
continue
|
||||
if operation and operation not in (cap.get("operations") or [operation]):
|
||||
continue
|
||||
return bool(cap.get("supported")) and bool(cap.get("installed", True))
|
||||
return False
|
||||
return cap
|
||||
return None
|
||||
|
||||
def supports(self, engine: str, model_id: str, operation: str) -> bool:
|
||||
"""Can this worker run this work at all?
|
||||
|
||||
``supported`` alone is not enough — an engine whose weights are not on
|
||||
disk cannot start without a download, and one that is not installed
|
||||
cannot start at all. Both are capability mismatches, not failures.
|
||||
"""
|
||||
cap = self._capability_for(engine, model_id, operation)
|
||||
if cap is None:
|
||||
return False
|
||||
return bool(cap.get("supported")) and bool(cap.get("installed", True))
|
||||
|
||||
def execution_device(self, engine: str, model_id: str, operation: str) -> str:
|
||||
"""Device used by the exact capability selected for this task."""
|
||||
for cap in self.record.capabilities:
|
||||
if cap.get("engine") != engine:
|
||||
continue
|
||||
if model_id and cap.get("model_id") not in (model_id, "", None):
|
||||
continue
|
||||
if operation and operation not in (cap.get("operations") or [operation]):
|
||||
continue
|
||||
if cap.get("cpu_fallback"):
|
||||
return "cpu"
|
||||
backend = str(cap.get("backend") or "").lower()
|
||||
return backend if backend in _KNOWN_EXECUTION_DEVICES else "cpu"
|
||||
return "cpu"
|
||||
cap = self._capability_for(engine, model_id, operation)
|
||||
if cap is None:
|
||||
return "cpu"
|
||||
if cap.get("cpu_fallback"):
|
||||
return "cpu"
|
||||
backend = str(cap.get("backend") or "").lower()
|
||||
return backend if backend in _KNOWN_EXECUTION_DEVICES else "cpu"
|
||||
|
||||
def under_provisioned(self, engine: str, model_id: str, operation: str) -> bool:
|
||||
"""Is this worker's GPU below the engine's declared VRAM floor?
|
||||
|
||||
The remote half of #1804. A card under the floor pages to system RAM and
|
||||
renders slower than a CPU, so it must not be given the shorter
|
||||
accelerated deadline. Decided from the two figures the WORKER itself
|
||||
advertises (``free_memory_bytes`` / ``min_memory_bytes``, both set in
|
||||
``worker/capabilities.py``): the control plane's own VRAM says nothing
|
||||
about the machine that will run the job, so
|
||||
``engine_routing.under_provisioned_vram`` — which probes THIS host —
|
||||
cannot answer for a remote worker.
|
||||
|
||||
Same rules as that predicate otherwise: dedicated-VRAM devices only
|
||||
(unified memory is not a comparable pool), and a zero on either side
|
||||
means "unknown", never "too small".
|
||||
"""
|
||||
cap = self._capability_for(engine, model_id, operation)
|
||||
if cap is None or cap.get("cpu_fallback"):
|
||||
return False
|
||||
if str(cap.get("backend") or "").lower() not in (
|
||||
"cuda", "rocm", "vulkan",
|
||||
):
|
||||
return False
|
||||
floor = int(cap.get("min_memory_bytes") or 0)
|
||||
have = int(cap.get("free_memory_bytes") or 0)
|
||||
return floor > 0 and 0 < have < floor
|
||||
|
||||
def is_warm(self, engine: str, model_id: str) -> bool:
|
||||
return self.capacity.is_resident(engine, model_id)
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -93,7 +93,7 @@ class HostInfo(_message.Message):
|
||||
def __init__(self, hostname: _Optional[str] = ..., os: _Optional[str] = ..., arch: _Optional[str] = ..., worker_version: _Optional[str] = ..., cpu_count: _Optional[int] = ..., system_memory_bytes: _Optional[int] = ..., gpus: _Optional[_Iterable[_Union[GpuInfo, _Mapping]]] = ...) -> None: ...
|
||||
|
||||
class ModelCapability(_message.Message):
|
||||
__slots__ = ("engine", "model_id", "operations", "supported", "installed", "downloaded", "resident", "min_memory_bytes", "precision", "derived_concurrency", "cpu_fallback", "repo_ids", "display_name")
|
||||
__slots__ = ("engine", "model_id", "operations", "supported", "installed", "downloaded", "resident", "min_memory_bytes", "precision", "derived_concurrency", "cpu_fallback", "repo_ids", "display_name", "backend", "free_memory_bytes")
|
||||
ENGINE_FIELD_NUMBER: _ClassVar[int]
|
||||
MODEL_ID_FIELD_NUMBER: _ClassVar[int]
|
||||
OPERATIONS_FIELD_NUMBER: _ClassVar[int]
|
||||
@@ -107,6 +107,8 @@ class ModelCapability(_message.Message):
|
||||
CPU_FALLBACK_FIELD_NUMBER: _ClassVar[int]
|
||||
REPO_IDS_FIELD_NUMBER: _ClassVar[int]
|
||||
DISPLAY_NAME_FIELD_NUMBER: _ClassVar[int]
|
||||
BACKEND_FIELD_NUMBER: _ClassVar[int]
|
||||
FREE_MEMORY_BYTES_FIELD_NUMBER: _ClassVar[int]
|
||||
engine: str
|
||||
model_id: str
|
||||
operations: _containers.RepeatedScalarFieldContainer[str]
|
||||
@@ -120,7 +122,9 @@ class ModelCapability(_message.Message):
|
||||
cpu_fallback: bool
|
||||
repo_ids: _containers.RepeatedScalarFieldContainer[str]
|
||||
display_name: str
|
||||
def __init__(self, engine: _Optional[str] = ..., model_id: _Optional[str] = ..., operations: _Optional[_Iterable[str]] = ..., supported: _Optional[bool] = ..., installed: _Optional[bool] = ..., downloaded: _Optional[bool] = ..., resident: _Optional[bool] = ..., min_memory_bytes: _Optional[int] = ..., precision: _Optional[str] = ..., derived_concurrency: _Optional[int] = ..., cpu_fallback: _Optional[bool] = ..., repo_ids: _Optional[_Iterable[str]] = ..., display_name: _Optional[str] = ...) -> None: ...
|
||||
backend: str
|
||||
free_memory_bytes: int
|
||||
def __init__(self, engine: _Optional[str] = ..., model_id: _Optional[str] = ..., operations: _Optional[_Iterable[str]] = ..., supported: _Optional[bool] = ..., installed: _Optional[bool] = ..., downloaded: _Optional[bool] = ..., resident: _Optional[bool] = ..., min_memory_bytes: _Optional[int] = ..., precision: _Optional[str] = ..., derived_concurrency: _Optional[int] = ..., cpu_fallback: _Optional[bool] = ..., repo_ids: _Optional[_Iterable[str]] = ..., display_name: _Optional[str] = ..., backend: _Optional[str] = ..., free_memory_bytes: _Optional[int] = ...) -> None: ...
|
||||
|
||||
class RegisterRequest(_message.Message):
|
||||
__slots__ = ("envelope", "protocol_version_min", "protocol_version_max", "enrollment_token", "worker_id", "public_key", "challenge_signature", "challenge", "host", "capabilities", "max_concurrent_tasks", "in_flight", "completed_unacked", "key_id", "nonce", "labels", "features")
|
||||
|
||||
@@ -168,6 +168,11 @@ message ModelCapability {
|
||||
// Human-readable UI label. Never use this as a scheduling or residency key;
|
||||
// unlike model_id it may change with ordinary copy edits.
|
||||
string display_name = 13;
|
||||
// Per-engine runtime routing. Native engines can select a provider that is
|
||||
// independent of the worker's global torch device.
|
||||
string backend = 14;
|
||||
// Memory measured for that exact selected provider/device; zero is unknown.
|
||||
uint64 free_memory_bytes = 15;
|
||||
}
|
||||
|
||||
message RegisterRequest {
|
||||
|
||||
@@ -652,7 +652,11 @@ class Scheduler:
|
||||
execution_device=worker.execution_device(
|
||||
task.engine, task.model_id, task.operation
|
||||
),
|
||||
under_provisioned=worker.under_provisioned(
|
||||
task.engine, task.model_id, task.operation
|
||||
),
|
||||
)
|
||||
attempt.deadlines = budget
|
||||
attempt.renew_lease(budget.accept_seconds, now=now)
|
||||
self._save(task, now=now)
|
||||
self._emit("assigned", task)
|
||||
@@ -1295,7 +1299,13 @@ class Scheduler:
|
||||
|
||||
def _budget_for(self, task: Task) -> deadline_policy.Deadlines:
|
||||
attempt = task.active_attempt
|
||||
if attempt is not None and attempt.deadlines is not None:
|
||||
return attempt.deadlines
|
||||
# A legacy attempt has no recorded device once its worker is absent.
|
||||
# Cover both configured device classes instead of assuming the shorter
|
||||
# CPU budget; for_task floors an under-provisioned GPU at max(CPU, GPU).
|
||||
worker = self.pool.get(attempt.worker_id) if attempt else None
|
||||
unknown_legacy_worker = attempt is not None and worker is None
|
||||
return deadline_policy.for_task(
|
||||
task.operation,
|
||||
text=task.params.get("text"),
|
||||
@@ -1303,7 +1313,13 @@ class Scheduler:
|
||||
input_seconds=float(task.params.get("input_seconds") or 0.0),
|
||||
execution_device=(
|
||||
worker.execution_device(task.engine, task.model_id, task.operation)
|
||||
if worker else None
|
||||
if worker else ("cuda" if unknown_legacy_worker else None)
|
||||
),
|
||||
under_provisioned=unknown_legacy_worker or bool(
|
||||
worker
|
||||
and worker.under_provisioned(
|
||||
task.engine, task.model_id, task.operation
|
||||
)
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
@@ -33,6 +33,7 @@ from core.db import db_conn
|
||||
from core.path_security import UnsafePath, resolve_within, safe_filename
|
||||
from worker.clock import resolve
|
||||
from worker.errors import ErrorClass, WorkerError
|
||||
from worker.deadlines import Deadlines
|
||||
from worker.lifecycle import Attempt, AttemptState, PriorityClass, Task, TaskState
|
||||
|
||||
logger = logging.getLogger("omnivoice.worker")
|
||||
@@ -67,6 +68,8 @@ def _row_to_attempt(row) -> Attempt:
|
||||
state=AttemptState(row["state"]),
|
||||
created_at=float(row["created_at"]),
|
||||
)
|
||||
if row["deadlines_json"]:
|
||||
attempt.deadlines = Deadlines(**json.loads(row["deadlines_json"]))
|
||||
attempt.accepted_at = row["accepted_at"]
|
||||
attempt.started_at = row["started_at"]
|
||||
attempt.finished_at = row["finished_at"]
|
||||
@@ -131,6 +134,32 @@ INPUT_PARAM_KEYS: tuple[str, ...] = (
|
||||
# task records what was staged for it. The record is what makes the purge
|
||||
# exact: an input is deletable only when no surviving task still refers to it.
|
||||
INPUTS_DIRNAME = "inputs"
|
||||
|
||||
|
||||
def artifact_id_for(name: str) -> str:
|
||||
"""The id a staged input is known by, everywhere.
|
||||
|
||||
This is a PROTOCOL identifier, not a local path: it is persisted in
|
||||
``params_json``, handed to remote workers over gRPC, and matched against
|
||||
what a later sweep finds on disk. ``os.path.join`` made it OS-specific, so
|
||||
a Windows control plane stored and shipped ``inputs\\<sha>.wav`` — which a
|
||||
Linux worker cannot resolve, and which stops matching the moment the same
|
||||
data directory is opened on another OS. Always ``/``; ``resolve_within``
|
||||
already treats both separators as structural, so resolution is unaffected.
|
||||
"""
|
||||
return f"{INPUTS_DIRNAME}/{name}"
|
||||
|
||||
|
||||
def normalize_artifact_id(artifact_id: str) -> str:
|
||||
"""Compare ids written by any host on equal terms.
|
||||
|
||||
Rows staged by a Windows control plane before this was canonicalised carry
|
||||
a backslash. The sweeper decides whether a file on disk is still
|
||||
referenced by comparing ids, so without this an upgraded install would
|
||||
read every legacy row as unreferenced and delete inputs that surviving
|
||||
tasks still point at.
|
||||
"""
|
||||
return (artifact_id or "").replace("\\", "/")
|
||||
INPUTS_PARAM_KEY = "inputs"
|
||||
|
||||
_HASH_CHUNK_BYTES = 1024 * 1024
|
||||
@@ -289,7 +318,7 @@ def stage_input(
|
||||
f"Could not read the task input {source!r}: {exc}"
|
||||
) from exc
|
||||
|
||||
artifact_id = os.path.join(INPUTS_DIRNAME, f"{digest}{_extension(source)}")
|
||||
artifact_id = artifact_id_for(f"{digest}{_extension(source)}")
|
||||
try:
|
||||
destination = resolve_within(base, artifact_id)
|
||||
except UnsafePath as exc: # pragma: no cover — the id is ours, hex only
|
||||
@@ -470,7 +499,7 @@ def _referenced_artifacts(conn) -> set[str]:
|
||||
continue
|
||||
for entry in entries:
|
||||
if isinstance(entry, dict) and entry.get("artifact_id"):
|
||||
referenced.add(str(entry["artifact_id"]))
|
||||
referenced.add(normalize_artifact_id(str(entry["artifact_id"])))
|
||||
return referenced
|
||||
|
||||
|
||||
@@ -557,8 +586,7 @@ def purge_artifacts(
|
||||
except OSError:
|
||||
return removed
|
||||
for name in names:
|
||||
artifact_id = os.path.join(INPUTS_DIRNAME, name)
|
||||
if artifact_id in referenced:
|
||||
if artifact_id_for(name) in referenced:
|
||||
continue
|
||||
path = os.path.join(inputs_dir, name)
|
||||
try:
|
||||
@@ -635,12 +663,13 @@ def _upsert_attempts(conn, task: Task) -> None:
|
||||
"INSERT INTO remote_task_attempts "
|
||||
"(id, task_id, worker_id, session_epoch, attempt_number, state, progress, stage, "
|
||||
" error_json, created_at, accepted_at, started_at, finished_at, lease_expires_at, "
|
||||
" grace_expires_at) "
|
||||
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) "
|
||||
" grace_expires_at, deadlines_json) "
|
||||
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?) "
|
||||
"ON CONFLICT(id) DO UPDATE SET state=excluded.state, progress=excluded.progress, "
|
||||
" stage=excluded.stage, error_json=excluded.error_json, accepted_at=excluded.accepted_at, "
|
||||
" started_at=excluded.started_at, finished_at=excluded.finished_at, "
|
||||
" lease_expires_at=excluded.lease_expires_at, grace_expires_at=excluded.grace_expires_at",
|
||||
" lease_expires_at=excluded.lease_expires_at, grace_expires_at=excluded.grace_expires_at, "
|
||||
" deadlines_json=excluded.deadlines_json",
|
||||
(
|
||||
attempt.attempt_id,
|
||||
attempt.task_id,
|
||||
@@ -657,6 +686,7 @@ def _upsert_attempts(conn, task: Task) -> None:
|
||||
attempt.finished_at,
|
||||
attempt.lease_expires_at,
|
||||
attempt.grace_expires_at,
|
||||
json.dumps(attempt.deadlines.to_dict()) if attempt.deadlines else None,
|
||||
),
|
||||
)
|
||||
|
||||
@@ -888,6 +918,8 @@ def purge_finished(
|
||||
|
||||
__all__ = [
|
||||
"INPUTS_DIRNAME",
|
||||
"artifact_id_for",
|
||||
"normalize_artifact_id",
|
||||
"INPUTS_PARAM_KEY",
|
||||
"INPUT_PARAM_KEYS",
|
||||
"InputStagingError",
|
||||
|
||||
@@ -229,10 +229,27 @@ def capability_to_pb(cap: dict) -> pb.ModelCapability:
|
||||
cpu_fallback=bool(cap.get("cpu_fallback")),
|
||||
repo_ids=list(cap.get("repo_ids") or []),
|
||||
display_name=str(cap.get("display_name") or ""),
|
||||
backend=str(cap.get("backend") or ""),
|
||||
free_memory_bytes=int(cap.get("free_memory_bytes") or 0),
|
||||
)
|
||||
|
||||
|
||||
def capability_from_pb(message: pb.ModelCapability) -> dict:
|
||||
def capability_from_pb(
|
||||
message: pb.ModelCapability, *, fallback_backend: str = ""
|
||||
) -> dict:
|
||||
"""Decode a capability, including protocol-v2 peers from before backend.
|
||||
|
||||
``backend`` was added to the existing protocol-v2 message, so an older
|
||||
peer legitimately sends its protobuf default (the empty string). The
|
||||
host-level GPU backend is the only compatible execution-device signal in
|
||||
that payload. A capability explicitly marked as a CPU fallback must stay
|
||||
on CPU even when its host also has a GPU.
|
||||
"""
|
||||
backend = str(message.backend or "").strip().lower()
|
||||
if message.cpu_fallback:
|
||||
backend = "cpu"
|
||||
elif not backend:
|
||||
backend = str(fallback_backend or "").strip().lower()
|
||||
return {
|
||||
"engine": message.engine,
|
||||
"model_id": message.model_id,
|
||||
@@ -249,6 +266,8 @@ def capability_from_pb(message: pb.ModelCapability) -> dict:
|
||||
"cpu_fallback": message.cpu_fallback,
|
||||
"repo_ids": list(message.repo_ids),
|
||||
"display_name": message.display_name,
|
||||
"backend": backend,
|
||||
"free_memory_bytes": message.free_memory_bytes,
|
||||
}
|
||||
|
||||
|
||||
|
||||
@@ -851,10 +851,12 @@ class WorkerServicer(pb_grpc.WorkerServiceServicer):
|
||||
epoch: int,
|
||||
) -> pb.RegisterResponse:
|
||||
session = identity.issue_session(worker_id=worker.id, key_id=worker.key_id, epoch=epoch)
|
||||
capabilities = [codec.capability_from_pb(c) for c in request.capabilities]
|
||||
host = codec.host_from_pb(request.host)
|
||||
|
||||
backend = host["gpus"][0].get("backend", "") if host.get("gpus") else ""
|
||||
capabilities = [
|
||||
codec.capability_from_pb(c, fallback_backend=backend)
|
||||
for c in request.capabilities
|
||||
]
|
||||
claimed_refs = {
|
||||
ref.attempt_id: codec.task_ref(
|
||||
ref.task_id, ref.attempt_id, ref.session_epoch
|
||||
@@ -2044,7 +2046,14 @@ class WorkerServicer(pb_grpc.WorkerServiceServicer):
|
||||
session.worker_id,
|
||||
)
|
||||
return
|
||||
caps = [codec.capability_from_pb(c) for c in update.capabilities]
|
||||
worker = self.pool.get(session.worker_id)
|
||||
fallback_backend = (
|
||||
worker.capacity.backend if worker is not None else ""
|
||||
)
|
||||
caps = [
|
||||
codec.capability_from_pb(c, fallback_backend=fallback_backend)
|
||||
for c in update.capabilities
|
||||
]
|
||||
self._queue_capability_update(session, caps)
|
||||
return
|
||||
|
||||
|
||||
@@ -14,12 +14,21 @@ cloning, and cinematic video dubbing — fully local, with no cloud API keys or
|
||||
|
||||

|
||||
|
||||
VoiceStudio runs entirely on your own hardware (CUDA / MPS / ROCm / CPU
|
||||
VoiceStudio runs entirely on your own hardware (CUDA / ROCm / CPU
|
||||
auto-detect) — nothing is sent to the cloud. This image is the **headless
|
||||
web-server build**: a FastAPI backend serving a pre-built React UI over HTTP, so
|
||||
you can run it on a homelab box, a GPU server, or anywhere Docker runs and open
|
||||
you can run it on an AMD64 homelab box or GPU server and open
|
||||
the UI in a browser.
|
||||
|
||||
**Architecture:** published images are **`linux/amd64` only**; there is no
|
||||
native ARM64 image. On Apple Silicon, use the
|
||||
[native macOS app](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/macos.md)
|
||||
for Apple GPU acceleration; the Linux container cannot access the Mac's Apple
|
||||
GPU through MPS or MLX. Other ARM64 hosts need an AMD64 server or CPU emulation,
|
||||
which can be much slower. See the
|
||||
[architecture requirements](https://github.com/debpalash/VoiceStudio/blob/main/docs/install/docker.md#architecture)
|
||||
before pulling an image.
|
||||
|
||||
> The Tauri desktop app's auto-updater and update-channel toggle are
|
||||
> **desktop-only** and do not apply to this image — to update, pull a newer tag
|
||||
> and recreate the container.
|
||||
@@ -111,12 +120,12 @@ publishing the web UI — see the [Docker install guide](https://github.com/debp
|
||||
|-----|--------------|
|
||||
| `:latest` | **Rolling preview** — latest commit on `main`, at or ahead of the last release. This is the preview channel; pin `:stable` for production. |
|
||||
| `:stable` | Most recent versioned release (updated on every `v*` git tag) |
|
||||
| `:0.5.1` | Exact release version |
|
||||
| `:0.5.2` | Exact release version |
|
||||
| `:0.5` | Latest patch within the `0.5` minor |
|
||||
| `:main` | Alias of the same rolling `main` build as `:latest` |
|
||||
| `:sha-xxxxxxx` | A specific commit (produced by manual workflow dispatch) |
|
||||
| `:rocm` | **AMD GPU (ROCm) build** of the rolling preview — the ROCm analogue of `:latest` |
|
||||
| `:stable-rocm`, `:0.5.1-rocm`, `:0.5-rocm`, `:sha-xxxxxxx-rocm` | ROCm builds of the corresponding tags above |
|
||||
| `:stable-rocm`, `:0.5.2-rocm`, `:0.5-rocm`, `:sha-xxxxxxx-rocm` | ROCm builds of the corresponding tags above |
|
||||
|
||||
Preview builds always come from `main` and never version-sort below `:stable`,
|
||||
so upgrades flow naturally. The same images and tags
|
||||
@@ -136,7 +145,7 @@ are mirrored on GHCR at
|
||||
- **📦 Batch Queue** — drop 50 videos and walk away; per-job progress.
|
||||
- **🤖 MCP Server** — drive VoiceStudio from Claude, Cursor, or any MCP client.
|
||||
- **🛡️ AI Watermark** — invisible AudioSeal (Meta) marking that survives compression.
|
||||
- **⚡ GPU Auto-Detect** — CUDA · MPS · ROCm · CPU, with auto-offload on ≤8 GB cards.
|
||||
- **⚡ GPU Auto-Detect** — CUDA · ROCm · CPU, with auto-offload on ≤8 GB cards.
|
||||
- **🧩 Extensible** — subclass `TTSBackend` to add any engine in ~50 lines.
|
||||
|
||||
Multiple TTS engines ship out of the box (IndexTTS, CosyVoice, Supertonic-3, and
|
||||
|
||||
+10
-1
@@ -96,7 +96,7 @@ bug to fix immediately, not backlog.
|
||||
| Channel | Source | Produced by | How to verify |
|
||||
|---|---|---|---|
|
||||
| GitHub Release: installers + signed `latest.json` (**Stable** updater channel) | the `vX.Y.Z` tag | `release.yml` on tag push | Release page has dmg (arm+intel), msi/exe, AppImage/deb, `latest.json`; body = the CHANGELOG section (not the auto-generated fallback), followed by per-platform checksums and a **Contributors** avatar strip (owner + every PR author for the tag — the `contributors-strip` job) |
|
||||
| **Preview** updater channel (rolling `preview` prerelease) | **`main` only** | `release.yml` nightly cron / manual dispatch | preview `latest.json` stamps `X.Y.Z-N` and semver-sorts above stable |
|
||||
| **Preview** updater channel (rolling `preview` prerelease) | **`main` only** | `release.yml` nightly cron / manual dispatch | preview `latest.json` uses main's version when it is ahead; otherwise it advances the stable patch, then appends `-N` so it semver-sorts above stable |
|
||||
| GHCR CUDA image: `:X.Y.Z`, `:X.Y`, `:stable` | the tag | `docker.yml` on tag push | `docker manifest inspect ghcr.io/debpalash/omnivoice-studio:X.Y.Z` |
|
||||
| GHCR ROCm image: `:X.Y.Z-rocm`, `:X.Y-rocm`, `:stable-rocm` | the tag | `docker.yml` on tag push | same, with `-rocm` suffix |
|
||||
| Docker Hub mirror of **all** the above tags | the tag | `docker.yml` (gated on `DOCKERHUB_*` secrets) | tag list at hub.docker.com/r/palashdeb/omnivoice-studio/tags |
|
||||
@@ -147,3 +147,12 @@ There's no "revert update" flow for clients — they'll only see a *newer* versi
|
||||
3. Clients auto-update to the "new" v0.2.1 which is actually the old code.
|
||||
|
||||
Ugly but it works. Better plan: test with Option B above before publishing the draft.
|
||||
|
||||
## Retrying a partially published build
|
||||
|
||||
Use GitHub Actions **Re-run failed jobs** for the same release run. On retries,
|
||||
the workflow removes only the current version's installers for that job's target
|
||||
before Tauri uploads them again. A macOS retry also replaces that architecture's
|
||||
versionless updater archive. Other versions, sibling platforms, and updater
|
||||
manifests remain intact. Inventory or deletion permission/network failures stop
|
||||
the job instead of hiding an upload collision.
|
||||
|
||||
+4
-4
@@ -18,7 +18,7 @@ Phase 5 · Productisation ░░░░░░░░░░ 0 / 5
|
||||
|
||||
Design track ▓▓▓▓▓▓▓▓▓░ ongoing · 14 primitives + ~67 migrated inline styles · DubTab/Header/Sidebar/CloneDesignTab drained
|
||||
Performance track ▓▓▓░░░░░░░ underway · profiling, preload, isolated engines + cache-remix I/O
|
||||
Feature-magic track ░░░░░░░░░░ not started
|
||||
Feature-magic track ▓▓░░░░░░░░ underway · project-level casting board shipped
|
||||
Quality track ▓▓░░░░░░░░ 12 smoke tests, 10 error messages rewritten
|
||||
```
|
||||
|
||||
@@ -205,15 +205,15 @@ None on the critical path to world-class. All are answers to real demand.
|
||||
| Interaction budgets (<50 ms UI, <200 ms preview, <4 s first seg) | 🟡 | `/ws/tts` reports real TTFA, total generation time and RTF; frontend responsiveness instrumentation exists, but no cross-surface budget gate yet. |
|
||||
| Dedicated dev-week per quarter | ⏳ | Cadence not yet booked. |
|
||||
|
||||
### ✨ Feature-magic track _(⏳ not started)_
|
||||
### ✨ Feature-magic track _(🟡 underway)_
|
||||
|
||||
| Feature | Status | Phase gate |
|
||||
|------|:---:|------|
|
||||
| Project-level casting view (drag voices to speakers) | ⏳ | After Phase 3 |
|
||||
| Project-level casting view (drag voices to speakers) | ✅ | Shipped (#1767): the dub CAST strip expands into a casting board — drag voice chips onto speaker rows, keyboard listbox included, same fields as the dropdowns. |
|
||||
| Voice memory across projects | ⏳ | After Phase 4 |
|
||||
| Context-aware pipeline (video frames → pipeline decisions) | ⏳ | After Phase 4 |
|
||||
| On-device learning from corrections (user edits → LoRA) | ⏳ | Research only; possibly Phase 5+ |
|
||||
| Real-time dub preview (stream TTS as you edit) | ⏳ | After Phase 4.1 |
|
||||
| Real-time dub preview (stream TTS as you edit) | ✅ | Shipped 2026-09-02 (#1769) — opt-in "Live preview" toggle on the dub segment table streams the edited line over `/ws/tts` with its CAST voice; export path unchanged. |
|
||||
|
||||
### 🧪 Quality track _(🟡 underway)_
|
||||
|
||||
|
||||
+109
-48
@@ -7,39 +7,71 @@ Every folder has a single job. Every file at the root earns its place.
|
||||
```
|
||||
VoiceStudio/
|
||||
│
|
||||
├── README.md ⟵ user-facing overview
|
||||
├── CHANGELOG.md ⟵ release history
|
||||
├── LICENSE
|
||||
├── README.md / README_CN.md ⟵ user-facing overview (English / Chinese)
|
||||
├── CHANGELOG.md ⟵ release history; release.yml extracts the tag's section verbatim
|
||||
├── CLAUDE.md / AGENTS.md ⟵ the working contract for AI agents — keep the two in sync
|
||||
├── LICENSE, LICENSE-NOTICE.md, SPONSORS.md
|
||||
│
|
||||
├── pyproject.toml ⟵ Python project manifest
|
||||
├── pyproject.toml ⟵ Python project manifest (+ pytest / lint config)
|
||||
├── uv.lock ⟵ Python lockfile
|
||||
├── package.json ⟵ monorepo manifest (Bun workspaces + Turborepo)
|
||||
├── bun.lock ⟵ JS lockfile
|
||||
├── bun.lock ⟵ JS lockfile — repo-root, covers frontend/ too
|
||||
├── turbo.json ⟵ turborepo pipeline
|
||||
│
|
||||
├── .coderabbit.yaml ⟵ CodeRabbit PR review config (fed CLAUDE.md)
|
||||
├── greptile.json ⟵ Greptile PR review config (fed CLAUDE.md)
|
||||
├── skills-lock.json ⟵ pins the sources + hashes of .agents/skills/
|
||||
├── .gitleaks.toml ⟵ secret-scan config
|
||||
├── .gitmodules ⟵ omnivoice-gallery submodule
|
||||
├── .python-version
|
||||
├── .dockerignore ⟵ Docker build context filter
|
||||
├── backend.spec ⟵ pyinstaller spec (stays at root by pyinstaller convention)
|
||||
├── alembic.ini ⟵ DB migration config (stays at root by alembic convention)
|
||||
│
|
||||
├── .env ⟵ user config; gitignored, .env.example is the template
|
||||
├── .gitignore
|
||||
├── .gitignore ⟵ a repo-local .env stays ignored, but user config is NOT
|
||||
│ kept here: the durable env file is ~/.config/omnivoice/env
|
||||
│ (backend/core/user_env.py), written by the Settings panel
|
||||
│
|
||||
├── backend/ ⟵ FastAPI server
|
||||
│ ├── main.py
|
||||
│ ├── api/routers/ HTTP endpoints (thin)
|
||||
│ ├── core/ config, db, task queue, metrics
|
||||
│ ├── services/ business logic
|
||||
│ └── schemas/ pydantic request/response shapes
|
||||
│ ├── main.py the one entry point; its boot order is load-bearing —
|
||||
│ │ read the comments before reordering anything
|
||||
│ ├── api/routers/ 39 routers, auto-included; thin HTTP/WS surface
|
||||
│ │ └── setup/ first-run wizard, model download
|
||||
│ ├── core/ config, db, job queue, event bus, auth/CSRF, path security,
|
||||
│ │ opt-in analytics, version, diagnostics
|
||||
│ ├── services/ 78 modules of business logic — TTS, dubbing pipeline,
|
||||
│ │ audio DSP, GPU gateway, engine routing, model lifecycle
|
||||
│ ├── engines/ per-engine adapters: indextts, supertonic3, confucius4,
|
||||
│ │ dots_tts, moss_tts_v15, pockettts, audiocpp,
|
||||
│ │ omnivoice_gguf, omnivoice_subprocess, _asr_sidecar, _echo
|
||||
│ ├── worker/ remote / distributed workers — scheduler, pool, routing,
|
||||
│ │ breaker, capacity, plus protocol/ and inbound/
|
||||
│ ├── mcp_shim/ MCP server entry point (docs/mcp.md)
|
||||
│ ├── speech_client/ speech sidecar client entry point
|
||||
│ ├── schemas/ pydantic request/response shapes
|
||||
│ ├── migrations/versions/ alembic revisions — every schema change goes through here
|
||||
│ ├── plugins/ plugin drop-in point (see services/plugin_sdk.py)
|
||||
│ ├── hooks/ pyinstaller runtime hooks
|
||||
│ ├── config/models.yaml model catalogue
|
||||
│ └── tests/ the isolated pytest session — see "Where tests live"
|
||||
│
|
||||
├── frontend/ ⟵ React 19 + Vite + Tauri desktop
|
||||
│ ├── package.json THE app version — every other version file mirrors it
|
||||
│ ├── src/
|
||||
│ │ ├── pages/ one file per top-level view
|
||||
│ │ ├── components/ reusable UI
|
||||
│ │ ├── api/ typed API clients
|
||||
│ │ ├── store/ Zustand slices
|
||||
│ │ ├── components/ reusable UI (+ audiobook/ clone/ dub/ gallery/ settings/ …)
|
||||
│ │ ├── ui/, lib/ shared primitives and helpers
|
||||
│ │ ├── api/ typed API clients, one per router group
|
||||
│ │ ├── store/ Zustand slices (+ persisted-state migrations)
|
||||
│ │ ├── hooks/ custom React hooks
|
||||
│ │ └── utils/
|
||||
│ ├── src-tauri/ Rust desktop shell
|
||||
│ │ ├── i18n/locales/ the ONLY home for user-facing strings
|
||||
│ │ ├── config/, data/, assets/, utils/
|
||||
│ │ └── test/ vitest setup + visual-test helpers
|
||||
│ ├── e2e/, e2e-perf/, e2e-prod/ Playwright suites: functional, perf, packaged bundle
|
||||
│ ├── src-tauri/ Rust desktop shell — backend spawn/bootstrap, updater
|
||||
│ │ │ channel, dictation shortcut, crash/reset/uninstall
|
||||
│ │ ├── capabilities/, icons/, wix/, debian/, appimage/ packaging inputs
|
||||
│ │ └── tests/
|
||||
│ └── public/
|
||||
│
|
||||
├── omnivoice/ ⟵ the underlying TTS model package
|
||||
@@ -51,46 +83,59 @@ VoiceStudio/
|
||||
│ ├── training/
|
||||
│ └── utils/
|
||||
│
|
||||
├── tests/ ⟵ all tests live here, no exceptions
|
||||
│ ├── conftest.py
|
||||
│ ├── test_api.py
|
||||
│ ├── test_dub_*.py
|
||||
│ ├── test_job_queue.py
|
||||
│ ├── test_segmentation.py
|
||||
│ └── frontend/ Node-based frontend tests
|
||||
├── tests/ ⟵ the main pytest session (testpaths in pyproject.toml)
|
||||
│ ├── conftest.py hermetic OMNIVOICE_DATA_DIR — never touches real app state
|
||||
│ ├── backend/, scripts/ mirrors of the source trees they cover
|
||||
│ ├── smoke/ fast end-to-end checks (own CI job, HF_HUB_OFFLINE=1)
|
||||
│ ├── evals/ quality evals (evals.yml)
|
||||
│ ├── probe/, fixtures/
|
||||
│ └── frontend/ Node-based frontend tests (legacy; vitest is the default)
|
||||
│
|
||||
├── scripts/ ⟵ dev / build / release shell + python scripts
|
||||
│ ├── install.sh universal installer (macOS/Linux/WSL)
|
||||
│ ├── install.ps1 universal installer (Windows)
|
||||
│ ├── run.sh universal launcher
|
||||
├── scripts/ ⟵ dev / build / release scripts (shell, python, mjs)
|
||||
│ ├── install.sh / install.ps1 universal installers
|
||||
│ ├── desktop-*.mjs dev, prod and fresh desktop launchers
|
||||
│ ├── smoke-test.sh end-to-end validation
|
||||
│ └── desktop-prod.sh production desktop build
|
||||
|
||||
│ ├── check-docs-drift.py the docs-drift.yml checker (docs/features.yaml is canonical)
|
||||
│ └── build-omnivoice-tts.sh builds the bin/ sidecars
|
||||
│
|
||||
├── bin/ ⟵ prebuilt omnivoice-tts sidecars, one per platform
|
||||
│
|
||||
├── .agents/skills/ ⟵ canonical skill copies (vite, fastapi-python), pinned by
|
||||
│ skills-lock.json — followed by path, never symlinked
|
||||
├── skills/ ⟵ skills this repo publishes (omnivoice, oss-maintainer)
|
||||
│
|
||||
├── infra/ ⟵ edge/deploy workers (not the Docker deploy path)
|
||||
│ └── install-redirect/ voicestudio.sh/install — UA-sniffing installer worker
|
||||
│
|
||||
├── deploy/ ⟵ Docker deployment configs
|
||||
│ ├── Dockerfile single-stage CUDA image
|
||||
│ └── docker-compose.yml one-click local deployment
|
||||
│ ├── Dockerfile CUDA by default; CI builds the ROCm variant from the same
|
||||
│ │ file via BASE_IMAGE / GPU_FLAVOR overrides
|
||||
│ ├── docker-compose.yml one-click local deployment
|
||||
│ ├── torch-constraints.txt pinned torch resolution for the image
|
||||
│ └── dockerhub-overview.md synced to the Docker Hub overview page at release
|
||||
│
|
||||
├── docs/ ⟵ developer docs, screenshots, branding
|
||||
│ ├── ROADMAP.md where this project is going
|
||||
│ ├── STRUCTURE.md you are here
|
||||
│ ├── mcp.json MCP config template
|
||||
│ ├── preview.png README hero image
|
||||
│ ├── logo.png, logo.svg branding assets
|
||||
│ ├── screenshot-*.png feature screenshots
|
||||
│ ├── languages.md
|
||||
│ ├── training.md
|
||||
│ ├── data_preparation.md
|
||||
│ ├── evaluation.md
|
||||
│ └── voice-design.md
|
||||
│ ├── RELEASING.md the release checklist — every deployment channel
|
||||
│ ├── features.yaml canonical feature inventory (drives docs-drift.yml)
|
||||
│ ├── adr/ architecture decision records
|
||||
│ ├── agents/ agent-facing docs (issue tracker, triage labels, domain)
|
||||
│ ├── engines/, dubbing/, install/, setup/, migration/, features/, playbooks/, specs/
|
||||
│ ├── media/, screenshot-*.png, preview.png, logo.*
|
||||
│ └── languages.md, training.md, data_preparation.md, evaluation.md, voice-design.md
|
||||
│
|
||||
├── examples/ ⟵ runnable demos + sample inputs
|
||||
├── examples/ ⟵ runnable demos + sample inputs (agentic/, speech-platform/)
|
||||
│
|
||||
├── notebooks/ ⟵ OmniVoice_Studio_Colab.ipynb
|
||||
│
|
||||
├── omnivoice-gallery/ ⟵ git submodule — the published voice gallery
|
||||
│
|
||||
├── omnivoice_data/ ⟵ Docker bind-mount target (gitignored)
|
||||
│ DB + HF cache live here when running via compose
|
||||
│
|
||||
├── .github/workflows/ ⟵ ci, docker, release, security, docs-drift, evals,
|
||||
│ install-smoke, build-omnivoice-tts
|
||||
└── .git/
|
||||
```
|
||||
|
||||
@@ -98,11 +143,22 @@ VoiceStudio/
|
||||
|
||||
1. **Nothing at the root is a runtime artifact.** Outputs, temp files, local DBs, crash logs — all go to `~/Library/Application Support/OmniVoice/` (or the OS equivalent), *never* into the repo. The one exception is `omnivoice_data/`, which exists as a bind-mount anchor for Docker.
|
||||
|
||||
2. **No ad-hoc scripts at the root.** One-off debug scripts live in `scripts/`. Tests live in `tests/`. Benchmarks live in `scripts/benchmarks/` (when we create them).
|
||||
2. **No ad-hoc scripts at the root.** One-off debug scripts live in `scripts/`. Tests live in one of the three homes below, never at the root.
|
||||
|
||||
3. **Each subdirectory owns one concern.** If you can't describe what goes in a directory in one sentence, it's wrong.
|
||||
|
||||
4. **Every package has a manifest.** `backend/`, `frontend/`, `omnivoice/` each have their own deps declared via `pyproject.toml` / `package.json` — they are independently testable.
|
||||
4. **Every package has a manifest.** `backend/`, `frontend/`, `omnivoice/` each have their own deps declared via `pyproject.toml` / `package.json` — they are independently testable. The JS lockfile is the **repo-root** `bun.lock` (Bun workspace), and `deploy/Dockerfile` installs from it with `--frozen-lockfile`.
|
||||
|
||||
## Where tests live
|
||||
|
||||
Three homes, each with its own runner. CI runs all three inside the single `test` job in
|
||||
`ci.yml`, as separate steps. The split is deliberate, not drift:
|
||||
|
||||
| Home | Runner | Why it's separate |
|
||||
|---|---|---|
|
||||
| `tests/` | `pytest tests/` — the `testpaths` default | The main suite. Its `conftest.py` points `OMNIVOICE_DATA_DIR` at a throwaway dir so a run can never touch the developer's real app state (#878). |
|
||||
| `backend/tests/` | `pytest backend/tests/` — its own pytest session (the `Run pytest (backend/tests, isolated)` step) | Runs as an isolated session against `backend/`'s bare imports. Its `conftest.py` sets the same hermetic data dir; **never** reintroduce module-level `sys.modules` stubs there — they leak process-wide at collection time and poison mixed runs. |
|
||||
| `frontend/src/**/*.test.{js,jsx,ts,tsx}` | `bun run test` (vitest, jsdom) | Co-located with the component under test. `frontend/e2e*/` hold the Playwright suites; `tests/frontend/` is the older `node:test` set. |
|
||||
|
||||
## What lives where
|
||||
|
||||
@@ -110,10 +166,14 @@ VoiceStudio/
|
||||
|---|---|
|
||||
| User-facing product code | `backend/`, `frontend/` |
|
||||
| The TTS model (independent of the studio) | `omnivoice/` |
|
||||
| A new TTS/ASR engine adapter | `backend/engines/<engine>/` |
|
||||
| Everything executable but not user-facing | `scripts/` |
|
||||
| Tests | `tests/` |
|
||||
| Prebuilt platform sidecars | `bin/` |
|
||||
| Python tests | `tests/` (or `backend/tests/` when the isolated session is required) |
|
||||
| Frontend unit tests | next to the component, as `*.test.jsx` |
|
||||
| Developer + user docs (Markdown) | `docs/` |
|
||||
| Architecture decision records (ADRs) | `docs/adr/` |
|
||||
| Agent-facing docs | `docs/agents/` |
|
||||
| Runnable demos and sample data | `examples/` |
|
||||
| Runtime data (never committed) | `~/Library/Application Support/OmniVoice/` on Mac |
|
||||
|
||||
@@ -137,10 +197,10 @@ Removed in the 2026-07-12 cleanup pass (all preserved in git history):
|
||||
| Dir | Why it was there | Where it went |
|
||||
|---|---|---|
|
||||
| `.planning/` (74 files) | GSD-era planning archive: phases, quick plans, issue clusters. The GSD workflow was retired 2026-07-08. | Deleted; the four load-bearing decision docs moved to `docs/adr/`. |
|
||||
| `specs/` | spec-kit specs for features 001–007 — all shipped. | Deleted. |
|
||||
| `specs/` | spec-kit specs for features 001–007 — all shipped. | Deleted; `docs/specs/` is the current home. |
|
||||
| `design/` | ASCII mockups of the pre-React target UX, superseded by the shipped app. | Deleted. |
|
||||
| `research/` | Archived legacy Gradio UI + April-2026 competitor notes. | Deleted. |
|
||||
| `.agents/` | Rules for a third-party agent tool no longer in use. | Deleted. |
|
||||
| `.agents/` | Rules for a third-party agent tool no longer in use. | Deleted — then reintroduced with a different job: `.agents/skills/` now holds the canonical skill copies pinned by `skills-lock.json`. |
|
||||
|
||||
## Scaling path (proposed, not yet executed)
|
||||
|
||||
@@ -169,12 +229,13 @@ VoiceStudio/
|
||||
- `backend.spec` (`['backend/main.py']`, `pathex=['.']`)
|
||||
- `frontend/src-tauri/tauri.*.conf.json` sidecar paths
|
||||
- every import that reads `from backend.main import …` (tests, scripts)
|
||||
- `frontend/package.json` as the version source of truth, and the mirrors that track it
|
||||
|
||||
Migrate when adding the second `apps/*` or the second `packages/*`. Not before.
|
||||
|
||||
## Conventions
|
||||
|
||||
- **Filenames:** snake_case for Python, kebab-case or PascalCase for JS/TS components, lowercase for Markdown.
|
||||
- **Tests mirror source paths.** `backend/services/dub_pipeline.py` → `tests/services/test_dub_pipeline.py`.
|
||||
- **Tests mirror source paths** where a mirror exists: `tests/backend/` mirrors `api/ core/ engines/ services/`, so `backend/services/ffmpeg_utils.py` → `tests/backend/services/test_ffmpeg_utils.py`. Everything else stays flat — `tests/backend/test_*.py` for backend-wide cases, `tests/test_*.py` for cross-cutting ones. A React component's test sits next to the component.
|
||||
- **One-off scripts** go into `scripts/` with a descriptive name, not `test_*.py` at the root.
|
||||
- **New top-level directories** require a PR that updates *this file*.
|
||||
|
||||
+4
-2
@@ -204,7 +204,8 @@ ws://gpu-box:3900/ws/transcribe?api_key=<key>
|
||||
That URL form is retained for non-browser compatibility only. The first-party
|
||||
UI never constructs it. A bearer administrator session first calls
|
||||
`POST /api/auth/ws-ticket` and puts only the returned `ws_ticket` in the URL.
|
||||
Tickets are scoped to `/ws/transcribe` or `/ws/events`, expire after 30 seconds,
|
||||
Tickets are scoped to one of `/ws/transcribe`, `/ws/events` or `/ws/tts` (the
|
||||
live dub preview stream), expire after 30 seconds,
|
||||
return the same bounded `expires_in`/`expires_at` pair, and are consumed
|
||||
atomically at most once. Same-origin UI WebSockets use the
|
||||
HttpOnly session cookie and must pass exact `Origin` validation; `null`, missing,
|
||||
@@ -338,7 +339,8 @@ drives exact-Origin checks and the session cookie's `Secure` attribute.
|
||||
For a public path prefix such as `/studio`, either strip that prefix before
|
||||
forwarding or configure the ASGI `root_path` to the same value. WebSocket ticket
|
||||
validation removes only that trusted, configured prefix; it never accepts an
|
||||
arbitrary path merely because it ends in `/ws/events` or `/ws/transcribe`.
|
||||
arbitrary path merely because it ends in `/ws/events`, `/ws/transcribe` or
|
||||
`/ws/tts`.
|
||||
|
||||
## Status codes
|
||||
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user