diff --git a/.github/assets/social-preview.png b/.github/assets/social-preview.png new file mode 100644 index 00000000..15799716 Binary files /dev/null and b/.github/assets/social-preview.png differ diff --git a/CHANGELOG.md b/CHANGELOG.md index fb341055..6c3168f4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,18 @@ The format is loosely based on [Keep a Changelog](https://keepachangelog.com/). Versions track the desktop app (`tauri.conf.json` + `frontend/src-tauri/Cargo.toml`). The bundled TTS model package (`pyproject.toml`) is versioned independently. +## [0.2.7] — Unreleased + +### Added +- **Frameless dictation widget.** Global dictation upgraded from an in-app FAB to a true OS-level floating widget that hovers over any application. Transparent, decorations-free, always-on-top secondary Tauri window activated by `⌘+⇧+Space`. Auto-hides 2.5 s after a successful paste. +- **Standalone `CaptureWidget` component.** Refactored `CaptureButton` into `CaptureWidget`, running on an isolated route (`/?window=widget`). +- **Social preview image.** Added `social-preview.png` for GitHub SEO. + +### Changed +- **README overhaul.** Compact 3-column feature grid, reorganized Quickstart (one-command install, Docker, Desktop App tips), updated comparison table, roadmap, and footer CTA. + +--- + ## [0.2.6] — Unreleased ### License diff --git a/README.md b/README.md index 45292856..1af8436d 100644 --- a/README.md +++ b/README.md @@ -1,8 +1,9 @@
- OmniVoice Logo + OmniVoice Logo

OmniVoice Studio

-

The open-source ElevenLabs alternative.

-

Voice cloning · Voice design · Video dubbing — 646 languages, runs 100% locally, forever free.

+

The open-source ElevenLabs alternative.

+

Real-time dictation, zero-shot voice cloning, and cinematic video dubbing — all on your desktop.
Open-source, no API keys, fully local. 646 languages.

+

Stars Release @@ -10,31 +11,176 @@ Issues Discord

+

- Download · - Features · Quickstart · - Why Open Source? · - Roadmap + Features · + Why OmniVoice Studio? · + TTS Engines · + Contributing · + Discord

+

- Download macOS DMG - Download Windows MSI - Download Linux AppImage - Download Debian .deb + Download macOS DMG + Download Windows MSI + Download Linux AppImage + Download Debian .deb


- OmniVoice Studio — Launchpad -
- Launchpad — Voice Clone · Voice Design · Video Dubbing, all in one place. + OmniVoice Studio — The open-source ElevenLabs alternative
+> [!WARNING] +> **OmniVoice Studio is in active beta.** Things may break between releases. For the latest features and fixes, clone the repo and run from source rather than using pre-built installers. Bug reports and PRs are very welcome — [open an issue](https://github.com/debpalash/OmniVoice-Studio/issues) or [join Discord](https://discord.gg/aRRdVj3de7). +
+## Features + + + + + + + + + + + + + + + + + + + + + + +
+

🎙️ Voice Cloning

+

3-second clip → mirror any voice.
646 languages, zero-shot.

+
+

🎨 Voice Design

+

Gender, age, accent, pitch, speed,
emotion, dialect — dial it in.

+
+

🎬 Video Dubbing

+

YouTube URL or file → transcribe →
translate → re-voice → MP4.

+
+

⌨️ Dictation Widget

+

⌘+⇧+Space from any app.
Transcribes, auto-pastes, disappears.

+
+

🔊 Vocal Isolation

+

Demucs-powered. Splits speech
from music, keeps the background.

+
+

👥 Speaker Diarization

+

Pyannote + WhisperX.
Auto-identifies who said what.

+
+

📦 Batch Queue

+

Drop 50 videos, walk away.
Progress bars per job.

+
+

🤖 MCP Server

+

Use OmniVoice from Claude,
Cursor, or any MCP client.

+
+

🛡️ AI Watermark

+

AudioSeal (Meta). Invisible,
survives compression.

+
+

🔐 100% Local

+

No keys, no cloud, no accounts.
Your machine only.

+
+

⚡ GPU Auto-Detect

+

CUDA · MPS · ROCm · CPU.
≤8 GB? Auto-offloads.

+
+

🧩 Extensible

+

Subclass TTSBackend,
add any engine in ~50 lines.

+
+ +--- + +## Quickstart + +### One-command install + +```bash +git clone https://github.com/debpalash/OmniVoice-Studio.git && cd OmniVoice-Studio && bun install && bun run dev +``` + +That's it. Open [localhost:3901](http://localhost:3901) and start cloning voices. + +### Docker + +```bash +# CPU mode +docker compose up --build -d + +# Or with NVIDIA GPU +docker compose --profile gpu up --build -d +``` + +Open [http://localhost:3900](http://localhost:3900) once the health check passes. First run downloads ~4 GB of model weights — progress is shown in `docker compose logs -f`. + +> **Network access:** the container binds to `127.0.0.1` only. To reach OmniVoice from another machine on your LAN, change the port mapping in `docker-compose.yml` to `"0.0.0.0:3900:3900"`. OmniVoice ships no built-in authentication — when exposing it beyond your machine, put it behind a reverse proxy with auth (Caddy `basic_auth`, nginx + htpasswd, Tailscale, etc.). + +### Desktop App + +Pre-built installers (~6–8 MB) are available on the [**Releases**](https://github.com/debpalash/OmniVoice-Studio/releases/latest) page. On first launch, the app bootstraps a Python environment and downloads model weights automatically — the splash screen shows progress. + +```bash +bun run desktop # Build from source (macOS / Windows / Linux) +``` + +
+macOS — "app is damaged and can't be opened" +
+ +macOS quarantines apps downloaded outside the App Store. After dragging to `/Applications`: + +```bash +xattr -cr /Applications/OmniVoice\ Studio.app +``` + +Open normally after. One-time fix. +
+ +
+Windows — first launch takes 5–10 minutes +
+ +The app bootstraps a Python virtual environment, installs dependencies, and downloads ffmpeg on first run. The splash screen shows each step. Subsequent launches start in seconds. +
+ +
+Linux — AppImage needs FUSE +
+ +If FUSE isn't available, use the `.deb` package or extract-and-run: + +```bash +chmod +x OmniVoice.Studio_*.AppImage +./OmniVoice.Studio_*.AppImage --appimage-extract-and-run +``` +
+ +> [!NOTE] +> First run downloads model weights (~2.4 GB). This works out of the box — no account needed. For faster downloads, optionally set `HF_TOKEN=hf_...` in your environment ([get a free token here](https://huggingface.co/settings/tokens)). +> +> **Having issues?** Join our [Discord](https://discord.gg/aRRdVj3de7) for setup help and troubleshooting. + +| Service | URL | Stack | +|---------|-----|-------| +| **Backend** | `localhost:3900` | FastAPI · 97 endpoints · WhisperX · Demucs · OmniVoice | +| **Frontend** | `localhost:3901` | React · Vite · Waveform timeline · Glassmorphism UI | + +--- + +## Screenshots +
@@ -83,7 +229,7 @@ --- -## Why Open Source? +## Why OmniVoice Studio? ElevenLabs charges **$5–$330/mo** and processes your audio on their servers. OmniVoice Studio runs **on your hardware, with no usage limits.** @@ -100,152 +246,7 @@ ElevenLabs charges **$5–$330/mo** and processes your audio on their servers. O | **Desktop App** | ❌ | ✅ macOS · Windows · Linux | | **Customizable** | ❌ Closed | ✅ Fork it, extend it, ship it | -Built on the [OmniVoice](https://github.com/k2-fsa/OmniVoice) 600-language zero-shot diffusion TTS model. Upload a video, get broadcast-quality dubs in any language with the original speaker's voice preserved. - -## Features - -### Core Pipeline -- **Video Dubbing** — Transcribe → translate → synthesize → mux back to MP4. One-click end-to-end. -- **Vocal Isolation** — Demucs-powered speech/music separation. Background audio preserved automatically. -- **Voice Cloning** — Clone any voice from a 3-second clip. Zero-shot, 600+ languages. -- **Multi-Speaker Diarization** — Pyannote + WhisperX fusion auto-identifies speakers and assigns unique voice profiles. - -### Studio Tools -- **Voice Capture** — Press `⌘+⇧+Space` **from any app** to dictate. Global system-wide hotkey records, transcribes, and auto-pastes into the active text field. Live partial results stream via WebSocket while you speak. -- **Speaker Casting** — Visual speaker-to-voice assignment grid. Auto-cast from video clones or assign saved profiles. -- **Voice Preview** — Floating widget for instant 8-step TTS testing. Try voices without leaving the workspace. -- **Real-time Dub Preview** — Edit a segment's text, preview the audio instantly without full re-render. -- **Multi-Language Batch** — Select multiple target languages, dub to all in one pass. -- **Batch Queue** — Drag-and-drop bulk video processing. Full pipeline: extract → transcribe → translate → generate → mix → export. Real-time progress bars per job. -- **Voice Library** — Browse, favorite, tag, and convert gallery clips into permanent voice profiles. -- **A/B Comparison** — Side-by-side voice audition for casting decisions. - -### Production Export -- **Selective Track Export** — Choose which language tracks to include in the final MP4. -- **Subtitle Export** — SRT and VTT generation alongside dubbed video. -- **Stem Export** — Separate vocals and background audio as individual files. -- **Per-Segment Mixing** — 0–200% gain control per segment for broadcast-quality balancing. - -### Technical -- **Cross-Platform GPU** — Auto-detects CUDA, Apple Silicon (MPS), ROCm, or CPU. Includes automatic cuDNN 8/9 compatibility handling. -- **VRAM-Aware** — Automatically offloads TTS to CPU during transcription on ≤8 GB GPUs. Zero config. -- **Streaming ASR** — WebSocket-based speech-to-text (`/ws/transcribe`) delivers live partial results during recording. 2s buffer interval, configurable. -- **Auto-Paste** — Dictated text is automatically pasted into the active app via system keyboard simulation (macOS Accessibility / Windows SendInput). -- **Live Telemetry** — Real-time CPU/RAM/VRAM stats with model warm-up indicator. -- **Keyboard-First** — `⌘+Enter` generate, `⌘+S` save, `⌘+Z`/`⌘+⇧+Z` undo/redo. - -### AI Provenance -- **Invisible Watermark** — AudioSeal-powered (Meta) neural watermark embedded in every generated audio. Imperceptible, survives compression/editing. -- **Detection API** — Upload any audio to `/watermark/detect` to verify OmniVoice origin with confidence score. -- **Video Branding** — Optional logo overlay on exported MP4s (5s fade-out, bottom-right). -- **Configurable** — Toggle invisible/visible watermarks independently in Settings → Privacy. - -### MCP Server (AI Agent Integration) -- **Model Context Protocol** — Expose OmniVoice as an AI agent tool for Claude, Cursor, and any MCP-compatible client. -- **5 Tools** — `generate_speech`, `list_voices`, `list_personalities`, `list_languages`, `check_health`. -- **stdio + SSE** — Works locally (Claude Desktop) or remotely (networked agents). -- **Zero config** — Drop `mcp.json` into your client config and go. See [`mcp.json`](docs/mcp.json). - -### Audio Effects Chain -- **6 presets** — Broadcast 📻, Cinematic 🎬, Podcast 🎙️, Warm ☀️, Bright ✨, Raw 🔇. -- **Pedalboard-powered** — Spotify's production-grade DSP (EQ, compressor, reverb, noise gate, limiter). -- **API-driven** — `GET /tools/effects` returns presets; custom chains via `apply_effects_chain()`. - -### Plugin SDK (Third-Party TTS Engines) -- **Abstract interface** — Subclass `TTSPlugin` to add any TTS engine in ~50 lines. -- **Built-in plugins** — ElevenLabs (cloud) and Bark (local) ship out of the box. -- **Auto-discovery** — Drop a `.py` file in `backend/plugins/`, it registers automatically. -- **API** — `GET /tools/plugins` lists all engines and their availability status. - -### GPU Safety -- **Crash sandbox** — GPU-intensive ops can run in subprocess isolation. A CUDA OOM or driver crash kills the worker, not the server. -- **6 color themes** — Gruvbox (default), Midnight Blue, Nord, Solarized, Rosé Pine, Catppuccin Mocha. - ---- - -## Quickstart - -### Docker (recommended) - -```bash -git clone https://github.com/debpalash/OmniVoice-Studio.git -cd OmniVoice-Studio - -# CPU mode -docker compose up --build -d - -# Or with NVIDIA GPU -docker compose --profile gpu up --build -d -``` - -Open [http://localhost:3900](http://localhost:3900) once the health check passes. First run downloads ~4 GB of model weights — progress is shown in `docker compose logs -f`. - -> **Network access:** the container binds to `127.0.0.1` only. To reach OmniVoice from another machine on your LAN, change the port mapping in `docker-compose.yml` to `"0.0.0.0:3900:3900"`. OmniVoice ships no built-in authentication — when exposing it beyond your machine, put it behind a reverse proxy with auth (Caddy `basic_auth`, nginx + htpasswd, Tailscale, etc.). - -### Local Development - -**Prerequisites:** [ffmpeg](https://ffmpeg.org/), [Bun](https://bun.sh/), [uv](https://docs.astral.sh/uv/) - -```bash -git clone https://github.com/debpalash/OmniVoice-Studio.git -cd OmniVoice-Studio -bun install -bun run dev -``` - -This boots both services: - -| Service | URL | Stack | -|---------|-----|-------| -| **Backend** | `localhost:3900` | FastAPI · 97 endpoints · WhisperX · Demucs · OmniVoice | -| **Frontend** | `localhost:3901` | React · Vite · Waveform timeline · Glassmorphism UI | - -> [!NOTE] -> First run downloads model weights (~2.4 GB). This works out of the box — no account needed. For faster downloads, optionally set `HF_TOKEN=hf_...` in your environment ([get a free token here](https://huggingface.co/settings/tokens)). -> -> **Having issues?** Join our [Discord](https://discord.gg/aRRdVj3de7) for setup help and troubleshooting. - -### Desktop App - -Pre-built installers (~6–8 MB) are available on the [**Releases**](https://github.com/debpalash/OmniVoice-Studio/releases/latest) page. On first launch, the app bootstraps a Python environment and downloads model weights automatically — the splash screen shows progress. - -To build from source instead: - -```bash -bun run desktop # Launches Tauri native app (macOS / Windows / Linux) -``` - -
-macOS — "app is damaged and can't be opened" -
- -macOS quarantines apps downloaded outside the App Store. After dragging to `/Applications`: - -```bash -xattr -cr /Applications/OmniVoice\ Studio.app -``` - -Open normally after. One-time fix. -
- -
-Windows — first launch takes 5–10 minutes -
- -The app bootstraps a Python virtual environment, installs dependencies, and downloads ffmpeg on first run. The splash screen shows each step. Subsequent launches start in seconds. -
- -
-Linux — AppImage needs FUSE -
- -If FUSE isn't available, use the `.deb` package or extract-and-run: - -```bash -chmod +x OmniVoice.Studio_*.AppImage -./OmniVoice.Studio_*.AppImage --appimage-extract-and-run -``` -
+OmniVoice Studio gives you professional-grade AI tools without the subscription or the cloud. --- @@ -317,18 +318,15 @@ OmniVoice ships a multi-engine TTS backend. The default engine (OmniVoice) is al | **State Management** | Zustand store migration — `uiSlice`, `pillSlice`, `dubSlice`, `generateSlice`, `prefsSlice`, `glossarySlice` | | **Desktop** | Cross-platform Tauri installers (macOS DMG, Windows MSI, Linux deb/AppImage), auto-update infrastructure | | **Windows Hardening** | Cross-platform log paths, Triton workaround, HF symlink bypass, 300s health check timeout | -| **Dictation** | Global system-wide hotkey (`⌘+⇧+Space`), streaming ASR via WebSocket, auto-paste into active app | +| **Dictation** | Global system-wide hotkey (`⌘+⇧+Space`), frameless floating widget, streaming ASR via WebSocket, auto-paste | | **Batch Pipeline** | Full batch TTS: extract → transcribe → translate → generate → mix → export, with live progress tracking | -### 🔜 Roadmap — completed ✅ +### 🔜 Up Next -**All planned features have been shipped.** - -- ~~Onboarding sample clip~~ · ~~Docker DX~~ · ~~Auto-updater~~ · ~~Deferred disk writes~~ -- ~~MCP server~~ · ~~Voice personalities~~ · ~~Audio effects chain~~ · ~~i18n framework~~ -- ~~Global hotkey dictation~~ · ~~Real-time dub preview~~ · ~~Speaker casting view~~ -- ~~Theme system~~ · ~~Plugin SDK~~ · ~~GPU crash sandbox~~ · ~~Waveform v2~~ -- ~~Batched TTS~~ · ~~Cold start optimization~~ · ~~Audiobook editor~~ · ~~Context-aware pipeline~~ +- 🎬 **Lip-sync v2** — visual speech timing with wav2lip +- 📖 **Audiobook Editor** — chapter-aware long-form narration +- 🌐 **Hosted Demo** — try OmniVoice without installing anything +- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects --- @@ -412,7 +410,12 @@ OmniVoice Studio is built on the shoulders of exceptional open-source work:
-**[⭐ Star on GitHub](https://github.com/debpalash/OmniVoice-Studio)** to follow updates. +
+ +If you read this far, you're our kind of person.
+**[⭐ Star this repo](https://github.com/debpalash/OmniVoice-Studio)** so others can find it too. + +
diff --git a/bun.lock b/bun.lock index d30776a1..5e1a296a 100644 --- a/bun.lock +++ b/bun.lock @@ -15,7 +15,7 @@ }, "frontend": { "name": "omnivoice-studio", - "version": "0.2.6", + "version": "0.2.7", "dependencies": { "@fontsource-variable/inter": "^5.2.8", "@fontsource-variable/source-serif-4": "^5.2.9", @@ -29,40 +29,40 @@ "@radix-ui/react-tabs": "^1.1.13", "@radix-ui/react-toggle-group": "^1.1.11", "@radix-ui/react-tooltip": "^1.2.8", - "@tailwindcss/vite": "4", - "@tanstack/react-query": "^5.100.4", + "@tailwindcss/vite": "^4.2.4", + "@tanstack/react-query": "^5.100.8", "@tanstack/react-table": "^8.21.3", "@tanstack/react-virtual": "^3.13.24", - "@tauri-apps/plugin-dialog": "^2.7.0", - "@tauri-apps/plugin-opener": "^2.5.3", + "@tauri-apps/plugin-dialog": "^2.7.1", + "@tauri-apps/plugin-opener": "^2.5.4", "@tauri-apps/plugin-process": "^2.3.1", "@tauri-apps/plugin-updater": "^2.10.1", "@tauri-apps/plugin-window-state": "^2.4.1", "i18next": "^26.0.8", "i18next-browser-languagedetector": "^8.2.1", - "lucide-react": "^1.8.0", + "lucide-react": "^1.14.0", "react": "^19.2.5", "react-dom": "^19.2.5", "react-hot-toast": "^2.6.0", "react-i18next": "^17.0.6", "react-window": "^2.2.7", - "tailwindcss": "4", + "tailwindcss": "^4.2.4", "wavesurfer.js": "^7.12.6", "zustand": "^5.0.12", }, "devDependencies": { "@eslint/js": "^10.0.1", - "@tauri-apps/api": "^2.10.1", - "@tauri-apps/cli": "^2.10.1", + "@tauri-apps/api": "^2.11.0", + "@tauri-apps/cli": "^2.11.0", "@types/react": "^19.2.14", "@types/react-dom": "^19.2.3", "@vitejs/plugin-react": "^6.0.1", - "eslint": "^10.2.1", + "eslint": "^10.3.0", "eslint-plugin-react-hooks": "^7.1.1", "eslint-plugin-react-refresh": "^0.5.2", - "globals": "^17.5.0", + "globals": "^17.6.0", "typescript": "^6.0.3", - "vite": "^8.0.9", + "vite": "^8.0.10", }, }, }, diff --git a/frontend/src-tauri/Cargo.lock b/frontend/src-tauri/Cargo.lock index 1017985f..dbf1fe71 100644 --- a/frontend/src-tauri/Cargo.lock +++ b/frontend/src-tauri/Cargo.lock @@ -2757,7 +2757,7 @@ dependencies = [ [[package]] name = "omnivoice-studio" -version = "0.2.6" +version = "0.2.7" dependencies = [ "enigo", "libc", diff --git a/frontend/src-tauri/src/lib.rs b/frontend/src-tauri/src/lib.rs index d4eb58ee..99dc9f5d 100644 --- a/frontend/src-tauri/src/lib.rs +++ b/frontend/src-tauri/src/lib.rs @@ -126,7 +126,7 @@ pub fn run() { .with_handler(move |app_handle, _shortcut, event| { if event.state == ShortcutState::Pressed { log::info!("Global shortcut triggered: dictation"); - if let Some(win) = app_handle.get_webview_window("main") { + if let Some(win) = app_handle.get_webview_window("widget") { let _ = win.show(); let _ = win.set_focus(); } diff --git a/frontend/src-tauri/tauri.conf.json b/frontend/src-tauri/tauri.conf.json index 61d9bfab..09e703ea 100644 --- a/frontend/src-tauri/tauri.conf.json +++ b/frontend/src-tauri/tauri.conf.json @@ -23,6 +23,20 @@ "fullscreen": false, "titleBarStyle": "Overlay", "hiddenTitle": true + }, + { + "label": "widget", + "title": "Dictation Widget", + "url": "/?window=widget", + "width": 350, + "height": 220, + "resizable": false, + "fullscreen": false, + "transparent": true, + "decorations": false, + "alwaysOnTop": true, + "visible": false, + "skipTaskbar": true } ], "security": { diff --git a/frontend/src/App.jsx b/frontend/src/App.jsx index 787a26e7..413c37b6 100644 --- a/frontend/src/App.jsx +++ b/frontend/src/App.jsx @@ -28,7 +28,7 @@ import Header from './components/Header'; import NavRail from './components/NavRail'; import ErrorBoundary from './components/ErrorBoundary'; import FloatingPill from './components/FloatingPill'; -import CaptureButton from './components/CaptureButton'; + import useRealtimeEvents from './hooks/useRealtimeEvents'; import { BootstrapSplash, useBootstrapStage } from './components/BootstrapSplash'; @@ -1767,7 +1767,7 @@ function App() { }}/> - +
* { @@ -69,8 +67,11 @@ border: 1px solid color-mix(in srgb, var(--chrome-border, #3c3836) 60%, transparent); border-radius: 16px; padding: 14px 16px; - min-width: 260px; - max-width: 320px; + width: 100%; + height: 100%; + display: flex; + flex-direction: column; + box-sizing: border-box; box-shadow: 0 8px 40px rgba(0, 0, 0, 0.35); animation: capture-slide-up 0.3s cubic-bezier(0.34, 1.56, 0.64, 1); } diff --git a/frontend/src/components/CaptureButton.jsx b/frontend/src/components/CaptureWidget.jsx similarity index 94% rename from frontend/src/components/CaptureButton.jsx rename to frontend/src/components/CaptureWidget.jsx index 0c4d3cd3..c4eee8ca 100644 --- a/frontend/src/components/CaptureButton.jsx +++ b/frontend/src/components/CaptureWidget.jsx @@ -2,7 +2,7 @@ import React, { useCallback, useEffect, useRef, useState } from 'react'; import { Mic, MicOff, Clipboard, X, Loader, Zap, Target, Check } from 'lucide-react'; import { toast } from 'react-hot-toast'; import { useAppStore } from '../store'; -import './CaptureButton.css'; +import './CaptureWidget.css'; import { API as API_BASE } from '../api/client'; import { addTranscription } from '../pages/Transcriptions'; @@ -33,11 +33,10 @@ const LS_AUTO_COPY = 'omni_capture_auto_copy'; * * Auto-copies to clipboard so users can immediately ⌘V into any app. */ -export default function CaptureButton() { +export default function CaptureWidget() { const [state, setState] = useState('idle'); // idle | recording | transcribing | done | error const [transcript, setTranscript] = useState(''); const [duration, setDuration] = useState(0); - const [expanded, setExpanded] = useState(false); const [captureMode, setCaptureMode] = useState(() => localStorage.getItem(LS_CAPTURE_MODE) || 'fast' ); @@ -137,6 +136,19 @@ export default function CaptureButton() { } catch { toast.success('Copied to clipboard — paste with ⌘V', { duration: 2000 }); } + + // Auto-dismiss the floating widget after 2.5 seconds so it gets out of the way + setTimeout(async () => { + setState('idle'); + setTranscript(''); + setDuration(0); + setCopied(false); + try { + const { getCurrentWindow } = await import('@tauri-apps/api/window'); + await getCurrentWindow().hide(); + } catch { /* not in Tauri */ } + }, 2500); + } catch { /* clipboard API may fail in some contexts */ } } }, [autoCopy]); @@ -371,12 +383,15 @@ export default function CaptureButton() { }); }, [transcript]); - const dismiss = () => { + const dismiss = async () => { setState('idle'); setTranscript(''); - setExpanded(false); setDuration(0); setCopied(false); + try { + const { getCurrentWindow } = await import('@tauri-apps/api/window'); + await getCurrentWindow().hide(); + } catch { /* not in Tauri */ } }; const toggleCapture = () => { @@ -395,11 +410,9 @@ export default function CaptureButton() { }; return ( -
- {/* Expanded panel */} - {expanded && ( +
-
+
{state === 'recording' && '🎙️ Listening…'} {state === 'transcribing' && '📝 Transcribing…'} @@ -491,22 +504,10 @@ export default function CaptureButton() {
-
+
{navigator.platform?.includes('Mac') ? '⌘' : 'Ctrl'}++Space
- )} - - {/* Main FAB button */} -
); } diff --git a/frontend/src/main-app.jsx b/frontend/src/main-app.jsx index 8b96b82c..404b3c24 100644 --- a/frontend/src/main-app.jsx +++ b/frontend/src/main-app.jsx @@ -27,11 +27,22 @@ const queryClient = new QueryClient({ }, }); +import { Suspense, lazy } from 'react'; +const CaptureWidget = lazy(() => import('./components/CaptureWidget.jsx')); + export function bootstrapApp() { + const isWidget = window.location.search.includes('window=widget'); + createRoot(document.getElementById('root')).render( - + {isWidget ? ( + Loading...
}> + + + ) : ( + + )} , ); diff --git a/pyproject.toml b/pyproject.toml index 9a389510..b8e9eaaf 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "omnivoice" -version = "0.2.4" +version = "0.2.7" description = "OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models" readme = "README.md" license = "Apache-2.0"