Files
VoiceStudio/docs/STRUCTURE.md
T
Palash Debnath 545b39c912 feat: Scalar API docs, community health files, Quickstart cards (#41)
* feat: Scalar API docs, community health files, Quickstart cards, GHCR Docker

Backend:
- Replace Swagger UI with Scalar at /docs (scalar-fastapi)
- Add OpenAI-compatible /v1/audio endpoints (openai_compat router)
- Add TTS streaming endpoint (tts_stream router)
- Add voice marketplace router (marketplace)
- Update TTS backend registry

Frontend:
- Refine CaptureWidget, WaveformTimeline, App layout
- CSS polish and index.css updates

Community health:
- SECURITY.md — vulnerability reporting policy
- CODE_OF_CONDUCT.md — Contributor Covenant v2.1
- .github/FUNDING.yml — GitHub Sponsors
- .github/ISSUE_TEMPLATE/ — bug report + feature request
- .github/pull_request_template.md — PR checklist

README:
- Quickstart redesigned as 3-column progressive cards
- Docker section updated with GHCR pull instructions
- API Docs row added to service table

Infra:
- scalar-fastapi added to pyproject.toml + uv.lock
- research/ added to .gitignore

* refactor: clean up documentation and logging while enhancing desktop packaging dependencies and capture UI performance.

* fix: address CodeRabbit review — streaming, escaping, thresholds

Backend:
- marketplace: stream zip entries via ZipFile.open()/copyfileobj, add 100MB
  upload cap, fix raise-from exception chaining (OOM prevention)
- openai_compat: _encode_audio returns actual file ext so Content-Disposition
  matches real format; forward non-profile voices when DB row not found
- tts_stream: send 'start' frame after generation so sample_rate is real;
  forward non-profile voices on DB miss
- capture_ws: split MIN_BUFFER_BYTES into separate partial/final thresholds
  so short utterances (<2s) still get transcribed

Frontend (Tauri):
- lib.rs: tray 'dictate' now toggles start/stop based on widget visibility
- commands.rs: XML-escape exe path in LaunchAgent plist, shell-quote in
  .desktop Exec line to prevent injection from special-char paths
- CaptureWidget.css: fix Stylelint violations (empty lines, font-family quotes)
2026-05-04 10:57:37 +05:30

180 lines
8.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Project Structure
Every folder has a single job. Every file at the root earns its place.
## Layout
```
OmniVoice/
│
├── README.md ⟵ user-facing overview
├── CHANGELOG.md ⟵ release history
├── LICENSE
│
├── pyproject.toml ⟵ Python project manifest
├── uv.lock ⟵ Python lockfile
├── package.json ⟵ monorepo manifest (Bun workspaces + Turborepo)
├── bun.lock ⟵ JS lockfile
├── turbo.json ⟵ turborepo pipeline
│
├── .dockerignore ⟵ Docker build context filter
├── backend.spec ⟵ pyinstaller spec (stays at root by pyinstaller convention)
├── alembic.ini ⟵ DB migration config (stays at root by alembic convention)
│
├── .env ⟵ user config; gitignored, .env.example is the template
├── .gitignore
│
├── backend/ ⟵ FastAPI server
│ ├── main.py
│ ├── api/routers/ HTTP endpoints (thin)
│ ├── core/ config, db, task queue, metrics
│ ├── services/ business logic
│ └── schemas/ pydantic request/response shapes
│
├── frontend/ ⟵ React 19 + Vite + Tauri desktop
│ ├── src/
│ │ ├── pages/ one file per top-level view
│ │ ├── components/ reusable UI
│ │ ├── api/ typed API clients
│ │ ├── store/ Zustand slices
│ │ ├── hooks/ custom React hooks
│ │ └── utils/
│ ├── src-tauri/ Rust desktop shell
│ └── public/
│
├── omnivoice/ ⟵ the underlying TTS model package
│ ├── models/
│ ├── cli/ CLI entry points (omnivoice-infer, etc.)
│ ├── data/ data utilities used by the model
│ ├── eval/ evaluation scripts
│ ├── scripts/ one-off utilities that ship with the package
│ ├── training/
│ └── utils/
│
├── tests/ ⟵ all tests live here, no exceptions
│ ├── conftest.py
│ ├── test_api.py
│ ├── test_dub_*.py
│ ├── test_job_queue.py
│ ├── test_segmentation.py
│ └── frontend/ Node-based frontend tests
│
├── scripts/ ⟵ dev / build / release shell + python scripts
│ ├── install.sh universal installer
│ ├── run.sh universal launcher
│ ├── smoke-test.sh end-to-end validation
│ └── desktop-prod.sh production desktop build
│
├── deploy/ ⟵ Docker deployment configs
│ ├── Dockerfile single-stage CUDA image
│ └── docker-compose.yml one-click local deployment
│
├── docs/ ⟵ developer docs, screenshots, branding
│ ├── ROADMAP.md where this project is going
│ ├── STRUCTURE.md you are here
│ ├── mcp.json MCP config template
│ ├── preview.png README hero image
│ ├── logo.png, logo.svg branding assets
│ ├── screenshot-*.png feature screenshots
│ ├── languages.md
│ ├── training.md
│ ├── data_preparation.md
│ ├── evaluation.md
│ └── voice-design.md
│
├── design/ ⟵ ASCII mockups of the target UX
│ ├── README.md
│ └── 00–08-*.md per-feature specs
│
├── research/ ⟵ reference material, competitor analysis, archived code
│ ├── LEARNINGS.md competitive analysis, what to absorb
│ ├── TheWhisper/ vendored reference (read-only)
│ ├── voice-pro/ vendored reference
│ └── legacy_gradio/ archived Gradio UI (pre-React rewrite)
│
├── examples/ ⟵ runnable demos + sample inputs
│
├── omnivoice_data/ ⟵ Docker bind-mount target (gitignored)
│ DB + HF cache live here when running via compose
│
└── .git/
```
## Rules of the root
1. **Nothing at the root is a runtime artifact.** Outputs, temp files, local DBs, crash logs — all go to `~/Library/Application Support/OmniVoice/` (or the OS equivalent), *never* into the repo. The one exception is `omnivoice_data/`, which exists as a bind-mount anchor for Docker.
2. **No ad-hoc scripts at the root.** One-off debug scripts live in `scripts/`. Tests live in `tests/`. Benchmarks live in `scripts/benchmarks/` (when we create them).
3. **Each subdirectory owns one concern.** If you can't describe what goes in a directory in one sentence, it's wrong.
4. **Every package has a manifest.** `backend/`, `frontend/`, `omnivoice/` each have their own deps declared via `pyproject.toml` / `package.json` — they are independently testable.
## What lives where
| Kind of thing | Goes in |
|---|---|
| User-facing product code | `backend/`, `frontend/` |
| The TTS model (independent of the studio) | `omnivoice/` |
| Everything executable but not user-facing | `scripts/` |
| Tests | `tests/` |
| Developer + user docs (Markdown) | `docs/` |
| Target-state mockups | `design/` |
| Competitor clones, legacy code, ref material | `research/` |
| Runnable demos and sample data | `examples/` |
| Runtime data (never committed) | `~/Library/Application Support/OmniVoice/` on Mac |
## What *doesn't* live at the root anymore
Removed in the cleanup pass:
| File | Why it was there | Where it went |
|---|---|---|
| `test_crash.py`, `test_server.py`, `test_whisper.py`, `test_mock.py`, `test_pyannote.py` | One-off debug scripts from an April 14 crash investigation. Imported symbols that no longer exist after the router refactor. | Deleted (already gitignored, referenced dead code). |
| `benchmark.py` | Another stale debug script; imported `backend.main._get_db` which no longer exists. | Deleted. |
| `output.wav`, `test.wav` | Runtime artifacts. | Deleted / moved out. |
| `crash_log.txt` | Runtime log. Now written to `$DATA_DIR/crash_log.txt`. | Deleted. |
| `omnivoice.zip` (148 MB) | Offline reference archive of the project itself. | Moved out of the repo to `../omnivoice.zip.bak`. |
| `data/` | Only contained `.DS_Store`. | Deleted. |
| `legacy_gradio/` | The pre-React Gradio UI. Kept for historical reference. | Archived to `research/legacy_gradio/`. |
| Scattered `.DS_Store` files | macOS Finder droppings. | Deleted from every non-ignored directory. |
## Scaling path (proposed, not yet executed)
The current flat layout works fine for the current size. If the project grows to include additional apps (a mobile companion, a plugin SDK, multiple backends), migrate to a Turborepo-style monorepo:
```
OmniVoice/
├── apps/
│ ├── api/ ← was backend/
│ ├── web/ ← was frontend/
│ └── desktop/ ← could extract src-tauri/ here later
├── packages/
│ ├── omnivoice-model/ ← was omnivoice/
│ └── tts-adapters/ ← new; the pluggable TTS interface from ROADMAP phase 3
├── config/
│ ├── docker/
│ └── pyinstaller/
├── tests/
├── docs/
├── design/
└── research/
```
**Do not execute this migration without a dedicated PR.** It breaks:
- `pyproject.toml` `[tool.hatch.build.targets.{sdist,wheel}]` paths
- `package.json` workspaces and scripts
- `turbo.json`, `Dockerfile`, `docker-compose.yml` paths
- `backend.spec` (`['backend/main.py']`, `pathex=['.']`)
- `frontend/src-tauri/tauri.*.conf.json` sidecar paths
- every import that reads `from backend.main import …` (tests, scripts)
Migrate when adding the second `apps/*` or the second `packages/*`. Not before.
## Conventions
- **Filenames:** snake_case for Python, kebab-case or PascalCase for JS/TS components, lowercase for Markdown.
- **Tests mirror source paths.** `backend/services/dub_pipeline.py` → `tests/services/test_dub_pipeline.py`.
- **One-off scripts** go into `scripts/` with a descriptive name, not `test_*.py` at the root.
- **New top-level directories** require a PR that updates *this file*.