Files
VoiceStudio/CONTRIBUTING.md
T
2574fccaf6 docs: community docs refresh — README, CONTRIBUTING, SECURITY, SUPPORT, Docker/macOS install (#341)
* docs: refresh community docs to match the project's current reality

- README: download badges now point to releases/latest (were frozen at
  v0.2.7); Intel-Mac note (pre-built bundle is Apple Silicon; source
  works on Intel; pre-built Intel tracked in #279)
- SECURITY: supported-versions table 0.2.x -> 0.3.x + 0.2.7 legacy row
- docs/install/docker.md: tag mapping matches docker.yml after #338 —
  :latest is the rolling main preview, :stable (new) pins releases
- PR template: removed the abolished two-RC/48h-soak ceremony; documents
  continuous-to-main
- CONTRIBUTING: new sections — what bot review looks like (CodeRabbit +
  Greptile), conventional-commit + issue-link expectations, the quality
  gates (cross-platform parity, 21-locale i18n + CJK allowlist, alembic,
  engine back-compat, local-first, loopback security posture), and a
  contribution-licensing grant that keeps the AGPL + commercial
  dual-license viable
- SUPPORT.md: new — channels, before-you-file checklist, expectations
- docs/install/macos.md: Intel caveat aligned with reality

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: codify the docs-sync hard rule — behavior changes update their docs in the same PR

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(agents): rtk rules for Antigravity — token-compressed tool output

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 13:58:43 +05:30

8.8 KiB

Contributing to OmniVoice Studio

Thanks for your interest in improving OmniVoice Studio! This guide covers everything you need to get started.

💬 Chat Discord
🐛 Bugs GitHub Issues
🏷️ Good First Issues Filtered list
📋 Roadmap README → Roadmap

Development Setup

Prerequisites

  • Git
  • Bun (frontend package manager)
  • uv (Python environment manager)
  • ffmpeg (audio/video processing)
  • Python 3.10+ (managed automatically by uv)

Clone & Run

git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio
bun install
bun run dev

This starts both services:

Service URL What it does
Backend localhost:3900 FastAPI server — TTS, ASR, diarization, dubbing pipeline
Frontend localhost:3901 React + Vite UI

Desktop App (Tauri)

bun run desktop

Requires Rust and platform-specific Tauri dependencies — see the Tauri prerequisites.


Project Structure

OmniVoice-Studio/
├── backend/                 # Python FastAPI server
│   ├── api/                 # Route handlers
│   ├── core/                # Config, prefs, constants
│   └── services/            # TTS engines, ASR, dubbing, audio DSP
│       └── tts_backend.py   # ← Multi-engine TTS registry
├── frontend/                # React + Vite
│   ├── src/
│   │   ├── components/      # UI components
│   │   ├── hooks/           # Custom React hooks
│   │   ├── stores/          # Zustand state slices
│   │   └── utils/           # Shared utilities
│   └── src-tauri/           # Rust/Tauri desktop shell
├── deploy/                  # Docker, CI configs
├── docs/                    # Screenshots, MCP config
└── scripts/                 # Build & release scripts

How to Contribute

Bug Reports

Open an issue with:

  1. What happened vs what you expected
  2. Steps to reproduce
  3. OS, GPU, and Python version (find in Settings → Logs)
  4. Error logs (Settings → Logs → copy relevant lines)

Pull Requests

  1. Fork the repo and create a branch from main
  2. Keep PRs focused — one feature or fix per PR
  3. Run tests before pushing:
    # Backend tests
    uv run pytest backend/ -x -q
    
    # Frontend build check
    cd frontend && npx vite build --mode development
    
  4. Write a clear PR title — it becomes the squash-merge commit message
  5. Don't include local machine stats, file paths, or private system info in PR descriptions

Adding a New TTS Engine

OmniVoice's TTS backend is a plugin registry. Adding a new engine takes ~50 lines:

  1. Open backend/services/tts_backend.py
  2. Create a class extending TTSBackend:
class MyEngineBackend(TTSBackend):
    id = "my-engine"
    display_name = "My Engine (description)"

    @classmethod
    def is_available(cls) -> tuple[bool, str]:
        try:
            import my_engine  # noqa: F401
            return True, "ready"
        except ImportError:
            return False, "my_engine not installed. pip install my-engine"

    @property
    def sample_rate(self) -> int:
        return 24000

    @property
    def supported_languages(self) -> list[str]:
        return ["en", "zh"]

    def generate(self, text: str, **kw) -> torch.Tensor:
        # ... call your engine, return [1, num_samples] tensor
  1. Register it in _REGISTRY at the bottom of the file
  2. That's it — it auto-appears in Settings → TTS Engine

Code Style

Python (Backend)

  • Formatter: We don't enforce one globally — match the style of the file you're editing
  • Logging: Use logger.warning() / logger.error(), never bare print()
  • Exceptions: Avoid bare except: pass — catch specific exceptions
  • Type hints: Use them for public API functions and class methods

JavaScript/React (Frontend)

  • Components: Functional components with hooks
  • State: Zustand stores in src/stores/, organized by slice
  • CSS: Vanilla CSS in component-level files — no Tailwind
  • Naming: PascalCase for components, camelCase for hooks and utils

Rust (Tauri)

  • Format: cargo fmt before committing
  • Modules: One concern per file (bootstrap.rs, tools.rs, config.rs, commands.rs)

Commit Messages

Write clear, concise messages. The PR title becomes the squash-merge commit.

good: fix: prevent CUDA OOM during concurrent transcription + TTS
good: feat: add CosyVoice 3 TTS backend adapter
good: docs: add platform compatibility matrix to README

bad:  fixed stuff
bad:  update
bad:  WIP

Testing

# Run all backend tests
uv run pytest backend/ -x -q

# Run a specific test file
uv run pytest backend/tests/test_api.py -x -q

# Frontend build validation (no test suite yet)
cd frontend && npx vite build --mode development

# Tauri shell check (requires Rust)
cd frontend/src-tauri && cargo check


What code review looks like

Every PR is reviewed by two AI reviewers before a human looks at it:

  • CodeRabbit posts a walkthrough (with a sequence diagram, and an ASCII before/after sketch for UI changes), inline findings, and warning-mode pre-merge checks against the project's hard rules.
  • Greptile reviews with the same project rubrics and learns from 👍/👎 reactions on its comments — react to train it.

Both are advisory, not gating: CI and the maintainer's approval decide. Don't be surprised by detailed bot comments minutes after you open a PR — address what's right, push back (in a reply) on what's wrong.

Commit & PR conventions: conventional-commit style with a scope (fix(dub): …, feat(setup): …) and link the issue (Closes #N / Refs #N) in the title or body.

Quality gates your PR must pass

  • Cross-platform parity (hard rule): anything that ships in default mode must behave identically on macOS, Windows, and Linux. Platform-specific implementation is fine; platform-divergent default behavior is a P0. Platform-only features go behind an explicit opt-in (Settings toggle, env var, or CLI flag).
  • i18n — all 21 locales (hard rule): every user-facing string goes through t('...') and the key must exist in all 21 files under frontend/src/i18n/locales/. Translate; don't copy English into non-English locales. CI fails on hardcoded CJK outside the allowlist in tests/test_no_hardcoded_cjk.py (extend _ALLOWED_FILES with a justification for legitimate functional CJK).
  • DB schema changes go through an alembic migration with a tested upgrade path — existing omnivoice_data/ must keep working with no manual steps.
  • Engine back-compat: already-installed engines (model weights on disk) must not require reinstall or re-download.
  • Local-first: no new outbound calls except GitHub Issues (opt-in reporting) and HuggingFace model downloads. Never log or persist secrets or absolute home paths.
  • Security posture: the backend serves loopback HTTP — treat every query/path/form parameter as hostile. User-chosen filesystem destinations are authorized in the Tauri process (save dialog), never via HTTP params.

Contribution licensing

OmniVoice Studio is AGPL-3.0-only, and the maintainer also offers a commercial license (see LICENSE). By submitting a contribution you agree that:

  1. you have the right to submit it (your own work, or compatibly licensed);
  2. it is licensed to the project under AGPL-3.0; and
  3. you grant the project maintainer a perpetual, worldwide, non-exclusive right to also distribute your contribution under the project's commercial license terms.

This inbound grant is what keeps the dual-license model viable. If you can't agree to (3) for a particular contribution, say so in the PR and we'll discuss before merging. Adding a Signed-off-by: line (DCO) to your commits is appreciated but not required.


Need Help?

Thank you for contributing! 🎙️