Skip to content

Backend Overview

Understanding the Subtide backend server.


The Subtide backend is a Flask-based REST API that handles:

  1. Video Processing - Downloading and extracting audio
  2. Transcription - Converting speech to text using Whisper
  3. Translation - Translating text using LLM APIs
  4. Streaming - Real-time translation for Tier 4
graph TB
    subgraph Backend[Subtide Backend]
        API[Flask API] --> Whisper[Whisper Service]
        Whisper --> Translation[Translation Service]

        Whisper --> MLX[MLX / Faster Whisper]
        Translation --> LLM[OpenAI / OpenRouter]

        API --> YouTube[YouTube Service<br/>yt-dlp]
    end

    style API fill:#4a90e2,stroke:#2c5aa0,color:#fff
    style Whisper fill:#7b68ee,stroke:#5a4fcf,color:#fff
    style Translation fill:#50c878,stroke:#3a9d5f,color:#fff
    style MLX fill:#ffa726,stroke:#f57c00,color:#fff
    style LLM fill:#26c6da,stroke:#00acc1,color:#fff
    style YouTube fill:#ff6b6b,stroke:#d63031,color:#fff

Option Best For Setup Difficulty
Binary Personal use, quick start Easy
Python Source Development, customization Medium
Docker Production, teams Medium
RunPod GPU acceleration, cloud Medium

Terminal window
# Download from releases
chmod +x subtide-backend-macos
./subtide-backend-macos
Terminal window
cd backend
docker-compose up subtide-tier2
Terminal window
cd backend
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
./run.sh

Subtide supports multiple Whisper implementations:

Optimized for M1/M2/M3/M4 Macs:

Terminal window
WHISPER_BACKEND=mlx ./subtide-backend
  • Uses unified memory efficiently
  • No separate GPU memory needed
  • Best for macOS users

CUDA-accelerated for NVIDIA GPUs:

Terminal window
WHISPER_BACKEND=faster ./subtide-backend
  • Requires CUDA toolkit
  • Significant speedup on supported GPUs
  • Best for Linux/Windows with NVIDIA

Original Whisper implementation:

Terminal window
WHISPER_BACKEND=openai ./subtide-backend
  • CPU-based (slower)
  • Works everywhere
  • Fallback option

Model Size Speed Quality
tiny ~39 MB Fastest Basic
base ~74 MB Fast Good
small ~244 MB Medium Better
medium ~769 MB Slow Great
large-v3 ~1.5 GB Slowest Best
large-v3-turbo ~800 MB Fast Excellent

Recommended: large-v3-turbo for best speed/quality ratio.


Mac Memory Recommended Model
8 GB Limited tiny, base
16 GB Good small, base
32 GB Excellent large-v3
64 GB+ Optimal Any model
GPU VRAM Recommended Model
RTX 3060 12 GB medium
RTX 3090/4080 16-24 GB large-v3
RTX 4090 24 GB Any model

Core configuration:

Variable Description Default
PORT Server port 5001
WHISPER_MODEL Model size base
WHISPER_BACKEND Backend type Auto
CORS_ORIGINS Allowed origins *

See Configuration for full list.


Verify the backend is running:

Terminal window
curl http://localhost:5001/health

Expected response:

{"status": "healthy"}