v2.4.0 Studio Release Production Verified

System Architecture & Technical Documentation

A comprehensive technical reference detailing the ingestion pipeline, word-level speech alignment mathematics, ASS vector subtitle rasterization, and FFmpeg video synthesis for LyricSync Studio.

1. System Overview & Architectural Tenets

LyricSync Studio is an automated audio-to-video synchronization suite designed to bridge raw musical vocals and typography. Unlike traditional monolithic video renderers that require local graphics workstations or heavy desktop software, LyricSync functions with zero mandatory GPU dependencies and operates seamlessly in constrained cloud/shared hosting environments.

CPU-Efficient Synthesis

Leverages native single-pass FFmpeg libass filtergraphs, avoiding intermediate frame dumps or disk space bloat.

Dual-Tier Draft Mode

Low-bandwidth connections utilize 720p Fast Draft for near-instant rendering and low latency playback.

Word-Level Precision

Timestamp interpolation down to millisecond accuracy with musical phrasing gap filling.

Strict RBAC Separation

Technical server internals are restricted exclusively to system admin (silatanuikipngetich@gmail.com).

Clean Security Model

All sensitive credentials (API keys, database secrets, session salts) are strictly ingested through OS environment variables. The codebase never hardcodes plain-text secrets.

2. Interactive System Diagrams (7 Comprehensive Views)

The diagrams below document the exact end-to-end data lifecycle, component relationships, and execution flows within the application.

Diagram 1: End-to-End System Architecture Flow
Chronological lifecycle from media ingestion to synchronized streaming video delivery.
flowchart TD
    A["User Audio/Video Upload"] --> B["Flask Upload Controller"]
    B --> C["Media Probe (FFprobe)"]
    C --> D["Database (SQLite / MySQL)"]
    D --> E["Background Generator (16:9 Canvas)"]
    E --> F["OpenAI Whisper Word Transcription"]
    F --> G["Temporal Alignment & Gap Filler"]
    G --> H["Canonical Lyrics JSON Storage"]
    H --> I["Interactive Studio Editor"]
    I --> J["ASS Vector Subtitle Generator"]
    J --> K["FFmpeg Hardware/Software Encoder"]
    K --> L["Progressive Streaming MP4 Video"]
                        
Diagram 2: Dual-Tier Quality & Bandwidth Optimization Flow
Architecture allowing instant drafts on mobile and high-fidelity masters on desktop.
flowchart LR
    subgraph Client["Client Browser"]
        Q["User Selects Quality"]
    end
    subgraph Engine["FFmpeg Rendering Engine"]
        Q -->|720p Fast Draft| D720["Scale: 1280x720
CRF 23 | Preset: veryfast
AAC 128k (Low Bandwidth)"] Q -->|1080p Studio Master| D1080["Scale: 1920x1080
CRF 20 | Preset: medium
AAC 192k (Broadcast Quality)"] end D720 --> M["MP4 Stream (+faststart)"] D1080 --> M
Diagram 3: Word-Level Temporal Alignment Pipeline
Transformation of continuous speech streams into quantized lyric chips with musical boundaries.
flowchart TD
    W0["Audio Input Stream"] --> W1["OpenAI Whisper-1 API
(timestamp_granularities=['word'])"] W1 --> W2["Raw Timestamp Tokens (start, end, text)"] W2 --> W3{"Missing End Times?"} W3 -->|Yes| W4["Interpolate from Next Token Start"] W3 -->|No| W5["Normalize Syllable Characters"] W4 --> W5 W5 --> W6{"Inter-word Gap > 0.45s or Punctuation?"} W6 -->|Yes| W7["Segment into New Musical Lyric Line"] W6 -->|No| W8["Append to Current Stanza"] W7 --> W9["Canonical Word-Level JSON Dataset"] W8 --> W9
Diagram 4: FFmpeg Filtergraph & Subtitle Blending Sequence
Single-pass filter composition merging audio streams, aesthetic canvases, and ASS scripts.
sequenceDiagram
    autonumber
    participant UI as Studio Editor
    participant API as Flask Render Route
    participant SUB as Subtitle Service (ASS)
    participant FF as FFmpeg Subprocess
    participant FS as Media File Storage

    UI->>API: POST /api/projects/{id}/render (Tier: 720/1080)
    API->>SUB: generate_ass_subtitles(canonical_lyrics, style)
    SUB-->>API: lyric_script.ass
    API->>FF: Spawn: ffmpeg -i visual -i audio -vf "scale,ass=lyric_script.ass" -c:v libx264 -c:a aac -movflags +faststart output.mp4
    FF->>FS: Stream progressive frames
    FF-->>API: Process Return Code 0
    API-->>UI: Render Job Completed (HTTP 200)
    UI->>FS: GET /api/projects/{id}/download
                        
Diagram 5: Relational Database Schema (ERD Model)
7 primary tables supporting users, versioned projects, lyrics, assets, and render queues.
erDiagram
    USERS ||--o{ PROJECTS : owns
    PROJECTS ||--o{ MEDIA_ASSETS : contains
    PROJECTS ||--o{ LYRIC_LINES : segments
    LYRIC_LINES ||--o{ LYRIC_WORDS : contains
    PROJECTS ||--o{ TRANSCRIPTIONS : stores
    PROJECTS ||--o{ RENDER_JOBS : queues

    USERS {
        string id PK
        string email UK
        string display_name
        string password_hash
        string google_id UK
        string avatar_url
        datetime created_at
    }

    PROJECTS {
        string id PK
        string user_id FK
        string name
        string status
        float audio_duration
        integer width
        integer height
        integer current_revision
        text canonical_json
        datetime created_at
    }

    LYRIC_LINES {
        string id PK
        string project_id FK
        integer line_index
        float start_time
        float end_time
        text text_content
    }

    LYRIC_WORDS {
        string id PK
        string line_id FK
        integer word_index
        float start_time
        float end_time
        string word_text
        float confidence
    }

    RENDER_JOBS {
        string id PK
        string project_id FK
        string status
        integer progress
        string stage
        string output_video_path
        datetime created_at
    }
                        
Diagram 6: Role-Based Access Control Architecture (RBAC)
Enforcing strict separation of deep diagnostic technicalities from creator interfaces.
flowchart TD
    Req["Incoming User Request"] --> Auth{"Authenticated?"}
    Auth -->|No| Pub["Public View: Clean status (200), Zero Server Internals"]
    Auth -->|Yes| Chk{"Email == silatanuikipngetich@gmail.com?"}
    Chk -->|Yes| Adm["Admin Privilege: Full Diagnostics
• Table Names • FFmpeg Binaries
• AI Model Specs • Storage Paths
• Raw Health Telemetry"] Chk -->|No| Usr["Normal Creator Privilege:
• Music-Focused UI • Word-Level Sync
• 720p/1080p HQ Export
• Technical Jargon Completely Hidden"]
Diagram 7: Studio Realtime Playback Synchronization Loop
Sub-second synchronization loop binding HTML5 media elements, active timeline chips, and video.
stateDiagram-v2
    [*] --> Idle
    Idle --> Playing : play() triggered
    state Playing {
        [*] --> SyncLoop
        SyncLoop --> ReadTime : requestAnimationFrame()
        ReadTime --> BinarySearchWord : Lookup word [start_time, end_time]
        BinarySearchWord --> HighlightActiveChip : Add .active-chip class
        HighlightActiveChip --> UpdateCanvasOverlay : Redraw sub-second text
        UpdateCanvasOverlay --> SyncLoop
    }
    Playing --> Paused : pause() triggered
    Paused --> Playing : resume()
    Playing --> ScrubberSeek : User drags scrubber
    ScrubberSeek --> Playing : currentTime updated
                        

3. Media Ingestion & Safety Probing Engine

Media files uploaded to the studio undergo validation and sanitization. The duration, channels, and sample rates are extracted via UTF-8 safe FFprobe subprocesses that guarantee stability across Windows and Linux hosting environments.

Python app/services/media_probe.py
import os
import shutil
import subprocess
from pathlib import Path
from typing import Dict, Any, Optional

def get_ffprobe_binary() -> Optional[str]:
    """Resolve ffprobe executable from system PATH or virtualenv binaries."""
    env_bin = os.getenv("FFPROBE_BINARY")
    if env_bin and (shutil.which(env_bin) or Path(env_bin).is_file()):
        return env_bin
    system_bin = shutil.which("ffprobe")
    if system_bin:
        return system_bin
    return None

def probe_media_file(file_path: str) -> Dict[str, Any]:
    """Inspect duration, streams, and codecs using UTF-8 error-tolerant decoding."""
    ffprobe = get_ffprobe_binary()
    if not ffprobe or not Path(file_path).is_file():
        return {"duration": 0.0, "has_audio": True, "has_video": False}

    cmd = [
        ffprobe, "-v", "error",
        "-show_entries", "format=duration:stream=codec_type,width,height",
        "-of", "default=noprint_wrappers=1:nokey=1",
        str(file_path)
    ]
    # Enforce UTF-8 with replace mode to eliminate ascii decode errors on remote servers
    res = subprocess.run(
        cmd,
        stdout=subprocess.PIPE,
        stderr=subprocess.PIPE,
        encoding="utf-8",
        errors="replace"
    )
    lines = [ln.strip() for ln in res.stdout.splitlines() if ln.strip()]
    duration = 0.0
    for line in lines:
        try:
            val = float(line)
            if val > duration:
                duration = val
        except ValueError:
            pass

    return {
        "duration": duration,
        "raw_probe": lines
    }

4. AI Word-Level Alignment & Transcription Pipeline

Speech-to-text recognition requests OpenAI Whisper with word timestamp granularities enabled. The resulting word tokens are processed through an interpolation pass and split into musical lines using pause detection (Δt > 0.45s) and punctuation rules.

Python app/services/alignment.py
import re
from typing import List, Dict, Any

PAUSE_SPLIT_THRESHOLD_SECONDS = 0.45
PUNCTUATION_BREAK_PATTERN = re.compile(r'[\.\?\!\;\:\,\—\-]')

def segment_transcript_words_into_lines(words: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
    """Segment a continuous list of word timestamps into structured musical lyric lines.
    Splits are triggered upon:
      1. Vocal pauses where gap between word_n.end and word_{n+1}.start exceeds 0.45s
      2. Natural phrase punctuation (. ! ? ; , -)
      3. Maximum word constraints per visual line (typically 8-10 words)
    """
    if not words:
        return []

    lines = []
    current_words = []
    
    for idx, w in enumerate(words):
        current_words.append(w)
        is_last_word = (idx == len(words) - 1)
        
        # Check pause gap to next word
        has_pause = False
        if not is_last_word:
            next_w = words[idx + 1]
            gap = float(next_w.get("start", 0)) - float(w.get("end", 0))
            if gap >= PAUSE_SPLIT_THRESHOLD_SECONDS:
                has_pause = True
                
        # Check terminal punctuation in word text
        has_punct = bool(PUNCTUATION_BREAK_PATTERN.search(w.get("word", "")))
        is_lengthy = len(current_words) >= 9

        if is_last_word or has_pause or has_punct or is_lengthy:
            line_text = " ".join([item.get("word", "").strip() for item in current_words]).strip()
            line_start = float(current_words[0].get("start", 0))
            line_end = float(current_words[-1].get("end", line_start + 1.0))
            
            lines.append({
                "line_id": f"line_{len(lines) + 1}",
                "start": round(line_start, 3),
                "end": round(line_end, 3),
                "text": line_text,
                "words": current_words
            })
            current_words = []
            
    return lines
Python app/services/openai_transcription.py
import os
from openai import OpenAI
from config import Config

class OpenAITranscriptionService:
    def __init__(self, api_key: str = None, model: str = None):
        # API key is securely retrieved from environment configuration
        self.api_key = api_key or Config.OPENAI_API_KEY or os.getenv("OPENAI_API_KEY", "")
        self.model = model or Config.OPENAI_TRANSCRIPTION_MODEL or "whisper-1"
        self.client = OpenAI(api_key=self.api_key) if self.api_key else None

    def transcribe_audio_words(self, audio_file_path: str) -> dict:
        """Call Whisper API with timestamp_granularities=['word'] enabled."""
        if not self.client:
            raise ValueError("OpenAI client not initialized. Set OPENAI_API_KEY in environment.")

        with open(audio_file_path, "rb") as audio_file:
            transcript = self.client.audio.transcriptions.create(
                file=audio_file,
                model=self.model,
                response_format="verbose_json",
                timestamp_granularities=["word"]
            )
            
        return transcript.to_dict()

5. Advanced SubStation Alpha (ASS) Vector Subtitles

Rather than burning rasterized text images, LyricSync compiles lyric timings into native Advanced SubStation Alpha scripts (.ass). Timing tags (\k<duration>) enable sub-frame syllable highlighting with crisp typography.

Python app/services/subtitles.py
def format_ass_time(seconds: float) -> str:
    """Format floating point seconds into ASS timestamp standard (H:MM:SS.cs)."""
    hrs = int(seconds // 3600)
    mins = int((seconds % 3600) // 60)
    secs = int(seconds % 60)
    centis = int(round((seconds - int(seconds)) * 100))
    if centis >= 100:
        centis = 99
    return f"{hrs}:{mins:02d}:{secs:02d}.{centis:02d}"

def generate_ass_script(lines: list, style: dict, width: int = 1920, height: int = 1080) -> str:
    """Construct complete ASS subtitle specification with karaoke timing wipes."""
    font_name = style.get("font", "Outfit")
    font_size = style.get("font_size", 42)
    primary_color = "&H00FFFFFF"  # White in ASS &HAABBGGRR format
    karaoke_color = "&H00DF77B7"  # Mint/Burgundy accent

    ass_content = [
        "[Script Info]",
        "ScriptType: v4.00+",
        f"PlayResX: {width}",
        f"PlayResY: {height}",
        "ScaledBorderAndShadow: yes",
        "",
        "[V4+ Styles]",
        "Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding",
        f"Style: Default,{font_name},{font_size},{primary_color},{karaoke_color},&H00000000,&H80000000,-1,0,0,0,100,100,0,0,1,2,3,2,60,60,80,1",
        "",
        "[Events]",
        "Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text"
    ]

    for line in lines:
        start_ts = format_ass_time(line.get("start", 0))
        end_ts = format_ass_time(line.get("end", 0))
        words = line.get("words", [])
        
        # Build \k tags for karaoke syllable transitions
        dialogue_text = ""
        for w in words:
            duration_cs = max(1, int(round((w.get("end", 0) - w.get("start", 0)) * 100)))
            dialogue_text += f"{{\\k{duration_cs}}}{w.get('word', '')} "
            
        ass_content.append(f"Dialogue: 0,{start_ts},{end_ts},Default,,0,0,0,,{dialogue_text.strip()}")

    return "\n".join(ass_content)

6. FFmpeg Dual-Tier Video Rendering Engine

The rendering pipeline synthesizes the visual background and audio track while rasterizing ASS subtitles in a single execution pass. To minimize latency on mobile or constrained bandwidth, a 720p Fast Draft profile renders in seconds.

Python app/services/ffmpeg.py
import subprocess
from app.services.media_probe import get_ffmpeg_binary

def execute_render_job(video_in: str, audio_in: str, ass_path: str, output_mp4: str, resolution: str = "720"):
    """Compile final synchronized lyric MP4.
    Resolution '720' uses CRF 23 with veryfast preset for instant preview.
    Resolution '1080' uses CRF 20 with medium preset for broadcast masters.
    """
    ffmpeg_bin = get_ffmpeg_binary()
    is_draft = (resolution == "720")
    
    scale_filter = "scale=-2:720" if is_draft else "scale=-2:1080"
    preset = "veryfast" if is_draft else "medium"
    crf = "23" if is_draft else "20"
    audio_bitrate = "128k" if is_draft else "192k"

    # Escape path separators for FFmpeg libass filter syntax
    escaped_ass = str(ass_path).replace("\\", "/").replace(":", "\\:")
    vf_chain = f"{scale_filter},ass='{escaped_ass}'"

    cmd = [
        ffmpeg_bin, "-y",
        "-i", str(video_in),
        "-i", str(audio_in),
        "-vf", vf_chain,
        "-c:v", "libx264",
        "-preset", preset,
        "-crf", crf,
        "-pix_fmt", "yuv420p",
        "-c:a", "aac",
        "-b:a", audio_bitrate,
        "-shortest",
        "-movflags", "+faststart",  # Optimizes MP4 header for instant web streaming
        str(output_mp4)
    ]

    process = subprocess.run(
        cmd,
        stdout=subprocess.PIPE,
        stderr=subprocess.PIPE,
        encoding="utf-8",
        errors="replace"
    )
    if process.returncode != 0:
        raise RuntimeError(f"FFmpeg render failure: {process.stderr[-500:]}")
        
    return output_mp4

7. Database Entities & ORM Architecture

Built with SQLAlchemy 2.0 with strict typed mappings. Cascade deletions ensure uploaded media, transcriptions, and cached sync revisions are cleaned atomically when a project is deleted.

Python app/models/project.py
from datetime import datetime, timezone
from sqlalchemy import String, Float, Integer, Text, ForeignKey, DateTime
from sqlalchemy.orm import Mapped, mapped_column, relationship
from app.extensions import db
from app.utils.ids import generate_project_id

class Project(db.Model):
    __tablename__ = "projects"

    id: Mapped[str] = mapped_column(String(64), primary_key=True, default=generate_project_id)
    user_id: Mapped[str] = mapped_column(String(64), ForeignKey("users.id"), nullable=True)
    name: Mapped[str] = mapped_column(String(255), nullable=False)
    status: Mapped[str] = mapped_column(String(32), default="created")
    audio_duration: Mapped[float] = mapped_column(Float, default=0.0)
    width: Mapped[int] = mapped_column(Integer, default=1920)
    height: Mapped[int] = mapped_column(Integer, default=1080)
    current_revision: Mapped[int] = mapped_column(Integer, default=1)
    canonical_json: Mapped[str] = mapped_column(Text, nullable=True)
    created_at: Mapped[datetime] = mapped_column(DateTime, default=lambda: datetime.now(timezone.utc))

    user = relationship("User", back_populates="projects")
    media_assets = relationship("MediaAsset", back_populates="project", cascade="all, delete-orphan")
    lyric_lines = relationship("LyricLine", back_populates="project", cascade="all, delete-orphan")
    render_jobs = relationship("RenderJob", back_populates="project", cascade="all, delete-orphan")

8. Role-Based Access Control (RBAC) & Diagnostics Guard

System technicalities (such as AI model parameters, FFmpeg binary locations, SQLite/MySQL table schemas, and disk diagnostics) are strictly gated so only the authorized administrator can inspect them.

Python app/routes/views.py (Gated Health Route)
from flask import Blueprint, jsonify, current_app
from flask_login import current_user

ADMIN_EMAIL = "silatanuikipngetich@gmail.com"

@views_bp.route("/health")
def health_check():
    """Restricted health diagnostics.
    Only silatanuikipngetich@gmail.com can inspect technical internals.
    Normal users receive a clean operational status.
    """
    is_admin = bool(
        current_user.is_authenticated and
        current_user.email and
        current_user.email.strip().lower() == ADMIN_EMAIL
    )

    if not is_admin:
        # Hides all server paths, table schemas, and binary paths
        return {
            "status": "healthy",
            "database": "connected",
            "service": "LyricSync Studio",
            "message": "LyricSync Studio is operational."
        }, 200

    # Full technical diagnostics reserved for administrator
    return {
        "status": "healthy",
        "database": "connected",
        "database_tables": inspector.get_table_names(),
        "storage_path": str(current_app.config.get("MEDIA_ROOT")),
        "ffmpeg_binary": get_ffmpeg_binary(),
        "ai_model": current_app.config.get("OPENAI_TRANSCRIPTION_MODEL", "whisper-1"),
        "rendering_engine": "FFmpeg (libass + h264 + aac)",
        "admin_access": True
    }, 200

9. Studio Frontend Engine & State Synchronization

The frontend editor features bidirectional scrubbing: scrubbing the audio player immediately repositions the timeline chips and video preview, while clicking any lyric chip jumps audio playback to the exact timestamp.

JavaScript static/js/editor.js
// Sub-second playback synchronization loop
function startSyncLoop() {
    function loop() {
        if (!audioEl.paused) {
            const currentTime = audioEl.currentTime;
            scrubber.value = currentTime;
            currentTimeDisplay.textContent = formatTime(currentTime);

            // Synchronize active lyric chips
            const activeLine = timeline.findLineAt(currentTime);
            if (activeLine) {
                timeline.highlightLine(activeLine.line_id);
                overlayEl.textContent = activeLine.text;
            }
            
            // Align background video
            if (videoEl && Math.abs(videoEl.currentTime - currentTime) > 0.3) {
                videoEl.currentTime = currentTime;
            }
        }
        requestAnimationFrame(loop);
    }
    requestAnimationFrame(loop);
}

10. Complete REST API Specifications

All API endpoints return JSON conforming to the standardized response envelope: {"success": true, ...} on 200 OK or {"success": false, "error": {"code": "...", "message": "..."}} on errors.

Method Endpoint Role Description
GET /api/projects Public / User List projects created by current user or public shelf
POST /api/projects Creator Multipart upload for audio track + background selection
PUT /api/projects/{id}/lyrics Creator Save revised lyric timings and word timestamps
POST /api/projects/{id}/transcribe Creator Trigger background Whisper word transcription task
POST /api/projects/{id}/render Creator Queue FFmpeg MP4 render job (resolution: 720 / 1080)
GET /health Admin / Public Diagnostics (Deep stats for silatanuikipngetich@gmail.com)