System Architecture & Technical Documentation
A comprehensive technical reference detailing the ingestion pipeline, word-level speech alignment mathematics, ASS vector subtitle rasterization, and FFmpeg video synthesis for LyricSync Studio.
1. System Overview & Architectural Tenets
LyricSync Studio is an automated audio-to-video synchronization suite designed to bridge raw musical vocals and typography. Unlike traditional monolithic video renderers that require local graphics workstations or heavy desktop software, LyricSync functions with zero mandatory GPU dependencies and operates seamlessly in constrained cloud/shared hosting environments.
CPU-Efficient Synthesis
Leverages native single-pass FFmpeg libass filtergraphs, avoiding intermediate frame dumps or disk space bloat.
Dual-Tier Draft Mode
Low-bandwidth connections utilize 720p Fast Draft for near-instant rendering and low latency playback.
Word-Level Precision
Timestamp interpolation down to millisecond accuracy with musical phrasing gap filling.
Strict RBAC Separation
Technical server internals are restricted exclusively to system admin (silatanuikipngetich@gmail.com).
All sensitive credentials (API keys, database secrets, session salts) are strictly ingested through OS environment variables. The codebase never hardcodes plain-text secrets.
2. Interactive System Diagrams (7 Comprehensive Views)
The diagrams below document the exact end-to-end data lifecycle, component relationships, and execution flows within the application.
flowchart TD
A["User Audio/Video Upload"] --> B["Flask Upload Controller"]
B --> C["Media Probe (FFprobe)"]
C --> D["Database (SQLite / MySQL)"]
D --> E["Background Generator (16:9 Canvas)"]
E --> F["OpenAI Whisper Word Transcription"]
F --> G["Temporal Alignment & Gap Filler"]
G --> H["Canonical Lyrics JSON Storage"]
H --> I["Interactive Studio Editor"]
I --> J["ASS Vector Subtitle Generator"]
J --> K["FFmpeg Hardware/Software Encoder"]
K --> L["Progressive Streaming MP4 Video"]
flowchart LR
subgraph Client["Client Browser"]
Q["User Selects Quality"]
end
subgraph Engine["FFmpeg Rendering Engine"]
Q -->|720p Fast Draft| D720["Scale: 1280x720
CRF 23 | Preset: veryfast
AAC 128k (Low Bandwidth)"]
Q -->|1080p Studio Master| D1080["Scale: 1920x1080
CRF 20 | Preset: medium
AAC 192k (Broadcast Quality)"]
end
D720 --> M["MP4 Stream (+faststart)"]
D1080 --> M
flowchart TD
W0["Audio Input Stream"] --> W1["OpenAI Whisper-1 API
(timestamp_granularities=['word'])"]
W1 --> W2["Raw Timestamp Tokens (start, end, text)"]
W2 --> W3{"Missing End Times?"}
W3 -->|Yes| W4["Interpolate from Next Token Start"]
W3 -->|No| W5["Normalize Syllable Characters"]
W4 --> W5
W5 --> W6{"Inter-word Gap > 0.45s or Punctuation?"}
W6 -->|Yes| W7["Segment into New Musical Lyric Line"]
W6 -->|No| W8["Append to Current Stanza"]
W7 --> W9["Canonical Word-Level JSON Dataset"]
W8 --> W9
sequenceDiagram
autonumber
participant UI as Studio Editor
participant API as Flask Render Route
participant SUB as Subtitle Service (ASS)
participant FF as FFmpeg Subprocess
participant FS as Media File Storage
UI->>API: POST /api/projects/{id}/render (Tier: 720/1080)
API->>SUB: generate_ass_subtitles(canonical_lyrics, style)
SUB-->>API: lyric_script.ass
API->>FF: Spawn: ffmpeg -i visual -i audio -vf "scale,ass=lyric_script.ass" -c:v libx264 -c:a aac -movflags +faststart output.mp4
FF->>FS: Stream progressive frames
FF-->>API: Process Return Code 0
API-->>UI: Render Job Completed (HTTP 200)
UI->>FS: GET /api/projects/{id}/download
erDiagram
USERS ||--o{ PROJECTS : owns
PROJECTS ||--o{ MEDIA_ASSETS : contains
PROJECTS ||--o{ LYRIC_LINES : segments
LYRIC_LINES ||--o{ LYRIC_WORDS : contains
PROJECTS ||--o{ TRANSCRIPTIONS : stores
PROJECTS ||--o{ RENDER_JOBS : queues
USERS {
string id PK
string email UK
string display_name
string password_hash
string google_id UK
string avatar_url
datetime created_at
}
PROJECTS {
string id PK
string user_id FK
string name
string status
float audio_duration
integer width
integer height
integer current_revision
text canonical_json
datetime created_at
}
LYRIC_LINES {
string id PK
string project_id FK
integer line_index
float start_time
float end_time
text text_content
}
LYRIC_WORDS {
string id PK
string line_id FK
integer word_index
float start_time
float end_time
string word_text
float confidence
}
RENDER_JOBS {
string id PK
string project_id FK
string status
integer progress
string stage
string output_video_path
datetime created_at
}
flowchart TD
Req["Incoming User Request"] --> Auth{"Authenticated?"}
Auth -->|No| Pub["Public View: Clean status (200), Zero Server Internals"]
Auth -->|Yes| Chk{"Email == silatanuikipngetich@gmail.com?"}
Chk -->|Yes| Adm["Admin Privilege: Full Diagnostics
• Table Names • FFmpeg Binaries
• AI Model Specs • Storage Paths
• Raw Health Telemetry"]
Chk -->|No| Usr["Normal Creator Privilege:
• Music-Focused UI • Word-Level Sync
• 720p/1080p HQ Export
• Technical Jargon Completely Hidden"]
stateDiagram-v2
[*] --> Idle
Idle --> Playing : play() triggered
state Playing {
[*] --> SyncLoop
SyncLoop --> ReadTime : requestAnimationFrame()
ReadTime --> BinarySearchWord : Lookup word [start_time, end_time]
BinarySearchWord --> HighlightActiveChip : Add .active-chip class
HighlightActiveChip --> UpdateCanvasOverlay : Redraw sub-second text
UpdateCanvasOverlay --> SyncLoop
}
Playing --> Paused : pause() triggered
Paused --> Playing : resume()
Playing --> ScrubberSeek : User drags scrubber
ScrubberSeek --> Playing : currentTime updated
3. Media Ingestion & Safety Probing Engine
Media files uploaded to the studio undergo validation and sanitization. The duration, channels, and sample rates are extracted via UTF-8 safe FFprobe subprocesses that guarantee stability across Windows and Linux hosting environments.
import os
import shutil
import subprocess
from pathlib import Path
from typing import Dict, Any, Optional
def get_ffprobe_binary() -> Optional[str]:
"""Resolve ffprobe executable from system PATH or virtualenv binaries."""
env_bin = os.getenv("FFPROBE_BINARY")
if env_bin and (shutil.which(env_bin) or Path(env_bin).is_file()):
return env_bin
system_bin = shutil.which("ffprobe")
if system_bin:
return system_bin
return None
def probe_media_file(file_path: str) -> Dict[str, Any]:
"""Inspect duration, streams, and codecs using UTF-8 error-tolerant decoding."""
ffprobe = get_ffprobe_binary()
if not ffprobe or not Path(file_path).is_file():
return {"duration": 0.0, "has_audio": True, "has_video": False}
cmd = [
ffprobe, "-v", "error",
"-show_entries", "format=duration:stream=codec_type,width,height",
"-of", "default=noprint_wrappers=1:nokey=1",
str(file_path)
]
# Enforce UTF-8 with replace mode to eliminate ascii decode errors on remote servers
res = subprocess.run(
cmd,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
encoding="utf-8",
errors="replace"
)
lines = [ln.strip() for ln in res.stdout.splitlines() if ln.strip()]
duration = 0.0
for line in lines:
try:
val = float(line)
if val > duration:
duration = val
except ValueError:
pass
return {
"duration": duration,
"raw_probe": lines
}
4. AI Word-Level Alignment & Transcription Pipeline
Speech-to-text recognition requests OpenAI Whisper with word timestamp granularities enabled. The resulting word tokens are processed through an interpolation pass and split into musical lines using pause detection (Δt > 0.45s) and punctuation rules.
import re
from typing import List, Dict, Any
PAUSE_SPLIT_THRESHOLD_SECONDS = 0.45
PUNCTUATION_BREAK_PATTERN = re.compile(r'[\.\?\!\;\:\,\—\-]')
def segment_transcript_words_into_lines(words: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Segment a continuous list of word timestamps into structured musical lyric lines.
Splits are triggered upon:
1. Vocal pauses where gap between word_n.end and word_{n+1}.start exceeds 0.45s
2. Natural phrase punctuation (. ! ? ; , -)
3. Maximum word constraints per visual line (typically 8-10 words)
"""
if not words:
return []
lines = []
current_words = []
for idx, w in enumerate(words):
current_words.append(w)
is_last_word = (idx == len(words) - 1)
# Check pause gap to next word
has_pause = False
if not is_last_word:
next_w = words[idx + 1]
gap = float(next_w.get("start", 0)) - float(w.get("end", 0))
if gap >= PAUSE_SPLIT_THRESHOLD_SECONDS:
has_pause = True
# Check terminal punctuation in word text
has_punct = bool(PUNCTUATION_BREAK_PATTERN.search(w.get("word", "")))
is_lengthy = len(current_words) >= 9
if is_last_word or has_pause or has_punct or is_lengthy:
line_text = " ".join([item.get("word", "").strip() for item in current_words]).strip()
line_start = float(current_words[0].get("start", 0))
line_end = float(current_words[-1].get("end", line_start + 1.0))
lines.append({
"line_id": f"line_{len(lines) + 1}",
"start": round(line_start, 3),
"end": round(line_end, 3),
"text": line_text,
"words": current_words
})
current_words = []
return lines
import os
from openai import OpenAI
from config import Config
class OpenAITranscriptionService:
def __init__(self, api_key: str = None, model: str = None):
# API key is securely retrieved from environment configuration
self.api_key = api_key or Config.OPENAI_API_KEY or os.getenv("OPENAI_API_KEY", "")
self.model = model or Config.OPENAI_TRANSCRIPTION_MODEL or "whisper-1"
self.client = OpenAI(api_key=self.api_key) if self.api_key else None
def transcribe_audio_words(self, audio_file_path: str) -> dict:
"""Call Whisper API with timestamp_granularities=['word'] enabled."""
if not self.client:
raise ValueError("OpenAI client not initialized. Set OPENAI_API_KEY in environment.")
with open(audio_file_path, "rb") as audio_file:
transcript = self.client.audio.transcriptions.create(
file=audio_file,
model=self.model,
response_format="verbose_json",
timestamp_granularities=["word"]
)
return transcript.to_dict()
5. Advanced SubStation Alpha (ASS) Vector Subtitles
Rather than burning rasterized text images, LyricSync compiles lyric timings into native Advanced SubStation Alpha scripts (.ass). Timing tags (\k<duration>) enable sub-frame syllable highlighting with crisp typography.
def format_ass_time(seconds: float) -> str:
"""Format floating point seconds into ASS timestamp standard (H:MM:SS.cs)."""
hrs = int(seconds // 3600)
mins = int((seconds % 3600) // 60)
secs = int(seconds % 60)
centis = int(round((seconds - int(seconds)) * 100))
if centis >= 100:
centis = 99
return f"{hrs}:{mins:02d}:{secs:02d}.{centis:02d}"
def generate_ass_script(lines: list, style: dict, width: int = 1920, height: int = 1080) -> str:
"""Construct complete ASS subtitle specification with karaoke timing wipes."""
font_name = style.get("font", "Outfit")
font_size = style.get("font_size", 42)
primary_color = "&H00FFFFFF" # White in ASS &HAABBGGRR format
karaoke_color = "&H00DF77B7" # Mint/Burgundy accent
ass_content = [
"[Script Info]",
"ScriptType: v4.00+",
f"PlayResX: {width}",
f"PlayResY: {height}",
"ScaledBorderAndShadow: yes",
"",
"[V4+ Styles]",
"Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding",
f"Style: Default,{font_name},{font_size},{primary_color},{karaoke_color},&H00000000,&H80000000,-1,0,0,0,100,100,0,0,1,2,3,2,60,60,80,1",
"",
"[Events]",
"Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text"
]
for line in lines:
start_ts = format_ass_time(line.get("start", 0))
end_ts = format_ass_time(line.get("end", 0))
words = line.get("words", [])
# Build \k tags for karaoke syllable transitions
dialogue_text = ""
for w in words:
duration_cs = max(1, int(round((w.get("end", 0) - w.get("start", 0)) * 100)))
dialogue_text += f"{{\\k{duration_cs}}}{w.get('word', '')} "
ass_content.append(f"Dialogue: 0,{start_ts},{end_ts},Default,,0,0,0,,{dialogue_text.strip()}")
return "\n".join(ass_content)
6. FFmpeg Dual-Tier Video Rendering Engine
The rendering pipeline synthesizes the visual background and audio track while rasterizing ASS subtitles in a single execution pass. To minimize latency on mobile or constrained bandwidth, a 720p Fast Draft profile renders in seconds.
import subprocess
from app.services.media_probe import get_ffmpeg_binary
def execute_render_job(video_in: str, audio_in: str, ass_path: str, output_mp4: str, resolution: str = "720"):
"""Compile final synchronized lyric MP4.
Resolution '720' uses CRF 23 with veryfast preset for instant preview.
Resolution '1080' uses CRF 20 with medium preset for broadcast masters.
"""
ffmpeg_bin = get_ffmpeg_binary()
is_draft = (resolution == "720")
scale_filter = "scale=-2:720" if is_draft else "scale=-2:1080"
preset = "veryfast" if is_draft else "medium"
crf = "23" if is_draft else "20"
audio_bitrate = "128k" if is_draft else "192k"
# Escape path separators for FFmpeg libass filter syntax
escaped_ass = str(ass_path).replace("\\", "/").replace(":", "\\:")
vf_chain = f"{scale_filter},ass='{escaped_ass}'"
cmd = [
ffmpeg_bin, "-y",
"-i", str(video_in),
"-i", str(audio_in),
"-vf", vf_chain,
"-c:v", "libx264",
"-preset", preset,
"-crf", crf,
"-pix_fmt", "yuv420p",
"-c:a", "aac",
"-b:a", audio_bitrate,
"-shortest",
"-movflags", "+faststart", # Optimizes MP4 header for instant web streaming
str(output_mp4)
]
process = subprocess.run(
cmd,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
encoding="utf-8",
errors="replace"
)
if process.returncode != 0:
raise RuntimeError(f"FFmpeg render failure: {process.stderr[-500:]}")
return output_mp4
7. Database Entities & ORM Architecture
Built with SQLAlchemy 2.0 with strict typed mappings. Cascade deletions ensure uploaded media, transcriptions, and cached sync revisions are cleaned atomically when a project is deleted.
from datetime import datetime, timezone
from sqlalchemy import String, Float, Integer, Text, ForeignKey, DateTime
from sqlalchemy.orm import Mapped, mapped_column, relationship
from app.extensions import db
from app.utils.ids import generate_project_id
class Project(db.Model):
__tablename__ = "projects"
id: Mapped[str] = mapped_column(String(64), primary_key=True, default=generate_project_id)
user_id: Mapped[str] = mapped_column(String(64), ForeignKey("users.id"), nullable=True)
name: Mapped[str] = mapped_column(String(255), nullable=False)
status: Mapped[str] = mapped_column(String(32), default="created")
audio_duration: Mapped[float] = mapped_column(Float, default=0.0)
width: Mapped[int] = mapped_column(Integer, default=1920)
height: Mapped[int] = mapped_column(Integer, default=1080)
current_revision: Mapped[int] = mapped_column(Integer, default=1)
canonical_json: Mapped[str] = mapped_column(Text, nullable=True)
created_at: Mapped[datetime] = mapped_column(DateTime, default=lambda: datetime.now(timezone.utc))
user = relationship("User", back_populates="projects")
media_assets = relationship("MediaAsset", back_populates="project", cascade="all, delete-orphan")
lyric_lines = relationship("LyricLine", back_populates="project", cascade="all, delete-orphan")
render_jobs = relationship("RenderJob", back_populates="project", cascade="all, delete-orphan")
8. Role-Based Access Control (RBAC) & Diagnostics Guard
System technicalities (such as AI model parameters, FFmpeg binary locations, SQLite/MySQL table schemas, and disk diagnostics) are strictly gated so only the authorized administrator can inspect them.
from flask import Blueprint, jsonify, current_app
from flask_login import current_user
ADMIN_EMAIL = "silatanuikipngetich@gmail.com"
@views_bp.route("/health")
def health_check():
"""Restricted health diagnostics.
Only silatanuikipngetich@gmail.com can inspect technical internals.
Normal users receive a clean operational status.
"""
is_admin = bool(
current_user.is_authenticated and
current_user.email and
current_user.email.strip().lower() == ADMIN_EMAIL
)
if not is_admin:
# Hides all server paths, table schemas, and binary paths
return {
"status": "healthy",
"database": "connected",
"service": "LyricSync Studio",
"message": "LyricSync Studio is operational."
}, 200
# Full technical diagnostics reserved for administrator
return {
"status": "healthy",
"database": "connected",
"database_tables": inspector.get_table_names(),
"storage_path": str(current_app.config.get("MEDIA_ROOT")),
"ffmpeg_binary": get_ffmpeg_binary(),
"ai_model": current_app.config.get("OPENAI_TRANSCRIPTION_MODEL", "whisper-1"),
"rendering_engine": "FFmpeg (libass + h264 + aac)",
"admin_access": True
}, 200
9. Studio Frontend Engine & State Synchronization
The frontend editor features bidirectional scrubbing: scrubbing the audio player immediately repositions the timeline chips and video preview, while clicking any lyric chip jumps audio playback to the exact timestamp.
// Sub-second playback synchronization loop
function startSyncLoop() {
function loop() {
if (!audioEl.paused) {
const currentTime = audioEl.currentTime;
scrubber.value = currentTime;
currentTimeDisplay.textContent = formatTime(currentTime);
// Synchronize active lyric chips
const activeLine = timeline.findLineAt(currentTime);
if (activeLine) {
timeline.highlightLine(activeLine.line_id);
overlayEl.textContent = activeLine.text;
}
// Align background video
if (videoEl && Math.abs(videoEl.currentTime - currentTime) > 0.3) {
videoEl.currentTime = currentTime;
}
}
requestAnimationFrame(loop);
}
requestAnimationFrame(loop);
}
10. Complete REST API Specifications
All API endpoints return JSON conforming to the standardized response envelope: {"success": true, ...} on 200 OK or {"success": false, "error": {"code": "...", "message": "..."}} on errors.
| Method | Endpoint | Role | Description |
|---|---|---|---|
| GET | /api/projects |
Public / User | List projects created by current user or public shelf |
| POST | /api/projects |
Creator | Multipart upload for audio track + background selection |
| PUT | /api/projects/{id}/lyrics |
Creator | Save revised lyric timings and word timestamps |
| POST | /api/projects/{id}/transcribe |
Creator | Trigger background Whisper word transcription task |
| POST | /api/projects/{id}/render |
Creator | Queue FFmpeg MP4 render job (resolution: 720 / 1080) |
| GET | /health |
Admin / Public | Diagnostics (Deep stats for silatanuikipngetich@gmail.com) |