- Initialize npm project with
package.json(name: voicebridge, type: module) - Create
tsconfig.jsonwith strict mode, ES2022 target, bundler module resolution - Create
vite.config.tswith multi-entry build for all extension contexts (service-worker, offscreen, content-script, widget, popup, sidepanel, options, onboarding, audio-processor.worklet) - Create directory structure:
src/background/,src/offscreen/,src/content/,src/popup/,src/sidepanel/,src/options/,src/onboarding/,src/lib/,src/worklets/,src/styles/ - Install dependencies:
@elevenlabs/elevenlabs-js,lucide-static,vite,vitest,fast-check,typescript - Create
src/manifest.jsonwith Manifest V3 config, permissions, content scripts, commands, CSP
- Create
src/styles/tokens.csswith all CSS custom properties: color system (dark + light mode), typography scale, spacing scale, motion tokens - Create
src/styles/shared.csswith base component styles: buttons (primary/secondary/ghost pill), toggles, inputs (underline), cards, segmented progress bars, dot-grid motif - Create
src/styles/widget.csswith floating widget styles: collapsed circle, expanded card, ghost mode, roulette mode, opacity fade, draggable positioning - Load Google Fonts: Doto (variable), Space Grotesk (300-700), Space Mono (400, 700)
- Create
src/lib/message-bus.tswith typedExtensionMessageinterface,MessageTypeunion,MessagePayloadMap, andExtensionContexttype - Implement
sendMessage(type, payload)with automatic timestamp and source context detection - Implement
onMessage(type, handler)with sender validation (sender.id === chrome.runtime.id) - Implement content script ↔ page bridge using
window.postMessagewithsource: 'voicebridge'and origin checking - Implement
MessageChannelport pair setup for high-frequency audio data transfer between offscreen and content script
- Create
src/lib/settings-store.tswith fullSettingsSchemainterface - Implement AES-GCM-256 encryption/decryption for API keys using Web Crypto API with PBKDF2 key derivation from
chrome.runtime.id+ per-install salt - Implement
get<K>(key)andset<K>(key, value)with type-safe access tochrome.storage.localandchrome.storage.sync - Implement settings migration for version upgrades
- Implement
exportSettings()andimportSettings()(excluding API keys) - Store install ID via
crypto.randomUUID()on first run
- Create
src/worklets/audio-processor.worklet.tswithAudioWorkletProcessorthat converts Float32 → Int16 PCM and posts viaMessagePortwithTransferable - Create
src/lib/audio-capture.tswithAudioCaptureModuleclass implementing mic access, AudioWorklet setup, ring buffer chunking (250ms / 4000 samples) - Implement energy-based VAD with state machine: SILENCE → SPEECH_PENDING (300ms onset) → SPEECH → SILENCE_PENDING (800ms offset) → SILENCE
- Implement noise gate with configurable threshold (default -40dB, Ghost Mode -55dB)
- Implement Ghost Mode gain boost (+20dB) and high-pass filter (100Hz)
- Implement mute/unmute for echo cancellation integration
- Emit
onAudioChunk,onVADStateChange,onSpeechEndcallbacks
- Create
src/lib/echo-cancellation.tswith discriminated unionEchoStatetype andEchoEventtype - Implement pure
transitionEchoState(current, event)function with exhaustive switch - Implement
EchoCancellationModuleclass that coordinates mic muting (SPEAKING/TRANSITIONING) and TTS stopping (barge-in) - Implement 200ms transition timer after TTS ends before re-enabling mic
- Implement barge-in detection: VAD speech during SPEAKING → fade out TTS (50ms), cancel TTS, switch to LISTENING within 100ms
- Implement Ghost Mode override: disable barge-in detection, full mic mute during SPEAKING
- Create
src/lib/stt-client.tswithSTTClientclass andSTTConnectionStatediscriminated union - Implement WebSocket connection to
wss://api.elevenlabs.io/v1/speech-to-text/streamwith config message on open - Implement single-use token acquisition via REST API (
POST /v1/speech-to-text/stream/token) - Implement binary audio frame sending (raw PCM Int16 ArrayBuffer)
- Implement
commit()method sending{ type: "commit" }on VAD speech-end - Handle
transcript.partialandtranscript.finalmessages with callbacks - Implement exponential backoff reconnection: 500ms base, 2× multiplier, 10s max, 5 attempts
- Implement 15-second heartbeat ping for silent disconnection detection
- Implement 10-second audio buffer queue for brief disconnections
- Create
src/lib/translation-engine.tswithTranslationEngineclass supporting OpenAI and Anthropic providers - Implement streaming translation via
AsyncGenerator<string>yielding tokens as they arrive - Implement system prompt with instructions: translate naturally, preserve tone, handle idioms, output only translated text, preserve proper nouns/technical terms
- Implement sliding context window (last N finalized pairs, configurable 5-20, default 10)
- Implement short utterance buffering: segments < 3 words wait 1.5s for continuation
- Implement preservation markers for URLs, emails, code snippets, numbers
- Implement length guard: flag translations > 3× source length, re-request with "be concise"
- Implement custom glossary injection into system prompt
- Implement formal/informal tone setting
- Handle errors: retry once after 200ms, skip on second failure, queue on 429 with Retry-After
- Create
src/lib/tts-client.tswithTTSClientclass andTTSConnectionStatediscriminated union - Implement WebSocket connection to
wss://api.elevenlabs.io/v1/text-to-speech/{voiceId}/stream-input?model_id=eleven_multilingual_v2 - Send initial config with voice settings, API key, and
output_format: "pcm_24000"on open - Implement token-by-token text streaming with
try_trigger_generation: true - Implement
flush()for end-of-utterance andcancel()for barge-in (flush + discard queue) - Handle binary audio frames (PCM Int16 24kHz) and JSON base64 audio responses
- Implement long sentence splitting at clause boundaries (commas, semicolons, conjunctions) for sentences > 50 words
- Implement exponential backoff reconnection matching STT client strategy
- Implement 15-second heartbeat ping
- Create
src/lib/audio-output.tswithAudioOutputModuleclass - Implement PCM Int16 24kHz → Float32 conversion and resampling to 48kHz via
AudioContext({ sampleRate: 48000 }) - Create
MediaStreamDestinationnode for virtual audio track generation - Implement
GainNodefor volume normalization (match user's average mic level from first 5s) - Implement 100ms playback buffer to prevent audio underruns
- Implement 50ms fade-out for barge-in scenarios
- Implement
getVirtualTrack()returning theMediaStreamTrackfor WebRTC injection - Implement
destroy()to release all audio resources (tracks stopped, context closed)
- Create
src/lib/meeting-detector.tswith URL pattern matching for Google Meet, Zoom, Teams, Discord - Implement
getInjectionStrategy(platform)returning the appropriate strategy type - Create
src/content/content-script.tswith platform detection on load and audio injection coordination - Implement Google Meet strategy: intercept
getUserMediavia main-world script injection (document_start) - Implement Teams/Discord strategy: monitor
RTCPeerConnectionconstructor,replaceTrackon audio sender - Implement Zoom fallback:
chrome.tabCaptureAPI for audio mixing - Implement generic "Force Enable" mode attempting
replaceTrackon any detectedRTCPeerConnection - Implement track lifecycle: store original → replace with virtual → restore on session end
- Create
src/lib/voice-profile.tswithVoiceProfileStatediscriminated union andVoiceProfileclass - Implement voice sample recording with duration tracking (30s min, 2min max)
- Implement sample validation: duration check, RMS noise analysis (> -30dB average)
- Implement upload to ElevenLabs Voice Cloning API (
POST /v1/voices/add) - Implement voice deletion via
DELETE /v1/voices/{voiceId} - Implement voice preview: synthesize test phrase via REST TTS endpoint
- Store voice_id encrypted in
chrome.storage.local
- Create
src/lib/latency-monitor.tswithLatencyMeasurementinterface andLatencyMonitorclass - Implement per-stage timing marks: captureStart, captureEnd, sttEnd, translationStart, translationFirstToken, ttsFirstByte, playbackStart
- Calculate per-utterance total latency and per-stage breakdown
- Implement rolling average calculation over configurable window
- Emit
onLatencyUpdatecallback for UI display - Store measurement history (last 100 measurements)
- Implement monotonically increasing sequence ID assignment in audio capture
- Implement
PipelineUtterancetracking through stages: CAPTURED → TRANSCRIBED → TRANSLATED → SYNTHESIZED → PLAYED / DROPPED - Implement strict ordering: never play utterance N+1 before N
- Implement backpressure: drop oldest unprocessed when queue > 3 utterances
- Implement failure handling: skip failed utterances, log reason, continue pipeline
- Wire all pipeline components together in offscreen document: AudioCapture → STT → Translation → TTS → AudioOutput
- Create
src/background/service-worker.tsas the session lifecycle orchestrator - Implement offscreen document creation/management (
chrome.offscreen.createDocumentwithhasDocumentcheck) - Implement session state persistence to
chrome.storage.session - Implement meeting detection via tab URL monitoring (
chrome.tabs.onUpdated) - Implement keyboard command handling (
chrome.commands.onCommand) for Alt+T, Ctrl+Space, Ctrl+Shift+X, Alt+G, Alt+R - Implement
chrome.runtime.onUpdateAvailableto defer updates during active sessions - Implement alarm-based periodic tasks (quota check, heartbeat monitoring)
- Handle service worker wake-up: re-attach to existing offscreen document
- Create
src/offscreen/offscreen.htmlandsrc/offscreen/offscreen.ts - Initialize all pipeline components on creation: AudioCapture, STTClient, TranslationEngine, TTSClient, AudioOutput, EchoCancellation, LatencyMonitor
- Wire message bus handlers for SESSION_START, SESSION_STOP, LANGUAGE_CHANGED, SETTINGS_UPDATED
- Implement audio data relay to content script via MessageChannel
- Implement quota tracking: characters sent to TTS, audio seconds to STT, LLM tokens
- Implement demo voice-time tracking (VAD-active time only)
- Create
src/content/widget.tswith Shadow DOM isolation for style encapsulation - Implement collapsed state: 48×48 circle with monoline status icon (mic/globe/speaker/pause/ghost)
- Implement expanded state on hover: 240×120 card with latency (Space Mono 36px), language pair, session duration, dot-grid background
- Implement click-to-toggle translation pipeline
- Implement draggable positioning with per-domain persistence via
chrome.storage.local - Implement opacity fade: 30% after 5s inactivity, 100% on hover (300ms/150ms transitions)
- Implement status color coding: green (<1500ms), yellow (1500-2500ms), red (>2500ms)
- Implement accent red dot (6px) for active recording state
- Implement error display: inline
[OFFLINE],[RECONNECTING...],[ERROR]text - Implement Ghost Mode display: ghost icon at 60% opacity with pulse animation
- Implement Language Roulette display: language name in Doto with fade transitions, segmented progress bar
- Create
src/popup/popup.htmlandsrc/popup/popup.ts - Implement main toggle switch (Nothing mechanical toggle style)
- Implement source language selector with auto-detect default, search/filter, recently-used (top 3)
- Implement target language selector excluding source language, same organization
- Implement voice profile status display (Not Set Up / Recording / Processing / Ready / Error)
- Implement real-time latency indicator with color coding (green/yellow/red)
- Implement connection status indicators for STT, TTS, LLM
- Implement session duration and estimated cost display
- Implement demo voice-time remaining indicator (
VOICE: 1:42 LEFT) - Implement demo limit reached card with reset countdown and BYO key option
- Implement Language Roulette button
- Implement Ghost Mode toggle
- Render within 100ms, max 400×500px
- Create
src/sidepanel/sidepanel.htmlandsrc/sidepanel/sidepanel.ts - Implement two-column transcript view: original (left) + translated (right)
- Implement partial transcript display with italic styling and pulsing indicator
- Implement final transcript pairs with timestamps (Space Mono caption)
- Implement auto-scroll to latest unless user has scrolled up
- Implement "Copy All" button (formatted text to clipboard)
- Implement search/filter input highlighting matches in both columns
- Implement "Export Transcript" button (.txt and .srt formats)
- Implement Demo Mode pipeline visualization (flow diagram with per-stage latency)
- Implement Language Roulette stacked list with current-language highlighting
- Create
src/options/options.htmlandsrc/options/options.ts - Implement API key inputs (ElevenLabs, LLM) with validation test requests and success/failure status
- Implement LLM provider selection dropdown (OpenAI / Anthropic)
- Implement voice profile management section: view, record, delete, preview
- Implement voice tuning sliders: stability, similarity boost, style (0.0-1.0)
- Implement audio settings: noise gate threshold, VAD sensitivity, echo cancellation mode
- Implement translation settings: context window size, preserve technical terms toggle, custom glossary editor
- Implement performance settings: latency/quality priority slider, max concurrent requests
- Implement Demo Mode section with limit explanation and "Get Credits" link
- Implement usage statistics: ElevenLabs characters, LLM tokens, daily history chart
- Implement Export/Import settings buttons
- Implement Debug Log section: scrollable, filterable, exportable (JSON)
- Implement Language Roulette sequence customization
- Implement Push-to-Translate hotkey configuration
- Create
src/onboarding/onboarding.htmlandsrc/onboarding/onboarding.ts - Implement Step 1 (Welcome): explain VoiceBridge, demo limitations (2 min / 24h), privacy notice
- Implement Step 2 (API Keys): ElevenLabs key input with validation, LLM key input, "Get free credits" card linking to hackathon offer
- Implement Step 3 (Voice Recording): real-time audio level meter, countdown timer, 3 reading prompts, sample validation
- Implement Step 4 (Language Selection): source + target language pickers
- Implement Step 5 (Test & Confirm): full pipeline test (5s capture → transcribe → translate → synthesize → playback)
- Implement step validation before progression, retry/skip on failure
- Store
onboardingComplete: trueon finish, auto-launch on first install
- Implement voice-time tracking: accumulate only during VAD SPEECH state, not wall-clock
- Implement 2-minute (120s) per-install limit within rolling 24-hour window
- Implement quota state machine: Available → Active → Warning30s → Warning10s → Exhausted → Cooldown → Available
- Implement BYO key detection: validate user key → switch to Unlimited mode, remove restrictions
- Implement embedded demo key assembly (obfuscated split base64 segments)
- Implement embedded key exhaustion detection (HTTP 402) with 6-hour recheck cache
- Implement reset timer: 24h from first voice-time usage in current window
- Wire quota state to UI: popup indicator, widget progress bar, limit-reached card
- Implement Language Roulette activation (button + Alt+R shortcut)
- Implement sentence capture (next spoken sentence or pre-loaded demo sample)
- Implement 10-language synthesis cycle: EN → JA → ES → AR → FR → ZH → DE → KO → PT → HI (customizable)
- Implement sequential playback with 200ms silence gaps between languages
- Implement widget display: language name in Doto with 150ms fade transitions
- Implement side panel display: stacked translations with current-language highlighting
- Implement segmented progress bar (10 segments, one per language)
- Implement completion state:
[COMPLETE], Replay button, Share button (clipboard text) - Implement "Record Roulette" option capturing full output as .webm
- Ensure roulette audio plays locally only (not sent to meeting)
- Implement Ghost Mode toggle (popup + Alt+G shortcut)
- Implement whisper capture: lower VAD threshold to -55dB, +20dB gain, high-pass filter 100Hz
- Implement TTS output at full volume regardless of whisper input
- Adjust voice settings in Ghost Mode: stability 0.7
- Implement aggressive echo cancellation: no barge-in detection, full mic mute during SPEAKING
- Implement sensitivity meter: 5-segment bar showing input level relative to whisper threshold
- Implement "too loud" warning:
[TOO LOUD — WHISPER]when input > -20dB - Implement first-time tooltip explaining Ghost Mode
- Implement widget ghost icon with pulse animation (opacity 40%→60%→40%, 2s cycle)
- Implement ElevenLabs subscription check at session start (
GET /v1/user/subscription) - Implement quota warning levels: 80% (yellow), 95% (urgent red), 100% (stop TTS, text-only mode)
- Implement per-session usage tracking: TTS characters, STT seconds, LLM tokens, estimated cost USD
- Implement daily usage history storage (last 30 days) in
chrome.storage.local - Display quota percentage bar and absolute numbers in popup
- Display estimated session cost in popup footer
- Implement connection loss detection within 2 seconds (heartbeat timeout)
- Implement automatic reconnection for all WebSocket connections within 3 seconds on network restore
- Implement STT audio buffer queue (10 seconds) during brief disconnections
- Implement graceful degradation cascade: Full → Text-Only → Transcription-Only → Passthrough
- Ensure original microphone always remains available regardless of failures
- Implement 30-second reconnection timeout → pause session, prompt user
- Implement panic button (Ctrl+Shift+X): stop all capture, close all connections, mute extension
- Implement language list fetching from ElevenLabs API (models endpoint for TTS, cached 24h)
- Implement language tier classification: 'full' (STT + TTS) vs 'text-only' (STT only)
- Implement BCP 47 language tag usage throughout
- Implement "Text Translation Mode" for languages without TTS support
- Implement language validation before session start (target language TTS-supported check)
- Cache language lists in
chrome.storage.localwith 24-hour TTL
- Register manifest commands: toggle-translation (Alt+T), push-to-translate (Ctrl+Space), panic-stop (Ctrl+Shift+X), toggle-ghost-mode (Alt+G), language-roulette (Alt+R)
- Implement full keyboard navigation in popup (Tab, Enter/Space, Escape)
- Implement keyboard navigation in side panel (arrows, Ctrl+C, Ctrl+F)
- Add ARIA labels to floating widget:
aria-label="VoiceBridge translation status: [status]" - Implement ARIA live regions for status change announcements
- Ensure visible focus indicators on all interactive elements
- Implement Push-to-Translate mode (audio only while Ctrl+Space held)
- Implement API key encryption with AES-GCM-256 + PBKDF2 key derivation (100,000 iterations, SHA-256)
- Ensure API keys never sent to content scripts (offscreen/service-worker only)
- Implement message sender validation on all
chrome.runtime.onMessagehandlers - Implement content script ↔ page message validation (origin check + source marker)
- Implement CSP:
script-src 'self'; object-src 'none' - Ensure no audio data persisted to disk — streaming buffers only
- Implement resource cleanup within 1 second of session end
- Implement privacy notice in onboarding
- Set up vitest configuration with TypeScript support
- Create mock utilities for
chrome.*APIs (storage, runtime, tabs, offscreen) - Create mock WebSocket server for STT/TTS protocol testing
- Implement Property 1: Audio Format Round-Trip (Float32 ↔ Int16 within ±1/32768)
- Implement Property 2: Audio Chunking Completeness (no samples lost/duplicated)
- Implement Property 3: Noise Gate Correctness
- Implement Property 4: VAD State Machine Hysteresis
- Implement Property 5: Echo Cancellation State Machine
- Implement Property 6: Exponential Backoff Bounds
- Implement Property 7: Translation Context Window
- Implement Property 8: Short Utterance Buffering
- Implement Property 9: Voice Sample Validation
- Implement Property 10: Long Sentence Clause Splitting
- Implement Property 11: Sample Rate Conversion Output Length
- Implement Property 12: Volume Normalization
- Implement Property 13: Meeting Platform URL Detection
- Implement Property 14: API Key Encryption Round-Trip
- Implement Property 15: Preservation Marker Detection
- Implement Property 16: Translation Length Ratio Guard
- Implement Property 17: Pipeline Ordering and Failure Resilience
- Implement Property 18: Pipeline Backpressure
- Implement Property 19: Circular Debug Log Buffer
- Implement Property 20: Language Tier Classification
- Implement Property 21: Voice-Time Accumulator and Demo Limit