This is the Part 2 specification for VoiceBridge, the real-time voice translation Chrome Extension. Part 1 (voice-translate-chrome-extension) defined all core modules and UI surfaces. Part 2 focuses on hardening the pipeline into a production-quality system: wiring all modules together with strict utterance ordering, eliminating double-voice through reliable track replacement, ensuring cross-platform meeting compatibility, adding network resilience with graceful degradation, streamlining the demo experience with embedded keys, and guaranteeing deterministic resource cleanup.
The existing codebase has all individual modules built (AudioCaptureModule, STTClient, TranslationEngine, TTSClient, AudioOutputModule, EchoCancellationModule, LatencyMonitor, MeetingDetector, MessageBus). The offscreen document (offscreen.ts) has a basic wiring that connects these modules but lacks: sequence-tracked utterance lifecycle, strict playback ordering, backpressure management, failure isolation, proper track replacement coordination with the content script, platform-specific injection execution, WebSocket reconnection buffering during pipeline operation, degradation cascading, and deterministic cleanup.
- Pipeline_Orchestrator: The central coordination logic in the offscreen document that tracks each utterance through stages (CAPTURED → TRANSCRIBED → TRANSLATED → SYNTHESIZED → PLAYED / DROPPED) using monotonic sequence IDs
- Utterance: A single speech segment identified by a unique sequence ID, tracked from audio capture through TTS playback
- Backpressure: The condition where unprocessed utterances accumulate faster than the pipeline can process them, requiring the oldest unprocessed utterances to be dropped
- Track_Replacement: The WebRTC mechanism (
RTCRtpSender.replaceTrack()) used to swap the user's original microphone track with the virtual TTS audio track in the meeting's peer connection - Virtual_Track: A
MediaStreamTrackgenerated by theAudioOutputModuleviaMediaStreamAudioDestinationNode, carrying either TTS audio or silence - Degradation_Cascade: The ordered fallback sequence when services fail: Full Pipeline → Text-Only → Transcription-Only → Passthrough
- Degradation_Level: A discriminated union representing the current pipeline capability:
full,text-only,transcription-only, orpassthrough - Platform_Adapter: A platform-specific implementation that handles audio injection for a particular meeting platform (Google Meet, Teams, Discord, Zoom, generic WebRTC)
- Mic_Passthrough: A mode where the virtual track carries the user's original microphone audio unmodified, used during silence and when translation is paused
- Session_Cleanup: The deterministic process of releasing all audio resources, closing WebSocket connections, restoring original tracks, and clearing buffers when a session ends
- Demo_Bootstrap: The process of auto-populating API keys from build-time environment variables into
chrome.storageon first run, enabling zero-configuration demo usage
User Story: As a user, I want every spoken utterance to be tracked through the full translation pipeline with strict ordering, so that translated speech plays back in the correct sequence without gaps, duplicates, or out-of-order playback.
- WHEN the Audio_Capture_Module emits an audio chunk with a new sequence ID, THE Pipeline_Orchestrator SHALL create a PipelineUtterance record with state CAPTURED and the current timestamp
- THE Pipeline_Orchestrator SHALL assign monotonically increasing sequence IDs starting from 1 per session, incrementing by 1 for each new utterance detected by the VAD speech-end event
- WHEN the STT_Client emits a final transcript for a sequence ID, THE Pipeline_Orchestrator SHALL transition that utterance to TRANSCRIBED and immediately forward the transcript to the Translation_Engine
- WHEN the Translation_Engine completes translation for a sequence ID, THE Pipeline_Orchestrator SHALL transition that utterance to TRANSLATED
- WHEN the TTS_Client emits audio chunks for a sequence ID, THE Pipeline_Orchestrator SHALL transition that utterance to SYNTHESIZED
- WHEN the Audio_Output_Module finishes playing audio for a sequence ID, THE Pipeline_Orchestrator SHALL transition that utterance to PLAYED
- THE Pipeline_Orchestrator SHALL enforce strict playback ordering: utterance N+1 SHALL NOT begin playback until utterance N has reached PLAYED or DROPPED state
- WHEN the pipeline queue contains more than 3 unprocessed utterances (state is CAPTURED or TRANSCRIBED), THE Pipeline_Orchestrator SHALL drop the oldest unprocessed utterances until the queue size is 3 or fewer, transitioning each dropped utterance to DROPPED with reason "backpressure"
- IF an utterance fails at any stage (STT timeout after 5 seconds, translation error, TTS failure), THEN THE Pipeline_Orchestrator SHALL transition that utterance to DROPPED with the failure reason, log the failure, and continue processing subsequent utterances without blocking
- THE Pipeline_Orchestrator SHALL emit SESSION_STATE_CHANGED messages with updated totalUtterances, droppedUtterances, and currentSequenceId counts after each state transition
- THE Pipeline_Orchestrator SHALL maintain a bounded map of active utterances (maximum 10 entries), evicting completed (PLAYED or DROPPED) utterances older than 30 seconds
- WHEN a new session starts, THE Pipeline_Orchestrator SHALL reset the sequence counter to 0 and clear all utterance tracking state
User Story: As a user in a meeting, I want other participants to hear ONLY my translated TTS voice when translation is active, never my original language voice, so that the experience is seamless and professional.
- WHEN a translation session starts, THE Content_Script SHALL locate the active RTCPeerConnection audio sender and store a reference to the original microphone MediaStreamTrack
- WHEN a translation session starts, THE Content_Script SHALL replace the original microphone track on the RTCPeerConnection audio sender with the Virtual_Track from the Audio_Output_Module using
RTCRtpSender.replaceTrack() - WHILE translation is active and the Echo_Cancellation_Module is in LISTENING state (no TTS playing), THE Virtual_Track SHALL carry silence (zero-filled audio frames) so that the user's original voice is not transmitted to the meeting
- WHILE translation is active and the Echo_Cancellation_Module is in SPEAKING state (TTS playing), THE Virtual_Track SHALL carry the TTS audio being played by the Audio_Output_Module
- WHEN a translation session ends, THE Content_Script SHALL restore the original microphone track on the RTCPeerConnection audio sender within 200 milliseconds
- IF the RTCPeerConnection is renegotiated or a new peer connection is created during an active session, THEN THE Content_Script SHALL detect the new connection and re-apply the Virtual_Track replacement within 500 milliseconds
- THE Content_Script SHALL monitor the RTCPeerConnection for
trackevents andconnectionstatechangeevents to detect when track replacement needs to be re-applied - IF
replaceTrack()fails (throws an error or the sender is null), THEN THE Content_Script SHALL log the failure, send a CONNECTION_STATE_CHANGED message with domain "meeting" and code "track-replace-failed", and fall back to the platform-specific alternative strategy - THE Audio_Output_Module SHALL provide a method
getMixedTrack()that returns a MediaStreamTrack which dynamically switches between silence and TTS audio based on the Echo_Cancellation_Module state, ensuring zero gaps in the audio stream sent to the meeting - WHEN barge-in is detected, THE Pipeline_Orchestrator SHALL transition the Virtual_Track from TTS audio to Mic_Passthrough within 100 milliseconds, allowing the user's original voice to be heard while they interrupt
User Story: As a user, I want VoiceBridge to work reliably on Google Meet, Microsoft Teams, Discord, Zoom Web, and generic WebRTC applications, so that I can use it regardless of which meeting platform my team uses.
- THE Content_Script SHALL implement a Platform_Adapter interface with methods:
initialize(),injectVirtualTrack(track: MediaStreamTrack),restoreOriginalTrack(),isInjected(), anddestroy() - FOR Google Meet, THE Content_Script SHALL inject a main-world script at
document_startthat interceptsnavigator.mediaDevices.getUserMediaand stores a reference to the original audio track, then replaces the track on the RTCPeerConnection audio sender when the session starts - FOR Microsoft Teams, THE Content_Script SHALL inject a main-world script that monitors the
RTCPeerConnectionconstructor, captures created peer connections, and usesRTCRtpSender.replaceTrack()on the audio sender to inject the Virtual_Track - FOR Discord, THE Content_Script SHALL inject a main-world script that monitors the
RTCPeerConnectionconstructor, captures created peer connections, and usesRTCRtpSender.replaceTrack()on the audio sender to inject the Virtual_Track - FOR Zoom Web Client, THE Content_Script SHALL use
chrome.tabCapture.capture()to obtain the tab audio stream, create an AudioContext mixing node that combines the TTS audio with the tab output, and route the mixed audio back to the tab - FOR generic WebRTC applications (Force Enable mode), THE Content_Script SHALL inject a main-world script that monitors the
RTCPeerConnectionconstructor and attemptsreplaceTrack()on any detected audio sender - WHEN the Meeting_Detector identifies the platform, THE Content_Script SHALL instantiate the corresponding Platform_Adapter and call
initialize()before any session can start - IF the Platform_Adapter
initialize()fails, THEN THE Content_Script SHALL send an ERROR message with domain "meeting" and code "injection-failed", and display[INJECTION FAILED]in the floating widget - WHEN a session ends, THE Content_Script SHALL call
restoreOriginalTrack()on the active Platform_Adapter and thendestroy()to clean up all injected scripts and event listeners - THE Content_Script SHALL communicate with the main-world injected script exclusively via
window.postMessagewith source markervoicebridgeand strict origin checking againstwindow.location.origin
User Story: As a user, I want the extension to handle brief network interruptions transparently, so that a momentary WiFi hiccup does not disrupt my meeting translation.
- WHEN the STT WebSocket connection drops during an active session, THE STT_Client SHALL buffer incoming audio chunks in a queue holding up to 10 seconds of audio (160,000 PCM Int16 samples at 16kHz) and replay them upon successful reconnection
- WHEN the TTS WebSocket connection drops during an active session, THE TTS_Client SHALL buffer pending text tokens and replay them upon successful reconnection
- WHEN any WebSocket connection drops, THE Pipeline_Orchestrator SHALL detect the disconnection within 2 seconds via heartbeat timeout and send a CONNECTION_STATE_CHANGED message with status "connecting"
- WHEN network connectivity is restored, THE STT_Client and TTS_Client SHALL each reconnect using exponential backoff (500ms base, 2x multiplier, 10-second maximum, 5 attempts maximum) and resume pipeline operation within 3 seconds of network restoration
- WHILE the STT WebSocket is disconnected, THE Audio_Capture_Module SHALL continue capturing audio and buffering chunks so that no speech is lost during brief outages
- WHILE the TTS WebSocket is disconnected, THE Translation_Engine SHALL continue processing transcripts and buffering translated text for synthesis upon reconnection
- IF reconnection fails after 5 attempts on any WebSocket, THEN THE Pipeline_Orchestrator SHALL transition to the next Degradation_Level and send a CONNECTION_STATE_CHANGED message with status "error" and retryable false
- THE Pipeline_Orchestrator SHALL attempt reconnection again after 30 seconds if the initial 5 attempts fail, providing a second chance before requiring user intervention
- IF reconnection fails after the second attempt cycle (30 seconds + 5 more attempts), THEN THE Pipeline_Orchestrator SHALL pause the session and send an ERROR message prompting the user to manually resume
- WHEN a WebSocket reconnects successfully, THE Pipeline_Orchestrator SHALL send a CONNECTION_STATE_CHANGED message with status "connected" and resume normal operation, flushing any buffered audio or text
User Story: As a user, I want the extension to continue providing value even when some services fail, degrading gracefully from full voice translation to text-only to transcription-only to passthrough, so that I am never blocked from speaking in my meeting.
- THE Pipeline_Orchestrator SHALL maintain a current Degradation_Level with four possible values:
full(STT + Translation + TTS + Audio Output),text-only(STT + Translation, display in Side_Panel only),transcription-only(STT only, display original transcript in Side_Panel),passthrough(no processing, original mic flows to meeting) - WHEN the TTS service becomes unavailable (connection failed, quota exhausted via HTTP 402, or voice not found), THE Pipeline_Orchestrator SHALL transition from
fulltotext-onlyand send a CONNECTION_STATE_CHANGED message indicating TTS is offline - WHEN the Translation service becomes unavailable (LLM connection failed, rate limited with no recovery, or timeout on 3 consecutive utterances), THE Pipeline_Orchestrator SHALL transition from
text-onlytotranscription-only - WHEN the STT service becomes unavailable (connection failed after all retry attempts), THE Pipeline_Orchestrator SHALL transition from
transcription-onlytopassthrough - WHILE in any degradation level, THE original microphone audio SHALL remain available to the meeting platform — the Pipeline_Orchestrator SHALL restore the original mic track when degrading below
full - WHEN a previously unavailable service recovers (WebSocket reconnects, rate limit expires), THE Pipeline_Orchestrator SHALL automatically upgrade to the highest available Degradation_Level within 5 seconds
- THE floating widget SHALL display the current Degradation_Level using inline status text:
[FULL](no indicator needed),[TEXT ONLY]in--warningcolor,[TRANSCRIPT ONLY]in--warningcolor,[PASSTHROUGH]in--accentcolor - THE Pipeline_Orchestrator SHALL log every degradation transition with the trigger reason, previous level, and new level to the debug log
- WHEN transitioning between degradation levels, THE Pipeline_Orchestrator SHALL complete the transition within 500 milliseconds to avoid perceptible gaps in service
- THE Pipeline_Orchestrator SHALL never transition directly from
fulltopassthrough— degradation SHALL always follow the cascade order: full → text-only → transcription-only → passthrough
User Story: As a first-time user or hackathon judge, I want the extension to work immediately after installation without entering any API keys, so that I can experience the full translation pipeline with zero setup friction.
- WHEN the extension is installed and opened for the first time, THE Service_Worker SHALL check for embedded demo API keys from build-time environment variables (
VITE_DEMO_ELEVENLABS_KEY,VITE_DEMO_LLM_KEY,VITE_DEMO_LLM_PROVIDER,VITE_DEMO_OPENROUTER_MODEL) - IF embedded demo keys are present and no user-provided keys exist in
chrome.storage, THEN THE Service_Worker SHALL write the embedded keys tochrome.storage.localusing the Settings_Store encryption for API keys - WHEN demo keys are populated, THE Onboarding_Wizard SHALL skip the API Keys step (Step 2) and proceed directly from Welcome (Step 1) to Voice Recording (Step 3)
- THE Onboarding_Wizard SHALL display a notice during the Welcome step when demo keys are active: "Demo mode active — 5 minutes of voice translation included. Enter your own API key in Settings for unlimited usage."
- WHEN the user enters their own API key in the Settings page, THE Settings_Store SHALL overwrite the demo key and THE Extension SHALL immediately switch to unlimited mode, removing all voice-time restrictions
- IF the embedded demo key is exhausted (ElevenLabs API returns HTTP 402), THEN THE Extension SHALL set
embeddedKeyExhausted: trueinchrome.storage.local, disable demo mode, and display a prompt to enter a personal API key - THE Extension SHALL cache the embedded key exhaustion state and recheck only once every 6 hours to avoid repeated failed API calls
- THE embedded demo keys SHALL be injected at build time via Vite environment variables and SHALL NOT be stored as plaintext string literals in the source code
User Story: As a user, I want all audio resources to be properly released when I stop translation, close the meeting tab, or encounter an error, so that my browser does not accumulate leaked resources or stale connections.
- WHEN a session ends (user toggle, meeting tab closed, error, or demo limit reached), THE Pipeline_Orchestrator SHALL execute cleanup in the following deterministic order: (a) stop Audio_Capture_Module, (b) disconnect STT_Client, (c) destroy Translation_Engine, (d) disconnect TTS_Client, (e) destroy Audio_Output_Module, (f) destroy Echo_Cancellation_Module, (g) clear Latency_Monitor
- THE Pipeline_Orchestrator SHALL complete all cleanup operations within 1 second of the session end trigger
- WHEN the Audio_Capture_Module is stopped, THE Audio_Capture_Module SHALL stop all MediaStream tracks, disconnect the AudioWorkletNode, disconnect the GainNode, and close the AudioContext
- WHEN the Audio_Output_Module is destroyed, THE Audio_Output_Module SHALL stop all MediaStreamDestination tracks, disconnect the GainNode, stop any playing AudioBufferSourceNode, and close the AudioContext
- WHEN the STT_Client is disconnected, THE STT_Client SHALL send an
end_of_streammessage if the WebSocket is open, close the WebSocket, clear the heartbeat timer, clear the reconnection timer, and clear the audio buffer queue - WHEN the TTS_Client is disconnected, THE TTS_Client SHALL close the WebSocket, clear the heartbeat timer, clear the reconnection timer, and clear the pending text buffer
- WHEN the meeting tab is closed during an active session, THE Content_Script SHALL detect the
beforeunloadevent and send a SESSION_STOP message with reason "tab-closed" to trigger cleanup in the offscreen document - IF any individual cleanup step fails (throws an error), THEN THE Pipeline_Orchestrator SHALL log the error, continue with the remaining cleanup steps, and report the partial cleanup failure in the debug log
- WHEN the Content_Script detects session end, THE active Platform_Adapter SHALL call
restoreOriginalTrack()to return the original microphone track to the RTCPeerConnection before the tab unloads - THE Pipeline_Orchestrator SHALL set all module references to null after cleanup to prevent use-after-destroy errors and enable garbage collection
- WHEN the offscreen document is destroyed (extension disabled or updated), THE Pipeline_Orchestrator SHALL execute the full cleanup sequence before the document unloads
User Story: As a user, I want the audio routing between my microphone and the meeting to be managed by a clear state machine, so that there is never a moment where both my original voice and the TTS voice are heard simultaneously, and there is never a moment of dead silence when I should be heard.
- THE Pipeline_Orchestrator SHALL implement an audio routing state machine with four states: PASSTHROUGH (original mic → meeting), MUTED (silence → meeting, mic captured for STT), TTS_PLAYING (TTS audio → meeting), and BARGE_IN (original mic → meeting, TTS fading out)
- WHEN translation is active and the user is not speaking (VAD state is silence), THE routing state SHALL be MUTED, sending silence to the meeting while the mic remains captured for STT processing
- WHEN translation is active and TTS audio is playing, THE routing state SHALL be TTS_PLAYING, sending TTS audio to the meeting
- WHEN the user begins speaking while TTS is playing (barge-in detected), THE routing state SHALL transition to BARGE_IN, fading out TTS over 50 milliseconds and switching to the user's original mic audio within 100 milliseconds total
- WHEN translation is inactive or the pipeline is in PASSTHROUGH degradation level, THE routing state SHALL be PASSTHROUGH, sending the user's original mic audio directly to the meeting
- THE routing state machine SHALL transition between states within 50 milliseconds to prevent audible gaps or overlaps
- THE routing state machine SHALL emit state change events that the floating widget can use to update the status icon (microphone for PASSTHROUGH, muted icon for MUTED, speaker for TTS_PLAYING, microphone with pulse for BARGE_IN)
- THE routing state machine SHALL coordinate with the Echo_Cancellation_Module: MUTED and TTS_PLAYING states correspond to echo cancellation SPEAKING state, PASSTHROUGH and BARGE_IN correspond to LISTENING state
User Story: As a user, I want the end-to-end translation latency to stay under 2 seconds consistently, so that the conversation feels natural and responsive.
- THE Pipeline_Orchestrator SHALL measure end-to-end latency for each utterance from audio capture start to TTS playback start using the Latency_Monitor
- THE Pipeline_Orchestrator SHALL enforce a per-stage timeout: STT must produce a final transcript within 5 seconds of speech-end, Translation must produce the first token within 3 seconds of transcript receipt, TTS must produce the first audio byte within 3 seconds of first text token
- IF any stage exceeds its timeout, THEN THE Pipeline_Orchestrator SHALL drop that utterance, transition it to DROPPED with reason "timeout-{stage}", and continue with the next utterance
- IF end-to-end latency exceeds 3000 milliseconds for 5 consecutive utterances, THEN THE Pipeline_Orchestrator SHALL send an ERROR message with a user-facing suggestion to check network connection or reduce quality settings
- THE Latency_Monitor SHALL report per-utterance latency breakdown (capture, STT, translation, TTS, routing) via LATENCY_UPDATE messages for display in the floating widget and Side_Panel
- THE Pipeline_Orchestrator SHALL prioritize latency over completeness: if translation is still streaming when the next utterance arrives and the queue has 2 or more pending utterances, THE Pipeline_Orchestrator SHALL flush the current translation and proceed
User Story: As a developer, I want a reliable, low-latency communication channel between the offscreen document (where TTS audio is generated) and the content script (where WebRTC injection happens), so that audio reaches the meeting without additional delay or data loss.
- WHEN a session starts, THE Service_Worker SHALL establish a MessageChannel port pair between the offscreen document and the content script, bypassing the service worker for audio data transfer
- THE offscreen document SHALL send TTS audio chunks as
TransferableArrayBuffer objects through the MessageChannel port to avoid memory copying - THE Content_Script SHALL receive audio chunks from the MessageChannel port and route them to the Platform_Adapter for injection into the meeting's RTCPeerConnection
- IF the MessageChannel port is disconnected (tab navigation, content script reload), THEN THE Service_Worker SHALL detect the disconnection and re-establish the port pair within 1 second
- THE MessageChannel SHALL carry typed messages with a discriminated
typefield:audio-chunk(PCM data + sequence ID),track-command(inject/restore/status), andstate-sync(echo state, routing state) - THE Content_Script SHALL acknowledge receipt of
track-commandmessages by sending a response through the MessageChannel port, allowing the offscreen document to confirm track replacement succeeded - WHEN the session ends, THE Service_Worker SHALL close both MessageChannel ports to release resources