Skip to content

Commit dbaa57a

Browse files
authored
feat(ai): add transcription across inline, streamed, and queued routes (#50910)
1 parent b6bc557 commit dbaa57a

56 files changed

Lines changed: 2838 additions & 284 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎packages/ai/AGENTS.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -100,7 +100,9 @@ Media does not fit the SSE-frames-to-event-state-machine LLM route. `MediaRoute.
100100

101101
`MediaProtocol.queued` is the submit-then-poll kind every video route uses: `start` (body + decode into `{ token, snapshot }`), `status`, `result`, and optional `cancel`, each addressed by a route-owned `token` whose `Schema.Codec` makes it serializable. `MediaRoute.inline` and `MediaRoute.queued` compose the two kinds with `Endpoint` and `Auth`; the queued route decodes the token once at the boundary (`start` output or `resume` input) and closes over it in a token-free `GenerationRoute` (`status`/`result`/`cancel` are plain Effects), so `Generation` never sees the token's shape and only carries the encoded JSON for persistence. Polls reuse the route's auth and deployment headers plus the request's `http` overlay after `start`, and resolve relative paths against the route base URL (provider-issued absolute URLs such as fal's `status_url` pass through). `result` is always its own GET even when the provider returns output inside the status document, so `Generation.await` behaves the same after `start` and after `resume`. `PollContext.auth` carries only what `Auth` added so protocols can hand download credentials to output assets as transient `Media.Asset.headers` (Veo) — never part of `source` or JSON. Status strings map through a per-protocol `STATUS` table via `MediaProtocol.status`; terminal generations without output fail through `output.ended` / `output.contentPolicy` with the provider document on `reason.body`. `Generation.AwaitOptions` (`{ poll?: Poll }`) is the one options type for `await`, `events`, `Video.generate`, and `Video.stream`.
102102

103-
`MediaProtocol.stream` is the incremental kind every speech route uses, with the same discipline as LLM protocols. `MediaRoute.stream` submits the caller's request as `MediaProtocol.Addressed<Request>` (`{ ...request, mode }`, `mode: "generate" | "stream"`), so one provider stays one protocol: `body.from`, the endpoint path, and `frames` read `request.mode` to pick the body, path, and framing. `frames(bytes, context)` returns frames — `Framing.sse`, `Framing.lines`, `Framing.document` (a single-document response shaped like a streamed record), or the raw `bytes` for chunked audio. `initial()` is fresh per-response parser state; `step` folds each frame into it and emits modality events; `finish(state, context)` runs once after the last frame with the request, body, and observed `http` (header-only usage lives there) and emits exactly one terminal event or fails with `MediaProtocol.incomplete`. Keep parser state to real accumulators and derive anything the request or body determines in `finish`. `generate` runs the same stream and folds it with the modality's `collect`. Request-derived URL parameters go on the JSON body's `query`, applied before route and caller `http.query`. Decode frames with `MediaProtocol.decodeFrame` and raise stream-time failures with `MediaProtocol.frameError` (the frame stays on `reason.body`); protocols never thread HTTP context, because the route fills `reason.http` on stream errors that lack it. Speech protocols share `protocols/utils/speech-stream.ts` for deltas, timestamps, voice ids, PCM and container descriptions, and the terminal asset.
103+
`MediaProtocol.stream` is the incremental kind every speech route uses, with the same discipline as LLM protocols. `MediaRoute.stream` submits the caller's request as `MediaProtocol.Addressed<Request>` (`{ ...request, mode }`, `mode: "generate" | "stream"`), so one provider stays one protocol: `body.from`, the endpoint path, and `frames` read `request.mode` to pick the body, path, and framing. `frames(bytes, context)` returns frames — `Framing.sse`, `Framing.lines`, `Framing.document` (a single-document response shaped like a streamed record), or the raw `bytes` for chunked audio. `initial()` is fresh per-response parser state; `step` folds each frame into it and emits modality events; `finish(state, context)` runs once after the last frame with the request, body, and observed `http` (header-only usage lives there) and emits exactly one terminal event or fails with `MediaProtocol.incomplete`. Keep parser state to real accumulators and derive anything the request or body determines in `finish`. `generate` runs the same stream and folds it with the modality's `collect`. Request-derived URL parameters go on the body's `query` (array values repeat the parameter), applied before route and caller `http.query`. Decode frames with `MediaProtocol.decodeFrame` and raise stream-time failures with `MediaProtocol.frameError` (the frame stays on `reason.body`); protocols never thread HTTP context, because the route fills `reason.http` on stream errors that lack it. Speech protocols share `protocols/utils/speech-stream.ts` for deltas, timestamps, voice ids, PCM and container descriptions, and the terminal asset.
104+
105+
Transcription uses all three kinds (OpenAI and Gemini stream, Deepgram is inline, AssemblyAI is queued): every route carries its `kind` and `TranscriptionClient` dispatches on it. Bodies are `json`, `multipart`, or `binary` (a raw upload), and a queued protocol that must upload media before submitting implements `start.prepare` (`MediaProtocol.Prepare`; AssemblyAI `/v2/upload`).
104106

105107
### URL Construction
106108

‎packages/ai/README.md‎

Lines changed: 65 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -766,6 +766,69 @@ for await (const event of ai.speech.stream({ model, text: "Hello from OpenCode."
766766
}
767767
```
768768

769+
## Transcription
770+
771+
Transcription (speech-to-text) is the one modality whose providers use every route kind: OpenAI and Gemini stream,
772+
Deepgram answers inline, and AssemblyAI is queued. `Transcription.generate` and `Transcription.stream` work on all of
773+
them; `Transcription.start` / `resume` return a `Generation` on queued routes and fail with `UnsupportedOperation`
774+
elsewhere. Models come from `.transcription(...)` selectors on the `OpenAI`, `Google`, `Deepgram`, and `AssemblyAI`
775+
facades. Common fields (`language`, `prompt`, `timestamps: "none" | "segment" | "word"`, `diarize`, `speakers`) lower
776+
natively or fail with a typed `AIError` before any network call; a route may return more than asked.
777+
778+
```ts
779+
import { Media, Transcription, TranscriptionEvent } from "@opencode/ai"
780+
import { AssemblyAI, Deepgram, OpenAI } from "@opencode/ai/providers"
781+
782+
const openai = OpenAI.configure({ apiKey: process.env.OPENAI_API_KEY })
783+
784+
const program = Effect.gen(function* () {
785+
const audio = yield* Media.file("./call.mp3")
786+
787+
// Speaker-labelled segments; labels are provider-native strings ("A", "0", "spk:0").
788+
const response = yield* Transcription.generate({
789+
model: Deepgram.configure({ apiKey }).transcription("nova-3"),
790+
audio,
791+
diarize: true,
792+
timestamps: "word",
793+
})
794+
response.text // "Hello from OpenCode."
795+
response.segments // [{ text, startSeconds, endSeconds, speaker: "0" }]
796+
response.words // [{ text, startSeconds, endSeconds, speaker, confidence }]
797+
response.language // the provider's own value, lowercased ("en", "english", "en_us")
798+
799+
// Text deltas as the model transcribes, then one finish carrying the whole transcript.
800+
yield* Transcription.stream({ model: openai.transcription("gpt-4o-mini-transcribe"), audio }).pipe(
801+
Stream.tap((event) => (TranscriptionEvent.is.textDelta(event) ? Console.log(event.delta) : Effect.void)),
802+
Stream.runDrain,
803+
)
804+
805+
// Queued: persist the token, resume from another process, and await.
806+
const model = AssemblyAI.configure({ apiKey }).transcription("universal-3-5-pro")
807+
const generation = yield* Transcription.start({ model, audio })
808+
const resumed = yield* Transcription.resume(model, JSON.parse(JSON.stringify(generation.token)))
809+
const transcript = yield* resumed.await({ poll: { interval: "3 seconds" } })
810+
})
811+
```
812+
813+
Inline routes emit only `finish` from `stream` (no faked deltas); queued routes emit `generation-queued` /
814+
`generation-progress` before it. `TranscriptionClient.layer` needs `RequestExecutor.Service`.
815+
816+
Provider notes:
817+
818+
- **OpenAI** takes inline audio only; `diarize` needs `gpt-4o-transcribe-diarize`, timestamps need `whisper-1`, and `whisper-1` does not stream.
819+
- **Gemini** needs a transcribe model (`gemini-3.5-transcribe`); `prompt` and `speakers` fail typed.
820+
- **Deepgram** detects the language unless `language` is set; vocabulary goes in `providerOptions.keyterm`.
821+
- **AssemblyAI** uploads inline audio before submitting and is the only route that accepts `speakers`.
822+
823+
The promise client mirrors the Effect API:
824+
825+
```ts
826+
const text = (await ai.transcription.generate({ model, audio })).text
827+
for await (const event of ai.transcription.stream({ model, audio })) if (event.type === "text-delta") write(event.delta)
828+
const generation = await ai.transcription.start({ model: assemblyai, audio })
829+
const transcript = await generation.await({ poll: { interval: 3_000 } })
830+
```
831+
769832
## Public API
770833

771834
- **`LLM.request({...})`** — build a provider-neutral `LLMRequest`. Accepts ergonomic inputs (`system: string`, `prompt: string`) that normalize into the canonical Schema classes.
@@ -778,7 +841,8 @@ for await (const event of ai.speech.stream({ model, text: "Hello from OpenCode."
778841
- **`Media`** — the shared asset type (`Media.Asset`, `Media.Source`) and constructors used by messages, tool results, and media requests.
779842
- **`Generation`** — provider-neutral handle for an in-flight media generation (`await`, `refresh`, `cancel`, `events`) used by queued media routes.
780843
- **`Speech.request` / `Speech.generate` / `Speech.stream`** — text-to-speech through a provider-neutral request; `SpeechClient` is its Effect service and layer.
781-
- **`@opencode/ai/promise`** — `AI.make({ layer? })` and a default `ai` client exposing `llm`, `image`, `video`, and `speech` as Promise / `AsyncIterable` APIs.
844+
- **`Transcription.request` / `generate` / `stream` / `start` / `resume`** — speech-to-text over inline, streaming, and queued routes; `TranscriptionClient` is its Effect service and layer.
845+
- **`@opencode/ai/promise`** — `AI.make({ layer? })` and a default `ai` client exposing `llm`, `image`, `video`, `speech`, and `transcription` as Promise / `AsyncIterable` APIs.
782846

783847
## Testing
784848

0 commit comments

Comments
 (0)