ModelInfo
English
All models

Gemini 3.5 Transcribe

Speech-to-text model with language detection, speaker diarization, word-level timestamps and custom vocabulary.

Googlegemini-3.5-transcribe2026-08-26Report incorrect data

Specs

Context
—
Max output
—
Input price
$2/ 1M tokens
Output price
$12/ 1M tokens
Released
2026-08-26

Context

Not published

Pricing · USD / 1M tokens

Input
$2
Output
$12
Cache read
—
Cache write
—

Modalities

Input
Audio
Output
Text

API

API types
audio

Reasoning

Reasoning
No

Info

Status
Active
Released
2026-08-26
Knowledge cutoff
2025-01

Features

Not published
Source · official docsVerified 2026-10-11

How to call

1 provider

googlemodel = gemini-3.5-transcribe

JSON

Standard format
gemini-3.5-transcribe.json
{
  "id": "gemini-3.5-transcribe",
  "object": "model",
  "created": 1787702400,
  "owned_by": "google",
  "name": "Gemini 3.5 Transcribe",
  "api": {
    "types": [
      "audio"
    ]
  },
  "limits": {
    "context": null,
    "input": null,
    "output": null
  },
  "modalities": {
    "input": [
      "audio"
    ],
    "output": [
      "text"
    ]
  },
  "reasoning": {
    "supported": false,
    "efforts": []
  },
  "pricing": {
    "currency": "USD",
    "unit": "1M_tokens",
    "input": 2,
    "output": 12,
    "cache_read": null,
    "cache_write": null
  },
  "features": [],
  "info": {
    "status": "active",
    "release_date": "2026-08-26",
    "knowledge_cutoff": "2025-01",
    "description": "Speech-to-text model with language detection, speaker diarization, word-level timestamps and custom vocabulary.",
    "docs": "https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe",
    "verified_at": "2026-10-11"
  }
}

Official sources

6