Voxtral TTS
文字转语音模型,支持零样本声音克隆和 9 种语言,流式输出的首个音频延迟约 90ms。
规格
- 上下文
- —
- 最大输出
- —
- 输入价
- $0/ 百万字符
- 输出价
- $16/ 百万字符
- 发布日期
- 2026-03-23
上下文
- 官方未公布
价格 · USD / 百万字符
- 输入
- $0
- 输出
- $16
- 缓存读取
- —
- 缓存写入
- —
输入输出
- 输入
- 文本
- 输出
- 音频
API
- 接口类型
思考
- 官方未公布
源信息
- 状态
- 可用
- 发布日期
- 2026-03-23
- 知识截止
- —
能力
- 官方未公布
调用方式
1 个 Provider
mistralmodel = voxtral-mini-tts-2603
audio
POSThttps://api.mistral.ai/v1/audio/speech
JSON
标准格式
{
"id": "voxtral-mini-tts-2603",
"object": "model",
"created": 1774224000,
"owned_by": "mistral",
"name": "Voxtral TTS",
"api": {
"types": [
"audio"
]
},
"limits": {
"context": null,
"input": null,
"output": null
},
"modalities": {
"input": [
"text"
],
"output": [
"audio"
]
},
"reasoning": {
"supported": null,
"efforts": []
},
"pricing": {
"currency": "USD",
"unit": "1M_characters",
"input": 0,
"output": 16,
"cache_read": null,
"cache_write": null
},
"features": [],
"info": {
"status": "active",
"release_date": "2026-03-23",
"knowledge_cutoff": null,
"description": "Text-to-speech model with zero-shot voice cloning, 9 languages, and streaming with about 90ms time-to-first-audio.",
"docs": "https://docs.mistral.ai/models/voxtral-tts-26-03",
"verified_at": "2026-10-11"
}
}官方来源
3