ModelInfo
English
All models

StepAudio 2.5 Chat

Conversational model that understands speech, including hesitation and laughter, and outputs text.

StepFunstepaudio-2.5-chatReport incorrect data

Specs

Context
—
Max output
—
Input price
$1.5/ 1M tokens
Output price
$3.5/ 1M tokens
Released
—

Context

Not published

Pricing · USD / 1M tokens

Input
$1.5
Output
$3.5
Cache read
$0.3
Cache write
—

Modalities

Input
AudioText
Output
Text

API

API types
chat

Reasoning

Not published

Info

Status
Active
Released
—
Knowledge cutoff
—

Features

Not published
Source · official docsVerified 2026-10-10

How to call

1 provider

stepfunmodel = stepaudio-2.5-chat

chat
POSThttps://api.stepfun.ai/v1/chat/completions

JSON

Standard format
stepaudio-2.5-chat.json
{
  "id": "stepaudio-2.5-chat",
  "object": "model",
  "created": null,
  "owned_by": "stepfun",
  "name": "StepAudio 2.5 Chat",
  "api": {
    "types": [
      "chat"
    ]
  },
  "limits": {
    "context": null,
    "input": null,
    "output": null
  },
  "modalities": {
    "input": [
      "audio",
      "text"
    ],
    "output": [
      "text"
    ]
  },
  "reasoning": {
    "supported": null,
    "efforts": []
  },
  "pricing": {
    "currency": "USD",
    "unit": "1M_tokens",
    "input": 1.5,
    "output": 3.5,
    "cache_read": 0.3,
    "cache_write": null
  },
  "features": [],
  "info": {
    "status": "active",
    "release_date": null,
    "knowledge_cutoff": null,
    "description": "Conversational model that understands speech, including hesitation and laughter, and outputs text.",
    "docs": "https://platform.stepfun.ai/docs/en/guides/models/stepaudio-2.5-chat",
    "verified_at": "2026-10-10"
  }
}

Official sources

4