ModelInfo
English
All models

Qwen3-VL-Flash

Small vision-language model with thinking and non-thinking modes and 2D/3D localization.

Alibabaqwen3-vl-flash2025-10-14Report incorrect data

Specs

Context
262.1K
Max output
32.8K
Input price
$0.05/ 1M tokens
Output price
$0.4/ 1M tokens
Released
2025-10-14

Context

Context window
262.1K
Max input
260.1K
Max output
32.8K

Pricing · USD / 1M tokens

Input
$0.05
Output
$0.4
Cache read
$0.005
Cache write
$0.0625

Modalities

Input
TextImageVideo
Output
Text

API

API types
chat

Reasoning

Reasoning
Yes
Levels
—

Info

Status
Retired
Released
2025-10-14
Knowledge cutoff
—

Features

Confirmed
structured_outputstreamingcaching
Source · official docsVerified 2026-10-11

How to call

1 provider

alibabamodel = qwen3-vl-flashreasoning = enable_thinking

chat
POSThttps://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions

JSON

Standard format
qwen3-vl-flash.json
{
  "id": "qwen3-vl-flash",
  "object": "model",
  "created": 1760400000,
  "owned_by": "alibaba",
  "name": "Qwen3-VL-Flash",
  "api": {
    "types": [
      "chat"
    ]
  },
  "limits": {
    "context": 262144,
    "input": 260096,
    "output": 32768
  },
  "modalities": {
    "input": [
      "text",
      "image",
      "video"
    ],
    "output": [
      "text"
    ]
  },
  "reasoning": {
    "supported": true,
    "efforts": []
  },
  "pricing": {
    "currency": "USD",
    "unit": "1M_tokens",
    "input": 0.05,
    "output": 0.4,
    "cache_read": 0.005,
    "cache_write": 0.0625
  },
  "features": [
    "structured_output",
    "streaming",
    "caching"
  ],
  "info": {
    "status": "retired",
    "release_date": "2025-10-14",
    "knowledge_cutoff": null,
    "description": "Small vision-language model with thinking and non-thinking modes and 2D/3D localization.",
    "docs": "https://www.alibabacloud.com/help/en/model-studio/text-generation-model",
    "verified_at": "2026-10-11"
  }
}

Official sources

5