ModelInfo
简体中文
全部模型

VILA

视觉语言模型,可对来自物理或虚拟世界的图像和视频进行查询与摘要。

NVIDIAvila报告数据错误

规格

上下文
—
最大输出
—
输入价
—
输出价
—
发布日期
—

上下文

官方未公布

价格

官方未公布

输入输出

输入
文本图像视频
输出
文本

API

接口类型
chat

思考

官方未公布

源信息

状态
可用
发布日期
—
知识截止
—

能力

官方未公布
来源 · 官方文档核验 2026-10-11

调用方式

1 个 Provider

nvidiamodel = nvidia/vila

JSON

标准格式
vila.json
{
  "id": "vila",
  "object": "model",
  "created": null,
  "owned_by": "nvidia",
  "name": "VILA",
  "api": {
    "types": [
      "chat"
    ]
  },
  "limits": {
    "context": null,
    "input": null,
    "output": null
  },
  "modalities": {
    "input": [
      "text",
      "image",
      "video"
    ],
    "output": [
      "text"
    ]
  },
  "reasoning": {
    "supported": null,
    "efforts": []
  },
  "pricing": null,
  "features": [],
  "info": {
    "status": "active",
    "release_date": null,
    "knowledge_cutoff": null,
    "description": "Vision language model that queries and summarizes images and video from the physical or virtual world.",
    "docs": "https://docs.api.nvidia.com/nim/reference/nvidia-vila",
    "verified_at": "2026-10-11"
  }
}

官方来源

1