HomeSkills › Audio Video To Text
Developer Tools

Audio Video To Text

1.8K downloads 1 stars Version 1.0.0 Rank #7658 of 10,000+

What this skill does

音视频转文字技能,使用 Whisper 进行语音识别。支持多种音视频格式,可输出纯文本、SRT/VTT 字幕或 JSON 格式。适用于会议记录、视频字幕生成、采访整理、播客转录等场景。

Audio Video To Text is part of the Developer Tools category — developer tools that help your agent write, review, debug, and ship code. You can install it on its own or alongside other developer tools skills from the OpenClaw catalog.

Audio Video To Text is ranked #7658 by downloads in the OpenClaw skill catalog (1.8K total downloads, 1 stars). It belongs to the Developer Tools category alongside 3305 other top-10000 skills.

How to install Audio Video To Text

The easiest path is via the OpenClaw Easy desktop app — one click, no terminal required:

  1. Download OpenClaw Easy for macOS or Windows (free, one-click installer, ~30 seconds).
  2. Open the in-app Skills panel.
  3. Search for audio-video-to-text and click Install.
  4. The skill activates automatically when an incoming message matches its description.

Install from the command line

If you already run the OpenClaw CLI, add Audio Video To Text with a single command:

openclaw skills add audio-video-to-text

This pulls audio-video-to-text from ClawHub and installs it into ~/.openclaw/skills/audio-video-to-text/. Restart the OpenClaw gateway afterwards so the new skill is discovered.

How to use Audio Video To Text

Once installed, Audio Video To Text activates on its own: when an incoming message on WhatsApp, Telegram, Slack, Discord, Feishu or Line matches the skill's description, your OpenClaw agent loads it and runs the workflow. You can also trigger it explicitly by describing the task in chat. No extra configuration is required after install.

Turning video into a finished short

Audio Video To Text works with video. Getting from there to something postable is a separate job: ViralMint — an open-source video pipeline — reframes to vertical, cleans the audio, cuts dead air, burns animated captions, and exports a version sized for TikTok, Reels, and Shorts.

It runs as an MCP server, so an OpenClaw agent can drive it from the same chat you already use: ask for a short, and the render comes back finished. See how to connect a video pipeline to your agent.

Manual install (advanced)

If you prefer manual installation:

  1. Click the Download .zip button above to grab audio-video-to-text-1.0.0.zip directly from our S3 mirror.
  2. Unzip into ~/.openclaw/skills/audio-video-to-text/ (create the directory if it does not exist).
  3. Restart OpenClaw Easy (or the OpenClaw CLI gateway) so the new skill is discovered.

Frequently asked questions

How do I install Audio Video To Text?

Install Audio Video To Text in the OpenClaw Easy desktop app by opening the Skills panel, searching for audio-video-to-text, and clicking Install. From a terminal you can run: openclaw skills add audio-video-to-text. Either way the skill is placed in ~/.openclaw/skills/audio-video-to-text/.

Is Audio Video To Text free?

Yes. Audio Video To Text is free and open-source, distributed under the Apache-2.0 license through ClawHub. No account or payment is required to download or run it.

What does Audio Video To Text do?

音视频转文字技能,使用 Whisper 进行语音识别。支持多种音视频格式,可输出纯文本、SRT/VTT 字幕或 JSON 格式。适用于会议记录、视频字幕生成、采访整理、播客转录等场景。

Related: more developer tools skills

If Audio Video To Text looks useful, you may also want to check out other developer tools skills in the OpenClaw catalog:

Browse the full OpenClaw skill catalog

This page covers just one skill. The OpenClaw skill hub has 10,000+ more — search, sort by downloads or stars, and install any of them in one click. There is also a curated awesome-openclaw-skills list grouped by use case.

Get OpenClaw Easy — Free

Install Audio Video To Text and 10,000+ other OpenClaw skills in one click. Free, open-source, runs locally on macOS & Windows.

Free, open-source · Apache-2.0 · Works with Claude, ChatGPT, Gemini, or local Ollama models