transcription
GitHubChatCut转录技能,指导如何识别媒体资产、触发/检查转录进度、查找字幕及处理失败或无音频状态。
Trigger Scenarios
Install
npx skills add ChatCut-Inc/agent-plugin --skill transcription -g -y
SKILL.md
Frontmatter
{
"name": "transcription",
"description": "Hosted ChatCut plugin sessions only (the `chatcut` MCP server). If the conversation is driving ChatCut Desktop (a `chatcut_desktop*` MCP server), load this skill only when the user explicitly chooses the plugin\/web surface — desktop sessions otherwise ship their own instructions and tools. Use for ChatCut transcription, transcript readiness, captions, subtitles, transcript repair, filler removal, and speech-led editing setup."
}
Transcription
- Use
browse_assetsto identify the video/audio asset and its transcription state. - For newly imported local media, complete the
asset-importworkflow first. - Check
track_progresswith targettranscription. It returns current status; follow the returned check-back guidance and do not busy-loop. - Use
find_transcriptfor timestamped text lookup. - Use the current caption tools to enable, inspect, translate, or style captions only after transcription is ready.
Do not call a transcription stuck from one pending status. Treat an explicit failed terminal state immediately; otherwise allow at least max(5 minutes, min(60 minutes, 2 x asset duration)), or at least 10 minutes across multiple checks when duration is unknown.
no_audio / no-speech is a successful terminal analysis result: the media is usable, but there is no transcript to read or caption. Do not report it as a transcription failure or retry it unless the user says the media contains speech that should have been detected.
Use trigger_transcript when the user explicitly wants transcription started or restarted. It is idempotent: only idle and error start work; ready, no-audio, and in-progress assets are left unchanged. If prepared transcription audio exists it is reused; otherwise the open Web/Desktop editor is asked to prepare and upload audio through its native pipeline. Then inspect readiness with track_progress.
For an idle or explicitly failed (error) run, call trigger_transcript with the asset id, then check transcription progress again. If an in-progress run has genuinely exceeded the stuck threshold, use manage_transcript action retry_transcription because trigger_transcript intentionally leaves in-progress work unchanged. The manage tool also remains available for low-level recovery where an external MCP host must supply prepared audioBase64. Repair source words with the transcript-fix action instead of rewriting visible captions when the source transcript itself is wrong.
For semantic speech edits, load talking-head-guide and use the Script workflow. Use mechanical cleanup only for fixed fillers and pauses; do not replace transcript-aware editing with destructive physical timeline cuts.
Version History
-
014586d
Current 2026-08-27 10:43
新增对闲置(idle)和错误(error)状态的显式处理逻辑,引入trigger_transcript工具以幂等方式启动转录,并明确区分机械清理与语义编辑场景。
-
926f56b
2026-08-12 11:06
更新内容:移除了对显式失败状态的立即处理逻辑,增加了 `no_audio` 作为成功终态的处理说明,细化了重试和修复的操作指引。
-
e1867a8
2026-08-02 23:36
重构了流程步骤,明确区分新导入媒体与现有资产的转录准备;细化了等待时间计算逻辑(基于资产时长);新增对语义语音编辑和机械清理的指导;优化了故障重试策略。
-
f39cdae
2026-07-30 21:53
更新获取资产ID的方法,由read_project改为browse_assets;优化本地资源处理指引,明确通过asset-import技能重新导入或先下载再导入,避免手动重链。
- 5e9afe0 2026-07-22 10:58


