transcription
GitHub处理音视频转录、字幕及修复。涵盖资产导入、状态检查、触发转录、进度追踪及失败重试,支持语义编辑与填充词清理,确保准确识别无语音状态。
Trigger Scenarios
Install
npx skills add ChatCut-Inc/agent-plugin --skill transcription -g -y
SKILL.md
Frontmatter
{
"name": "transcription",
"description": "Hosted ChatCut plugin sessions only (the `chatcut` MCP server). If the conversation is driving ChatCut Desktop (a `chatcut_desktop*` MCP server), load this skill only when the user explicitly chooses the plugin\/web surface — desktop sessions otherwise ship their own instructions and tools. Use for ChatCut transcription, transcript readiness, captions, subtitles, transcript repair, filler removal, and speech-led editing setup."
}
Transcription
- Use
browse_assetsto identify the video/audio asset and its transcription state. - For newly imported local media, complete the
asset-importworkflow first. - Check
track_progresswith targettranscription. It returns current status; follow the returned check-back guidance and do not busy-loop. - Use
find_transcriptfor timestamped text lookup. - Use the current caption tools to enable, inspect, translate, or style captions only after transcription is ready.
Do not call a transcription stuck from one pending status. Treat an explicit failed terminal state immediately; otherwise allow at least max(5 minutes, min(60 minutes, 2 x asset duration)), or at least 10 minutes across multiple checks when duration is unknown.
no_audio / no-speech is a successful terminal analysis result: the media is usable, but there is no transcript to read or caption. Do not report it as a transcription failure or retry it unless the user says the media contains speech that should have been detected.
Use trigger_transcript when the user explicitly wants transcription started or restarted. It is idempotent: only idle and error start work; ready, no-audio, and in-progress assets are left unchanged. If prepared transcription audio exists it is reused; otherwise the open Web/Desktop editor is asked to prepare and upload audio through its native pipeline. Then inspect readiness with track_progress.
For an idle or explicitly failed (error) run, call trigger_transcript with the asset id, then check transcription progress again. If an in-progress run has genuinely exceeded the stuck threshold, use manage_transcript action retry_transcription because trigger_transcript intentionally leaves in-progress work unchanged. The manage tool also remains available for low-level recovery where an external MCP host must supply prepared audioBase64. Repair source words with the transcript-fix action instead of rewriting visible captions when the source transcript itself is wrong.
For semantic speech edits, load talking-head-guide and use the Script workflow. Use mechanical cleanup only for fixed fillers and pauses; do not replace transcript-aware editing with destructive physical timeline cuts.
Version History
- 0dd9c5e Current 2026-09-22 02:33


