meeting-video-grounding
GitHub将会议视频转换为音频优先的结构化输出。提取音频后,复用现有技能生成转录本,再基于转录本生成结构化会议笔记,适用于以语音为主要信息源的视频分析场景。
Trigger Scenarios
Install
npx skills add gaotiexinqu/OneResearchClaw --skill meeting-video-grounding -g -y
SKILL.md
Frontmatter
{
"name": "meeting-video-grounding",
"description": "Convert a meeting video into an audio-first transcript bundle, then use meeting-grounding to produce structured meeting grounding outputs."
}
Meeting Video Grounding
Convert a meeting video into structured meeting grounding outputs by:
- extracting audio from the video
- reusing the existing
audio_structuringskill to producemeeting_transcript.txt - reusing the existing
meeting-groundingskill to turn that transcript into meeting grounding outputs
This skill is for meeting videos where the primary information comes from speech. It is intentionally audio-first. It does not attempt full visual understanding of the video.
When to Use
Use this skill when:
- the input is a meeting video or discussion video
- the main information is expected to come from spoken content
- you want to reuse the existing audio transcription and meeting grounding workflow
Do not use this skill when:
- the task requires visual analysis of slides, demos, whiteboards, or screen content as first-class evidence
- the video has little or no speech
- the task is to write a polished final report directly from the raw video
Input
A single meeting video file.
Typical examples:
.mp4.mkv.mov.webm
Optional input:
transcription_language: language code such asenorzh
Output Bundle
For each input video, create one bundle directory:
data/grounded_notes/<ground_id>/
Inside that bundle, the expected outputs are always:
<bundle_dir>/
├─ extracted.md
├─ extracted_meta.json
├─ grounded.md
├─ audio/
│ └─ meeting_audio.wav
└─ transcript/
└─ meeting_transcript.txt
If the meeting contains multiple independent topics that should be researched separately downstream, the bundle may also contain:
<bundle_dir>/
├─ topic_manifest.json
└─ child_outputs/
├─ topic_01/
│ └─ grounded.md
├─ topic_02/
│ └─ grounded.md
└─ ...
Important separation of responsibilities
scripts/run.shis responsible for:- extracting audio from the input video
- calling the existing
audio_structuringskill - creating the bundle files:
audio/meeting_audio.wavtranscript/meeting_transcript.txtextracted.mdextracted_meta.json
- The agent is responsible for:
- reading the transcript bundle
- applying the existing
meeting-groundingskill - always writing the meeting-level:
grounded.md
- and, when appropriate, also writing:
topic_manifest.jsonchild_outputs/topic_xx/grounded.md
grounded.md must be a real grounding note.
It must not remain a placeholder scaffold.
Required Agent Workflow
- Run the existing
scripts/run.shentrypoint for this skill. - Confirm that the bundle exists and that these files are present:
extracted.mdextracted_meta.jsonaudio/meeting_audio.wavtranscript/meeting_transcript.txt
- Read
transcript/meeting_transcript.txtas the primary grounding evidence. - Reuse the existing
meeting-groundingskill on that transcript. - Always save the meeting-level structured note to
grounded.mdinside the same bundle directory. - If the transcript clearly contains multiple independent topics, also save:
topic_manifest.jsonchild_outputs/topic_xx/grounded.md
- Do not stop after confirming that the transcript bundle exists.
Important scope rule
This skill currently treats meeting videos as audio-first inputs.
That means:
- the extracted transcript is the primary evidence for grounding
- absence of visual analysis is not, by itself, a failure
- do not invent slide content, visual details, or screen evidence that were not captured in the transcript
Output Format
The final grounded.md must follow the existing meeting-grounding schema exactly.
If topic children are created, each child grounded note must also follow the same meeting-grounding schema.
Do not invent a new schema here.
Instructions
- Do not implement a new ASR pipeline.
- Do not implement a new meeting summarizer.
- Do not directly summarize the raw video without first running the existing workflow.
- Reuse the existing
audio_structuringskill for transcription. - Reuse the existing
meeting-groundingskill for transcript grounding. - Treat the transcript as the primary evidence.
- Keep the workflow simple and stable.
Failure Handling
- If the input video file does not exist, fail clearly.
- If the video has no audio stream, fail clearly.
- If audio extraction fails, fail clearly.
- If
meeting_transcript.txtis not produced, fail clearly. - Do not pretend the task succeeded if only part of the workflow completed.
Example Invocation
/meeting-video-grounding
Version History
- 37e86c6 Current 2026-07-24 12:30


