maple-chatcut-story-first-editing
GitHub基于ChatCut的中文语音主导视频编辑工作流,涵盖钩子、字幕、章节、动态图形及品牌视觉设计。通过MCP连接管理项目状态,确保可编辑性并执行最终QA,适用于教程和演示视频制作。
Trigger Scenarios
Install
npx skills add MapleShaw/codex-chatcut-workflow --skill maple-chatcut-story-first-editing -g -y
SKILL.md
Frontmatter
{
"name": "maple-chatcut-story-first-editing",
"description": "Run a content-first, review-driven ChatCut workflow for Chinese speech-led videos, demos, tutorials, talking-head footage, and similar edits. Use when planning or executing hooks, captions, chapter progress, Motion Graphics, logos, sound cues, reframing, summary cards, cover insertion, cover-image prompts, timeline retiming, brand-led visual direction, or final edit QA while preserving an editable ChatCut project. Also use when the user asks to repeat the established premium editing style or workflow without locking every video to one color palette or visual template."
}
ChatCut Story-First Editing
Create a polished speech-led edit by aligning every visual and sound cue to the spoken idea. Preserve flexibility: standardize the workflow, timing discipline, and verification—not a fixed visual style.
Compose the ChatCut skills
Use the active ChatCut skills that match the work:
- Use the ChatCut basics skill to target, read, and mutate the live project.
- Use the talking-head guide for speech rhythm and subject-safe placement.
- Use the transcription skill for transcript or caption work.
- Use the Motion Graphics or shader skill only when the requested design needs it.
- Use the verification skill before reporting an edit complete.
These are capability roles, not mandatory installed skill names. If a companion skill is absent, discover the live tool schema and use its supported equivalent. Never invent a tool or install unrelated dependencies. This workflow skill does not itself provide the ChatCut connection.
For installation recovery, structural insertion, source/timeline mapping, PiP synchronization, audio continuity, or a claim that an edit is missing, read references/sync-and-recovery.md.
For a new edit, hook design, or substantial visual pass, read references/editorial-playbook.md. For any visual-direction, Motion Graphics, card, progress, persistent-topic, or caption-style pass, also read references/visual-language.md. For a narrow mechanical timing fix, follow this workflow without loading the references unless judgment is needed.
Non-negotiable operating rules
- Validate the ChatCut MCP connection first, target the exact project, and read the current project state before every mutation batch.
- Never infer timeline state from the web UI and never use UI automation to bypass an unavailable or timed-out MCP connection.
- If the MCP becomes unreadable or times out, stop further mutations and ask the user to refresh or reconnect it.
- Keep source video, timeline clips, captions, and overlays editable. Do not flatten locally, replace the project with an exported render, or modify unrelated repositories.
- Do not export unless the user explicitly requests export. Stop at a reviewable project state when approval is pending.
- Treat structural timeline changes as high risk. Obtain approval before choosing or inserting a hook, removing content, or materially reordering the main story unless the user already approved that exact operation.
- When a full-frame Motion Graphic or bridge card is meant to play by itself, insert a real timeline gap for its full duration and shift every affected video, audio, caption, overlay, effect, and cue group together. Never place the card on an otherwise empty track while speech or screen recording continues underneath it.
- Treat captions as the final locked layer: after any structural timing, content, reframing, transition, or audio change, refresh or rebuild the caption plan and spot-check its boundaries before reporting completion.
Workflow
1. Establish a fresh baseline
Read and record:
- project, active timeline, canvas size, frame rate, and duration;
- main A-roll track and item order;
- caption project, card count, first/last card, text alignment, and layout;
- audio cues, Motion Graphics, logos, images, effects, and their timing;
- current chapter, hook, summary, and cover structures.
- the video's named subject or product, any official brand assets already present, and whether an exact official theme color has been verified.
Check the main A-roll track for gaps, overlaps, and unexpected offsets before editing. Inspect representative frames when placement or composition matters.
2. Build a semantic map
Map the transcript and timeline into:
- opening promise or tension;
- chapters and their actual spoken boundaries;
- demonstrations, mistakes, corrections, and reveals;
- summary or payoff;
- visual whitespace, face position, and subtitle-safe areas.
For a hook, identify candidates with source time range, exact copy, why each works, and a proposed cut structure. A stumble comparison needs enough lead-in and the actual restart, even when that exceeds a short highlight's usual 5–10 seconds. Preserve gesture-linked repetitions. If the user already selected a candidate or authorized the exact change, implement it without asking again; otherwise obtain their choice.
3. Propose the visual direction
Offer a small number of clearly distinct directions when the treatment is subjective. Explain scale, hierarchy, animation behavior, and semantic role—not just color.
Prefer bold, immediately legible visual communication in unused space. Avoid decorative labels, tiny corner treatments, ambiguous marks, or repeated branding that does not improve comprehension.
Before choosing a palette for a named product, company, tool, or theme, verify its current official colors from primary sources such as the official site, product CSS, documentation, or supplied brand assets. Record the exact value and its source. If no official color can be verified, derive a restrained accent from the footage or provided assets and label it as an inference.
Treat the verified theme color as an accent system, not a fill-everything instruction. Build hierarchy with neutral white, off-white, charcoal, and muted gray; reserve the brand color for the subject name, active state, connector, selected result, or one key metric. Keep error/success colors semantic and instantly understandable.
Use the established premium visual-language defaults in references/visual-language.md: large readable type, strong hierarchy, generous spacing, disciplined alignment, and layouts derived from the actual negative space. Do not ship a dense first draft and wait for the user to request basic legibility.
4. Implement in bounded batches
Group dependent items before moving them. A semantic event may include:
- card or large text;
- icon, logo, glow, or supporting graphic;
- sound cue;
- subject resize/reposition;
- chapter/progress state;
- captions affected by structural insertion.
Move the whole group together. After each batch, re-read the project before starting the next batch.
Do not assume a ripple operation moves all tracks. Capture a before/after item ledger, compute each item's expected position, and repair only the unmoved dependents; never shift an already-moved item twice. Split or extend media crossing the insertion boundary only with valid source handles, and preserve its gain, fades, and routing. Handle caption timing through caption tools, never item operations. Use Script for spoken-content selection when the current tool contract requires it.
For an insertion at the beginning, convert the requested duration to integer frames. Shift every time-bound downstream element by exactly that frame count, including captions, audio cues, overlays, progress graphics, effects, and A-roll. Ensure the inserted cover or cold-open is visually clean until its final frame unless an overlay was explicitly requested.
For a standalone full-frame bridge card, use the same insertion discipline at any timeline position: create the card's actual gap, ripple all dependent tracks by the same delta, and verify the card's start, settled frame, last frame, and the first resumed frame. The card should have a clear reading task and normally hold for roughly 1–3 seconds; extend it when the copy or diagram cannot be read comfortably.
5. Lock the edit before captions
Captions are the final editorial layer, not part of the first rough cut. Do not enable, generate, or style captions while the A-roll selection, B-roll placement, Motion Graphics, reframing, transitions, and sound cues are still changing. First lock the spoken structure and visual/audio layers, then perform a structural and visual check; only after that enable or refresh captions and verify them against the locked cut. If the timeline later changes, disable or rebuild/refresh captions as part of the final pass so they never become stale.
6. Tune captions and timing
Keep captions horizontally centered unless the composition requires a deliberate exception. Read the current layout before changing position; move in measured increments and inspect actual frames instead of applying assumed coordinates.
For this user's Chinese speech-led videos, prefer Smiley Sans for captions when search_fonts confirms it is available; use Noto Sans SC as the fallback. Preserve the user's established high-legibility treatment unless the footage requires a change: bold white text, restrained dark backing, no more than two lines, and a 1080p baseline around 56–60 px. Scale proportionally for other canvases and verify real pixels rather than copying raw numbers blindly. For Motion Graphics, prefer renderer-supported Chinese fonts such as Noto Sans SC unless a verified brand font is both available and cloud-safe.
Align cues to meaning:
- show a prompt when its idea is introduced, not merely near the same timestamp;
- land a ding or error sound on the visible transition, mistake reveal, or correction beat;
- make wrong/right symbols correspond unambiguously to the text they judge;
- resize the speaker before a side card needs the space, and restore the speaker when the card's role ends;
- when one member of a cue group moves, re-check every member.
7. Verify after every material change
Perform both structural and visual verification.
Structural checks:
- captions still exist and expected count/content remains;
- the A-roll track has no gaps, overlaps, or accidental source offsets;
- intended standalone cards and playback segments explain any main-track gaps, with no accidental empty composition;
- total duration and boundary frames are coherent;
- every dependent cue moved by the intended delta;
- deleted overlays are absent only from the intended ranges;
- no unsupported font remains in Motion Graphics before export.
Visual checks:
- inspect frames immediately before, at, and after every changed boundary;
- inspect the first visible frame and a settled frame of entrance animations;
- verify caption visibility, face clearance, hierarchy, icon meaning, and cue synchronization;
- inspect representative chapter and summary transitions, not only the opening.
Do not claim success from mutation responses alone. Report completion only after the read-back and frame inspection agree.
Reusable and public-package hygiene
When packaging or sharing this workflow, keep the skill abstract and portable:
- include any referenced lightweight playbooks or visual-language files with the package, and keep their links relative;
- remove project IDs, timeline IDs, asset/item IDs, local filesystem paths, private chat excerpts, transient timestamps, and one-off media references;
- describe durable editorial judgments and verification checks instead of copying a frozen timeline recipe;
- treat optional companion skills as capabilities with a graceful fallback, so the workflow remains useful when a specific helper skill is unavailable.
8. Hand off for review
Summarize what changed, the exact timing delta where relevant, and the verification result. Keep the ChatCut editor available for review. Mention any deliberate remaining choice or known limitation. Do not export during a review handoff.
9. Deliver a cover-image prompt at the end
When a cover is requested and has not already been accepted, provide a production-ready cover-image generation prompt based on the actual video—not a generic template. Skip this step for timing-only fixes or a video with an approved cover. The prompt must include:
- the video's clearest promise, conflict, proof point, or result;
- the exact overlay copy to appear on the cover, written verbatim inside the prompt (brand name, main headline, supporting line, and any optional proof label);
- hierarchy and placement for every text line, including which words are largest, boldest, or accent-colored;
- the verified theme color and restrained neutral palette, with the source already recorded during visual-direction planning;
- the intended aspect ratio and composition derived from the footage, including a clean area reserved for compositing the real person and protection for the headline;
- a negative prompt covering unwanted people, fake logos, garbled text, tiny type, clipping, clutter, and overuse of effects.
Do not omit the text with instructions such as “add title later.” The user explicitly needs to paste the prompt into an image generator, so include the exact Chinese wording and tell the generator to render it clearly and accurately. Keep the visual treatment premium: one focal statement, generous spacing, strong contrast, disciplined alignment, and the brand color as an accent rather than a full-frame fill. Adapt the layout to the video's negative space; do not force the current video's left-side layout onto a different subject.
Before handing over the prompt, check that the brand spelling is exact, the headline is readable at thumbnail size, no line is likely to clip, the reserved person area is unoccupied, and the title explains the video's topic even when the cover is viewed without audio.
Preference hierarchy
When tradeoffs arise, prioritize:
- spoken meaning and narrative clarity;
- timing and synchronization;
- legibility and visual hierarchy;
- continuity and editability;
- stylistic novelty.
Adapt palette, typography, layout, iconography, and animation language to each video's subject and footage. Reuse the decision process, not the surface treatment.
Version History
- e82aa01 Current 2026-09-27 10:26


