paper-collage-explainer-generator
GitHub将叙述或知识转化为半调子纸拼贴风格的解释性视觉内容,包括静态提示、分镜或定格动画。遵循特定美学政策,规划隐喻与运动,生成具有触觉质感的最终构图及组装动画。
触发场景
安装
npx skills add vllm-project/vllm-omni --skill paper-collage-explainer-generator -g -y
SKILL.md
Frontmatter
{
"name": "paper-collage-explainer-generator",
"metadata": {
"compatibility": "File-based workflow with sibling h3-prompt-writing; media creation uses available image\/video tools and an editor when needed."
},
"description": "Translate narration, ideas, or knowledge points into editorial halftone paper-collage metaphors, still prompts, storyboards, or stop-motion clips. Use for photographic cutout collage and tactile assembly, distinct from paper-puppet dioramas or generic cartoons."
}
Paper Collage Explainer Generator
Read the portable workflow and collage design when constructing scenes. Honor prompt-only and asset-only requests as well as finished video requests.
Meaning and media policy
Extract each source line's meaning, emotion, action verb, visual metaphor, and three to six readable object groups. Preserve the user's topic and wording as context; do not put narration on screen merely because it is available.
Default to tactile paper SFX, with no BGM, voiceover, or subtitles unless requested. An explainer does not inherently require speech. Existing user requests for those layers already establish the choice; do not ask again. A short 16:9 clip of roughly four seconds is a useful default for one assembly beat, subject to backend limits.
Plan the image and motion together
Create a concise plan of metaphor, palette, object groups, assembly order, duration, and sound cues. For a series, keep paper texture, halftone treatment, keylines, shadows, and rhythmic language consistent; vary the metaphor and color field only where the meaning benefits. A user's no-text or no-music requirement applies to every downstream asset, prompt, clip, and edit.
Design each still as the completed final composition. Use black-and-white photographic halftone cutouts, selective colored cardstock, warm cream keylines, subtle fiber/torn edges, and soft physical layer shadows on a bold color field. Keep layered depth and negative space; avoid excessive age, dirt, wrinkles, or automatic brown/kraft backgrounds. Exact editable layers require a compositing workflow, not a claim that a generated raster video is a layered project.
Generate final-composition stills when requested/needed and inspect them. Follow any requested approval gate; otherwise continue within the authorized workflow.
Animate the assembly
Start from a clean paper field matching the planned final color. Base pieces enter, then the main metaphor, then supporting objects. Each piece slides or pops in, bounces slightly, presses flat, pauses, and locks. End with a brief hold on the completed composition. Use stop-motion assembly rather than global fades, smooth digital drifting, fast spinning, or liquid morphs unless explicitly requested.
Use the supplied final still as a last-frame anchor only when the backend supports it. Otherwise identify it honestly as composition/style guidance; do not describe a reference image as a guaranteed final-frame constraint. Compile the actual H3 mode with the sibling guide and bind only real inputs.
Map short paper slides, pops, taps, rustles, and snaps to motion. Preserve those SFX during assembly. When the user adds music or narration later, mix with the useful SFX unless they request full replacement.
Review and deliver
Check the metaphor without relying on text, object count/readability, matching opening color, visible piece-by-piece construction, final-frame resemblance, controlled paper texture, and no unwanted UI/letters/music/voiceover. If the assembly is weak, simplify the groups and make their arrival order concrete.
For multiple clips, preserve the planned order and SFX. Deliver prompts, current stills, and actual clips/assembly as requested. Explain any missing generation capability or remaining drift; keep raw output separate from edited fixes.
版本历史
- f7e2834 当前 2026-09-22 16:47


