dashscope
GitHub集成阿里云百炼DashScope,提供图像生成、语音合成及带词级时间戳的自动语音识别功能。支持通过原生API调用Qwen系列模型,适用于多媒体内容创作与处理场景。
Trigger Scenarios
Install
npx skills add calesthio/OpenMontage --skill dashscope -g -y
SKILL.md
Frontmatter
{
"name": "dashscope",
"description": "DashScope (Alibaba Cloud Bailian \/ 阿里云百炼) integration — image generation (qwen-image-2.0-pro), text-to-speech (qwen3-tts-flash), and ASR with word-level timestamps (qwen3-asr-flash-filetrans). Use when generating images via Qwen-Image, narrating via Qwen-TTS, or transcribing with word-level timestamps via Qwen-ASR."
}
DashScope
Requires DASHSCOPE_API_KEY in .env. Get one at https://dashscope.aliyun.com/.
Current API
CRITICAL: DashScope's /compatible-mode/v1/ only supports /chat/completions and /embeddings. Image generation, TTS, and ASR all use DashScope-native endpoints — not OpenAI-compatible paths.
All three tools use Authorization: Bearer $DASHSCOPE_API_KEY.
Image Generation
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
- Model:
qwen-image-2.0-pro(default),qwen-image-max,wan2.7-image,z-image-turbo - Body:
{model, input: {messages: [{role: "user", content: [{text: "prompt"}]}]}, parameters: {size: "W*H", n, prompt_extend, watermark}} - Size format uses asterisk:
"1024*1024"not"1024x1024" - Response:
output.choices[0].message.content[0].image(URL, valid ~24h) — must download separately
Text-to-Speech
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
Same endpoint as image gen, different body.
- Model:
qwen3-tts-flash(default),qwen3-tts-instruct-flash,qwen-tts-2025-05-22 - Body:
{model, input: {text, voice: "Cherry", language_type: "Auto"}} - Response:
output.audio.url(WAV, valid ~24h) — must download separately
ASR with Word-Level Timestamps
POST https://dashscope.aliyuncs.com/api/v1/services/audio/asr/transcription
Header: X-DashScope-Async: enable
- Model:
qwen3-asr-flash-filetrans(NOTqwen3-asr-flash— the sync version has no word timestamps) - Body:
{model, input: {file_url: "https://public-url/audio.mp3"}, parameters: {enable_words: true, language_hints: ["zh","en"]}} - Returns
task_id→ pollGET /api/v1/tasks/{task_id}untilSUCCEEDED→ downloadoutput.result.transcription_url→ JSON withtranscripts[].sentences[].words[] - Timestamps in
begin_time/end_timeare in milliseconds — the tool normalizes to seconds
OpenMontage Usage
Image via selector
from tools.graphics.image_selector import ImageSelector
result = ImageSelector().execute({
"preferred_provider": "dashscope",
"prompt": "一只猫坐在沙发上",
"output_path": "projects/my-video/assets/images/cat.png",
})
TTS via selector
from tools.audio.tts_selector import TTSSelector
result = TTSSelector().execute({
"preferred_provider": "dashscope",
"text": "如果 AI 真的会改变未来,普通人到底该怎么参与?",
"voice": "Cherry",
"output_path": "projects/my-video/assets/audio/narration.wav",
})
ASR directly (word timestamps for subtitles)
from tools.analysis.dashscope_asr import DashscopeAsr
result = DashscopeAsr().execute({
"audio_url": "https://example.com/narration.wav",
"output_path": "projects/my-video/assets/audio/transcription.json",
})
# result.data["words"] is a flat list of {text, begin_time_seconds, end_time_seconds}
Recommended Workflow
- Image: Generate a sample first. Check
prompt_extend: true(default) — DashScope rewrites your prompt for better results. Disable if you need literal prompt adherence. - TTS: Generate a 10-15 second sample before full narration. Approve voice and pacing before committing to full generation.
- ASR: Audio must be at a publicly accessible URL. Upload to any public host (S3, etc.) first. Local paths are rejected with a clear error.
- Subtitles: Build from
result.data["words"]— each word hasbegin_time_secondsandend_time_seconds. Group words into caption phrases by language semantics, not fixed character count.
Parameters
Image (dashscope_image)
prompt(required): text promptmodel: defaultqwen-image-2.0-prosize: default"1024*1024"— asterisk separator, not "x"n: 1-6 imagesnegative_prompt: things to avoid (max 500 chars)prompt_extend: defaulttrue— auto-rewrite prompt for better resultswatermark: defaultfalseseed: for reproducibility
TTS (dashscope_tts)
text(required): text to synthesize (max 600 chars for qwen3-tts-flash)model: defaultqwen3-tts-flashvoice: default"Cherry"— other voices:"Ethan","Chelsie", etc.language_type: default"Auto"—"Chinese","English","Japanese","Korean"instructions: natural language delivery instructions (only forqwen3-tts-instruct-flash)
ASR (dashscope_asr)
audio_url(required): must be publicly accessible URLmodel:qwen3-asr-flash-filetrans(only model that supports word timestamps)language_hints: default["zh", "en"]enable_words: defaulttrue— required for word-level timestampspoll_interval_seconds: default5.0timeout_seconds: default300
Troubleshooting
- Image size error: Use
"W*H"with asterisk, not"WxH". Example:"2048*2048". - TTS no audio URL: Check
output.audio.url— if empty, the model name or voice may be wrong. - ASR "file not accessible":
audio_urlmust be publicly reachable. DashScope servers fetch the file; local paths and auth-gated URLs don't work. - ASR poll timeout: Increase
timeout_seconds(default 300). Long audio files take longer to transcribe. - ASR no word timestamps: Ensure
enable_words: trueand model isqwen3-asr-flash-filetrans(not the syncqwen3-asr-flash). - Auth error (401): Verify
DASHSCOPE_API_KEYis set. UseAuthorization: Bearer $KEYheader.
Safety
Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names. The tool's _safe_error() method redacts the key from error messages.
Version History
- 0af32ce Current 2026-07-24 22:20


