DashScope
Requires DASHSCOPEAPIKEY in .env. Get one at https://dashscope.aliyun.com/.
Current API
CRITICAL: DashScope's /compatible-mode/v1/ only supports /chat/completions and /embeddings. Image generation, TTS, and ASR all use DashScope-native endpoints — not OpenAI-compatible paths.
All three tools use Authorization: Bearer $DASHSCOPEAPIKEY.
Image Generation
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
- Model:
qwen-image-2.0-pro (default), qwen-image-max, wan2.7-image, z-image-turbo
- Body:
{model, input: {messages: [{role: "user", content: [{text: "prompt"}]}]}, parameters: {size: "W*H", n, prompt_extend, watermark}}
- Size format uses asterisk:
"1024*1024" not "1024x1024"
- Response:
output.choices[0].message.content[0].image (URL, valid ~24h) — must download separately
Text-to-Speech
POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
Same endpoint as image gen, different body.
- Model:
qwen3-tts-flash (default), qwen3-tts-instruct-flash, qwen-tts-2025-05-22
- Body:
{model, input: {text, voice: "Cherry", language_type: "Auto"}}
- Response:
output.audio.url (WAV, valid ~24h) — must download separately
ASR with Word-Level Timestamps
POST https://dashscope.aliyuncs.com/api/v1/services/audio/asr/transcription
Header: X-DashScope-Async: enable
- Model:
qwen3-asr-flash-filetrans (NOT qwen3-asr-flash — the sync version has no word timestamps)
- Body:
{model, input: {fileurl: "https://public-url/audio.mp3"}, parameters: {enablewords: true, language_hints: ["zh","en"]}}
- Returns
taskid → poll GET /api/v1/tasks/{taskid} until SUCCEEDED → download output.result.transcription_url → JSON with transcripts[].sentences[].words[]
- Timestamps in
begintime/endtime are in milliseconds — the tool normalizes to seconds
OpenMontage Usage
Image via selector
from tools.graphics.image_selector import ImageSelector
result = ImageSelector().execute({
"preferred_provider": "dashscope",
"prompt": "一只猫坐在沙发上",
"output_path": "projects/my-video/assets/images/cat.png",
})
TTS via selector
from tools.audio.tts_selector import TTSSelector
result = TTSSelector().execute({
"preferred_provider": "dashscope",
"text": "如果 AI 真的会改变未来,普通人到底该怎么参与?",
"voice": "Cherry",
"output_path": "projects/my-video/assets/audio/narration.wav",
})
ASR directly (word timestamps for subtitles)
from tools.analysis.dashscope_asr import DashscopeAsr
result = DashscopeAsr().execute({
"audio_url": "https://example.com/narration.wav",
"output_path": "projects/my-video/assets/audio/transcription.json",
})
# result.data["words"] is a flat list of {text, begin_time_seconds, end_time_seconds}
Recommended Workflow
- Image: Generate a sample first. Check
prompt_extend: true (default) — DashScope rewrites your prompt for better results. Disable if you need literal prompt adherence.
- TTS: Generate a 10-15 second sample before full narration. Approve voice and pacing before committing to full generation.
- ASR: Audio must be at a publicly accessible URL. Upload to any public host (S3, etc.) first. Local paths are rejected with a clear error.
- Subtitles: Build from
result.data["words"] — each word has begintimeseconds and endtimeseconds. Group words into caption phrases by language semantics, not fixed character count.
Parameters
Image (dashscope_image)
prompt (required): text prompt
model: default qwen-image-2.0-pro
size: default "1024*1024" — asterisk separator, not "x"
n: 1-6 images
negative_prompt: things to avoid (max 500 chars)
prompt_extend: default true — auto-rewrite prompt for better results
watermark: default false
seed: for reproducibility
TTS (dashscope_tts)
text (required): text to synthesize (max 600 chars for qwen3-tts-flash)
model: default qwen3-tts-flash
voice: default "Cherry" — other voices: "Ethan", "Chelsie", etc.
language_type: default "Auto" — "Chinese", "English", "Japanese", "Korean"
instructions: natural language delivery instructions (only for qwen3-tts-instruct-flash)
ASR (dashscope_asr)
audio_url (required): must be publicly accessible URL
model: qwen3-asr-flash-filetrans (only model that supports word timestamps)
language_hints: default ["zh", "en"]
enable_words: default true — required for word-level timestamps
pollintervalseconds: default 5.0
timeout_seconds: default 300
Troubleshooting
- Image size error: Use
"WH" with asterisk, not "WxH". Example: "20482048".
- TTS no audio URL: Check
output.audio.url — if empty, the model name or voice may be wrong.
- ASR "file not accessible":
audio_url must be publicly reachable. DashScope servers fetch the file; local paths and auth-gated URLs don't work.
- ASR poll timeout: Increase
timeout_seconds (default 300). Long audio files take longer to transcribe.
- ASR no word timestamps: Ensure
enable_words: true and model is qwen3-asr-flash-filetrans (not the sync qwen3-asr-flash).
- Auth error (401): Verify
DASHSCOPEAPIKEY is set. Use Authorization: Bearer $KEY header.
Safety
Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names. The tool's safeerror() method redacts the key from error messages.