Skip to content

Commit c08de0c

Browse files
docs: close remaining Tier 1 gaps, mirror to zh/ja
1 parent 205c1bd commit c08de0c

3 files changed

Lines changed: 603 additions & 0 deletions

File tree

‎README.md‎

Lines changed: 193 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -93,6 +93,55 @@ async def main():
9393
asyncio.run(main())
9494
```
9595

96+
### Function Calling
97+
98+
Pass OpenAI-style tool definitions via `tools`; the model requests a call through `message.tool_calls`, which your code executes and feeds back:
99+
100+
```python
101+
from dashscope import Generation
102+
103+
tools = [{
104+
"type": "function",
105+
"function": {
106+
"name": "get_current_weather",
107+
"description": "Get the current weather for a city.",
108+
"parameters": {
109+
"type": "object",
110+
"properties": {"location": {"type": "string", "description": "The city name."}},
111+
"required": ["location"],
112+
},
113+
},
114+
}]
115+
response = Generation.call(
116+
model="qwen-plus",
117+
messages=[{"role": "user", "content": "What's the weather in Hangzhou?"}],
118+
tools=tools,
119+
result_format="message",
120+
)
121+
tool_call = response.output.choices[0].message.tool_calls[0]
122+
print(tool_call.function.name, tool_call.function.arguments)
123+
```
124+
125+
### Thinking Mode
126+
127+
Hybrid thinking models can expose their reasoning process separately from the final answer via `enable_thinking` (requires `stream=True`); the reasoning appears in `message.reasoning_content`, the final answer in `message.content`:
128+
129+
```python
130+
from dashscope import Generation
131+
132+
responses = Generation.call(
133+
model="qwen-plus",
134+
messages=[{"role": "user", "content": "Which is bigger, 1.1 or 0.9?"}],
135+
result_format="message",
136+
enable_thinking=True,
137+
incremental_output=True,
138+
stream=True,
139+
)
140+
for response in responses:
141+
message = response.output.choices[0].message
142+
print(message.get("reasoning_content") or message.content, end="")
143+
```
144+
96145
### Error Handling
97146

98147
Missing required arguments (e.g. no `model`, no `messages`/`prompt`, no API
@@ -321,6 +370,45 @@ async def main():
321370
asyncio.run(main())
322371
```
323372

373+
Video is passed as a list of frame image URLs/paths (not a single video file):
374+
375+
```python
376+
from dashscope import MultiModalConversation
377+
378+
messages = [{
379+
"role": "user",
380+
"content": [
381+
{"video": ["frame1.jpg", "frame2.jpg", "frame3.jpg", "frame4.jpg"]},
382+
{"text": "Describe what happens in this video."},
383+
],
384+
}]
385+
response = MultiModalConversation.call(model="qwen-vl-max-latest", messages=messages)
386+
print(response.output.choices[0].message.content[0]["text"])
387+
```
388+
389+
`qwen-vl-ocr` models accept an `ocr_options` parameter for structured extraction (e.g. filling a JSON schema from a document image):
390+
391+
```python
392+
from dashscope import MultiModalConversation
393+
394+
messages = [{
395+
"role": "user",
396+
"content": [
397+
{"image": "https://example.com/invoice.jpg"},
398+
{"text": "Extract fields from this document into the given JSON schema: {result_schema}"},
399+
],
400+
}]
401+
response = MultiModalConversation.call(
402+
model="qwen-vl-ocr-latest",
403+
messages=messages,
404+
ocr_options={
405+
"task": "key_information_extraction",
406+
"task_config": {"result_schema": {"invoice_number": "", "total_amount": ""}},
407+
},
408+
)
409+
print(response.output.choices[0].message.content[0]["text"])
410+
```
411+
324412
### Using Local Files
325413

326414
Every field that accepts a URL (`image`, `audio`, `video` in messages; `images` on `ImageSynthesis`, etc.) also accepts a local file path — the SDK uploads it to OSS automatically, no manual header needed:
@@ -365,6 +453,26 @@ resp = MultiModalEmbedding.call(
365453
print(resp.output)
366454
```
367455

456+
Use the explicit item classes to combine text/image/audio into a single fused vector (each requires a `factor` weight, and `enable_fusion` on `qwen3-vl-embedding`):
457+
458+
```python
459+
from dashscope import MultiModalEmbedding
460+
from dashscope.embeddings.multimodal_embedding import (
461+
MultiModalEmbeddingItemText,
462+
MultiModalEmbeddingItemImage,
463+
)
464+
465+
resp = MultiModalEmbedding.call(
466+
model="qwen3-vl-embedding",
467+
input=[
468+
MultiModalEmbeddingItemText(text="a red sports car", factor=1.0),
469+
MultiModalEmbeddingItemImage(image="https://dashscope.oss-cn-beijing.aliyuncs.com/images/256_1.png", factor=1.0),
470+
],
471+
enable_fusion=True,
472+
)
473+
print(resp.output)
474+
```
475+
368476
### Batch (Offline) Text Embedding
369477

370478
For large volumes of text, submit a file (one text per line) for asynchronous batch embedding instead of calling `TextEmbedding.call` per item:
@@ -381,6 +489,21 @@ if resp.output.task_status == "SUCCEEDED":
381489
print(resp.output.url) # download the result file from here
382490
```
383491

492+
Submit without blocking, then poll separately:
493+
494+
```python
495+
from dashscope import BatchTextEmbedding
496+
497+
task = BatchTextEmbedding.async_call(
498+
model=BatchTextEmbedding.Models.text_embedding_async_v2,
499+
url="https://example.com/texts.txt",
500+
)
501+
print(task.output.task_id)
502+
503+
result = BatchTextEmbedding.wait(task)
504+
print(result.output.task_status)
505+
```
506+
384507
### Text ReRank
385508

386509
```python
@@ -455,6 +578,64 @@ if rsp.status_code == HTTPStatus.OK:
455578
print(result.url)
456579
```
457580

581+
`sync_call` (currently only for `wan2.2-t2i-flash`/`wan2.2-t2i-plus`) returns the result directly instead of polling an async task:
582+
583+
```python
584+
from http import HTTPStatus
585+
from dashscope import ImageSynthesis
586+
587+
rsp = ImageSynthesis.sync_call(
588+
model="wan2.2-t2i-flash",
589+
prompt="a flower shop with delicate windows and a wooden door",
590+
n=1,
591+
size="1024*1024",
592+
)
593+
if rsp.status_code == HTTPStatus.OK:
594+
print(rsp.output)
595+
```
596+
597+
Use `AioImageSynthesis.sync_call` for the `async`/`await` form:
598+
599+
```python
600+
import asyncio
601+
from dashscope import AioImageSynthesis
602+
603+
async def main():
604+
rsp = await AioImageSynthesis.sync_call(
605+
model="wan2.2-t2i-flash",
606+
prompt="a flower shop with delicate windows and a wooden door",
607+
n=1,
608+
size="1024*1024",
609+
)
610+
print(rsp.output)
611+
612+
asyncio.run(main())
613+
```
614+
615+
### Sketch-to-Image and Image Editing
616+
617+
`ImageSynthesis.call` also accepts a hand-drawn sketch or an existing image to edit, via dedicated models and parameters:
618+
619+
```python
620+
from dashscope import ImageSynthesis
621+
622+
# Sketch to image
623+
rsp = ImageSynthesis.call(
624+
model=ImageSynthesis.Models.wanx_sketch_to_image_v1,
625+
prompt="a cute cat, watercolor style",
626+
sketch_image_url="https://example.com/sketch.png",
627+
)
628+
629+
# Edit an existing image with a text instruction
630+
rsp = ImageSynthesis.call(
631+
model=ImageSynthesis.Models.wanx_2_1_imageedit,
632+
prompt="change the background to a beach",
633+
function="description_edit",
634+
base_image_url="https://example.com/photo.png",
635+
)
636+
print(rsp.output)
637+
```
638+
458639
### Video Generation
459640

460641
Video generation runs as an async task; `call` blocks until it completes, or use `async_call` + `wait`/`fetch` to poll manually.
@@ -473,6 +654,18 @@ if rsp.status_code == HTTPStatus.OK:
473654
print(rsp.output.video_url)
474655
```
475656

657+
Submit without blocking, then poll separately:
658+
659+
```python
660+
from dashscope import VideoSynthesis
661+
662+
task = VideoSynthesis.async_call(model="wan2.7-t2v", prompt="a kitten running under the moonlight")
663+
print(task.output.task_id)
664+
665+
rsp = VideoSynthesis.wait(task)
666+
print(rsp.output.video_url)
667+
```
668+
476669
### Speech Synthesis (TTS)
477670

478671
Qwen-TTS models are called through `MultiModalConversation`, passing `text`/`voice` instead of `messages`:

0 commit comments

Comments
 (0)