Product launch · XiaoHu explains

Qwen releases Qwen-Image-3.0: one 3,000-character prompt, nine infographics and a full newspaper page in a single pass

The prompt cap jumps from 1k to 4.5k tokens, and 10px text stays legible. Bailian already serves qwen-image-3.0-pro, free for a limited time.
60-second summary
  • Newspaper layouts, nine-panel teaching charts, short-drama storyboards — text-heavy images like these used to require generating one cell at a time and stitching them together. Tongyi's Qwen-Image-3.0 raises the single-prompt cap to 4.5k tokens (roughly 3,000 Chinese characters), putting nine unrelated subjects into one image in a single pass.
  • The smallest legible type size reaches 10px, holding up across a full newspaper front page, an entire LaTeX paper, and a densely packed eight-section reference chart.
  • On coverage: native rendering in 12 languages, 100+ artistic styles, and faithful mockups of interfaces like developer tools and food-delivery apps. It can also pull live information to build an image, such as generating a same-day weather and sightseeing guide.
  • Generation and editing live in the same model: add red-pen annotations to an existing book page, restore a damaged classical painting, overlay scientific labels on a real photo. Give it a character sheet and a location, and it lays out a 12-panel storyboard complete with shot size, duration, and camera-move notes.
  • Alibaba Cloud's Bailian already lists qwen-image-3.0-pro, free for a limited time; the previous-generation pro was 0.5 RMB per image. Open weights still stop at Qwen-Image-2512 from 2025-12-30 — neither 3.0 nor February's 2.0 has been released.
Capability descriptions and sample images are from Qwen's official launch page. Pricing and open-source details were independently verified by this site.
What changed this generation

Each generation of the model has had one keyword — this one's is "substance"

Alibaba's Tongyi team released Qwen-Image-3.0 on July 21 — the third-generation foundation model in the Qwen-Image line.

Qwen-Image-3.0 banner: an undulating sea of text
The official banner is itself a demonstration: the whole image is built from undulating text, mixing the word "image" in Thai, Korean, Arabic, Russian, Chinese, and Vietnamese, with a corner caption reading "text becomes image, every character a seal." Source: Qwen launch page

Every generation of this line has summed itself up in a phrase. The first, released last August, was "accurate" — the whole point was getting text right. The second, released this February, was "accurate, versatile, consistent, beautiful, real" — it folded generation and editing into one model and supported 1k-token prompts at 2K resolution. This generation compresses down to one word: substance.

"Substance" breaks into three parts, and together they make up the entire capability list for this update.

Rich content
How complex a scene it can draw
  • 4.5k-token prompt cap
  • Lateral: multiple subjects side by side
  • Depth: interfaces nested layer within layer
  • Newspapers, storyboards, exam sheets, direct
Realistic detail
How realistic it can render
  • Precise rendering down to 10px text
  • Pore- and strand-level rendering
  • Material and paper texture
  • Small text holds up in editing too
Deep knowledge
How broad its coverage is
  • Native rendering in 12 languages
  • 100+ art styles
  • Simulates web, game, livestream UIs
  • Pulls current world knowledge online

Qwen frames this update's direction as a shift from "looking good" to "being useful," aimed at working scenarios like newspaper layouts, short-drama storyboards, and complex interface mockups. Here's each of the three broken down.

Capability 1 · Rich content

Prompt cap raised to 4.5k tokens — a whole complex layout, spelled out in one go

What this capability is

A single prompt can now hold up to 4.5k tokens — roughly 3,000 Chinese characters, up from 1k in the last generation. How long a prompt can run directly determines how many things can be arranged in one image.

Qwen-Image-2.0
1k
Bailian API's current cap
1300
Qwen-Image-3.0
4.5k

The middle bar is the number spelled out in Alibaba Cloud Bailian's API docs: the qwen-image-2.0 line caps prompts at 1300 tokens, other models at 800, with anything past that simply truncated.

When prompts were short, a complex layout had to be drawn in pieces and stitched together. Nine subjects, each with its own title, body text, formulas, and character actions — a thousand characters couldn't even cover half of that. Raising the cap to around three thousand is what turns this kind of job from piecemeal collage into a single pass.

Qwen's proof point here is a nine-panel grid, whose description ran to 3.7k tokens.

A nine-panel complex infographic generated by Qwen-Image-3.0 in one pass
All nine panels were generated from a single 3.7k-token prompt, with no stitching in between. The subjects: a tunnel safety-distance comic, spatial geometry, a stylistic analysis of the classical essay "Chu Shi Biao," projectile motion, parasitology, a medical diagram, group theory's Sylow theorems, a bank internal-controls infographic, and a cell/DNA structure comparison. Source: Qwen launch page

Pull out any single panel and it stands on its own. The upper-middle one is a complete spatial-geometry teaching page — problem, premises, theorem, and conclusion, all present.

One panel from the grid: a teaching diagram on line-plane perpendicularity in spatial geometry
The information density of a single panel matches that of a full teaching slide. Source: Qwen launch page

Worth noting: the derivations in these panels aren't just for show. The group-theory panel works out how many Sylow 3-subgroups a group of order 72 has — factoring 72 into 8 times 9, applying the constraint that the count must be congruent to 1 mod 3 and divide 8, and landing on an answer of 1 or 4. The whole chain holds up. The physics panel derives the angular velocity formula from an object's initial velocity at the point of release, and the landing-time and displacement substitutions in between check out too.

Capability 2 · Rich content

Spatial control, two directions: side-by-side layout, and layer-by-layer nesting

What this capability is

Lateral spread tests semantic juxtaposition: multiple unrelated concepts appear on the same canvas at once, each holding its own territory without content bleeding across. Depth tests logical nesting: one image draws several layers of interface from outside in, and each layer keeps its own distinct look.

Lateral: subjects side by side Depth: interfaces nested Spread across one plane Each holds its spot, no cross-talk VSCode Qwen chat WeChat chat Coffee poster Layer within layer Each layer keeps its own interface look
Diagram by this site

The nine-panel grid in the last section proved lateral spread. The depth example is a single prompt producing four nested layers: a VSCode editor on the outside, a Qwen chat window open inside it, a WeChat screenshot inside that chat, and a pour-over coffee process poster shared inside that WeChat conversation.

A picture within a picture within a picture: a Qwen chat window open in VSCode, containing a WeChat chat, containing a pour-over coffee poster
Four nested layers, and each layer's interface elements hold up on their own: VSCode has its file tree and status bar, WeChat has its contact name and timestamps, the poster has its four numbered steps. Source: Qwen launch page

These two directions serve different jobs. Lateral spread handles multi-column layouts — newspapers, exam sheets, teaching wall charts, product comparison pages. Depth handles product-demo decks and interface sequences, showing how a feature jumps from one piece of software to another. What used to take three screenshots stitched together now takes one prompt.

Capability 3 · Realistic detail

Smallest legible text hits 10px — enough to lay out a full newspaper page or a whole paper

What this capability is

10px is the floor for body-text-sized copy. At that scale, characters easily blur into a gray smear and strokes tend to merge. Holding up at this size is what makes dense text layouts possible at all.

Three kinds of work depend most on this capability.

1. Formula-dense academic layouts

The hard part of LaTeX layout is that superscripts, subscripts, braces, fraction bars, Greek letters, and multi-line alignment all show up at once — miss any one position by half a grid step and it shows.

A full page of a generated academic paper on algebraic geometry
A full two-column page where formula sub/superscript placement, multi-line alignment, and special symbols all follow LaTeX conventions. The header reads "Fake PDF" — this page was generated to spec for a fictional paper. Source: Qwen launch page

2. Newspaper layout

A newspaper front page has a lot that has to work at once: the calligraphic strokes of the masthead, the dividing lines in the date bar, the headline hierarchy across several stories, a chart card in the middle, a reading-guide column on the right, and the baseline alignment of body text in every column. Miss any one and it stops looking like a newspaper.

A generated newspaper front page
The whole page's content is fictional, per the prompt. Beyond the dense small text, the paper's creases, shadows, and print texture were generated too. Source: Qwen launch page

3. High-density science infographics

This kind of image crams body text, leader-line labels, a legend, and a scale bar all together, with text pushed down to the smallest size — and every label still has to point to the right spot.

A whale shark field guide, an infographic with eight densely labeled sections
Eight numbered sections, plus a world map of migration routes and a size-comparison scale, all with text at the smallest size tier. Source: Qwen launch page
Capability 4 · Realistic detail

Pores, hair strands, and material textures all hold up to a close look

What this capability is

Micro-level rendering: portraits get pore, fine-hair, and skin-highlight gradation; non-portrait objects get the structure and reflectivity of the material itself. This is what decides whether a generated image can be used as a finished asset.

Portraits: the harder the light, the harder to hide flaws

Hard light pushes every bump and pore on the skin into plain view, and shadow edges can't be allowed to blur. The image below also has a floral-branch shadow cast across the face — the detail right at the light-dark boundary is where the real test is.

A hard-light portrait with a floral-branch shadow across the face
A petal-shaped shadow falls across the face, while a row of ear cartilage piercings, the highlight on the lip gloss, and the print on a floral shirt all stay crisp at the same time. Source: Qwen launch page

The natural-light shot tests something else: hair strands. Wind-blown wisps have to separate strand by strand against backlight, with light visibly passing through them.

An outdoor natural-light portrait with wind-blown hair
The Swiss cheese plant and ferns in the background stay softly blurred by depth of field, while the loose strands of hair in the foreground separate individually. Source: Qwen launch page

Materials: four completely different surfaces

The non-portrait examples show the range better. For embroidery to work, three things have to land at once: feather layering has to come through in stitch direction, the thread needs the actual sheen of silk, and the linen weave underneath has to show through.

An embroidered warbler on an embroidery hoop
Thread direction, color gradation, and the base fabric's weave are three distinct textures, and each one holds up on its own within the same image. Source: Qwen launch page

Animal fur runs on a different logic: black fur has to keep its layering in the shadows, and white fur can't blow out under a flash.

Two dogs under flash at night
A nighttime flash shot where the black coat's layering and the white coat's downy texture both hold up at once, with streetlight bokeh in the background. Source: Qwen launch page

Rock art tests the relationship between pigment and stone: mineral color has to seep into the pores of sandstone, and it needs the mottling of weathering and flaking.

Ochre-red rock art on sandstone
Ochre-red pigment painted on sandstone, with the rock's cracks, grain, and pigment flaking all rendered together. Source: Qwen launch page

This last one is an extreme macro shot: the thick stacked strokes of oil paint, every individual bristle on the brush, and the scratches and reflections on its metal ferrule.

A macro close-up of an oil brush dipped in blue paint
The paint's thick-thin buildup, the separated strands of hog-bristle brush hair, and the canvas weave peeking through at the edge — three layers of texture, all near the same focal plane. Source: Qwen launch page
Capability 5 · Generation and editing, combined

One model both draws from scratch and edits the image you already have

What this capability is

Since the last generation, generation and editing have lived in the same model — no switching models or endpoints. All the detail capabilities above carry over into editing mode too: adding small text, patching gaps, layering on annotations. The three pairs below are all before-and-after shots of the same image.

Adding handwritten annotations to an existing book page

The hard part is how real the handwriting looks. Red-pen strokes need pressure variation, circled words need the irregular hand-drawn oval of an actual pen stroke, arrow tails need a slight hook back at the end, and it all has to sit on top of the printed text rather than float off to the side.

The original book page, before annotation
Before: a book page about Xu Xiake, with a travel-route map on the lower half
The same page, with red-pen annotations added
After: more than a dozen red-pen annotations, underlines, circles, and arrows layered on top

Restoring missing sections of a damaged painting

The requirement here is that the brushwork match the original after restoration, with no sense of a patch job.

The damaged classical painting Wanli Yingyang (Ten Thousand Miles, Eagle Soaring)
Before: white gaps of loss across the whole surface, horizontal creases, brown stains
Wanli Yingyang after restoration
After: even silk tone throughout, with the inscription seal and composition's brushwork held in their original places
The inscription in the upper left of the painting reads "Wanli Yingyang." Qwen's launch page refers to this piece as "Eagle Strike" in its body text; this article follows the inscription on the painting itself.

Overlaying a professional annotation layer on a real photo

The input is an ordinary macro photo, and the output has to be usable directly as an academic figure: the subject can't distort, and labels have to point to the right spots.

The original macro photo of damselflies
Before: two damselflies on a blade of grass
A damselfly infographic with an academic annotation layer added
After: taxonomic information, morphological labels, zoomed-in detail, and a scale bar added
Capability 6 · Producing complete sets

Give it a character and a scene, and it lays out a full storyboard

What this capability is

Short-drama storyboarding is one of the productivity use cases Qwen names explicitly. Pulling it off requires the model to hold three things at once: the same character looking consistent across a dozen-plus panels, a coherent scene, and clear shot type and camera movement in every panel.

Qwen's example set here is four images. The first three are source material: a character reference, a scene, and one key shot.

Character reference for the storyboard: a studio portrait
Character reference: black-frame glasses, a beige suit, a wine-red turtleneck. Source: Qwen launch page
Scene reference for the storyboard: an old Japanese mountain town
Scene: an old mountain town at dusk, with black tile roofs, timber construction, stone-paved streets, and layered distant mountains. Source: Qwen launch page
Key shot for the storyboard: a silhouette with arms spread in the rain
Key shot: black and white, a backlit silhouette with arms spread wide in a curtain of rain. Source: Qwen launch page

The fourth image strings the first three together into a complete storyboard.

A 12-panel storyboard with shot type, duration, and camera-movement labels
Twelve panels, each labeled with a shot number, duration (ranging from 3s to 1s), shot type, and action description — wide shot entering the town, medium shot walking, close-up reading a note, over-the-shoulder shot spotting a dark figure ahead, handheld turn, insert shot of a hand pushing a door open — with the camera-movement notation (hold, move, push in, follow) marked in the bottom-right corner. Source: Qwen launch page

This image draws on every capability covered so far: the same character carries through all 12 panels, the scene and weather advance with the narrative (entering the town, rain starting, the door pushed open), each panel's black label bar packs in dense small text, and the whole thing is a lateral-spread layout in its own right. Panel 10 is the rain silhouette shown above.

The old workflow for this went: generate the character reference, then the scene, then each panel one at a time, then pull everything into design software to lay out and annotate. Now the whole set comes out in one pass.

Capability 7 · Deep knowledge

12 languages, 100-plus styles, mainstream software interfaces — and it pulls in today's information online

What this capability is

This dimension of coverage breaks into four parts: native multilingual rendering, an art-style library, software and web interface simulation, and world knowledge — including pulling fresh information online and recognizing specific IP characters.

Multiple languages, each paired with its own local layout conventions

This dimension covers native rendering in 12 languages and more than 100 art styles. The three examples here each take a completely different route. The Japanese one is a manga title page — black-and-white screentone, vertical dialogue boxes, and sound effects all have to line up together as one visual language.

A Japanese manga title page in black-and-white screentone style
Vertical Japanese text, manga paneling, and screentone style all hold up at once. Source: Qwen launch page

The Korean example is an e-commerce product page: spec icons on the left, a model in the middle, a parameter table on the right, four lifestyle shots along the bottom — a complete product-page layout.

A Korean e-commerce product page for a dress
A dress product page: six line-drawing icons with labels, five color-swatch thumbnails, a spec table on the right, and four styled-outfit shots along the bottom. Source: Qwen launch page

The Spanish example switches to handwriting: whiteboard-marker lettering in three colors, plus a neural-network diagram and math formulas.

A Spanish-language whiteboard explaining multilayer perceptrons
A handwritten whiteboard explanation of multilayer perceptrons, in five numbered sections: basic structure, forward propagation, loss function, backpropagation, and characteristics — including a neuron-connection diagram and formulas. Source: Qwen launch page

Interface simulation

Web, game, and livestream interfaces each follow their own set of control conventions. This capability feeds directly into product prototypes and demo mockups.

A simulated PyCharm interface
A developer-tool interface: menu bar, project tree, syntax-highlighting colors, and the log format in the console at the bottom. Source: Qwen launch page
A simulated food-delivery app interface
A food-delivery app's listing page: phone status bar, search box, filter tags, four restaurant cards (each with a product photo, rating, monthly order count, delivery time, and discount coupon), a coupon banner and five navigation tabs at the bottom. The information density here is even higher than the developer-tool example. Source: Qwen launch page

Pulling today's information online

The model's built-in world knowledge has a cutoff date; going online is what fills in anything newer. Asked for today's Hangzhou weather forecast, it comes back with a complete weather-and-sightseeing guide.

A Hangzhou weather and sightseeing guide
Today's temperature range, weather, wind, and humidity, paired with a temperature curve — the lower half of the page has four attraction cards and a one-day route. Source: Qwen launch page

Recognizing specific IP characters and reworking them

This isn't just about recognizing a face — the model has to carry over the style that character is supposed to be drawn in, too.

Qi Baishi and Van Gogh in a livestream
Qi Baishi rendered with photorealistic texture, Van Gogh kept in thick oil-paint brushwork — two completely different rendering styles placed into the same livestream lighting setup. Source: Qwen launch page
Getting started

Where to use it now, what it costs, and whether it's open source

Qwen's launch page never mentions any of these three things. This site checked them directly.

In Alibaba Cloud Bailian's pricing table for Qwen text-to-image models, qwen-image-3.0-pro already sits in the top row, marked "free for a limited time" in both the mainland China and international regions. For comparison, the previous-gen qwen-image-2.0-pro cost ¥0.5 per image.

ModelMainland China price
qwen-image-3.0-proFree for a limited time
qwen-image-2.0-pro¥0.5/image (100 free images)
qwen-image-2.0¥0.2/image (100 free images)
qwen-image-max¥0.5/image
qwen-image / qwen-image-plus¥0.2–0.25/image

Bailian's API documentation hadn't been updated as of that day — the newest entry in the model list was still the June 22 version of qwen-image-2.0-pro. The pricing table is already live; the docs will need to catch up. On the web, it's available through Qwen Chat by selecting "image generation."

Open source is a separate story. In the Qwen organization's image-model repositories on Hugging Face, the last weight release was Qwen-Image-2512, dated December 30, 2025. February's Qwen-Image-2.0 never got a weights repository — the corresponding entry in the GitHub changelog only links to the blog post and Qwen Chat. 3.0 currently has neither.

Qwen-Image open-weight release history
  • 2025-08-02 Qwen-Image (first generation, 20B MMDiT, Apache 2.0 license)
  • 2025-08-17 Qwen-Image-Edit, 2025-09-22 Edit-2509, 2025-12-17 Edit-2511
  • 2025-12-17 Qwen-Image-Layered
  • 2025-12-30 Qwen-Image-2512 (the last update to image-generation weights)
  • 2.0 in February 2026 and 3.0 in July 2026: no weights released
🧰 Quick-start card · Qwen-Image-3.0
Accesschat.qwen.ai, select "image generation"
PriceFree on the Qwen Chat web app; the Bailian API's qwen-image-3.0-pro is currently marked free for a limited time, while the previous-gen qwen-image-2.0-pro costs ¥0.5 per image
Barrier to entryThe web app just needs a typed prompt; using the API requires an Alibaba Cloud Bailian account and API key, and the official API docs haven't listed 3.0 in the model table yet

When using it, write longer prompts. This generation raised the cap to 4.5k tokens precisely so you can spell out every word that needs to appear and where every block needs to sit.

Source
Qwen-Image-3.0: Rich content, realistic detail, deep knowledgeQwenTeam·qwen.ai·2026-07-21
Editor's note
All 30 sample images from the launch page have been included in this article. Pricing, open-weight release history, and the API's prompt-length cap come from Alibaba Cloud Bailian's pricing table and API docs, the Hugging Face repository listing, and the GitHub changelog — none of which the launch page itself mentions. The lateral-spread and depth comparison diagram was drawn by this site.