
Over the past week, Alibaba unveiled three new multimodal models capable of processing both image and audio. Designed to handle rich information, complex reasoning, agent tool integration, and lifelike dialogue, these models can be used to enhance real-world applications in short films, comics, interactive education, and immersive entertainment.
Alibaba has released its latest image generation model, Qwen-Image-3.0. The new model supports ultra-long inputs of up to 4.5k tokens, enabling the one-step generation of knowledge diagrams and complex UI interfaces that cover multiple elements such as mathematical symbols, geometric shapes, and logical deduction steps. Supporting native rendering in 12 languages and over 20 font styles, the model can significantly reduce the production costs for commercial materials development, such as multilingual product posters and storyboards for films or comics.
Compared to its predecessor, Qwen-Image-3.0 significantly enhances the model’s capacity to handle complex information, achieving a 4.5-fold increase in text input length. This means that in scenarios requiring ultra-long prompts — such as storyboarding for short films, creating complex infographics, or designing product description pages — users can input detailed instructions, which allow the model to generate more sophisticated content, reducing the costs associated with repeated generations and manual adjustments.

When a user inputs “Generate a pour-over coffee poster posted in a chat interface via the Qwen App, within the VSCode programming interface,” the model generates the visual accordingly.
Equipped with more comprehensive knowledge, Qwen-Image-3.0 can understand complex instructions and the logical relationships within images. Based on a single prompt, the new model can generate different visual expressions such as webpages, software interfaces, chat windows, and posters, all within a single image. For example, when a user inputs “Generate a pour-over coffee poster posted in a chat interface via the Qwen App, within the VSCode programming interface,” the model accurately understands the multi-layered UI nesting relationships, and generates the visual accordingly.
Alibaba has unveiled two new voice models aimed at advancing digital communication. The lineup includes Qwen-Audio-3.0-Realtime, a real-time conversational model capable of complex reasoning, agent tool integration, and lifelike dialogue, alongside Qwen-Audio-3.0-TTS, a text-to-speech model engineered for precise control, broader language coverage, and exceptional audio clarity even in noisy environments.
Optimized for latency-sensitive applications, Qwen-Audio-3.0-Realtime delivers millisecond-level responses for natural conversation. The model also supports tool calling to connect with external knowledge bases, MCP, and OpenAPIs, while maintaining memory for consistent, context-aware dialogue. To ensure lifelike interactions, it dynamically adjusts its tone, pitch, and emotion, while supporting duplex conversation events such as background noise handling, backchanneling, and interruptions.

The model is ideal for a wide range of real-world applications, including intelligent customer service, interactive education and training, immersive entertainment and digital companionship.
Alibaba’s latest text-to-speech model Qwen-Audio-3.0-TTS, features free-style natural language control and supports embedded emotional tags such as gasp, giggles, and angry to deliver lifelike micro-expressions. Optimized for global applications, the model supports 16 languages, including English, Chinese, Japanese, Korean, and German, alongside 20 Chinese dialects.
Delivering studio-grade audio that meets the professional requirements of recording and film dubbing, the model is ideal for premium video voiceovers and audiobooks, supporting up to three minutes of continuous synthesis per session. It is available in two versions: a Flash version optimized for real-time interaction and a Plus version designed for high-quality generation.
Among the two versions offered, Qwen-Audio-3.0-TTS-Plus features upgraded 48kHz studio-grade audio output. It has also topped the Speech Arena Leaderboard on Artificial Analysis.

This article was originally published on Alizila written by Crystal Liu and Shao Xiaoyi.
Alibaba Cloud Research Reveals Key Insights on AI Deployment Across Asian Companies
1,495 posts | 509 followers
FollowAlibaba Cloud Community - September 27, 2025
Alibaba Cloud Community - June 8, 2026
Alibaba Cloud Community - February 27, 2025
Alibaba Cloud Community - December 16, 2025
Alibaba Cloud Community - December 31, 2025
Alibaba Cloud Community - September 19, 2024
1,495 posts | 509 followers
Follow
Alibaba Cloud Model Studio
A one-stop generative AI platform to build intelligent applications that understand your business, based on Qwen model series such as Qwen-Max and other popular models
Learn More
Qwen
Full-range, open-source, multimodal, and multi-functional
Learn More
Token Plan
Build more, spend less. One plan, every modality.
Learn More
Alibaba Cloud for Generative AI
Accelerate innovation with generative AI to create new business success
Learn MoreMore Posts by Alibaba Cloud Community