Mage

A Lightweight, Research-Friendly Multimodal Model Family

Microsoft Mage Team
Mage family

Mage is a family of lightweight, research-friendly multimodal models built at a fixed 4B-parameter budget, sharing a codec-aligned efficiency philosophy — spend representation capacity where the signal is — across both visual understanding and generation. Both models are compact enough to train, fine-tune, and deploy on modest hardware, yet remain competitive with much larger open systems.

Vision–Language · Understanding
Mage-VL
A Codec-Native Proactive Streaming Multimodal Foundation Model
A codec-native, from-scratch VLM for image & video understanding — reads video the way a codec does (anchor/predicted frames, 16×16 patches), with bio-inspired proactive streaming.
Coming soon
Generation · Editing
Mage-Flow
An Efficient Native-Resolution Foundation Model for Image Generation and Editing
A compact 4B generative stack (Mage-VAE + Native-Resolution MMDiT) for text-to-image generation and instruction-based editing at native resolution, with Base / RL / 4-step Turbo variants.
Enter project page