Instruction Following
Visual instruction tuning aligns user intent, multimodal context, and model responses in one generative interface—from image dialogue to video, documents, and interleaved inputs.

A unified behavior-shaping view of how pretrained MLLMs become reliable, grounded, and task-oriented multimodal systems.
* Corresponding author · is.pengpengzeng@gmail.com





Multimodal pretraining builds broad perception and alignment. Post-training determines how those capabilities behave when the evidence is ambiguous, the task is complex, or the interaction has real consequences.

MLLMs Post-training has rapidly become a central mechanism for endowing pre-trained MLLMs with the ability to exhibit more aligned and reliable behaviors, marking significant progress in multimodal intelligence.
MLLMs post-training connects multimodal learning with digital AI and physical AI, representing a core step in the progression towards Artificial General Intelligence (AGI).



Visual instruction tuning aligns user intent, multimodal context, and model responses in one generative interface—from image dialogue to video, documents, and interleaved inputs.

Benchmarks define what is learned and what counts as progress—from instruction compliance and faithfulness to multimodal reasoning and domain reliability.

MME · MMBench · MM-Vet · SEED-Bench · MM-IFEval · MIA-Bench
POPE · MMHal-Bench · HallusionBench · SafeBench · Multimodal RewardBench · VL-RewardBench
MMMU · MMMU-Pro · MathVista · MathVerse · We-Math · LogicVista
ScreenSpot · OSWorld · DocVQA · OCRBench · SLAKE · DriveLM
Key directions for advancing MLLM post-training beyond current paradigms fall into three complementary themes: grounded behavior shaping, reliability-aware evaluation, and scaling for generalization.
Post-training should learn behavior directly from multimodal evidence instead of organizing every modality around language-style supervision.
Most pipelines eventually align visual, video, or audio inputs to text responses. Native signals should preserve spatial layout, temporal continuity, acoustic cues, and cross-modal correspondence, allowing models to learn from modality-specific structure rather than treating non-text inputs as auxiliary context.
Future agents must move beyond understanding digital images, videos, documents, and screens. Post-training should connect perception with action-oriented reasoning in continuous environments where actions change the world and feedback can be delayed or uncertain.
Evaluation must reveal whether improved scores correspond to grounded, calibrated, safe, and stable behavior.
High benchmark accuracy does not guarantee reliability. New metrics should diagnose visual grounding, calibration, consistency, safety, and robustness under distribution shift—exposing hallucination, overconfidence, shortcut reasoning, and unstable responses rather than rewarding only final-answer correctness.
Real applications contain ambiguous evidence, long-horizon context, changing environments, and interaction constraints. Benchmarks should move beyond static images and closed-form QA so models must maintain state, handle uncertainty, and adapt across continuous multimodal inputs.
Scaling should produce reusable behaviors that transfer across tasks, domains, modalities, and continuously changing worlds.
Progress requires more than additional data or compute. We need to understand how diverse supervision signals, model capacity, and optimization jointly drive generalization, enabling learned behaviors to be composed, reused, and adapted beyond the datasets, task formats, and reward functions seen during training.
MLLMs need persistent memory, temporal event abstraction, selective retrieval, online feedback integration, and continual adaptation. They should absorb new experiences while preserving previous capabilities, maintaining behavioral consistency, and avoiding catastrophic forgetting or uncontrolled model drift.
The companion repository continuously curates papers, code, datasets, and project pages across the MMPoT landscape.
Browse the reading list@article{zhang2026survey,
title={A Survey on Post-Training of Multimodal Large Language Models},
author={Haonan Zhang and Pengpeng Zeng and Libin Cao and Wenrui Lai and
Jinlong Li and Duo Peng and Yi Bin and Xuanhan Wang and Ji Zhang and
Jingkuan Song and Nicu Sebe and Yuchuan Wu and Yongbin Li and
Heng Tao Shen and Jieping Ye},
year={2026},
publisher={Preprints}
}