Pith. sign in

REVIEW 2 major objections 1 minor 87 references

Semantic Generative Tuning uses image segmentation as a generative proxy to align understanding and generation in unified multimodal models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 18:28 UTC pith:U6HXRAVQ

load-bearing objection SGT uses segmentation as a post-training generative proxy to align understanding and generation in UMMs, with reported benchmark gains and mechanistic support, but details are thin in the abstract. the 2 major comments →

arxiv 2605.18714 v2 pith:U6HXRAVQ submitted 2026-05-18 cs.CV cs.AI

Semantic Generative Tuning for Unified Multimodal Models

classification cs.CV cs.AI
keywords unified multimodal modelssemantic generative tuningimage segmentationgenerative proxiesrepresentation alignmentmultimodal comprehensiongenerative fidelity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Unified multimodal models seek to combine visual understanding and generation in one architecture, yet separate training creates misaligned representation spaces that limit mutual improvement. The paper systematically tests generative post-training and identifies high-level semantic tasks, particularly segmentation, as effective proxies because they supply structural information rather than low-level textures. Semantic Generative Tuning applies segmentation during post-training to bridge the gap. A sympathetic reader cares because the approach offers a concrete way to make the two capabilities reinforce each other instead of remaining isolated. Mechanistic checks confirm gains in feature separability and attention patterns, with consistent benchmark improvements in both comprehension and generation.

Core claim

The paper states that hierarchical visual tasks formulated as generative proxies, with image segmentation as the optimal choice, bridge the isolation between understanding and generation in unified multimodal models. Semantic Generative Tuning then applies this proxy to align representation spaces, producing improved feature linear separability, optimized visual-textual attention allocation, and measurable gains in both multimodal comprehension and generative fidelity across benchmarks.

What carries the argument

Semantic Generative Tuning (SGT), a post-training method that treats image segmentation as a generative proxy to align visual and textual representation spaces.

Load-bearing premise

High-level semantic tasks such as image segmentation supply structural semantics that optimally bridge understanding and generation without introducing distracting low-level texture signals.

What would settle it

If a low-level task such as texture or edge prediction produces larger gains than segmentation on the same unified multimodal model benchmarks and mechanistic metrics, the claim that segmentation is the optimal proxy would not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Multimodal comprehension improves across standard benchmarks
  • Generative layout fidelity increases without separate dense pixel objectives
  • Feature representations become more linearly separable
  • Visual-textual attention allocation becomes more optimized

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The proxy approach could be tested on video or 3D tasks to check whether other high-level semantic signals produce similar alignment effects.
  • Post-training with segmentation might allow existing unified models to close the performance gap between understanding and generation without full retraining.
  • The method implies that choosing the right generative proxy during alignment could reduce reliance on paired text-image data for joint optimization.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper claims to present the first systematic investigation of generative post-training for unified multimodal models (UMMs). It formulates hierarchical visual tasks as generative proxies to address misalignment between understanding (sparse text) and generation (dense pixels), empirically identifying high-level semantic tasks—particularly image segmentation—as optimal proxies that provide structural semantics without low-level texture distractions. Building on this, it introduces Semantic Generative Tuning (SGT) and reports that mechanistic analyses show improved feature linear separability and optimized visual-textual attention allocation, with extensive evaluations demonstrating consistent gains in multimodal comprehension and generative fidelity across benchmarks; code is released.

Significance. If the empirical results and mechanistic findings hold under scrutiny, the work could establish a practical post-training paradigm for aligning understanding and generation in UMMs via semantic proxies, with the segmentation insight and attention/separability analyses offering reusable mechanistic guidance. Code availability strengthens reproducibility and potential adoption.

major comments (2)
  1. [Abstract] Abstract: the central claim that 'high-level semantic tasks, particularly image segmentation, serve as optimal proxies' rests on an unspecified 'empirical investigation' into hierarchical tasks; without details on the task hierarchy tested, comparison metrics, or controls for low-level vs. high-level effects, it is impossible to verify whether segmentation is demonstrably optimal or if the reported gains are driven by post-hoc selection.
  2. [Abstract] Abstract: the mechanistic findings ('improved feature linear separability and optimized visual-textual attention allocation pattern') and the claim of 'consistent improvements... across mainstream benchmarks' are presented without any quantitative tables, ablation results, baseline comparisons, or analysis methods; these are load-bearing for the soundness of both the empirical and mechanistic contributions.
minor comments (1)
  1. [Abstract] The provided link ends with a trailing period ('https://song2yu.github.io/SGT/.') which appears to be a minor formatting artifact.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We appreciate the referee's feedback highlighting areas where the abstract could be more self-contained. We agree to revise the abstract to provide more context on the empirical investigation and to reference the quantitative support for the mechanistic and benchmark claims.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central claim that 'high-level semantic tasks, particularly image segmentation, serve as optimal proxies' rests on an unspecified 'empirical investigation' into hierarchical tasks; without details on the task hierarchy tested, comparison metrics, or controls for low-level vs. high-level effects, it is impossible to verify whether segmentation is demonstrably optimal or if the reported gains are driven by post-hoc selection.

    Authors: The details of the empirical investigation into hierarchical tasks are presented in the main text of the manuscript. To address this comment, we will revise the abstract to briefly outline the task hierarchy tested and the metrics used for comparison, ensuring the optimality claim is better supported within the abstract itself. revision: yes

  2. Referee: [Abstract] Abstract: the mechanistic findings ('improved feature linear separability and optimized visual-textual attention allocation pattern') and the claim of 'consistent improvements... across mainstream benchmarks' are presented without any quantitative tables, ablation results, baseline comparisons, or analysis methods; these are load-bearing for the soundness of both the empirical and mechanistic contributions.

    Authors: While the abstract summarizes the findings, the supporting quantitative tables, ablation results, baseline comparisons, and analysis methods are included in the main manuscript. We will revise the abstract to include key quantitative highlights and references to the analysis methods to make these claims more verifiable from the abstract. revision: yes

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper's core contribution is an empirical study of hierarchical visual tasks as generative proxies for unified multimodal models, with segmentation identified as optimal through experiments, followed by the introduction of SGT and reported benchmark gains plus attention analyses. No derivation chain, equations, or first-principles claims are present that reduce by construction to fitted parameters, self-definitions, or self-citation load-bearing premises; the claims rest on external benchmark results and mechanistic observations rather than internal tautologies.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

Abstract provides no explicit hyperparameters or invented entities; the core assumption that segmentation supplies structural semantics superior to low-level tasks is presented as an empirical discovery rather than an axiom.

axioms (1)
  • domain assumption High-level semantic tasks such as image segmentation provide structural semantics that enhance both perception and generative layout fidelity without the distraction of texture details.
    Invoked as the key empirical finding that justifies choosing segmentation over other proxies.

pith-pipeline@v0.9.1-grok · 5730 in / 1120 out tokens · 21546 ms · 2026-06-30T18:28:49.953172+00:00 · methodology

0 comments
read the original abstract

Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and generation through dense pixel objectives. Such a decoupled strategy yields misaligned representation spaces, isolating visual understanding from generation and hindering their mutual reinforcement. This work presents the first systematic investigation into generative post-training, where we formulate hierarchical visual tasks as generative proxies to bridge the isolation in UMMs. Our empirical investigation reveals that high-level semantic tasks, particularly image segmentation, serve as optimal proxies. Unlike low-level tasks that distract models with texture details, segmentation provides structural semantics that significantly enhance both vision-centric perception and generative layout fidelity. Building upon these insights, we introduce Semantic Generative Tuning (SGT), a novel paradigm that leverages segmentation as a generative proxy to align and synergize multimodal capabilities. Mechanistic analyses further demonstrate that SGT fundamentally improves feature linear separability and optimizes visual-textual attention allocation pattern. Extensive evaluations show that SGT consistently improves both multimodal comprehension and generative fidelity across mainstream benchmarks. Our code is available on the https://song2yu.github.io/SGT/.

Figures

Figures reproduced from arXiv: 2605.18714 by Songsong Yu, Yanwei Li, Ying Shan, Yuxin Chen.

Figure 1
Figure 1. Figure 1: Comparison of alignment strategies for UMMs. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the generative tuning paradigm. An RGB image and a concise textual instruction are processed by respective vision and text encoders to extract independent embeddings. UMMs then integrate these embeddings and map the rep￾resentations to the designated task. Because empirical evaluations demonstrate that visual generation targets at an advanced semantic level yield the most significant per￾forman… view at source ↗
Figure 3
Figure 3. Figure 3: Empirical evaluation of the hierarchical task ladder across diverse understand￾ing and generation dimensions. (a) High-level proxy tasks yield greater performance gains than low-level tasks in multimodal understanding. (b) Various generative objec￾tives consistently improve performance in the position dimension, yielding comparable overall gains. (From left to right): Position, Colors, Color Attributes, Co… view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison on compositional text-to-image generation. 4.3 More Explorations Optimal data recipe. While our analysis in Sec. 3.3 indicates that SGT independently enhances both understanding and generation, we posit that a comprehensive post-training regime must synergize SGT objectives with SFT data to maximize performance. Therefore, we conduct an ablation study to determine the optimal data sa… view at source ↗
Figure 5
Figure 5. Figure 5: Ablation studies on segmentation data integration. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Training dynamics with different SFT:Seg ratios. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Feature space analysis on fine-grained classes. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Analysis of attention patterns. (a) Layer-wise changes in attention to vi￾sual features for three proxy tasks relative to the BAGEL baseline, demonstrating a consistent increase in visual focus in deeper layers. (b) Attention distribution over text tokens. The segmentation objective effectively enhancing the focus on critical tokens (Object, Color, Relation). 4.4 Mechanistic Insights: Why Semantic Proxies … view at source ↗
Figure 9
Figure 9. Figure 9: Illustration of various computer vision tasks. Top row: RGB Image, Semantic Segmentation, Instance Segmentation, Panoptic Segmentation, Object Detection, and Depth Estimation. Bottom row: (This figure serves solely illustrative purposes and does not originate from the MS COCO dataset.) De-raining, De-hazing, Denoising, Image Super-Resolution (ISR), Deblurring, Edge Detection, Low-light Enhancement, and RGB… view at source ↗
Figure 10
Figure 10. Figure 10: Visualization of images generated by SGT, demonstrating high-quality and diverse generations across a wide range of prompts and scenes. compared to GenEval. The lack of substantial improvement on this benchmark suggests that the SGT framework does not inherently facilitate complex instruc￾tion parsing capabilities. Further enhancement of these editing proficiencies likely requires the integration of speci… view at source ↗
Figure 11
Figure 11. Figure 11: Pseudocode for t-SNE visualization pipeline. visualization, we first apply Principal Component Analysis (PCA) to reduce the feature dimensionality to 50, followed by t-SNE projection onto a 2D plane for visualization. For t-SNE, we adopt the default perplexity value of 30. The complete pipeline is summarized in [PITH_FULL_IMAGE:figures/full_fig_p022_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Pseudocode for keyword-level attention analysis during image generation. We extract keywords from the prompt, compute GQA attention maps at selected timesteps and layers, and aggregate attention scores for each keyword to quantify its influence on the generated image [PITH_FULL_IMAGE:figures/full_fig_p023_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Token-level attention distribution during image generation. [PITH_FULL_IMAGE:figures/full_fig_p024_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 46 canonical work pages · 24 internal anchors

  1. [1]

    Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.arXiv preprint arXiv:2301.08243, April 2023

    Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., Le- Cun, Y., Ballas, N.: Self-supervised learning from images with a joint-embedding predictive architecture. arXiv:2301.08243 (2023) 2, 7

  2. [2]

    Qwen2.5-VL Technical Report

    Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-vl technical report. arXiv:2502.13923 (2025) 9

  3. [3]

    Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., Ramesh, A.: Video gener- ation models as world simulators (2024),https://openai.com/research/video- generation-models-as-world-simulators1

  4. [4]

    Chen, F., Jing, M., Lu, W., Feng, Y., Li, X., Cao, X.: Unihetero: Could gen- eration enhance understanding for vision-language-model at large data scale? arXiv:2512.23512 (2025) 5

  5. [5]

    Chen, L., Li, J., Dong, X., Zhang, P., Zang, Y., Chen, Z., Duan, H., Wang, J., Qiao, Y., Lin, D., et al.: Are we on the right way for evaluating large vision-language models? NeurIPS37, 27056–27087 (2024) 6, 9, 11, 17

  6. [6]

    Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

    Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C.: Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv:2501.17811 (2025) 3, 5, 10

  7. [7]

    arXiv preprint arXiv:2401.14404 , year=

    Chen, X., Liu, Z., Xie, S., He, K.: Deconstructing denoising diffusion models for self-supervised learning. arXiv:2401.14404 (2024) 4

  8. [8]

    NeurIPS36, 49250–49267 (2023) 2

    Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P.N., Hoi, S.: Instructblip: Towards general-purpose vision-language models with instruction tuning. NeurIPS36, 49250–49267 (2023) 2

  9. [9]

    Emerging Properties in Unified Multimodal Pretraining

    Deng, C., Zhu, D., Li, K., Gou, C., Li, F., Wang, Z., Zhong, S., Yu, W., Nie, X., Song, Z., et al.: Emerging properties in unified multimodal pretraining. arXiv:2505.14683 (2025) 3, 5, 7, 9, 10

  10. [10]

    In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2

    Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., Yi, L.: DreamLLM: Synergistic multimodal comprehension and creation. In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2

  11. [11]

    Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction, 2025

    Du, S., Guo, J., Li, B., Cui, S., Xu, Z., Luo, Y., Wei, Y., Gai, K., Wang, X., Wu, K., et al.: Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction. arXiv:2511.23386 (2025) 4

  12. [12]

    In: ACMMM

    Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., et al.: Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In: ACMMM. pp. 11198–11201 (2024) 9

  13. [13]

    In: ICML (2024) 1, 4

    Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al.: Scaling rectified flow transformers for high-resolution image synthesis. In: ICML (2024) 1, 4

  14. [14]

    MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

    Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al.: Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394 (2023) 9, 11

  15. [15]

    In: ECCV

    Fu, X., Yin, W., Hu, M., Wang, K., Ma, Y., Tan, P., Shen, S., Lin, D., Long, X.: Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. In: ECCV. pp. 241–258. Springer (2024) 4 26 Songsong Yu et al

  16. [16]

    BLINK: Multimodal Large Language Models Can See but Not Perceive

    Fu, X., Hu, Y., Li, B., Feng, Y., Wang, H., Lin, X., Roth, D., Smith, N.A., Ma, W.C., Krishna, R.: Blink: Multimodal large language models can see but not per- ceive. arXiv:2404.12390 (2024) 9, 11

  17. [17]

    arXiv preprint arXiv:2407.00783 , year=

    Fuest, M., Ma, P., Gui, M., Schusterbauer, J., Hu, V.T., Ommer, B.: Diffusion models and representation learning: A survey. arXiv:2407.00783 (2024) 4

  18. [18]

    SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

    Ge,Y.,Zhao,S.,Zhu,J.,Ge,Y.,Yi,K.,Song,L.,Li,C.,Ding,X.,Shan,Y.:Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv:2404.14396 (2024) 3

  19. [19]

    NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14

    Ghosh, D., Hajishirzi, H., Schmidt, L.: Geneval: An object-focused framework for evaluating text-to-image alignment. NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14

  20. [20]

    In: CVPR

    Graikos,A.,Yellapragada,S.,Le,M.Q.,Kapse,S.,Prasanna,P.,Saltz,J.,Samaras, D.: Learned representation-guided diffusion models for large-image generation. In: CVPR. pp. 8532–8542 (2024) 4

  21. [21]

    In: CVPR

    Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., et al.: Hallusionbench: an advanced diagnostic suite for entangled lan- guage hallucination and visual illusion in large vision-language models. In: CVPR. pp. 14375–14385 (2024) 6, 9, 11, 17

  22. [22]

    In: CVPR

    Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., Liu, X.: Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis. In: CVPR. pp. 15733–15744 (2025) 1

  23. [23]

    In: CVPR

    Hudson, D.A., Zoran, D., Malinowski, M., Lampinen, A.K., Jaegle, A., McClelland, J.L., Matthey, L., Hill, F., Lerchner, A.: Soda: Bottleneck diffusion models for representation learning. In: CVPR. pp. 23115–23127 (2024) 4

  24. [24]

    arXiv preprint arXiv:2402.03161 , year=

    Jin, Y., Sun, Z., Xu, K., Chen, L., Jiang, H., Huang, Q., Song, C., Liu, Y., Zhang, D., Song, Y., et al.: Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization. arXiv:2402.03161 (2024) 2

  25. [25]

    arXiv preprint arXiv:2309.04669 , year=

    Jin, Y., Xu, K., Chen, L., Liao, C., Tan, J., Huang, Q., Chen, B., Lei, C., Liu, A., Song, C., et al.: Unified language-vision pretraining in llm with dynamic discrete visual tokenization. arXiv:2309.04669 (2023) 2

  26. [26]

    In: ICCV

    Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: ICCV. pp. 4015–4026 (2023) 9

  27. [27]

    LLaVA-OneVision: Easy Visual Task Transfer

    Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., Li, C.: Llava-onevision: Easy visual task transfer. arXiv:2408.03326 (2024) 9

  28. [28]

    In: EMNLP

    Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., Wen, J.R.: Evaluating object hallucination in large vision-language models. In: EMNLP. pp. 292–305 (2023) 6, 9, 11, 17

  29. [29]

    Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

    Li, Z., Liu, Z., Zhang, Q., Lin, B., Yuan, S., Yan, Z., Ye, Y., Yu, W., Niu, Y., Yuan, L.: Uniworld-v2: Reinforce image editing with diffusion negative-aware finetuning and mllm implicit feedback. arXiv:2510.16888 (2025) 2

  30. [30]

    arXiv:2512.19680 (2025) 5

    Liao, X., He, Q., Xu, K., Qu, X., Li, Y., Wei, W., Yao, A.: Va-π: Variational policy alignment for pixel-aware autoregressive generation. arXiv:2512.19680 (2025) 5

  31. [31]

    UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

    Lin, B., Li, Z., Cheng, X., Niu, Y., Ye, Y., He, X., Yuan, S., Yu, W., Wang, S., Ge, Y., et al.: Uniworld: High-resolution semantic encoders for unified visual understanding and generation. arXiv:2506.03147 (2025) 10

  32. [32]

    Transactions of the Association for Computational Linguistics (2023) 6, 9, 17

    Liu, F., Emerson, G.E.T., Collier, N.: Visual spatial reasoning. Transactions of the Association for Computational Linguistics (2023) 6, 9, 17

  33. [33]

    In: NeurIPS (2023) 1

    Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: NeurIPS (2023) 1

  34. [34]

    Step1X-Edit: A Practical Framework for General Image Editing

    Liu, S., Han, Y., Xing, P., Yin, F., Wang, R., Cheng, W., Liao, J., Wang, Y., Fu, H., Han, C., Li, G., Peng, Y., Sun, Q., Wu, J., Cai, Y., Ge, Z., Ming, R., Xia, L., Semantic Generative Tuning for Unified Multimodal Models 27 Zeng, X., Zhu, Y., Jiao, B., Zhang, X., Yu, G., Jiang, D.: Step1x-edit: A practical framework for general image editing. arXiv:2504...

  35. [35]

    Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al.: Mmbench: Is your multi-modal model an all-around player? In: ECCV. pp. 216–233. Springer (2024) 9

  36. [36]

    Science China Information Sciences67(12), 220102 (2024) 6, 17

    Liu, Y., Li, Z., Huang, M., Yang, B., Yu, W., Li, C., Yin, X.C., Liu, C.L., Jin, L., Bai, X.: Ocrbench: on the hidden mystery of ocr in large multimodal models. Science China Information Sciences67(12), 220102 (2024) 6, 17

  37. [37]

    arXiv:2312.17172 (2023) 3

    Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., Kembhavi, A.: Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. arXiv:2312.17172 (2023) 3

  38. [38]

    In: ICLR (2024) 6, 9, 17

    Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.W., Galley, M., Gao, J.: Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In: ICLR (2024) 6, 9, 17

  39. [39]

    In: NeurIPS (2022) 6, 17

    Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.W., Zhu, S.C., Tafjord, O., Clark, P., Kalyan, A.: Learn to explain: Multimodal reasoning via thought chains for science question answering. In: NeurIPS (2022) 6, 17

  40. [40]

    arXiv:2405.15232 (2024) 4

    Luo, R., Li, Y., Chen, L., He, W., Lin, T.E., Liu, Z., Zhang, L., Song, Z., Xia, X., Liu, T., et al.: Deem: Diffusion models serve as the eyes of large language models for image perception. arXiv:2405.15232 (2024) 4

  41. [41]

    Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,

    Ma, C., Jiang, Y., Wu, J., Yang, J., Yu, X., Yuan, Z., Peng, B., Qi, X.: Unitok: A unified tokenizer for visual generation and understanding. arXiv:2502.20321 (2025) 3

  42. [42]

    In: ICCV

    Ma, S., Ge, Y., Wang, T., Guo, Y., Ge, Y., Shan, Y.: Genhancer: Imperfect genera- tive models are secretly strong vision-centric enhancers. In: ICCV. pp. 24402–24412 (2025) 4, 5, 7

  43. [43]

    In: CVPR

    Ma, Y., Liu, X., Chen, X., Liu, W., Wu, C., Wu, Z., Pan, Z., Xie, Z., Zhang, H., Yu, X., et al.: Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation. In: CVPR. pp. 7739–7751 (2025) 2

  44. [44]

    In: WACV

    Mathew, M., Karatzas, D., Jawahar, C.: Docvqa: A dataset for vqa on document images. In: WACV. pp. 2200–2209 (2021) 6, 17

  45. [45]

    WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

    Niu, Y., Ning, M., Zheng, M., Jin, W., Lin, B., Jin, P., Liao, J., Ning, K., Feng, C., Zhu, B., Yuan, L.: Wise: A world knowledge-informed semantic evaluation for text-to-image generation. arXiv:2503.07265 (2025) 2

  46. [47]

    Transfer between Modalities with MetaQueries

    Pan, X., Shukla, S.N., Singh, A., Zhao, Z., Mishra, S.K., Wang, J., Xu, Z., Chen, J., Li, K., Juefei-Xu, F., et al.: Transfer between modalities with metaqueries. arXiv:2504.06256 (2025) 2

  47. [48]

    In: ECCV

    Parihar, R., Sachidanand, V., Mani, S., Karmali, T., Venkatesh Babu, R.: Pre- cisecontrol: Enhancing text-to-image diffusion models with fine-grained attribute control. In: ECCV. pp. 469–487. Springer (2024) 4

  48. [49]

    In: CVPR (2022) 14

    Peng, X., Wei, Y., Deng, A., Wang, D., Hu, D.: Balanced multimodal learning via on-the-fly gradient modulation. In: CVPR (2022) 14

  49. [50]

    In: CVPR

    Qu, L., Zhang, H., Liu, Y., Wang, X., Jiang, Y., Gao, Y., Ye, H., Du, D.K., Yuan, Z., Wu, X.: Tokenflow: Unified image tokenizer for multimodal understanding and generation. In: CVPR. pp. 2545–2555 (2025) 4

  50. [51]

    V., Zettlemoyer, L., and Yu, L

    Shi,W.,Han,X.,Zhou,C.,Liang,W.,Lin,X.V.,Zettlemoyer,L.,Yu,L.:Lmfusion: Adapting pretrained language models for multimodal generation. arXiv:2412.15188 (2024) 2 28 Songsong Yu et al

  51. [52]

    In: CVPR

    Shipard, J., Wiliem, A., Thanh, K.N., Xiang, W., Fookes, C.: Diversity is definitely needed: Improving model-agnostic zero-shot classification via stable diffusion. In: CVPR. pp. 769–778 (2023) 4

  52. [53]

    Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

    Su, Z., Wei, H., Cen, K., Wang, Y., Chen, G., Yuan, C., Chu, X.: Generation en- hances understanding in unified multimodal models via multi-representation gen- eration. arXiv:2601.21406 (2026) 4, 10

  53. [54]

    Unilip: Adapting clip for unified multimodal understanding, generation and editing

    Tang, H., Xie, C., Bao, X., Weng, T., Li, P., Zheng, Y., Wang, L.: Unilip: Adapting clip for unified multimodal understanding, generation and editing. arXiv:2507.23278 (2025) 10

  54. [55]

    Chameleon: Mixed-Modal Early-Fusion Foundation Models

    Team, C.: Chameleon: Mixed-modal early-fusion foundation models. arXiv:2405.09818 (2024) 4, 10

  55. [56]

    NeurIPS37, 84839–84865 (2024) 1

    Tian, K., Jiang, Y., Yuan, Z., Peng, B., Wang, L.: Visual autoregressive model- ing: Scalable image generation via next-scale prediction. NeurIPS37, 84839–84865 (2024) 1

  56. [57]

    NeurIPS 36, 48382–48402 (2023) 4

    Tian, Y., Fan, L., Isola, P., Chang, H., Krishnan, D.: Stablerep: Synthetic images from text-to-image models make strong visual representation learners. NeurIPS 36, 48382–48402 (2023) 4

  57. [58]

    NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17

    Tong, P., Brown, E., Wu, P., Woo, S., Iyer, A.J.V., Akula, S.C., Yang, S., Yang, J., Middepogu, M., Wang, Z., et al.: Cambrian-1: A fully open, vision-centric ex- ploration of multimodal llms. NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17

  58. [59]

    MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

    Tong, S., Fan, D., Zhu, J., Xiong, Y., Chen, X., Sinha, K., Rabbat, M., LeCun, Y., Xie, S., Liu, Z.: Metamorph: Multimodal understanding and generation via instruction tuning. arXiv:2412.14164 (2024) 4

  59. [60]

    In: CVPR

    Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., Xie, S.: Eyes wide shut? exploring the visual shortcomings of multimodal llms. In: CVPR. pp. 9568–9578 (2024) 6, 9, 11, 17

  60. [61]

    arXiv preprint arXiv:2410.09575 , year=

    Wang, H., Zheng, A., Zhao, Y., Wang, T., Ge, Z., Zhang, X., Zhang, Z.: Recon- structive visual instruction tuning. arXiv:2410.09575 (2024) 4, 5

  61. [62]

    arXiv:2508.03320 (2025) 2

    Wang, P., Peng, Y., Gan, Y., Hu, L., Xie, T., Wang, X., Wei, Y., Tang, C., Zhu, B., Li, C., et al.: Skywork unipic: Unified autoregressive modeling for visual un- derstanding and generation. arXiv:2508.03320 (2025) 2

  62. [63]

    arXiv:2407.20171 (2024) 4, 5

    Wang, W., Sun, Q., Zhang, F., Tang, Y., Liu, J., Wang, X.: Diffusion feedback helps clip see better. arXiv:2407.20171 (2024) 4, 5

  63. [64]

    Emu3: Next-Token Prediction is All You Need

    Wang, X., Zhang, X., Luo, Z., Sun, Q., Cui, Y., Wang, J., Zhang, F., Wang, Y., Li, Z., Yu, Q., et al.: Emu3: Next-token prediction is all you need. arXiv:2409.18869 (2024) 3, 10

  64. [65]

    In: ICML

    Wang, Y., Schiff, Y., Gokaslan, A., Pan, W., Wang, F., De Sa, C., Kuleshov, V.: Infodiffusion: Representation learning using information maximizing diffusion models. In: ICML. pp. 36336–36354. PMLR (2023) 4

  65. [66]

    arXiv:2510.22946 (2025) 3

    Wang, Z., Chen, Z., Gou, C., Li, F., Deng, C., Zhu, D., Li, K., Yu, W., Tu, H., Fan, H., et al.: Lightfusion: A light-weighted, double fusion framework for unified multimodal understanding and generation. arXiv:2510.22946 (2025) 3

  66. [67]

    In: ICCV

    Wei, C., Mangalam, K., Huang, P.Y., Li, Y., Fan, H., Xu, H., Wang, H., Xie, C., Yuille, A., Feichtenhofer, C.: Diffusion models as masked autoencoders. In: ICCV. pp. 16284–16294 (2023) 4

  67. [68]

    In: ECCV

    Weng, N., Pegios, P., Petersen, E., Feragen, A., Bigdeli, S.: Fast diffusion-based counterfactuals for shortcut removal and generation. In: ECCV. pp. 338–357. Springer (2024) 4

  68. [69]

    Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

    Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., et al.: Janus: Decoupling visual encoding for unified multimodal understanding and generation. arXiv:2410.13848 (2024) 2 Semantic Generative Tuning for Unified Multimodal Models 29

  69. [70]

    OmniGen2: Towards Instruction-Aligned Multimodal Generation

    Wu, C., Zheng, P., Yan, R., Xiao, S., Luo, X., Wang, Y., Li, W., Jiang, X., Liu, Y., Zhou, J., et al.: Omnigen2: Exploration to advanced multimodal generation. arXiv:2506.18871 (2025) 3, 5, 7, 9, 10

  70. [71]

    Representa- tion entanglement for generation: Training diffusion trans- formers is much easier than you think.arXiv preprint arXiv:2507.01467, 2025

    Wu, G., Zhang, S., Shi, R., Gao, S., Chen, Z., Wang, L., Chen, Z., Gao, H., Tang, Y., Yang, J., et al.: Representation entanglement for generation: Training diffusion transformers is much easier than you think. arXiv:2507.01467 (2025) 4

  71. [72]

    IJCV (2025) 3

    Wu, J., Jiang, Y., Ma, C., Liu, Y., Zhao, H., Yuan, Z., Bai, S., Bai, X.: Liquid: Language models are scalable and unified multi-modal generators. IJCV (2025) 3

  72. [73]

    arXiv preprint arXiv:2505.23661 , year=

    Wu, S., Wu, Z., Gong, Z., Tao, Q., Jin, S., Li, Q., Li, W., Loy, C.C.: Ope- nuni: A simple baseline for unified multimodal understanding and generation. arXiv:2505.23661 (2025) 5, 10

  73. [74]

    In: ICCV

    Wu, S., Zhang, W., Xu, L., Jin, S., Wu, Z., Tao, Q., Liu, W., Li, W., Loy, C.C.: Harmonizing visual representations for unified multimodal understanding and gen- eration. In: ICCV. pp. 17739–17750 (2025) 10

  74. [75]

    VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

    Wu, Y., Zhang, Z., Chen, J., Tang, H., Li, D., Fang, Y., Zhu, L., Xie, E., Yin, H., Yi, L., et al.: Vila-u: a unified foundation model integrating visual understanding and generation. arXiv:2409.04429 (2024) 3

  75. [76]

    xAI: Grok-1.5 vision preview.https://x.ai/news/grok-1.5v(2024) 9

  76. [77]

    Reconstruction Alignment Improves Unified Multimodal Models

    Xie, J., Darrell, T., Zettlemoyer, L., Wang, X.: Reconstruction alignment improves unified multimodal models. arXiv:2509.07295 (2025) 2, 4, 8, 10

  77. [78]

    Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

    Xie, J., Mao, W., Bai, Z., Zhang, D.J., Wang, W., Lin, K.Q., Gu, Y., Chen, Z., Yang, Z., Shou, M.Z.: Show-o: One single transformer to unify multimodal under- standing and generation. arXiv:2408.12528 (2024) 4, 10

  78. [79]

    Show-o2: Improved Native Unified Multimodal Models

    Xie,J.,Yang,Z.,Shou,M.Z.:Show-o2:Improvednativeunifiedmultimodalmodels. arXiv:2506.15564 (2025) 4

  79. [80]

    In: CVPR

    Yang, J., Yin, D., Zhou, Y., Rao, F., Zhai, W., Cao, Y., Zha, Z.J.: Mmar: Towards lossless multi-modal auto-regressive probabilistic modeling. In: CVPR. pp. 7974– 7985 (2025) 2

  80. [81]

    In: ICCV

    Yang, X., Wang, X.: Diffusion model as representation learner. In: ICCV. pp. 18938–18949 (2023) 4

Showing first 80 references.