REVIEW 2 major objections 1 minor 87 references
Semantic Generative Tuning uses image segmentation as a generative proxy to align understanding and generation in unified multimodal models.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 18:28 UTC pith:U6HXRAVQ
load-bearing objection SGT uses segmentation as a post-training generative proxy to align understanding and generation in UMMs, with reported benchmark gains and mechanistic support, but details are thin in the abstract. the 2 major comments →
Semantic Generative Tuning for Unified Multimodal Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper states that hierarchical visual tasks formulated as generative proxies, with image segmentation as the optimal choice, bridge the isolation between understanding and generation in unified multimodal models. Semantic Generative Tuning then applies this proxy to align representation spaces, producing improved feature linear separability, optimized visual-textual attention allocation, and measurable gains in both multimodal comprehension and generative fidelity across benchmarks.
What carries the argument
Semantic Generative Tuning (SGT), a post-training method that treats image segmentation as a generative proxy to align visual and textual representation spaces.
Load-bearing premise
High-level semantic tasks such as image segmentation supply structural semantics that optimally bridge understanding and generation without introducing distracting low-level texture signals.
What would settle it
If a low-level task such as texture or edge prediction produces larger gains than segmentation on the same unified multimodal model benchmarks and mechanistic metrics, the claim that segmentation is the optimal proxy would not hold.
If this is right
- Multimodal comprehension improves across standard benchmarks
- Generative layout fidelity increases without separate dense pixel objectives
- Feature representations become more linearly separable
- Visual-textual attention allocation becomes more optimized
Where Pith is reading between the lines
- The proxy approach could be tested on video or 3D tasks to check whether other high-level semantic signals produce similar alignment effects.
- Post-training with segmentation might allow existing unified models to close the performance gap between understanding and generation without full retraining.
- The method implies that choosing the right generative proxy during alignment could reduce reliance on paired text-image data for joint optimization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to present the first systematic investigation of generative post-training for unified multimodal models (UMMs). It formulates hierarchical visual tasks as generative proxies to address misalignment between understanding (sparse text) and generation (dense pixels), empirically identifying high-level semantic tasks—particularly image segmentation—as optimal proxies that provide structural semantics without low-level texture distractions. Building on this, it introduces Semantic Generative Tuning (SGT) and reports that mechanistic analyses show improved feature linear separability and optimized visual-textual attention allocation, with extensive evaluations demonstrating consistent gains in multimodal comprehension and generative fidelity across benchmarks; code is released.
Significance. If the empirical results and mechanistic findings hold under scrutiny, the work could establish a practical post-training paradigm for aligning understanding and generation in UMMs via semantic proxies, with the segmentation insight and attention/separability analyses offering reusable mechanistic guidance. Code availability strengthens reproducibility and potential adoption.
major comments (2)
- [Abstract] Abstract: the central claim that 'high-level semantic tasks, particularly image segmentation, serve as optimal proxies' rests on an unspecified 'empirical investigation' into hierarchical tasks; without details on the task hierarchy tested, comparison metrics, or controls for low-level vs. high-level effects, it is impossible to verify whether segmentation is demonstrably optimal or if the reported gains are driven by post-hoc selection.
- [Abstract] Abstract: the mechanistic findings ('improved feature linear separability and optimized visual-textual attention allocation pattern') and the claim of 'consistent improvements... across mainstream benchmarks' are presented without any quantitative tables, ablation results, baseline comparisons, or analysis methods; these are load-bearing for the soundness of both the empirical and mechanistic contributions.
minor comments (1)
- [Abstract] The provided link ends with a trailing period ('https://song2yu.github.io/SGT/.') which appears to be a minor formatting artifact.
Simulated Author's Rebuttal
We appreciate the referee's feedback highlighting areas where the abstract could be more self-contained. We agree to revise the abstract to provide more context on the empirical investigation and to reference the quantitative support for the mechanistic and benchmark claims.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'high-level semantic tasks, particularly image segmentation, serve as optimal proxies' rests on an unspecified 'empirical investigation' into hierarchical tasks; without details on the task hierarchy tested, comparison metrics, or controls for low-level vs. high-level effects, it is impossible to verify whether segmentation is demonstrably optimal or if the reported gains are driven by post-hoc selection.
Authors: The details of the empirical investigation into hierarchical tasks are presented in the main text of the manuscript. To address this comment, we will revise the abstract to briefly outline the task hierarchy tested and the metrics used for comparison, ensuring the optimality claim is better supported within the abstract itself. revision: yes
-
Referee: [Abstract] Abstract: the mechanistic findings ('improved feature linear separability and optimized visual-textual attention allocation pattern') and the claim of 'consistent improvements... across mainstream benchmarks' are presented without any quantitative tables, ablation results, baseline comparisons, or analysis methods; these are load-bearing for the soundness of both the empirical and mechanistic contributions.
Authors: While the abstract summarizes the findings, the supporting quantitative tables, ablation results, baseline comparisons, and analysis methods are included in the main manuscript. We will revise the abstract to include key quantitative highlights and references to the analysis methods to make these claims more verifiable from the abstract. revision: yes
Circularity Check
No significant circularity detected
full rationale
The paper's core contribution is an empirical study of hierarchical visual tasks as generative proxies for unified multimodal models, with segmentation identified as optimal through experiments, followed by the introduction of SGT and reported benchmark gains plus attention analyses. No derivation chain, equations, or first-principles claims are present that reduce by construction to fitted parameters, self-definitions, or self-citation load-bearing premises; the claims rest on external benchmark results and mechanistic observations rather than internal tautologies.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption High-level semantic tasks such as image segmentation provide structural semantics that enhance both perception and generative layout fidelity without the distraction of texture details.
read the original abstract
Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, prevailing training paradigms independently optimize understanding via sparse text signals and generation through dense pixel objectives. Such a decoupled strategy yields misaligned representation spaces, isolating visual understanding from generation and hindering their mutual reinforcement. This work presents the first systematic investigation into generative post-training, where we formulate hierarchical visual tasks as generative proxies to bridge the isolation in UMMs. Our empirical investigation reveals that high-level semantic tasks, particularly image segmentation, serve as optimal proxies. Unlike low-level tasks that distract models with texture details, segmentation provides structural semantics that significantly enhance both vision-centric perception and generative layout fidelity. Building upon these insights, we introduce Semantic Generative Tuning (SGT), a novel paradigm that leverages segmentation as a generative proxy to align and synergize multimodal capabilities. Mechanistic analyses further demonstrate that SGT fundamentally improves feature linear separability and optimizes visual-textual attention allocation pattern. Extensive evaluations show that SGT consistently improves both multimodal comprehension and generative fidelity across mainstream benchmarks. Our code is available on the https://song2yu.github.io/SGT/.
Figures
Reference graph
Works this paper leans on
-
[1]
Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., Le- Cun, Y., Ballas, N.: Self-supervised learning from images with a joint-embedding predictive architecture. arXiv:2301.08243 (2023) 2, 7
-
[2]
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-vl technical report. arXiv:2502.13923 (2025) 9
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[3]
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., Ramesh, A.: Video gener- ation models as world simulators (2024),https://openai.com/research/video- generation-models-as-world-simulators1
2024
- [4]
-
[5]
Chen, L., Li, J., Dong, X., Zhang, P., Zang, Y., Chen, Z., Duan, H., Wang, J., Qiao, Y., Lin, D., et al.: Are we on the right way for evaluating large vision-language models? NeurIPS37, 27056–27087 (2024) 6, 9, 11, 17
2024
-
[6]
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C.: Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv:2501.17811 (2025) 3, 5, 10
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[7]
arXiv preprint arXiv:2401.14404 , year=
Chen, X., Liu, Z., Xie, S., He, K.: Deconstructing denoising diffusion models for self-supervised learning. arXiv:2401.14404 (2024) 4
-
[8]
NeurIPS36, 49250–49267 (2023) 2
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P.N., Hoi, S.: Instructblip: Towards general-purpose vision-language models with instruction tuning. NeurIPS36, 49250–49267 (2023) 2
2023
-
[9]
Emerging Properties in Unified Multimodal Pretraining
Deng, C., Zhu, D., Li, K., Gou, C., Li, F., Wang, Z., Zhong, S., Yu, W., Nie, X., Song, Z., et al.: Emerging properties in unified multimodal pretraining. arXiv:2505.14683 (2025) 3, 5, 7, 9, 10
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[10]
In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., Yi, L.: DreamLLM: Synergistic multimodal comprehension and creation. In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2
2024
-
[11]
Du, S., Guo, J., Li, B., Cui, S., Xu, Z., Luo, Y., Wei, Y., Gai, K., Wang, X., Wu, K., et al.: Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction. arXiv:2511.23386 (2025) 4
-
[12]
In: ACMMM
Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., et al.: Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In: ACMMM. pp. 11198–11201 (2024) 9
2024
-
[13]
In: ICML (2024) 1, 4
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al.: Scaling rectified flow transformers for high-resolution image synthesis. In: ICML (2024) 1, 4
2024
-
[14]
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al.: Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394 (2023) 9, 11
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[15]
In: ECCV
Fu, X., Yin, W., Hu, M., Wang, K., Ma, Y., Tan, P., Shen, S., Lin, D., Long, X.: Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. In: ECCV. pp. 241–258. Springer (2024) 4 26 Songsong Yu et al
2024
-
[16]
BLINK: Multimodal Large Language Models Can See but Not Perceive
Fu, X., Hu, Y., Li, B., Feng, Y., Wang, H., Lin, X., Roth, D., Smith, N.A., Ma, W.C., Krishna, R.: Blink: Multimodal large language models can see but not per- ceive. arXiv:2404.12390 (2024) 9, 11
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[17]
arXiv preprint arXiv:2407.00783 , year=
Fuest, M., Ma, P., Gui, M., Schusterbauer, J., Hu, V.T., Ommer, B.: Diffusion models and representation learning: A survey. arXiv:2407.00783 (2024) 4
-
[18]
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Ge,Y.,Zhao,S.,Zhu,J.,Ge,Y.,Yi,K.,Song,L.,Li,C.,Ding,X.,Shan,Y.:Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv:2404.14396 (2024) 3
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[19]
NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14
Ghosh, D., Hajishirzi, H., Schmidt, L.: Geneval: An object-focused framework for evaluating text-to-image alignment. NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14
2023
-
[20]
In: CVPR
Graikos,A.,Yellapragada,S.,Le,M.Q.,Kapse,S.,Prasanna,P.,Saltz,J.,Samaras, D.: Learned representation-guided diffusion models for large-image generation. In: CVPR. pp. 8532–8542 (2024) 4
2024
-
[21]
In: CVPR
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., et al.: Hallusionbench: an advanced diagnostic suite for entangled lan- guage hallucination and visual illusion in large vision-language models. In: CVPR. pp. 14375–14385 (2024) 6, 9, 11, 17
2024
-
[22]
In: CVPR
Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., Liu, X.: Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis. In: CVPR. pp. 15733–15744 (2025) 1
2025
-
[23]
In: CVPR
Hudson, D.A., Zoran, D., Malinowski, M., Lampinen, A.K., Jaegle, A., McClelland, J.L., Matthey, L., Hill, F., Lerchner, A.: Soda: Bottleneck diffusion models for representation learning. In: CVPR. pp. 23115–23127 (2024) 4
2024
-
[24]
arXiv preprint arXiv:2402.03161 , year=
Jin, Y., Sun, Z., Xu, K., Chen, L., Jiang, H., Huang, Q., Song, C., Liu, Y., Zhang, D., Song, Y., et al.: Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization. arXiv:2402.03161 (2024) 2
-
[25]
arXiv preprint arXiv:2309.04669 , year=
Jin, Y., Xu, K., Chen, L., Liao, C., Tan, J., Huang, Q., Chen, B., Lei, C., Liu, A., Song, C., et al.: Unified language-vision pretraining in llm with dynamic discrete visual tokenization. arXiv:2309.04669 (2023) 2
-
[26]
In: ICCV
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: ICCV. pp. 4015–4026 (2023) 9
2023
-
[27]
LLaVA-OneVision: Easy Visual Task Transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., Li, C.: Llava-onevision: Easy visual task transfer. arXiv:2408.03326 (2024) 9
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[28]
In: EMNLP
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., Wen, J.R.: Evaluating object hallucination in large vision-language models. In: EMNLP. pp. 292–305 (2023) 6, 9, 11, 17
2023
-
[29]
Li, Z., Liu, Z., Zhang, Q., Lin, B., Yuan, S., Yan, Z., Ye, Y., Yu, W., Niu, Y., Yuan, L.: Uniworld-v2: Reinforce image editing with diffusion negative-aware finetuning and mllm implicit feedback. arXiv:2510.16888 (2025) 2
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[30]
Liao, X., He, Q., Xu, K., Qu, X., Li, Y., Wei, W., Yao, A.: Va-π: Variational policy alignment for pixel-aware autoregressive generation. arXiv:2512.19680 (2025) 5
-
[31]
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Lin, B., Li, Z., Cheng, X., Niu, Y., Ye, Y., He, X., Yuan, S., Yu, W., Wang, S., Ge, Y., et al.: Uniworld: High-resolution semantic encoders for unified visual understanding and generation. arXiv:2506.03147 (2025) 10
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[32]
Transactions of the Association for Computational Linguistics (2023) 6, 9, 17
Liu, F., Emerson, G.E.T., Collier, N.: Visual spatial reasoning. Transactions of the Association for Computational Linguistics (2023) 6, 9, 17
2023
-
[33]
In: NeurIPS (2023) 1
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: NeurIPS (2023) 1
2023
-
[34]
Step1X-Edit: A Practical Framework for General Image Editing
Liu, S., Han, Y., Xing, P., Yin, F., Wang, R., Cheng, W., Liao, J., Wang, Y., Fu, H., Han, C., Li, G., Peng, Y., Sun, Q., Wu, J., Cai, Y., Ge, Z., Ming, R., Xia, L., Semantic Generative Tuning for Unified Multimodal Models 27 Zeng, X., Zhu, Y., Jiao, B., Zhang, X., Yu, G., Jiang, D.: Step1x-edit: A practical framework for general image editing. arXiv:2504...
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[35]
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al.: Mmbench: Is your multi-modal model an all-around player? In: ECCV. pp. 216–233. Springer (2024) 9
2024
-
[36]
Science China Information Sciences67(12), 220102 (2024) 6, 17
Liu, Y., Li, Z., Huang, M., Yang, B., Yu, W., Li, C., Yin, X.C., Liu, C.L., Jin, L., Bai, X.: Ocrbench: on the hidden mystery of ocr in large multimodal models. Science China Information Sciences67(12), 220102 (2024) 6, 17
2024
-
[37]
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., Kembhavi, A.: Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. arXiv:2312.17172 (2023) 3
-
[38]
In: ICLR (2024) 6, 9, 17
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.W., Galley, M., Gao, J.: Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. In: ICLR (2024) 6, 9, 17
2024
-
[39]
In: NeurIPS (2022) 6, 17
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.W., Zhu, S.C., Tafjord, O., Clark, P., Kalyan, A.: Learn to explain: Multimodal reasoning via thought chains for science question answering. In: NeurIPS (2022) 6, 17
2022
-
[40]
Luo, R., Li, Y., Chen, L., He, W., Lin, T.E., Liu, Z., Zhang, L., Song, Z., Xia, X., Liu, T., et al.: Deem: Diffusion models serve as the eyes of large language models for image perception. arXiv:2405.15232 (2024) 4
-
[41]
Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,
Ma, C., Jiang, Y., Wu, J., Yang, J., Yu, X., Yuan, Z., Peng, B., Qi, X.: Unitok: A unified tokenizer for visual generation and understanding. arXiv:2502.20321 (2025) 3
-
[42]
In: ICCV
Ma, S., Ge, Y., Wang, T., Guo, Y., Ge, Y., Shan, Y.: Genhancer: Imperfect genera- tive models are secretly strong vision-centric enhancers. In: ICCV. pp. 24402–24412 (2025) 4, 5, 7
2025
-
[43]
In: CVPR
Ma, Y., Liu, X., Chen, X., Liu, W., Wu, C., Wu, Z., Pan, Z., Xie, Z., Zhang, H., Yu, X., et al.: Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation. In: CVPR. pp. 7739–7751 (2025) 2
2025
-
[44]
In: WACV
Mathew, M., Karatzas, D., Jawahar, C.: Docvqa: A dataset for vqa on document images. In: WACV. pp. 2200–2209 (2021) 6, 17
2021
-
[45]
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Niu, Y., Ning, M., Zheng, M., Jin, W., Lin, B., Jin, P., Liao, J., Ning, K., Feng, C., Zhu, B., Yuan, L.: Wise: A world knowledge-informed semantic evaluation for text-to-image generation. arXiv:2503.07265 (2025) 2
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[47]
Transfer between Modalities with MetaQueries
Pan, X., Shukla, S.N., Singh, A., Zhao, Z., Mishra, S.K., Wang, J., Xu, Z., Chen, J., Li, K., Juefei-Xu, F., et al.: Transfer between modalities with metaqueries. arXiv:2504.06256 (2025) 2
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[48]
In: ECCV
Parihar, R., Sachidanand, V., Mani, S., Karmali, T., Venkatesh Babu, R.: Pre- cisecontrol: Enhancing text-to-image diffusion models with fine-grained attribute control. In: ECCV. pp. 469–487. Springer (2024) 4
2024
-
[49]
In: CVPR (2022) 14
Peng, X., Wei, Y., Deng, A., Wang, D., Hu, D.: Balanced multimodal learning via on-the-fly gradient modulation. In: CVPR (2022) 14
2022
-
[50]
In: CVPR
Qu, L., Zhang, H., Liu, Y., Wang, X., Jiang, Y., Gao, Y., Ye, H., Du, D.K., Yuan, Z., Wu, X.: Tokenflow: Unified image tokenizer for multimodal understanding and generation. In: CVPR. pp. 2545–2555 (2025) 4
2025
-
[51]
V., Zettlemoyer, L., and Yu, L
Shi,W.,Han,X.,Zhou,C.,Liang,W.,Lin,X.V.,Zettlemoyer,L.,Yu,L.:Lmfusion: Adapting pretrained language models for multimodal generation. arXiv:2412.15188 (2024) 2 28 Songsong Yu et al
-
[52]
In: CVPR
Shipard, J., Wiliem, A., Thanh, K.N., Xiang, W., Fookes, C.: Diversity is definitely needed: Improving model-agnostic zero-shot classification via stable diffusion. In: CVPR. pp. 769–778 (2023) 4
2023
-
[53]
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
Su, Z., Wei, H., Cen, K., Wang, Y., Chen, G., Yuan, C., Chu, X.: Generation en- hances understanding in unified multimodal models via multi-representation gen- eration. arXiv:2601.21406 (2026) 4, 10
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[54]
Unilip: Adapting clip for unified multimodal understanding, generation and editing
Tang, H., Xie, C., Bao, X., Weng, T., Li, P., Zheng, Y., Wang, L.: Unilip: Adapting clip for unified multimodal understanding, generation and editing. arXiv:2507.23278 (2025) 10
-
[55]
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Team, C.: Chameleon: Mixed-modal early-fusion foundation models. arXiv:2405.09818 (2024) 4, 10
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[56]
NeurIPS37, 84839–84865 (2024) 1
Tian, K., Jiang, Y., Yuan, Z., Peng, B., Wang, L.: Visual autoregressive model- ing: Scalable image generation via next-scale prediction. NeurIPS37, 84839–84865 (2024) 1
2024
-
[57]
NeurIPS 36, 48382–48402 (2023) 4
Tian, Y., Fan, L., Isola, P., Chang, H., Krishnan, D.: Stablerep: Synthetic images from text-to-image models make strong visual representation learners. NeurIPS 36, 48382–48402 (2023) 4
2023
-
[58]
NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17
Tong, P., Brown, E., Wu, P., Woo, S., Iyer, A.J.V., Akula, S.C., Yang, S., Yang, J., Middepogu, M., Wang, Z., et al.: Cambrian-1: A fully open, vision-centric ex- ploration of multimodal llms. NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17
2024
-
[59]
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Tong, S., Fan, D., Zhu, J., Xiong, Y., Chen, X., Sinha, K., Rabbat, M., LeCun, Y., Xie, S., Liu, Z.: Metamorph: Multimodal understanding and generation via instruction tuning. arXiv:2412.14164 (2024) 4
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[60]
In: CVPR
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., Xie, S.: Eyes wide shut? exploring the visual shortcomings of multimodal llms. In: CVPR. pp. 9568–9578 (2024) 6, 9, 11, 17
2024
-
[61]
arXiv preprint arXiv:2410.09575 , year=
Wang, H., Zheng, A., Zhao, Y., Wang, T., Ge, Z., Zhang, X., Zhang, Z.: Recon- structive visual instruction tuning. arXiv:2410.09575 (2024) 4, 5
-
[62]
Wang, P., Peng, Y., Gan, Y., Hu, L., Xie, T., Wang, X., Wei, Y., Tang, C., Zhu, B., Li, C., et al.: Skywork unipic: Unified autoregressive modeling for visual un- derstanding and generation. arXiv:2508.03320 (2025) 2
-
[63]
Wang, W., Sun, Q., Zhang, F., Tang, Y., Liu, J., Wang, X.: Diffusion feedback helps clip see better. arXiv:2407.20171 (2024) 4, 5
-
[64]
Emu3: Next-Token Prediction is All You Need
Wang, X., Zhang, X., Luo, Z., Sun, Q., Cui, Y., Wang, J., Zhang, F., Wang, Y., Li, Z., Yu, Q., et al.: Emu3: Next-token prediction is all you need. arXiv:2409.18869 (2024) 3, 10
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[65]
In: ICML
Wang, Y., Schiff, Y., Gokaslan, A., Pan, W., Wang, F., De Sa, C., Kuleshov, V.: Infodiffusion: Representation learning using information maximizing diffusion models. In: ICML. pp. 36336–36354. PMLR (2023) 4
2023
-
[66]
Wang, Z., Chen, Z., Gou, C., Li, F., Deng, C., Zhu, D., Li, K., Yu, W., Tu, H., Fan, H., et al.: Lightfusion: A light-weighted, double fusion framework for unified multimodal understanding and generation. arXiv:2510.22946 (2025) 3
-
[67]
In: ICCV
Wei, C., Mangalam, K., Huang, P.Y., Li, Y., Fan, H., Xu, H., Wang, H., Xie, C., Yuille, A., Feichtenhofer, C.: Diffusion models as masked autoencoders. In: ICCV. pp. 16284–16294 (2023) 4
2023
-
[68]
In: ECCV
Weng, N., Pegios, P., Petersen, E., Feragen, A., Bigdeli, S.: Fast diffusion-based counterfactuals for shortcut removal and generation. In: ECCV. pp. 338–357. Springer (2024) 4
2024
-
[69]
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., et al.: Janus: Decoupling visual encoding for unified multimodal understanding and generation. arXiv:2410.13848 (2024) 2 Semantic Generative Tuning for Unified Multimodal Models 29
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[70]
OmniGen2: Towards Instruction-Aligned Multimodal Generation
Wu, C., Zheng, P., Yan, R., Xiao, S., Luo, X., Wang, Y., Li, W., Jiang, X., Liu, Y., Zhou, J., et al.: Omnigen2: Exploration to advanced multimodal generation. arXiv:2506.18871 (2025) 3, 5, 7, 9, 10
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[71]
Wu, G., Zhang, S., Shi, R., Gao, S., Chen, Z., Wang, L., Chen, Z., Gao, H., Tang, Y., Yang, J., et al.: Representation entanglement for generation: Training diffusion transformers is much easier than you think. arXiv:2507.01467 (2025) 4
-
[72]
IJCV (2025) 3
Wu, J., Jiang, Y., Ma, C., Liu, Y., Zhao, H., Yuan, Z., Bai, S., Bai, X.: Liquid: Language models are scalable and unified multi-modal generators. IJCV (2025) 3
2025
-
[73]
arXiv preprint arXiv:2505.23661 , year=
Wu, S., Wu, Z., Gong, Z., Tao, Q., Jin, S., Li, Q., Li, W., Loy, C.C.: Ope- nuni: A simple baseline for unified multimodal understanding and generation. arXiv:2505.23661 (2025) 5, 10
-
[74]
In: ICCV
Wu, S., Zhang, W., Xu, L., Jin, S., Wu, Z., Tao, Q., Liu, W., Li, W., Loy, C.C.: Harmonizing visual representations for unified multimodal understanding and gen- eration. In: ICCV. pp. 17739–17750 (2025) 10
2025
-
[75]
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Wu, Y., Zhang, Z., Chen, J., Tang, H., Li, D., Fang, Y., Zhu, L., Xie, E., Yin, H., Yi, L., et al.: Vila-u: a unified foundation model integrating visual understanding and generation. arXiv:2409.04429 (2024) 3
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[76]
xAI: Grok-1.5 vision preview.https://x.ai/news/grok-1.5v(2024) 9
2024
-
[77]
Reconstruction Alignment Improves Unified Multimodal Models
Xie, J., Darrell, T., Zettlemoyer, L., Wang, X.: Reconstruction alignment improves unified multimodal models. arXiv:2509.07295 (2025) 2, 4, 8, 10
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[78]
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Xie, J., Mao, W., Bai, Z., Zhang, D.J., Wang, W., Lin, K.Q., Gu, Y., Chen, Z., Yang, Z., Shou, M.Z.: Show-o: One single transformer to unify multimodal under- standing and generation. arXiv:2408.12528 (2024) 4, 10
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[79]
Show-o2: Improved Native Unified Multimodal Models
Xie,J.,Yang,Z.,Shou,M.Z.:Show-o2:Improvednativeunifiedmultimodalmodels. arXiv:2506.15564 (2025) 4
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[80]
In: CVPR
Yang, J., Yin, D., Zhou, Y., Rao, F., Zhai, W., Cao, Y., Zha, Z.J.: Mmar: Towards lossless multi-modal auto-regressive probabilistic modeling. In: CVPR. pp. 7974– 7985 (2025) 2
2025
-
[81]
In: ICCV
Yang, X., Wang, X.: Diffusion model as representation learner. In: ICCV. pp. 18938–18949 (2023) 4
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.