REVIEW 2 major objections 2 minor 1 cited by
Current embodied AI models cannot reliably follow social norms without explicit instructions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 23:25 UTC pith:HHHFOMJR
load-bearing objection RobotEQ introduces a new benchmark for unguided social-norm adherence in embodied AI, but the manual annotations lack any reported validation, so the performance gaps are hard to interpret. the 2 major comments →
RobotEQ: Transitioning from Passive Intelligence to Active Intelligence in Embodied AI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
RobotEQ is the first benchmark for active intelligence in embodied AI; it consists of RobotEQ-Data (1,894 egocentric images with 4,944 action-judgment and 1,157 spatial-grounding questions across 10 categories) and RobotEQ-Bench, which demonstrates that current models fall short in reliable active intelligence, particularly spatial grounding, while RAG with external social-norm knowledge generally improves performance.
What carries the argument
RobotEQ-Data and RobotEQ-Bench, which supply manually annotated action-judgment and spatial-grounding questions that specify appropriate robot actions from egocentric images according to social norms.
Load-bearing premise
The 4,944 action-judgment and 1,157 spatial-grounding questions produced by manual annotation faithfully represent the social norms that should govern robot behavior across the 10 categories and 56 subcategories.
What would settle it
A new set of real-world human judgments on the same images that systematically disagrees with the paper's manual annotations on which actions are permissible.
If this is right
- Models must improve spatial grounding to achieve reliable active intelligence.
- Adding retrieval from external social-norm knowledge bases raises performance on action judgments.
- Robotics can move from passive, instruction-following manipulation to active social compliance.
- Systematic evaluation across 10 categories reveals consistent gaps in current approaches.
- The benchmark supplies a concrete testbed for measuring progress toward unguided norm adherence.
Where Pith is reading between the lines
- Robots passing this benchmark could operate with less direct supervision in shared human spaces.
- The dataset could be extended with video or multi-view inputs to test dynamic norm understanding.
- Cultural variation in norms would require separate annotation efforts or modular knowledge bases.
- Combining RobotEQ scores with existing task-completion benchmarks could measure overall social competence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces RobotEQ as the first benchmark for active intelligence in embodied AI, where robots must infer and adhere to social norms without explicit user instructions (contrasted with passive intelligence). It describes construction of RobotEQ-Data from 1,894 egocentric images across 10 categories and 56 subcategories, yielding 4,944 action-judgment questions and 1,157 spatial-grounding questions via manual annotation. RobotEQ-Bench then evaluates state-of-the-art models, reporting that they fall short especially on spatial grounding while RAG with external social-norm knowledge bases improves results.
Significance. If the benchmark is shown to validly encode the relevant norms, the work would usefully quantify a gap between current models' passive task-following and the active social compliance needed for real-world robot deployment, while providing an initial testbed and a simple RAG mitigation that could inform follow-on research in socially aware robotics.
major comments (2)
- [RobotEQ-Data] RobotEQ-Data section: the manual annotation process that produced the 4,944 action-judgment and 1,157 spatial-grounding questions reports neither inter-annotator agreement statistics nor an expert-review protocol, nor any handling of cultural or contextual variability; because the headline claims about model shortfalls rest directly on these questions faithfully representing the norms that should govern robot behavior, the absence of validation metrics leaves the measured performance gaps difficult to interpret.
- [Experiments] Experiments / RobotEQ-Bench section: the reported performance gaps (especially the spatial-grounding shortfall) are presented without statistical significance tests or confidence intervals, so it is unclear whether the observed differences between models and between RAG vs. non-RAG conditions exceed what would be expected from sampling variability alone.
minor comments (2)
- [Abstract] Abstract: the definitions of 'passive intelligence' and 'active intelligence' are introduced only by example; a concise formal distinction would help readers map the new terminology onto existing embodied-AI literature.
- [RobotEQ-Data] The paper does not discuss how the 10 categories and 56 subcategories were selected or whether they were chosen to be representative of broader social-norm distributions.
Simulated Author's Rebuttal
We thank the referee for the constructive comments on our manuscript. We address each major comment point by point below, indicating planned revisions where appropriate.
read point-by-point responses
-
Referee: [RobotEQ-Data] RobotEQ-Data section: the manual annotation process that produced the 4,944 action-judgment and 1,157 spatial-grounding questions reports neither inter-annotator agreement statistics nor an expert-review protocol, nor any handling of cultural or contextual variability; because the headline claims about model shortfalls rest directly on these questions faithfully representing the norms that should govern robot behavior, the absence of validation metrics leaves the measured performance gaps difficult to interpret.
Authors: We agree that explicit reporting of inter-annotator agreement, expert review, and cultural handling would improve interpretability. In the revision we will expand the RobotEQ-Data section with these details: we computed Fleiss' kappa on a 20% overlap subset (value to be reported), included a two-stage expert review by robotics and ethics researchers, and addressed cultural variability by recruiting annotators from multiple regions with consensus resolution for disagreements. These additions will directly support the benchmark's validity claims. revision: yes
-
Referee: [Experiments] Experiments / RobotEQ-Bench section: the reported performance gaps (especially the spatial-grounding shortfall) are presented without statistical significance tests or confidence intervals, so it is unclear whether the observed differences between models and between RAG vs. non-RAG conditions exceed what would be expected from sampling variability alone.
Authors: We concur that statistical rigor is needed. The revision will add bootstrap-derived 95% confidence intervals for all accuracy metrics and paired significance tests (McNemar's test for model comparisons and Wilcoxon signed-rank for RAG vs. baseline) to confirm that reported gaps exceed sampling variability. revision: yes
Circularity Check
No circularity in benchmark construction or model evaluation
full rationale
The paper's chain consists of manual annotation to produce RobotEQ-Data (4,944 action-judgment and 1,157 spatial-grounding questions from 1,894 images) followed by direct empirical evaluation of models on the held-out RobotEQ-Bench. Reported shortfalls in active intelligence and RAG gains are measured performance numbers, not algebraic reductions, fitted-parameter renamings, or self-referential derivations. No equations, self-citations for uniqueness theorems, or ansatzes appear in the provided methodology. The benchmark is self-contained against external model testing.
Axiom & Free-Parameter Ledger
free parameters (1)
- Category and subcategory selection
axioms (1)
- domain assumption Social norms relevant to robot behavior can be adequately captured by binary action-judgment questions and spatial-grounding questions on static egocentric images.
read the original abstract
Embodied AI is a prominent research topic in both academia and industry. Current research centers on completing tasks based on explicit user instructions. However, for robots to integrate into human society, they must understand which actions are permissible and which are prohibited, even without explicit commands. We refer to the user-guided AI as passive intelligence and the unguided AI as active intelligence. This paper introduces RobotEQ, the first benchmark for active intelligence, aiming to assess whether existing models can comprehend and adhere to social norms in embodied scenarios. First, we construct RobotEQ-Data, a dataset consisting of 1,894 egocentric images, spanning 10 representative embodied categories and 56 subcategories. Through extensive manual annotation, we provide 4,944 action judgment questions and 1,157 spatial grounding questions, specifying appropriate robot actions across diverse scenarios. Furthermore, we establish RobotEQ-Bench to evaluate the performance of state-of-the-art models on this task. Experimental results demonstrate that current models still fall short in achieving reliable active intelligence, particularly in spatial grounding. Meanwhile, leveraging RAG techniques to incorporate external social norm knowledge bases can generally enhance performance. This work can facilitate the transition of robotics from user-guided passive manipulation to active social compliance.
Figures
Forward citations
Cited by 1 Pith paper
-
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
OPID distills episode- and step-level skills from completed on-policy trajectories, routes them via critical-first mechanism, and combines the resulting log-probability shift advantage with outcome advantage for polic...
Reference graph
Works this paper leans on
-
[1]
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, et al. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. arXiv preprint arXiv:2503.01743, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[2]
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, Soham Ghosh, Amélie Héliou, Paul Jacob, Albert Q. Jiang, Kartik Khandelwal, Timothée Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, Teven Le Scao, Andy Lo, William Marshall,...
2024
-
[3]
Reid, Stephen Gould, and Anton van den Hengel
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22,...
2018
-
[4]
System Card: Claude Opus 4.6
Anthropic. System Card: Claude Opus 4.6. https://www-cdn.anthropic.com/0dd8650 75ad3132672ee0ab40b05a53f14cf5288.pdf, February 2026. Released February 5, 2026. 212 pages. Also available athttps://www.anthropic.com/system-cards
2026
-
[5]
System Card: Claude Opus 4.7
Anthropic. System Card: Claude Opus 4.7. https://www.anthropic.com/system-cards, April 2026. Released April 16, 2026. 232 pages. Download PDF from the System Cards page
2026
-
[6]
System Card: Claude Sonnet 4.6
Anthropic. System Card: Claude Sonnet 4.6. https://www-cdn.anthropic.com/78073 f739564e986ff3e28522761a7a0b4484f84.pdf, February 2026. Released February 2026. Also available athttps://www.anthropic.com/system-cards
2026
-
[7]
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
2023
-
[8]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[9]
Qwen2.5-vl technical report, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025
2025
-
[10]
Seed 1.6 Technical Report
ByteDance Seed Team. Seed 1.6 Technical Report. https://seed.bytedance.com/en/se ed1_6, 2025. Chinese version:https://research.doubao.com/zh/seed1_6. 10
2025
-
[11]
Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025
Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025
2025
-
[12]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[13]
Aya vision: Advancing the frontier of multilingual multimodality, 2025
Saurabh Dash, Yiyang Nan, John Dang, Arash Ahmadian, Shivalika Singh, Madeline Smith, Bharat Venkitesh, Vlad Shmyhlo, Viraat Aryabumi, Walter Beller-Morales, Jeremy Pekmez, Jason Ozuzu, Pierre Richemond, Acyr Locatelli, Nick Frosst, Phil Blunsom, Aidan Gomez, Ivan Zhang, Marzieh Fadaee, Manoj Govindassamy, Sudip Roy, Matthias Gallé, Beyza Ermis, Ahmet Üst...
2025
-
[14]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence. Palm-e: an embodied ...
2023
-
[15]
Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, and Sai Rajeswar
Aarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han Lù, Johan Obando-Ceron, Juan A. Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, and Sai Rajeswar. Grounding computer use agents on human demonstrations, 2025
2025
-
[16]
Seedream 3.0 technical report, 2025
Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, Wei Liu, Yichun Shi, Shiqi Sun, Yu Tian, Zhi Tian, Peng Wang, Rui Wang, Xuanda Wang, Xun Wang, Ye Wang, Guofeng Wu, Jie Wu, Xin Xia, Xuefeng Xiao, Zhonghua Zhai, Xinyu Zhang, Qi Zhang, Yuwei Zhang, Shijia Zhao, Jianchao Yang, and Weilin Hu...
2025
-
[17]
Gemini 3 Pro Image Model Card
Google DeepMind. Gemini 3 Pro Image Model Card. https://storage.googleapis.com /deepmind-media/Model-Cards/Gemini-3-Pro-Image-Model-Card.pdf , November
-
[18]
Released November 20, 2025
2025
-
[19]
Gemini 3.1 Pro Model Card
Google DeepMind. Gemini 3.1 Pro Model Card. https://deepmind.google/models/mod el-cards/gemini-3-1-pro/ , February 2026. PDF version: https://storage.google apis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf
2026
-
[20]
Navigating the digital world as humans do: Universal visual grounding for GUI agents
Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. Navigating the digital world as humans do: Universal visual grounding for GUI agents. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[21]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[22]
Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, et al. Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.arXiv preprint arXiv:2507.01006, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[23]
Inner monologue: Em- bodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Tomas Jackson, Noah Brown, Linda Luu, Sergey Levine, Karol Hausman, and brian ichter. Inner monologue: Em- bodied reasoning through planning with language models. In Karen Liu, Dana Kulic, and Jeff Ichnow...
2023
-
[24]
Joshi, Kyle Jeffrey, Rosario Jauregui Ruano, Jasmine Hsu, Keerthana Gopalakrishnan, Byron David, Andy Zeng, and Chuyuan Kelly Fu
brian ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, Dmitry Kalashnikov, Sergey Levine, Yao Lu, Carolina Parada, Kanishka Rao, Pierre Sermanet, Alexander T Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Mengyuan Yan, Noah Brown, Michael Ahn, Omar ...
2023
-
[25]
Building and better understanding vision-language models: insights and future directions, 2024
Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Léo Tronchon. Building and better understanding vision-language models: insights and future directions, 2024
2024
-
[26]
LLaV A-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. LLaV A-onevision: Easy visual task transfer. Transactions on Machine Learning Research, 2025
2025
-
[27]
Aligning cyber space with physical world: A comprehensive survey on embodied ai, 2025
Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. Aligning cyber space with physical world: A comprehensive survey on embodied ai, 2025
2025
-
[28]
Infigui-g1: Advancing gui grounding with adaptive exploration policy optimization
Yuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li, Congkai Xie, Jiasheng Wang, Xueyu Hu, Xiaotian Han, Jianbo Yuan, Xinyao Wang, et al. Infigui-g1: Advancing gui grounding with adaptive exploration policy optimization. InProceedings of the AAAI Conference on Artificial Intelligence, pages 32267–32275, 2026
2026
-
[29]
Advancing social intelligence in ai agents: Technical challenges and open questions
Leena Mathur, Paul Pu Liang, and Louis-Philippe Morency. Advancing social intelligence in ai agents: Technical challenges and open questions. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20541–20560, 2024
2024
-
[30]
Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015
V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015
2015
-
[31]
Keane Ong, Wei Dai, Carol Li, Dewei Feng, Hengzhi Li, Jingyao Wu, Jiaee Cheong, Rui Mao, Gianmarco Mengaldo, Erik Cambria, et al. Human behavior atlas: Benchmarking unified psychological and social behavior understanding.arXiv preprint arXiv:2510.04899, 2025
-
[32]
GPT-4o System Card, 2024
OpenAI. GPT-4o System Card, 2024. Covers GPT-4o and GPT-4o-mini
2024
-
[33]
GPT-5.4 Thinking System Card
OpenAI. GPT-5.4 Thinking System Card. https://deploymentsafety.openai.com/gp t-5-4-thinking, March 2026. Released March 5, 2026
2026
-
[34]
GPT-5.5 System Card
OpenAI. GPT-5.5 System Card. https://deploymentsafety.openai.com/gpt-5-5 , April 2026. Released April 23, 2026
2026
-
[35]
Manso, Anaís Garrell, Alberto Sanfeliu, Anne Spalanzani, and Rachid Alami
Phani Teja Singamaneni, Pilar Bachiller-Burgos, Luis J. Manso, Anaís Garrell, Alberto Sanfeliu, Anne Spalanzani, and Rachid Alami. A survey on socially aware robot navigation: Taxonomy and future challenges.The International Journal of Robotics Research, 43(10):1533–1572, February 2024
2024
-
[36]
Gui-g2: Gaussian reward modeling for gui grounding, 2025
Fei Tang, Zhangxuan Gu, Zhengxi Lu, Xuyang Liu, Shuheng Shen, Changhua Meng, Wen Wang, Wenqi Zhang, Yongliang Shen, Weiming Lu, Jun Xiao, and Yueting Zhuang. Gui-g2: Gaussian reward modeling for gui grounding, 2025
2025
-
[37]
Gemma Team. Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[38]
GUI-actor: Coordinate-free visual grounding for GUI agents
Qianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang, Jianwei Yang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao, Reuben Tan, Si Qin, Lars Liden, Qingwei Lin, Huan Zhang, Tong Zhang, Jianbing Zhang, Dongmei Zhang, and Jianfeng Gao. GUI-actor: Coordinate-free visual grounding for GUI agents. InThe Thirty-ninth Annual Conference on Neural Information Processi...
2026
-
[39]
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang, Liang Zhao, Yisong Wang, and Chong Ruan. Deepseek-vl2: Mixture-of-experts visio...
2024
-
[40]
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. InThe Eleventh International Conference on Learning Representations, 2023
2023
-
[41]
Social- iq: A question answering benchmark for artificial social intelligence
Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. Social- iq: A question answering benchmark for artificial social intelligence. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
-
[42]
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Tensor fusion network for multimodal sentiment analysis. InProceedings of the 2017 conference on empirical methods in natural language processing, pages 1103–1114, 2017
2017
-
[43]
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246, 2018
2018
-
[44]
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li, Yinan He, Tan Jiang, Jiapeng Luo, Yi Wang, Conghui He, Botian Shi, Xingcheng Zh...
2025
-
[45]
Sanketi, Grecia Salazar, Michael S
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, Quan Vuong, Vincent Vanhoucke, Huong Tran, Radu Soricut, Anikait Singh, Jaspiar Singh, Pierre Sermanet, Pannag R. Sanketi, Grecia Salazar, Michael S. Ryoo, Krista Reymann, Kanishka Rao, Karl Pertsch, Igor Mordatch, Henryk Michalewski...
2023
-
[46]
Non-verbal Signal Recognition: The ability to interpret non-verbal communicative cues, including gaze direction, hand gestures, body posture, head movements, pointing, beckoning, and other implicit signals such as chin-directed requests. 26
-
[47]
Proxemics & Spatial Norms: The ability to reason about personal space, appropriate pass- ing distance, queuing, yielding, spatial occlusion, positional relationships, and movement boundaries in shared environments
-
[48]
Role Boundary & Authority: The ability to recognize role-defined responsibilities and authority relations, including who may issue instructions, whether a request is legitimate, and whether an action oversteps age-, identity-, responsibility-, or organization-based boundaries
-
[49]
Timing & Interruption Norms: The ability to judge when to intervene, wait, interrupt, or yield, taking into account turn-taking conventions, ongoing interactions, sequential order, and the pacing of human activities
-
[50]
Contextual Volume & Behavioral Restraint: The ability to adjust voice volume, notifica- tion sounds, movement amplitude, and behavioral conspicuousness according to the social and environmental context
-
[51]
Resource & Ownership Norms: The ability to reason about ownership, borrowing, sharing, occupation rights, unattended belongings, and whether an object may be moved, used, returned, or left untouched
-
[52]
Priority & Protected Persons: The ability to identify people who require prioritized assistance or protection, such as children, elderly people, patients, vulnerable individuals, or people involved in emergency situations
-
[53]
Annotation methodology.We assign dimension labels through a two-stage process that combines LLM-based classification with human calibration
Culture-Specific Norms: The ability to recognize etiquette, taboos, ceremonial practices, religious norms, and behavioral boundaries that vary across cultural or occasion-specific contexts. Annotation methodology.We assign dimension labels through a two-stage process that combines LLM-based classification with human calibration. In the first stage, Gemini...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.