On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li, Qingshan Xu, Bobo Li, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Qingyu Shi, Zhiyuan Zhou, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu, Liyu Jia, Wentao Hu, Meng Luo, Jiebo Luo, Tat-Seng Chua, Shuicheng YAN, Hanwang Zhang
Overview
This talk, presented by Tianjie Ju and a large team of co-authors from numerous institutions, introduces a novel framework for evaluating general-purpose multimodal foundation models. Titled "On Path to Multimodal Generalist: General-Level and General-Bench," the presentation addresses a critical gap in the rapidly evolving field of multimodal AI: the lack of a reliable, long-term benchmark capable of truly assessing the generalization and synergy capabilities of these complex models. The core contribution lies in two interconnected components: General-Level, a five-tier system designed to quantify multimodal intelligence, and General-Bench, a massive, comprehensive benchmark dataset built to facilitate this evaluation.

Key moments
- 0:00 Introduction: Towards a true multimodal generalist AI
- 2:00 Limitations of current narrow model evaluation methods
- 3:00 Proposing a five-level General-Level synergy system
- 4:00 Detailed explanation of General-Level 2 through 5
- 6:00 Clarifying synergy measurement: Generalist vs. Specialist
- 8:00 Criteria for advancing to higher General-Level tiers
- 9:00 Introducing General-Bench: A new comprehensive benchmark
- 9:40 General-Bench scale: 29 domains, 700+ tasks, 300k+ samples
On Path to Multimodal Generalist: General-Level and General-Bench
Speakers: Hao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li, Qingshan Xu, Bobo Li, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Qingyu Shi, Zhiyuan Zhou, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu, Liyu Jia, Wentao Hu, Meng Luo, Jiebo Luo, Tat-Seng Chua, Shuicheng YAN, Hanwang Zhang
Conference: ICML 2025
YouTube: https://slideslive.com/39044102
Overview
This talk, presented by Tianjie Ju and a large team of co-authors from numerous institutions, introduces a novel framework for evaluating general-purpose multimodal foundation models. Titled "On Path to Multimodal Generalist: General-Level and General-Bench," the presentation addresses a critical gap in the rapidly evolving field of multimodal AI: the lack of a reliable, long-term benchmark capable of truly assessing the generalization and synergy capabilities of these complex models. The core contribution lies in two interconnected components: General-Level, a five-tier system designed to quantify multimodal intelligence, and General-Bench, a massive, comprehensive benchmark dataset built to facilitate this evaluation.
The motivation behind this ambitious project stems from the observation that while large multimodal language models (LLMs) are booming and increasingly handle diverse modalities beyond just images, their underlying intelligence often remains heavily language-centric. Current evaluations tend to be narrow, focusing on incremental improvements on specific visual benchmarks rather than assessing a model's ability to generalize knowledge across tasks and modalities. The speakers argue that true Artificial General Intelligence (AGI) will manifest as a "modality generalist," akin to the powerful AIs depicted in science fiction, capable of synergistic reasoning and generation across diverse data types. This work aims to redefine how the community measures progress towards such a generalist AI, shifting the focus from mere accuracy to genuine cross-modal generalization.
The significance of this work is profound, as it challenges the prevailing methods of multimodal model comparison and proposes a more rigorous standard. By introducing a structured evaluation system and an expansive benchmark, the project seeks to guide the development of future multimodal generalists away from specialized, bolted-on modules towards architectures that foster deep, synergistic intelligence. The findings from their extensive evaluation of over 100 generalist models and 700 specialist models reveal that current multimodal AI, despite its advancements, is far from achieving true generalism, highlighting critical areas for future research and development.
Background
▶ Watch: Introduction: Towards a true multimodal generalist AI (0:00)
The landscape of AI research has witnessed an explosive growth in multimodal Large Language Models (LLMs), which are increasingly capable of processing and generating information across various data types. Initially confined to processing images alongside text, these models are now expanding to encompass more modalities. However, a fundamental challenge persists: most current multimodal LLMs (or MLMs) are designed with an LLM at their core, with additional vision or generation modules "bolted on." This architecture means the entire system heavily relies on the intelligence derived from language, with language intelligence often performing the "heavy lifting" for other modalities. While LLMs have demonstrated remarkable emergent abilities and generalization across numerous NLP tasks, the aspiration for a truly general multimodal AI demands more than just a language model augmented with other sensory inputs. It requires an AI that can generalize across tasks and modalities with the same fluidity and robustness that ChatGPT exhibits across diverse NLP challenges.
The prevailing methods for evaluating these MLMs have proven to be insufficient. The community largely focuses on benchmarking and comparing models based on their performance on a limited set of visual benchmarks, where gaining a few points of accuracy over a baseline is often touted as a significant advancement. The speakers critically question this approach, arguing that a marginal improvement in accuracy on specific tasks does not necessarily equate to stronger multimodal intelligence. Such narrow evaluations fail to capture the essence of true generalization and, crucially, synergy—the ability of a model to leverage knowledge learned in one task or modality to enhance its performance in others. Without a mechanism to measure this synergy, the community risks developing highly specialized models that, despite impressive individual task scores, remain far from the vision of a general-purpose AI.
To address this critical gap, the authors drew inspiration from established evaluation methodologies, particularly the capability levels used in self-driving technology to define performance tiers. This analogy highlights the need for a structured, hierarchical system that moves beyond simple accuracy metrics to assess deeper levels of intelligence, generalization, and synergistic reasoning. The problem, therefore, is not merely a lack of benchmarks, but a fundamental misunderstanding of what constitutes "multimodal generalist" intelligence and how to rigorously measure its emergence. This foundational shift in perspective underpins the development of both the General-Level evaluation system and the General-Bench benchmark.
Key Findings
▶ Watch: Proposing a five-level General-Level synergy system (3:00)
The core contribution of this work is the introduction of a sophisticated framework comprising two pivotal components: General-Level, a novel five-tier system for assessing multimodal intelligence, and General-Bench, an expansive benchmark designed to facilitate this evaluation. This framework allowed the researchers to conduct a comprehensive analysis of the current state of multimodal AI, yielding several critical findings that challenge conventional wisdom and highlight the significant distance remaining on the path to true multimodal generalism.
First, the General-Level system provides a clear, hierarchical pathway to evaluate a model's generalization and synergy. Moving from Level 1 (specialist) to Level 5 (full modality synergy), models are progressively challenged to demonstrate broader task coverage, outperform specialized models, and exhibit synergistic understanding and generation across modalities. This framework revealed a stark reality: no existing models have yet achieved Level 5, the ultimate goal of full cross-modal synergy supporting SOTA NLP specialists. Only three models demonstrated true understanding-generation synergy, qualifying for Level 4. Furthermore, many models that appear strong at Level 2 (handling multiple modalities/tasks) significantly drop out at Level 3, which demands outperforming specialists and thus proving genuine synergy. This indicates that while many models boast multimodal capabilities, few demonstrate the deep generalization required to surpass dedicated experts.
Second, the development of General-Bench provides an unprecedented scale and diversity for multimodal evaluation. Recognizing the limitations of existing benchmarks, General-Bench was constructed to cover 29 domains, five major modality types, over 145 modality skills, and more than 700 distinct tasks, encompassing over 300,000 samples. Crucially, it includes both understanding and generation paradigms, which is often a missing component in current benchmarks. This comprehensive nature allowed for a more holistic assessment of multimodal models, moving beyond the narrow confines of vision-only or understanding-only evaluations.
The experimental results from evaluating over 700 specialist models and more than 100 multimodal generalist models up to the end of 2024 painted a clear picture of the current state of the art:
- Specialization over Generalization: Most multimodal models are heavily specialized, focusing on specific tasks or modalities, and are far from being true generalists.
- Unbalanced Modality Support: Modality coverage is highly uneven, with the vast majority of models centered around vision, particularly images, and often neglecting other crucial modalities.
- Dominance of Understanding: Most models primarily support multimodal understanding, with generation performance across various tasks and skills proving to be quite uneven and inconsistent.
- Lack of Synergy: The ability to achieve synergy across different paradigms (e.g., understanding supporting generation) and across diverse modalities is largely absent in current models.
In essence, the key findings underscore that despite significant progress in multimodal AI, the community's current "generalist" models often lack the breadth of coverage, the depth of generalization, and the true cross-modal synergy necessary to advance towards AGI. The General-Level and General-Bench framework provides the much-needed tools to accurately measure and guide this progress.
Technical Deep Dive
▶ Watch: Clarifying synergy measurement: Generalist vs. Specialist (6:00)
The technical core of this work lies in the meticulously designed General-Level evaluation system and the expansive General-Bench benchmark. Together, these components provide a robust methodology for assessing the true generalist capabilities of multimodal AI models, moving beyond superficial performance metrics.
General-Level Evaluation System
The General-Level system proposes a five-tier hierarchy to quantify the level of generalization and synergy demonstrated by a multimodal model. The concept of a "multimodal generalist" is defined as a large multimodal foundation model or agent, typically built around an LLM core, capable of handling multiple tasks across various modalities, with users interacting via natural language prompts. In contrast, a "specialist model" is defined as a model that achieves state-of-the-art (SOTA) performance on one specific task, is usually fine-tuned on corresponding task training data, often has a smaller parameter scale, and typically does not use an LLM as a central reasoning core.
A crucial aspect of General-Level is its approach to measuring synergy. The authors initially considered pairwise comparisons between tasks (e.g., Task A vs. Task B) but found this impractical due to the extensive joint pre-training and fine-tuning that generalist models undergo, making it difficult to construct truly independent task distributions for fair comparison. Instead, they simplified the measurement: if a generalist model outperforms a SOTA specialist model on an individual task, it suggests synergy, implying the generalist is leveraging knowledge from other tasks or modalities. This comparison is intentionally "harsh" or "unfair" in that specialists are heavily fine-tuned for their specific task, while generalists are not allowed any task-specific fine-tuning within this evaluation framework. This strict condition is deemed necessary because a true multimodal generalist, by definition, should demonstrate strong generalization and even surpass task-specific experts without explicit fine-tuning for every single task.
The scoring for each general level reflects the expected performance across modalities, often as a weighted average, and importantly, the general level score decreases monotonically with increasing level, signifying the increasing difficulty and higher standards required for each successive tier.
The five levels are defined as follows:
- Level 1 (Specialist): This serves as a reference point. It encompasses task-specific models that are top performers in their narrow domains. These models establish the baseline performance that generalists must aim to surpass.
- Level 2 (Generalist): Models at this level are capable of handling multiple modalities and tasks. The primary criterion is broad coverage, even if no explicit synergy is yet demonstrated. Scoring is based on average performance across the covered tasks and modalities. To advance from Level 1 to Level 2, a model must support as many modalities, tasks, and capabilities as broadly as possible.
- Level 3 (Task-level Synergy): This level introduces the crucial condition of synergy. A model only receives credit on a task if it outperforms the specialist for that task. This demonstrates real generalization, suggesting that the model is leveraging knowledge across its broader capabilities. If a model achieves a non-zero score at Level 3, it indicates some form of task-level synergy. To move from Level 2 to Level 3, a model needs to demonstrate broad cross-task generalization by outperforming as many SOTA specialists as possible across different tasks.
- Level 4 (Understanding-Generation Synergy): This level goes further, requiring that understanding and generation capabilities mutually support each other. Instead of a simple average, a harmonic mean is used for scoring. This ensures that strong performance in one paradigm cannot compensate for weak performance in the other; a high score requires balanced strength in both understanding and generation. For instance, a model with high generation scores but low understanding scores would receive a very low Level 4 score. To move from Level 3 to Level 4, models must outperform SOTA specialists across tasks from different paradigms.
- Level 5 (Full Modality Synergy): This is the ultimate goal, representing full synergy across modalities. A model qualifies for Level 5 if other modalities help boost language tasks to the extent that it can support or even surpass SOTA NLP specialists. This implies a truly integrated multimodal intelligence where non-linguistic information profoundly enhances linguistic capabilities. To reach Level 5, the model should be able to surpass as many SOTA NLP specialists as possible.
General-Bench Benchmark
The motivation for building General-Bench was clear: existing multimodal LLM benchmarks are insufficient for evaluating generalist models due to gaps in limited modalities, restricted scales, or the absence of crucial paradigms like generation. General-Bench is designed to be a massive and comprehensive benchmark, addressing these limitations head-on.
Its scale is unprecedented:
- Domains: 29 distinct domains.
- Modality Types: 5 major modality types (e.g., implied vision, text, audio, etc., though specific types beyond vision were not detailed in the transcript).
- Modality Skills: Over 145 different modality skills.
- Tasks: More than 700 individual tasks.
- Samples: Over 300,000 samples.
Crucially, General-Bench includes both understanding and generation paradigms, offering a more balanced assessment of a model's capabilities. The datasets within General-Bench are derived from diverse sources, including existing benchmarks, AI-generated content, and human annotation. To address concerns about potential data leakage, the authors mentioned maintaining an "open set" for models to use and a "closed set" to which external users do not have direct access, ensuring a fair comparison while mitigating the risk of models being overly optimized for specific benchmark data.
Experimental Setup & Results
▶ Watch: Criteria for advancing to higher General-Level tiers (8:00)
The experimental evaluation leveraging the General-Level framework and the General-Bench benchmark was conducted up to the end of 2024, involving a substantial effort to assess the capabilities of current multimodal AI models.
Models Evaluated: The researchers conducted an extensive evaluation of over 700 specialist models across the 700+ tasks defined in General-Bench. This provided a robust set of SOTA baselines against which generalist models would be measured. For the generalist category, more than 100 multimodal generalist models were tested.
General-Bench Scale and Coverage: The benchmark itself serves as a crucial component of the experimental setup, providing the diverse and comprehensive testing ground. As detailed previously, General-Bench spans 29 domains, encompasses 5 major modality types, includes over 145 distinct modality skills, and features more than 700 tasks with a total of 300,000+ samples. Critically, it incorporates both multimodal understanding and generation paradigms, ensuring a holistic evaluation that extends beyond mere comprehension.
Metrics and Baselines: The evaluation employed the scoring mechanisms defined by the General-Level system. For Level 2, average performance across tasks and modalities was the metric. For Level 3 and above, the primary metric shifted to outperforming SOTA specialist models on individual tasks, with the "specialist models" serving as crucial baselines. Level 4 introduced the harmonic mean to ensure balanced performance across understanding and generation. Level 5's metric was the ability to boost SOTA NLP specialists, indicating deep cross-modal synergy.
Headline Results and Ablations:
The results provided a sobering assessment of the current state of multimodal generalists:
- Level 2 Performance: At Level 2, the initial evaluation revealed that many widely recognized multimodal models ranked lower than expected. This was primarily attributed to their insufficient coverage of the broad range of tasks or modalities included in General-Bench. The speaker noted that newer versions of models like GPT would likely perform much better, implying that ongoing rapid development could shift these rankings.
- Level 3 Drop-off: A significant drop-off occurred at Level 3. Most models that appeared strong at Level 2 disappeared from contention. This stark reduction highlighted the difficulty of meeting the Level 3 criterion, which demands outperforming specialists and thus demonstrating true task-level synergy. It underscored that simply handling multiple modalities does not equate to genuine generalization.
- Level 4 Scarcity: The challenge intensified at Level 4, where only three models out of the more than 100 evaluated generalists showed true understanding-generation synergy. This indicates a severe lack of models capable of robustly integrating and leveraging both comprehension and generative capabilities across modalities in a balanced and mutually supportive manner.
- Level 5 Unattained: Most notably, no models were found to have reached Level 5. This signifies that current multimodal generalists have not yet achieved the ultimate goal of full cross-modality synergy, where non-linguistic modalities significantly boost the performance of SOTA NLP tasks.
Deep Analysis Observations: Beyond the hierarchical level scores, a deeper analysis of the results revealed several overarching trends:
- High Specialization: The majority of evaluated models are heavily specialized, demonstrating proficiency in narrow domains rather than exhibiting broad generalist capabilities. They are far from the ideal of a true generalist AI.
- Vision-Centric Bias: Modality support across the models is highly unbalanced. The vast majority of models are centered around vision, specifically image processing, with other modalities receiving comparatively less attention or integration.
- Understanding Dominance, Generation Weakness: While many models excel at multimodal understanding, their performance in generation across diverse tasks and skills is notably uneven and inconsistent. This points to a significant gap in their generative capabilities when compared to their comprehension abilities.
- Limited Synergy: Overall, the ability to demonstrate synergy across different paradigms (e.g., understanding-generation) and across various modalities remains a major weakness, reinforcing the findings from the General-Level tiered evaluation.
These experimental results provide concrete evidence that despite the rapid advancements in multimodal AI, the field is still in its nascent stages regarding true generalist capabilities and synergistic intelligence.
Practical Implications
▶ Watch: General-Bench scale: 29 domains, 700+ tasks, 300k+ samples (9:40)
The findings from the "On Path to Multimodal Generalist" talk carry profound practical implications for various stakeholders in the AI/ML ecosystem, influencing how models are built, evaluated, and deployed.
For Practitioners and Model Builders:
The most significant implication is the urgent need to rethink model design and evaluation strategies. The current focus on narrowly beating benchmarks or achieving marginal accuracy gains on specific visual tasks is revealed as insufficient and potentially misleading. Model builders should shift their efforts from creating highly specialized, bolted-on modules to developing architectures that inherently foster cross-modal generalization and synergy. This means designing models where knowledge learned from one modality or task can genuinely enhance performance in others, rather than operating as independent silos. The General-Level framework provides a clear roadmap for this, urging developers to aim for higher tiers by improving broad task coverage, ensuring balanced understanding and generation, and integrating modalities more deeply. The "unfair" comparison against specialists, where generalists are not allowed task-specific fine-tuning, is a deliberate design choice to push for true generalization, suggesting that future generalist models should be robust enough to perform well on diverse tasks out-of-the-box.
For Infrastructure Teams and Deployers:
The insights from General-Bench highlight the need for infrastructure capable of supporting diverse modalities and paradigms. If models are to achieve higher General-Level tiers, they will require training and inference environments that can efficiently handle 29 domains, 5 major modality types, 145+ skills, and 700+ tasks. This implies significant demands on data pipelines, specialized hardware (e.g., for different sensor inputs), and robust software stacks that can manage complex multimodal inputs and outputs. The finding that most models are vision-centric and primarily support understanding suggests that current deployment strategies might be similarly skewed. Future infrastructure must be prepared for models that exhibit balanced understanding and generation capabilities across a broader spectrum of modalities, moving beyond image and text. The inconsistency in generation performance also implies a need for more robust and reliable generative output handling in deployed systems.
Tradeoffs and Limitations:
This work deliberately introduces a challenging evaluation paradigm, which inherently presents tradeoffs. A key tradeoff is the "unfair" comparison between generalists and specialists. While specialists, being fine-tuned on specific tasks, will often retain SOTA performance in their narrow domains, the General-Level framework deliberately penalizes generalists for not outperforming specialists without task-specific fine-tuning. This is a necessary design choice to push for true general intelligence, but it means that for certain critical, high-stakes applications, a specialized model might still be the optimal choice for maximum performance in its niche.
Another practical implication concerns data leakage. The talk acknowledges this challenge, mentioning the use of both "open" and "closed" sets within General-Bench. This highlights the ongoing tension between providing benchmarks for community development and preventing models from simply overfitting to the test data. For practitioners, this means that while the open sets can guide development, achieving high scores on the closed sets (or in real-world, unseen scenarios) will be the true test of generalizability.
Ultimately, the project serves as a wake-up call, indicating that despite the "boom" in multimodal LLMs, the AI community is still far from building a true "modality generalist" akin to the powerful AIs envisioned for AGI. The current generation of models is heavily specialized, vision-centric, and lacks the deep, synergistic intelligence across modalities that this framework demands. This necessitates a fundamental re-evaluation of research priorities and investment in developing more integrated and broadly capable multimodal architectures.
Key Takeaways
- Current evaluations for multimodal foundation models are largely insufficient, focusing on narrow accuracy gains rather than true generalization and synergy across tasks and modalities.
- The General-Level system introduces a novel, five-tier hierarchical framework to rigorously assess multimodal intelligence, pushing models beyond simple task performance to demonstrate deep cross-modal generalization.
- General-Bench is an expansive, comprehensive benchmark featuring 29 domains, 5 major modality types, 145+ skills, 700+ tasks, and over 300,000 samples, covering both understanding and generation paradigms, designed to facilitate this advanced evaluation.
- Experimental results reveal that current multimodal generalist models are predominantly specialized, vision-centric, and primarily focused on understanding, with uneven and inconsistent generation capabilities.
- A critical finding is the pervasive lack of true cross-modal synergy: no models have reached Level 5 (full modality synergy), and only three achieved Level 4 (understanding-generation synergy), indicating a significant gap on the path to AGI.
- The project provides a much-needed, reliable, and long-term benchmark to guide the multimodal AI community, prompting a fundamental shift in model design and evaluation towards truly integrated and synergistic generalist AI.
About the Speaker(s)
The talk was presented by Tianjie Ju, representing a large collaborative team of co-authors from numerous different institutions. While specific individual bios were not detailed in the transcript, the collective effort underscores a broad academic and industrial collaboration. The overarching goal of this diverse team is to establish a reliable, long-term benchmark for the multimodal AI community, thereby guiding the development of future general-purpose multimodal foundation models and advancing the collective journey towards achieving Artificial General Intelligence (AGI).
Reviews
Maya Iyer (Theoretical ML Researcher) — WEAK
This talk introduces General-Level, a five-tier evaluation hierarchy for multimodal foundation models, and General-Bench, a large-scale benchmark covering 700+ tasks across multiple modalities. The ambition is real and the community need is genuine — current evaluations are narrow and reward specialization over generalization. But the core theoretical apparatus is thin. The central notion of 'synergy' is operationalized in a way that collapses the concept rather than formalizes it, the tiered system conflates distinct properties without principled justification, and the benchmark construction raises standard data-leakage and saturation concerns that are acknowledged but not resolved. This…
Chen Zhao (Applied ML Researcher & Empiricist) — SOLID
General-Level and General-Bench is a well-motivated benchmark paper that identifies a real problem — the multimodal evaluation ecosystem is fragmented, vision-centric, and systematically ignores generation — and proposes a structured five-tier framework plus a large-scale benchmark to address it. The scope is genuinely impressive (700+ tasks, 300K+ samples, both understanding and generation), and the core finding — that essentially no current generalist model demonstrates true cross-modal synergy under a strict no-fine-tuning protocol — is a useful and probably correct diagnosis of the field. However, as reviewed from the article summary alone, the paper's experimental hygiene raises real…
→ Top-rated talks at International Conference on Machine Learning 2025
All talks from International Conference on Machine Learning 2025