~ similar to 2607.06560· 18 results
Sara Dorfman, Maya Vishnevsky, Omer Dahary, Or Patashnik +1 more
This paper introduces a method for controlled diversity in text-to-image models, enabling semantic browsing and creative exploration through structured image galleries.
Tianjiao Yu, Xinzhuo Li, Yifan Shen, Onkar Susladkar +3 more
The paper introduces ELSA3D, a unified 3D model that uses elastic semantic anchoring to improve interaction between text and 3D representations, achieving state-of-the-art performance with reduced FLO…
Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao +13 more
Introduce PerceptionRubrics, a rubric-based evaluation framework for addressing real-world brittleness of models using 1,038 images and over 12,000 instance-specific rubrics.
Fangzhou Lin, Peiran Li, Lingyu Xu, Wenjing Chen +11 more
The paper introduces CV-Arena, a large-scale open benchmark for instructional computer vision, demonstrating that professional-grade image editing requires advanced capabilities in physical reasoning…
Chong Bao, Shichen Liu, Lijun Yu, David Futschik +8 more
The paper introduces Archon, a unified, fully pretrained multimodal model that addresses the challenge of generating holistic digital humans by integrating seven modalities (including text, audio, mot…
Yang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun +4 more
The paper proposes CSMR, a cognitive scheduling framework that allows a language model to dynamically decide when to acquire task-relevant visual evidence, significantly improving multimodal reasoning…
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu +2 more
The paper proposes VLM3, a simple, scalable method that demonstrates standard Vision Language Models (VLMs) can natively learn 3D understanding by focusing on architectural simplicity and specific dat…
The paper introduces Staged Executable Inverse Graphics (SEIG), an agentic framework that uses general-purpose Vision-Language Models (VLMs) to reconstruct editable 3D scenes directly into executable…
Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu +8 more
Lumos-Nexus is a training-efficient framework that enhances video generation quality by progressively bridging generation from a lightweight model to a high-fidelity generator in a shared latent space…
The paper analyzes token reduction for efficient unified VLM training, finding that while task-specific acceleration saves computation, it destroys the mutual performance gains achieved through joint…
Kaixiong Gong, Xin Cai, Bin Lin, Hao Wang +8 more
This paper proposes Twins, a unified continuous token space for multimodal models using ViT and VAE features, and addresses optimization imbalance with a focal regression objective.
This paper introduces a method for generating diagnostic saliency maps for vision-language models using transformer models, revealing how the models allocate focus across visual elements during answer…
Wenjia Jiang, Zongyuan Cai, Yuanhang Shao, Chenru Wang +6 more
This paper introduces ManimAgent, a self-evolving multimodal agent that retains reflection experience across tasks using an Episodic Memory Bank, leading to improved performance on a code-generation t…
Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang +10 more
The paper introduces SceneActBench, a benchmark for evaluating vision-language model agents' ability to perform actions on multi-object 3D scenes.
The paper introduces UniKE, a benchmark showing that successful knowledge edits in text-only multimodal models do not reliably transfer to image generation, revealing a significant modality gap.
The paper introduces GPIC, a massive, permissively licensed, and safety-filtered image corpus of 28 trillion pixels, designed to serve as a stable and accessible benchmark for large-scale visual gener…