ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers8 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(3)

Learning to Read the Contextual Tokens in Diffusion Transformers

Oct 5, 2026

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al.

Text tokens in image-generating transformers develop interpretable semantic representations of the emerging image that can be read with an LLM probe—and explicitly training to strengthen these representations improves generation quality.

This paper reveals what text tokens learn during image generation in multimodal diffusion transformers. Researchers built a tool to 'read' these hidden representations by connecting them to a language model, discovering they encode rich scene information early in generation. They then used these insights to improve image quality through a new training technique.

multimodaltraining

Recursive Video In-Context Learning for Agentic Robot

Oct 5, 2026

Wenrui Bao, Xinxin Liu, Bingxin Xu et al.

By organizing demonstration videos into navigable hierarchies rather than static prompts, agents can access task details only when needed, reducing context overhead while improving learning from single examples.

This paper presents Recursive Video In-Context Learning (RV-ICL), a method that helps robot agents learn from demonstration videos more effectively. Instead of feeding entire videos as prompts, RV-ICL organizes a single demo video into a hierarchical structure of sub-events (like grasps and releases) that the agent can navigate on-demand.

Sep 28 – Oct 4(15)

VISTA: A Visual Harness for Reasoning in an Interactive World

Oct 1, 2026

Qiushi Han, Keya Hu, Linlu Qiu et al.

Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.

VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.

agentsmultimodalreasoning

Generative Cinematographer: Composing Camera and Object Motion in 3D

Oct 1, 2026

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.

By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.

GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.

Sep 21 – Sep 27(11)

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

Sep 25, 2026

Itzel Tlelo-Coyotecatl, Hugo Jair Escalante

Most hate speech detection datasets focus on English; this dataset enables models to learn culturally-specific patterns in Mexican Spanish video content, which is essential since hate speech is deeply tied to local context and language nuances.

MexHat is a new video dataset with ~1,000 annotated clips for detecting hate speech in Mexican Spanish. It addresses the lack of non-English resources by capturing linguistic and cultural context specific to Mexico, with annotations for both general categories (offensive vs. hate speech) and fine-grained hate speech subtypes.

datamultimodalsafety

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Sep 25, 2026

Kunxiong Zhu, Zhihao Shu, Hangyu Zheng et al.

Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.

EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.

Sep 14 – Sep 20(11)

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Sep 17, 2026

Kevin Qu, Tao Sun, Massimiliano Viola et al.

By processing multiple sparse observations together rather than single images, and training with a motion-range supervision objective, the model better understands articulation without needing large labeled datasets—it generates synthetic training data procedurally.

FAMOS is a feed-forward model that predicts how articulated objects (like doors, drawers) move and which parts are movable from multiple sparse 3D views.

architecturemultimodal

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Sep 17, 2026

Ji Xie, Dewei Zhou, Xinyu Huang et al.

By combining real image supervision with pure-color anchors and a shared hex-prompt interface, you can now control exact object colors in both image generation and editing tasks with professional-grade precision.

Paint-Anything enables precise color control in AI image generation and editing by letting users specify exact colors using hex codes (like #FF5733). The system learns to understand hex values through a training dataset of 500K images with color labels and pure-color reference images, achieving much better color accuracy than existing methods.

Sep 7 – Sep 13(11)

Can Edge-Deployable Vision-Language Models Identify Species?

Sep 10, 2026

William Zhou, Mayukha Siripuram, Xiao Yan et al.

Small deployable VLMs can identify species above chance but degrade sharply on real camera-trap images; specialized training data outweighs model scale, and all models occasionally hallucinate fake species names when uncertain.

This paper evaluates whether small vision-language models (2-8B parameters) suitable for edge deployment can accurately identify animal species from camera-trap photos.

evaluationmultimodalapplications

TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

Sep 10, 2026

Akshaj Gupta, Hwi Joo Park, Andrea Guzman et al.

This is the first guitar transcription system that produces both fingering positions and expressive techniques directly from audio, achieving significant improvements over prior work and handling real-world noisy recordings.

TART is a four-stage system that converts guitar audio into tablature (written notation showing which strings and frets to play). It tackles three real problems: recognizing guitar techniques like slides and bends, correctly identifying which string-fret combinations produce each note, and handling noisy real-world recordings.

Aug 31 – Sep 6(9)

UniMate: One Unified Model to Animate Diverse Skeletons

Sep 4, 2026

Linzhan Mou, Jiahui Lei, Zhiyang Dou et al.

For the first time, a single model can animate any skeleton topology (bipedal, quadrupedal, insects, etc.) from text alone—no per-character fine-tuning or reference motions needed at inference time.

UniMate is a foundation model that generates realistic motion for any 3D character skeleton from text descriptions, without needing to retrain for each new character type. It uses a specialized neural architecture that understands skeleton structure through graph-based attention mechanisms, and was trained on a diverse dataset of 13,000+ motion sequences across different creature types.

multimodalarchitecture

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

Sep 4, 2026

Ji Soo Lee, Xilun Chen, Pierce Chuang et al.

Most LLMs struggle with real-world wearable health reasoning (19.6%-72.9% accuracy), revealing a significant gap between general language abilities and the specialized reasoning needed to interpret longitudinal physiological data.

WearableQA is a benchmark with 4,084 multiple-choice questions testing whether AI models can reason about real health data from wearables. It uses 500 days of measurements from 200 actual users, including heart rate, sleep, and blood tests, organized into 16 question types that test both data computation and health interpretation skills.

Aug 24 – Aug 30(1)

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Aug 27, 2026

Kechen Liu, Ola Shorinwa

By training on cross-embodiment video data with a unified physics understanding, CLAP creates robot world models that match or beat single-embodiment models and generalize to new robots without retraining.

CLAP is a video world model that learns physics from diverse videos of humans and robots by treating physical laws as universal. It solves the challenge of different action representations across embodiments using end-effector poses, language, and learned latent actions, then applies a curriculum approach to train models that work zero-shot on real robot tasks.

multimodal

Aug 17 – Aug 23(11)

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

Aug 21, 2026

Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor et al.

State-of-the-art vision-language models fail at interpreting real scientific artifacts that domain experts find straightforward, highlighting a critical gap between general image understanding and specialized scientific visual reasoning needed for practical biotech applications.

VIALS is a benchmark of 161 visual question-answering tasks based on real scientific images (gel blots, microscopy, flow cytometry plots, etc.) from biotech workflows. Current vision-language models struggle with these domain-specific images despite excelling at natural images, revealing gaps in scientific reasoning that limit their usefulness in professional life sciences research.

evaluationmultimodalapplications

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Aug 21, 2026

Haonan Jia, Shichao Dong, Zenghui Sun et al.

Retrieval can guide reinforcement learning to improve vision-language models for image captioning by helping identify and correct errors, outperforming standard supervised fine-tuning approaches.

This paper proposes Re³Cap, a method that uses retrieval-guided reasoning to improve image captioning with reinforcement learning. Instead of just fine-tuning models, it retrieves similar images and captions to help identify and fix errors (hallucinations and omissions) in generated descriptions, achieving better results than supervised fine-tuning without needing extra labeled data.

Aug 10 – Aug 16(15)

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Aug 13, 2026

Bobo Li, Hao Fei, Tianjie Ju et al.

Direct perception of raw scientific data—not just text summaries—is critical for AI systems to conduct rigorous, evidence-grounded research. OmniScientist shows that multimodal input improves all aspects of automated scientific discovery.

OmniScientist is an AI system that conducts scientific research across multiple disciplines by directly processing raw data in many formats—images, videos, audio, 3D structures, tables, and more.

multimodalagentsreasoning

Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology

Aug 13, 2026

Yunsung Chung, Yingshuo Liu, Abboud F. Hassan et al.

Clinical prediction improves when models treat recovery as an evolving process rather than a static snapshot—incorporating asynchronous post-procedure events and imaging can significantly boost outcome forecasting accuracy for cardiac interventions.

This paper presents a clinical AI model that predicts post-surgery outcomes in heart rhythm procedures by tracking how a patient's condition evolves over time.

multimodal

Aug 3 – Aug 9(11)

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Aug 7, 2026

Youjun Zhao, Alex Warren, Gary K. L. Tam et al.

Mirror reflection generation requires modeling two distinct challenges—semantic consistency (what reflects) and geometric accuracy (how it's arranged)—which can be addressed by distilling relational knowledge and learning spatial transformations.

MirrorWorld tackles the problem of generating realistic mirror reflections in videos by teaching diffusion models to understand what scene content should appear in mirrors and how it should be spatially arranged. The method uses semantic guidance from visual foundation models and geometric transformation learning to ensure reflections are consistent with their surroundings.

architecturemultimodaltraining

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

Aug 7, 2026

Zixuan Lan, Luzhe Sun, Matthew R. Walter et al.

VLMs have a critical weakness: they often trust learned world knowledge over what's actually shown in images, and SABRE provides a reusable framework to systematically identify and measure such failures.

SABRE is an automated pipeline that creates stress tests for vision-language models by converting task designs into images and question-answer pairs. It tests whether VLMs rely on visual evidence or learned assumptions about the world, revealing that current models struggle significantly (17.8-31.3% accuracy) when images contradict expectations.

Jul 27 – Aug 2(2)

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Jul 31, 2026

Senyu Fei, Xiaopeng Yu, Siyin Wang et al.

By adding explicit world modeling to the critic function, WCM helps robot learning systems better understand how observations change over time, leading to more reliable value estimates and improved performance on both seen and unseen tasks.

This paper introduces World Critic Model (WCM), a new approach for training robot control systems that combines vision, language, and action learning. The key innovation is having the critic (value estimator) explicitly learn to predict future states alongside estimating action values, rather than just predicting scalar returns.

multimodal

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

Jul 31, 2026

Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino et al.

Multimodal AI models can match human performance on social inference tasks, but they use different reasoning strategies and don't benefit from visual information the way humans do, suggesting gaps in how they process social cues.

FriendBench is a benchmark that tests whether AI models and humans can tell if two people already know each other or are strangers by watching a 20-second conversation clip. Researchers compared 26 AI models against human judges across text, audio, and video formats.

agentsreasoningmultimodal

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Oct 5, 2026

Yaohui Zhang, Binxu Li, Haoyi Duan et al.

Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.

PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.

evaluationmultimodaldata
multimodalarchitectureapplications

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Oct 1, 2026

Hanchu Zhou, Dechen Gao, Hang Wang et al.

Using semantic communication between robots—where they exchange meaningful descriptions rather than sensor data—enables better coordination on long-horizon tasks while keeping each robot's execution independent and reliable.

DuoMind is a framework that enables multiple robots to coordinate and work together on complex tasks by combining vision-language models for high-level reasoning with vision-language-action models for precise execution. Robots communicate through semantic messages rather than raw data, allowing them to share understanding of the task and environment while maintaining independent control.

agentsmultimodalreasoning

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Oct 1, 2026

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al.

Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.

This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.

trainingmultimodaldata

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Oct 1, 2026

Yakun Zhu, Yi Bin, Yujuan Ding et al.

Separating geometry into common and residual components while routing visual learning through structured latents prevents representation collapse and improves 3D spatial reasoning from images.

This paper addresses 3D spatial reasoning from 2D images by introducing GeoLatent, which uses decomposed spatial representations (position, direction, geometry) with geometric supervision and routed optimization. The method prevents geometry collapse and ensures latents are actively used during learning, achieving state-of-the-art results on spatial reasoning benchmarks.

reasoningmultimodalarchitecture

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Sep 30, 2026

Tian Xia, Minghao Liu, Yiqing Liang et al.

For imbalanced clinical tasks, optimizing prompts for ranking metrics (AUROC) instead of accuracy can dramatically improve model performance—up to 16 percentage points—because accuracy-based optimization fails when one class dominates the data.

This paper addresses class imbalance in clinical diagnosis by optimizing multimodal language models for AUROC instead of accuracy. The authors introduce Ranking-PE, a prompt optimization method that evaluates candidate prompts based on how well they rank positive cases above negative cases, rather than raw correctness.

evaluationmultimodalapplications

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Sep 30, 2026

Xinghao Chen, Xiangbo Gao, Jiongze Yu et al.

Video text editing requires balancing three competing goals: correct text, smooth motion, and unchanged background. This benchmark and dataset help measure those trade-offs and establish baselines for the community.

ViTeX-Bench is a benchmark for video scene text editing—replacing text on surfaces like signs and labels while keeping the rest of the video unchanged. It includes 387 real videos, evaluation metrics for text accuracy and visual quality, and a baseline editor that achieves strong results. This addresses a gap where video text editing lags behind image editing.

evaluationmultimodalapplications

Image Classifiers are Efficient Self-Supervised Video Representation Learners

Sep 30, 2026

Owais Iqbal, Sudipta Sarkar, Shyam Marjit et al.

You can efficiently learn video representations by repurposing pretrained image models with clever masking strategies, avoiding expensive 3D architectures and reconstruction overhead.

VideoMSN uses standard image Vision Transformers to learn video representations without 3D models or reconstruction. It treats videos as grids of frames and masks either spatial patches or temporal frames, then aligns the two views using a Siamese loss. This achieves top results on video benchmarks while needing 32-160x fewer training epochs than prior methods.

efficiencymultimodal

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Sep 30, 2026

Ziyan Jiang, Jingbo Yang, Jiabao Ji et al.

Multimodal agents struggle to couple exploration and visual reasoning in 3D worlds—they can see anomalies or navigate, but struggle to do both together effectively, suggesting a fundamental gap in how these systems integrate action and perception.

WorldAuditBench is a benchmark for testing how AI agents find problems in 3D virtual worlds—like floating objects or walls you can walk through. It evaluates multimodal AI systems (vision-language models and vision-language-action models) on 213 anomaly detection tasks across 13 environments, measuring how well agents can explore systematically and visually identify issues.

evaluationmultimodalagents

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

Sep 30, 2026

Anson Y. Lam, Shuqing Li, Michael R. Lyu

Text-to-material generation can be more effective and controllable by outputting human-readable programs rather than raw images—users get both the final material and the explicit rules that construct it.

MatLoom generates realistic materials from text descriptions by creating compact, readable programs that define how layers combine to form appearance and physical properties.

multimodalapplications

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Sep 29, 2026

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.

Training MLLMs to internally reconstruct 3D scene geometry (even in compact form) improves spatial reasoning and 3D understanding, suggesting that learning to imagine scenes is more effective than explicit geometric supervision alone.

This paper teaches multimodal AI models to reason about 3D scenes by first imagining a compact 3D representation before answering questions. Instead of relying on detailed geometric details, the model learns to assemble a coarse 3D layout from multiple viewpoints—similar to how humans understand 3D space—then uses this mental model to answer spatial reasoning questions more accurately.

multimodalreasoningarchitecture

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Sep 29, 2026

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

Factorizing attention across different data dimensions improves crop segmentation, but dataset construction choices (tile size, class definitions) have outsized impact on results and must be standardized for meaningful model comparisons.

This paper introduces PAtteRNS, a transformer-convolutional model for crop segmentation in satellite imagery that separately applies self-attention to temporal, spectral, and spatial dimensions. The authors also highlight critical dataset issues—flawed class groupings and incompatible tile-size variants—that undermine fair model comparison and suggest standardization is needed.

architecturemultimodalevaluation

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Sep 29, 2026

Hui Ren, Lei Fan, Henry Pao et al.

For long-video question answering, visually grounding object identity across time—not just retrieving relevant clips—is critical. GEB shows that organizing observations into entity biographies improves accuracy by 4+ points on day-long and week-long videos.

This paper solves a key challenge in long-video understanding: tracking the same physical object across hours or days of footage. The authors introduce Grounded Entity Biographies (GEB), which groups visual observations of the same object into retrievable "biographies" that preserve context.

multimodalreasoning

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Sep 28, 2026

Yijia Fan, Ziqi Huang, Zhongang Cai et al.

Unified models can learn to self-correct their outputs by applying RL to complete reflection loops, where both the reasoning about what's wrong and the actual image fixes improve together without needing external verifiers.

This paper presents UMM-Reflection, a method that teaches unified multimodal models to critique and fix their own image generations through reinforcement learning. Instead of just generating images once, the model can now look at what it created, identify problems, revise the image, and repeat—all within a single model.

trainingreasoningmultimodal
efficiencymultimodalarchitecture

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Sep 24, 2026

Wenhao Li, Zhibin Wu, Chong Xiao et al.

Using LLM-generated semantics as a shared anchor point for aligning incomplete multimodal data is more robust than trying to reconstruct missing modalities or design complex fusion mechanisms.

SemMSA tackles multimodal sentiment analysis when some data is missing by using large language models to create rich semantic representations that ground all modalities together. Instead of reconstructing missing features, it aligns visual, acoustic, and text representations through spectral methods, achieving better results on standard benchmarks.

multimodalalignmenttraining

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

Sep 24, 2026

Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri

Audio language models have a significant blind spot: they verify written facts reliably but fail on identical claims when spoken. Retrieval-augmented approaches help only when combined with explicit reasoning, not retrieval alone.

VeriSpeak is a benchmark for fact-checking spoken claims using audio language models. It contains nearly 4,000 spoken statements about real-world facts and tests whether models can verify claims directly from speech, especially when given retrieved text evidence.

evaluationmultimodalreasoning

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

Sep 24, 2026

Yuncong Yang, Jinlong Li, Yulong Xue et al.

Object-centric predictive models trained on multi-view video can enable underwater robots to plan manipulation tasks by imagining future states, without requiring expensive contact sensors or reconstruction.

This paper presents Underwater C³-JEPA, a world model that predicts how objects move during underwater robot salvage operations. Using multiple camera views and robot control signals, it learns to predict future object positions in a compact latent space without needing contact sensors. The model transfers well to downstream tasks and works on real underwater video.

multimodal

The Alignment Illusion in Multimodal Large Language Models

Sep 24, 2026

Hong-Han Wang, Yuntao Wang, Hu Ding

Don't trust standard alignment scores as proof that multimodal models understand images—they can be fooled by the model's internal structure. Use task-based validation and the PA gap metric instead to verify genuine cross-modal integration.

This paper reveals that high visual-text similarity scores in multimodal AI models don't actually mean the model is properly integrating images and text. Using experiments across 13 models, the researchers show that replacing images with random noise barely changes similarity scores, even though task performance drops sharply.

evaluationmultimodal

Do Audio Language Models Hear and Read Distinctive Features Alike?

Sep 24, 2026

Yuanhao Chen, Peter Chin

Audio language models don't reliably represent the same phonetic features in the same direction across speech and text—suggesting these models may process the two modalities quite differently despite using a shared decoder.

This paper investigates whether audio language models represent phonetic features consistently across speech and text inputs. Using minimal pairs of phonemes that differ in single features, researchers measured whether the same distinctive features (like voicing) are encoded in the same direction in both modalities across 6 models, 7 features, and 15 languages.

multimodalevaluationarchitecture

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Sep 21, 2026

Yiran Wang, Xingyilang Yin, Junfu Pu et al.

This benchmark reveals that gameplay requires coordinating visual understanding, instruction following, and long-horizon planning—and shows clear performance gaps between model families, providing a standardized way to measure progress on embodied AI tasks.

GameHorizon is a comprehensive dataset and benchmark for evaluating AI models on video game tasks. It includes 5,000 hours of gameplay from 21 AAA games with aligned videos, actions, and instructions at multiple time scales, plus offline and online evaluation tracks to measure how well models understand and execute gameplay across different planning horizons.

evaluationmultimodalreasoning

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Sep 21, 2026

Wangbo Yu, Kunhao Liu, Wenbo Hu et al.

By storing observations in a viewpoint-aware implicit 3D memory rather than explicit depth maps, video world models can generate longer, more consistent videos across different camera angles without running out of token budget.

WorldCrafter is a video world model that maintains consistent 3D-aware memory across different camera viewpoints and long time horizons. It uses an implicit memory system that compresses multi-view observations into tokens optimized for the requested viewpoint, enabling realistic minute-long video generation from a single image or text prompt while respecting what was seen before.

architecturemultimodalreasoning

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Sep 21, 2026

Haoran Yuan, Zekai Wang, Boning Shao et al.

Tactile sensing matters for dexterous manipulation: modeling contact evolution as part of world state (not just as input) nearly triples performance, and pretrained vision models can efficiently extend to touch with minimal data.

DexTacWAM combines vision and touch sensing for robot hand manipulation by extending video prediction models with tactile data from fingertips. The system predicts both visual and contact dynamics together, achieving 70% success on complex multi-finger tasks—nearly double the best existing approach—while learning from minimal touch data by adapting a pretrained vision model.

multimodal

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Sep 21, 2026

Muzhe Wu, Zuchen Li, Xu Wang et al.

Contextualizing visual instructions to match a user's real workspace—through AI-generated images and videos—improves task performance and trust compared to generic pre-authored tutorials.

This paper presents a system that generates live visual instructions tailored to a user's specific workspace and task progress, rather than showing pre-recorded tutorials. Using AR and generative AI, it creates goal images and demo videos that match the user's actual environment, helping them complete physical tasks more accurately and confidently.

applicationsmultimodalevaluation
multimodal
training
applications

JEPA-Anything: Learning Predictive Models across Different Worlds

Sep 17, 2026

Taoyong Cui, Zhongyao Wang, Xinyue Xu et al.

A single learning principle based on factorized predictions can work across radically different domains, suggesting world modeling doesn't need domain-specific architectures—and can even guide real scientific discovery.

JEPA-Anything is a unified framework for building predictive models across completely different domains—from videos to molecules to weather—using a technique called orthogonal predictive factorization. Instead of training separate models for each domain, it learns to decompose predictions into independent factors that work across vision, biology, physics, and clinical data.

architecturereasoningmultimodal

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

Sep 17, 2026

Tica Lin, Deepak Chandran, Gauri Jagatap et al.

A shared, human-readable schema can simultaneously ground agent generation and enable human interpretation, making AI-generated content more transparent and controllable.

This paper introduces a semantic action graph—a structured representation of sports matches using performers, actions, recipients, moments, and states—that enables both AI agents to generate video highlights and humans to understand and control them.

agentsmultimodalapplications

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Sep 17, 2026

Hanchu Zhou, Brendan Lynch, Raman Goyal et al.

By treating vision and tactile signals at different timescales in a lightweight architecture, you can build fast tactile-aware robot controllers (11.9ms latency) that outperform larger models on real contact-rich tasks.

This paper presents Agile-WAM, a tactile world action model that predicts future robot states and actions for contact-rich manipulation tasks.

architecturemultimodalefficiency

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Sep 17, 2026

Haocheng Xi, Yiming Xie, Hexu Zhao et al.

Hybrid attention architectures that combine local Softmax with linear memory can dramatically accelerate video generation without sacrificing quality—enabling practical real-time video synthesis on modern hardware.

Video DeltaNet combines local attention with efficient linear memory to speed up video diffusion models. By mixing Softmax attention for fine details with a new Video Delta Attention mechanism for long-range context, it generates high-quality videos 14.5x faster than baseline models while maintaining visual quality.

efficiencyarchitecturemultimodal

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

Sep 17, 2026

Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi et al.

Multilingual healthcare AI needs explicit testing for how language, communication style, and missing information affect risk assessment—aggregate accuracy scores alone don't catch safety failures that emerge in specific languages or contexts.

HerHealthEval is an evaluation framework that tests how well AI language models understand women's health concerns across multiple languages (English, French, Arabic) and communication styles (clinical, casual, emotional, vague).

evaluationmultimodalsafety

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Sep 16, 2026

Sara Pieri, Evangelos Kazakos, Shizhe Chen et al.

Combining dense captioning with pixel-level grounding requires jointly optimizing text generation and mask selection—PANORAMA shows this can be done effectively by conditioning a segmenter on phrase representations and learning which masks correspond to each phrase.

This paper introduces PANORAMA, a vision-language model that generates detailed image captions while simultaneously grounding each phrase with pixel-level segmentation masks.

multimodalevaluationarchitecture

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Sep 16, 2026

Guanhua Ji, Tianyu Li, Dayoon Suh et al.

Combining generated video with generated audio allows robots to infer not just motion but also the forces needed for contact-heavy tasks—something video alone cannot provide.

This paper shows how robots can learn contact-rich manipulation tasks by generating both video and audio together. The system uses the loudness of generated contact sounds to create force profiles that guide the robot's movements, enabling tasks like pushing and grasping that require precise force control. The approach works without task-specific training data.

agentsmultimodaldata

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

Sep 15, 2026

Chuhao Chen, Peter Wonka, Chaoyang Wang et al.

This model enables mid-generation control of video generation using physics-grounded motion signals rather than pixel positions, making it possible to interactively steer object dynamics while the video is being generated.

PhysStream generates videos from images while letting you control object motion in real-time using physics-based signals. Instead of specifying exact positions, you provide velocity changes that the model interprets as physical dynamics. It uses scene memory (tracking where objects are) to maintain consistency across frames, enabling interactive control over multi-object scenes during generation.

multimodal

Tables Decoded: DELTA for Structure, TARQA for Understanding

Sep 15, 2026

Jahanvi Rajput, Dhruv Kudale, Saikiran Kasturi et al.

Processing tables as structured text instead of images makes them easier for LLMs to understand, scales better across languages, and achieves better performance on table QA tasks without needing visual encoders.

This paper presents DELTA and TARQA, a two-stage approach for table understanding that converts table images into structured text (OTSL format) rather than relying on vision-language models. DELTA handles table structure recognition and OCR, while TARQA is an LLM fine-tuned to answer questions about tables in this text format.

multimodal
applicationsmultimodaltraining

MindTopo: Can Foundation Models Reason in Topological Space?

Sep 10, 2026

Yunfei Ge, Anbang Liu, Qineng Wang et al.

Foundation models excel at metric spatial reasoning (distance, angles) but fail at topological reasoning—a more fundamental aspect of spatial understanding. This gap persists even with fine-tuning and reinforcement learning, suggesting current architectures lack core spatial intuition.

MindTopo is a benchmark testing whether AI models can reason about topological properties—spatial relationships that stay the same under stretching or bending. It evaluates 14 multimodal AI models on five topological concepts (continuity, separation, order, enclosure, knots) through reasoning tasks and embodied planning.

evaluationreasoningmultimodal

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Sep 10, 2026

Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi

Arabic speech AI was severely underrepresented in multilingual models—this project provides the dataset, trained models, and evaluation tools needed to build general-purpose Arabic speech systems from scratch.

Nuha-Speech addresses the lack of Arabic speech AI models by creating a comprehensive framework: a 1.5M-sample Arabic speech question-answering dataset, fine-tuned models based on Qwen-Omni, and evaluation benchmarks. This work establishes foundational infrastructure for building Arabic-capable speech AI systems despite limited available resources.

multimodaldataevaluation

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Sep 10, 2026

Bowen Zhang, Hsiu-Wen Cheng, Hongyu Yang et al.

Foundation models for time-series forecasting need task-specific fine-tuning to work well for glucose prediction, and multimodal dietary data (food images + nutrition) provides clinically useful signals that significantly improve postprandial glucose forecasting.

This study evaluates whether time-series foundation models can forecast blood glucose levels from continuous glucose monitoring (CGM) data, and whether adding dietary information improves predictions.

evaluationmultimodalapplications

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

Sep 10, 2026

Lisa Bylinina

Visual grounding helps language models learn concrete object properties, but this benefit doesn't show up in standard benchmarks—suggesting we need better evaluation methods to measure what models actually learn from multimodal information.

This paper tests whether giving a small language model visual grounding for words (like showing it images of 'banana') before training helps it learn language better. The author finds that visual initialization leaves a lasting imprint on the model, especially for object-property knowledge, but this advantage is invisible to most standard benchmarks that test grammar and abstract reasoning.

multimodalevaluationtraining

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

Sep 8, 2026

Anqi Li, Yuxin Chen, Zhaobo Li et al.

End-to-end vision-language models can control complex humanoid robots for real-world navigation by learning whole-body coordination in simulation, eliminating the need for modular planning pipelines and real-world data collection.

TANGO is a vision-language model that enables humanoid robots to navigate cluttered indoor spaces by predicting full-body joint movements directly from natural language instructions and camera images. Unlike traditional 2D path planning, it coordinates arm placement, torso adjustment, and walking patterns to move through complex 3D environments.

agentsmultimodal

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Sep 8, 2026

Siting Li, Zhengyang Wang, Simon Shaolei Du et al.

When building multimodal models, image tokenizer choice matters not just for image quality but for how well it helps the model learn text and vision together—and the best tokenizer for one task may not be best for another.

This paper studies how image tokenizers work in multimodal AI models by building a controlled training setup that tracks how well different tokenizers perform across text, image generation, and image understanding tasks.

multimodaltrainingarchitecture

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

Sep 8, 2026

Tobias Susetzky, Raphael Rehms, Dmitrii Seletkov et al.

This is the first generative model to handle the full complexity of real patient data (multiple modalities, irregular timing, uncertainty), enabling both predictive tasks and counterfactual reasoning for personalized medicine.

NOAH is a generative AI model that learns from complete patient medical histories to predict future health outcomes. It processes diverse data types—images, vital signs, lab results, and clinical notes—across irregular time intervals, enabling doctors to forecast patient trajectories, simulate treatment scenarios, and identify disease risks.

multimodalreasoning

Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs

Sep 8, 2026

Xiaofu Chen, Stella Frank, Yova Kementchedjhieva

Vision encoders learn and store conceptual semantic information like canonical colors independently of visual input, and this knowledge is tightly linked to object recognition—providing a measurable way to understand what abstract concepts models actually learn.

This paper investigates whether vision encoders in vision-language models (VLMs) learn conceptual information beyond what's visible in images. Using canonical color (like 'bananas are yellow') as a test case, researchers found that models can decode object colors from grayscale images, suggesting they learn abstract semantic knowledge.

evaluationmultimodal

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

Sep 8, 2026

Yankai Fu, Ning Chen, Junkai Zhao et al.

Adaptive tactile fusion and joint visual-tactile prediction enable VLA models to handle contact-rich manipulation where vision alone fails, achieving 71% success on dexterous tasks.

DeCAL is a vision-language-action model for robot hands that combines visual and tactile (touch) sensing to perform complex manipulation tasks. It uses specialized AI experts working together to understand scenes, imagine future states, and generate actions—all while explicitly modeling contact dynamics that cause visual occlusions during dexterous manipulation.

multimodalagentsarchitecture
evaluationmultimodalreasoning

Reflection-aware Generative Novel View Synthesis

Sep 4, 2026

GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

You can generate realistic novel views of mirror scenes by treating reflections as virtual views and using gated attention mechanisms—no retraining needed, just clever use of existing diffusion models.

This paper presents Ref-GeNVS, a method for generating novel views of scenes containing mirrors without requiring additional training. The key innovation is treating mirror reflections as complementary views by estimating the mirror plane and reflecting camera poses, then using a two-stage approach with special attention mechanisms to ensure reflections stay consistent during generation.

multimodalarchitectureapplications

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Sep 4, 2026

Vivek Chavan, Pengtao Xie, Yahuan Shi et al.

Visual distractors cause robot policies to select wrong objects not because they lose manipulation skills, but because they struggle to ground attention on the correct target—a problem that can be fixed by explicitly training better target selection.

This paper identifies why robot learning policies fail when similar-looking objects are present, even though they work well normally.

evaluationagentsmultimodal

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

Sep 3, 2026

Ye-Chan Kim, Seunghee Choi, SeungJu Cha et al.

Using a VLM to detect actual visual transitions in videos produces better event localization and captions than pre-written transition descriptions inserted at fixed positions.

This paper tackles dense video captioning—describing multiple events in long videos—by using a vision-language model to intelligently detect transition moments between events rather than blindly inserting captions everywhere.

multimodalevaluationtraining

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Sep 3, 2026

Sixu Yan, Shikang Wang, Binhua Huang et al.

Decoupling physical grasp synthesis from task-dependent reasoning allows robotic systems to leverage foundation model improvements without retraining, and enables cross-hand generalization through explicit kinematic and stability constraints.

AdaRoboVLG combines vision-language models with robotic grasping by separating physical grasp synthesis from task understanding. A base policy generates and evaluates grasp candidates using kinematics and stability checks, while foundation models provide task-specific context.

applicationsmultimodalreasoning

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

Sep 3, 2026

Tingyu Song, Mingxin Li, Yanzhao Zhang et al.

Distilling ranking knowledge from a reranker into embedding models significantly improves compositional reasoning—the ability to distinguish scenes by attribute-object relationships—achieving 82.7% on compositional benchmarks while maintaining standard retrieval performance.

This paper improves how multimodal AI models understand complex visual scenes by teaching embedding models to better distinguish between scenes with the same objects but different arrangements.

multimodaltrainingevaluation

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Sep 2, 2026

Vasileios Baltatzis, Mert Inan, Connor Gillis et al.

Sign language translation needs discourse context to maintain spatial consistency and entity tracking across multiple sentences—not just translating each sentence independently.

DiscoSign translates English text to American Sign Language glosses while tracking discourse-level phenomena like spatial references and entity consistency. Traditional sign language systems work sentence-by-sentence, missing how entities and concepts maintain meaning across longer conversations.

multimodalapplicationsevaluation

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Sep 1, 2026

Haoyuan Deng, Haichao Liu, Wenkai Guo et al.

By explicitly modeling contact forces alongside visual and kinematic information, this system achieves 82% success on precision assembly tasks—a massive jump from 15% for existing approaches—showing that understanding physical interaction is critical for real-world robot manipulation.

Facet-0 is a robotic foundation model that learns to perform precise assembly tasks by predicting how its actions will affect contact forces with objects. It combines vision, language, and force feedback to generate robot movements that maintain sub-millimeter accuracy during assembly, trained on 1,000 hours of real robot data across multiple manufacturing setups.

multimodal
trainingmultimodalreasoning

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Aug 20, 2026

Shiao Xie, Siyu Chen, Jianwei Lv et al.

Medical AI needs dual optimization: factual correctness (verifiable through evidence) and patient communication quality (context-dependent). G-CARL shows that structured checklists paired with retrieval-based verification can train models for both simultaneously better than standard approaches.

This paper introduces a new task where AI systems explain medical reports to patients in accurate, accessible language. The key innovation is G-CARL, a training method that uses retrieval-based fact-checking and customized checklists to ensure explanations are both medically accurate and responsive to what patients actually want to know, without limiting creative variation in responses.

multimodalalignmentevaluation

A comparison between ceiling-mounted FMCW, IR-UWB and Wi-Fi radar for in-bedroom human activity monitoring and sleep interruption detection

Aug 20, 2026

Anton Lambrecht, Reda El Hail, Xianjun Jiao et al.

Different RF sensing technologies excel at different things: IR-UWB wins on accuracy, FMCW on generalization to new spaces. For healthcare monitoring, choose based on whether you prioritize activity recognition or robustness to environmental changes.

This paper compares three radar technologies (FMCW, IR-UWB, and Wi-Fi) for detecting human activities and sleep patterns from ceiling-mounted sensors. Using the same neural network and test conditions across 20 people and different room layouts, the study shows IR-UWB performs best overall (89% accuracy), while FMCW adapts better to new environments.

evaluationmultimodalapplications

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

Aug 20, 2026

Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli et al.

Large language models can match or exceed traditional machine learning for travel behavior prediction without task-specific training, and adding visual context from survey images improves performance—showing that multimodal AI can enhance behavioral modeling when integrated with human-centered ...

This paper presents a three-agent workflow that combines chatbot surveys, data processing, and prediction to model how weather affects commuter mode choices. The system collected 454 survey responses about travel preferences across different weather scenarios, then compared traditional statistical models with nine different large language models (2-35B parameters) for predicting travel behavior.

agentsapplicationsmultimodal

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Aug 20, 2026

Yu Chen, Ting Lei, Yaoyi Li et al.

Breaking down rule-following tasks into separate perception, rule interpretation, and planning steps helps multimodal models generalize better to new rules and constraints than end-to-end approaches.

This paper introduces RuleMaze, a benchmark for testing whether multimodal AI models can navigate mazes while following natural-language rules. The authors propose a method called Disentangled Multimodal Planning that breaks down the task into separate steps—understanding the scene, interpreting rules, and planning actions—to help models follow complex constraints more reliably.

reasoningmultimodalevaluation

Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

Aug 20, 2026

Mosharof Hossain, Md Rabiul Islam, Limon Halder et al.

Prompts can be integrated deep into image segmentation networks through channel-wise attention, rather than just at the end, making models better at finding anatomical structures across different medical imaging modalities.

This paper introduces PCCA, a mechanism that uses text prompts to guide how neural networks process medical images at multiple levels, improving segmentation accuracy across different body parts and imaging types. The method adaptively adjusts which image features matter most based on the prompt, achieving 10-23% improvements over standard approaches.

architecturemultimodalapplications

Decoding silent reading from non-invasive EEG

Aug 20, 2026

Ingo Marquardt, Anthilia Alchanat, Priyanka Jain

Silent reading produces decodable word-level information in EEG that improves with more training data—this opens a scalable path to studying how the brain processes language without relying on unreliable self-reports of inner speech.

Researchers decoded which words people were silently reading from brain activity (EEG), using 49 hours of data from one participant. They trained a neural network to match EEG signals with word embeddings from a language model, achieving above-chance accuracy on 240,000 word presentations.

evaluationmultimodaldata

DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing

Aug 20, 2026

Haoxiang Cao, Jiajiong Cao, Xuanpu Zhang et al.

By decomposing credit assignment across both pipeline stages and within the planner's reasoning trace, DARS enables more efficient training of complex multi-stage image editing systems compared to naive joint optimization.

DARS is a reinforcement learning system for training instruction-based image editing that uses a two-stage pipeline (planner + renderer). It solves the credit assignment problem—figuring out whether failures come from bad planning or bad rendering—using structured reasoning outputs and multi-level reward analysis to provide targeted feedback to each component.

reasoningmultimodal

SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval

Aug 19, 2026

Zhenyao Cui, Siyuan Kan, Siyang Li et al.

Brain-based image retrieval can work for new users without retraining by learning to recover the geometric transformation between their brain's coordinate system and a shared visual space, using only unlabeled test-time alignment.

This paper tackles cross-subject EEG-to-image retrieval—retrieving images that match brain signals from new users without labeled training data. The key insight is that different people's brains organize visual concepts similarly but along different coordinate directions.

multimodalalignmentevaluation

Beyond Trial Averaging: Anchoring Neural and Visual Representations for Few-Repetition Brain-to-Image Retrieval

Aug 19, 2026

Zhenyao Cui, Siyuan Kan, Dingkun Liu et al.

Brain-to-image decoding can work with far fewer repetitions by anchoring both neural and visual representations to a shared reference point, rather than just denoising the query alone.

This paper tackles brain-to-image retrieval with limited neural data. Current methods require averaging 80+ brain scans per image, but the authors show that low-repetition queries fail not just due to noise, but because brain and image representations misalign.

multimodalevaluationefficiency

Intern-S2-Preview: Scientific Agentic Foundation Model

Aug 13, 2026

Lei Bai, Jiaqi Cao, Chiyu Chen et al.

This work demonstrates how to build AI agents for science by combining multimodal pre-training with agentic reinforcement learning and memory-augmented architectures, achieving strong performance on scientific reasoning and forecasting tasks without requiring task-specific model modifications.

Intern-S2-Preview is a large multimodal AI system designed to tackle scientific discovery tasks by reasoning over diverse data types, using scientific tools, and working across long-horizon problems.

agentsmultimodalreasoning

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

Aug 13, 2026

Yi-Chung Chen, Philip Jacobson, Tom Lampo et al.

Using physical trajectory data as training supervision helps video embedding models better understand motion-centric driving events, improving retrieval accuracy by 5-10% while keeping inference simple and efficient.

This paper tackles retrieving relevant driving video clips from large datasets by improving multimodal embedding models. The key innovation is TraVEL, which fine-tunes video embeddings using trajectory (vehicle motion) as training supervision, helping the model understand motion-centric events like turning or accelerating.

multimodaltrainingapplications

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Aug 13, 2026

Daniel Perkins, John Squires, Janou Milligan et al.

Instead of training separate models for each domain, you can use an MLLM router to dynamically select which vision backbone handles each image, getting both better generalization and easier updates without retraining.

ARMDIL uses a multimodal language model to intelligently route images to the best-suited vision model (CNNs, self-supervised learners, or vision-language models) within an ensemble.

multimodalarchitectureevaluation

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

Aug 13, 2026

Serli Kopar, Sam Gijsen, Abner Hernandez et al.

Speech-based disease detection models may not be learning genuine disease characteristics but rather dataset-specific patterns, raising serious concerns about their reliability for real-world clinical use.

This paper investigates whether speech models trained to detect Parkinson's disease actually learn disease-specific patterns or just exploit dataset quirks.

evaluationmultimodaldata

Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks

Aug 13, 2026

Muhammad Hannan Akram, Muhammad Abubakar Rashid, Wassi Haider Kabir et al.

Heterogeneous AI agents can efficiently synchronize their understanding in distributed networks by translating belief updates through edge-deployed models, without requiring shared training or identical architectures.

This paper proposes a framework for AI agents in 6G networks to synchronize their understanding (beliefs) despite using different AI models and operating under different constraints.

agentsmultimodalefficiency

AVA-Encoder: Towards Agent-Native Video Representation Learning

Aug 12, 2026

Chuyue Li, Jinpeng Yu, Haozhe Wang et al.

Agents can now learn from and generate high-quality videos by working with structured knowledge graph representations instead of raw pixels, improving video generation quality by 20.7% over existing methods.

AVA-Encoder learns video representations as knowledge graphs that agents can reason about and edit. It converts videos into structured text and asset layers, then reconstructs videos from these representations. A natural-language feedback loop optimizes the encoding, enabling agents to work with cinematic-quality videos more effectively.

multimodalagentsarchitecture

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Aug 12, 2026

Yan Deng, Fei Xu

For embodied AI agents navigating from visual instructions, explicitly modeling temporal context (past-only memory), multi-step planning with single-step execution, and decoupled termination detection significantly improves navigation success and efficiency.

DreamFly improves aerial drone navigation by combining three key techniques: a causal memory system that uses only past observations to avoid information leakage, a receding-horizon planning approach that predicts multiple future actions but executes one at a time, and explicit stop detection from action predictions.

agentsreasoningmultimodal

Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling

Aug 12, 2026

Pedro Sousa, Will Tebbutt, Sadiq Jaffer et al.

Foundation models trained on satellite imagery can capture persistent surface properties (terrain, vegetation, water) that explain local weather variations better than hand-crafted features, enabling more accurate probabilistic weather predictions at arbitrary locations.

This paper shows that Earth observation embeddings from satellite data can improve weather downscaling—predicting local weather at specific locations from coarse global weather models.

multimodalapplicationsefficiency

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Aug 12, 2026

Weihao Bo, Shan Zhang, Yanpeng Sun et al.

Current multimodal AI models excel at understanding diagrams visually but struggle significantly with converting them to executable code, suggesting this is a critical gap to address for scientific writing tools.

Diagram-MMU is a benchmark with 3.7k scientific diagrams and 18.3k questions that tests how well AI models can understand diagrams and convert them to code. The benchmark evaluates 12 models on three tasks: turning diagrams into LaTeX code, editing diagram code, and answering questions about diagrams—revealing that code generation is much harder for models than visual understanding.

multimodalevaluationapplications

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Aug 11, 2026

Changhao Xiang, Shangyu Xing, Zhen Wu et al.

By interleaving visual objects directly into text during pretraining, you can teach multimodal models object-level grounding 12x more efficiently than traditional image-text pair training.

This paper introduces MultiModal Code-Switching (MMCS), a new way to train vision-language models by replacing words in text with their corresponding visual objects. Instead of just pairing whole images with descriptions, MMCS explicitly shows the model which objects match which words, making training much more efficient—achieving the same performance with 12x less data.

multimodaltrainingdata

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Aug 10, 2026

Oluwanifemi Bamgbose, Simon Rosen, Jash Shah et al.

Automated TTS evaluators need to move beyond generic 'naturalness' scores and evaluate specific linguistic dimensions of speech quality, as current tools miss many errors that human listeners easily detect.

This paper reveals that current automated Text-to-Speech evaluation tools (MOS predictors and Audio-LLM judges) fail to capture the full range of speech quality issues that humans perceive.

evaluationmultimodal

Multimodal Model Diffing for Feature Discovery and Control

Aug 10, 2026

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar et al.

Sparse autoencoders can turn multimodal model internals into controllable feature interfaces: you can identify what changed during multimodal training, find features causing specific behaviors, and steer or remove them to improve safety or task performance.

This paper introduces MMDiff, a framework that uses sparse autoencoders to decompose multimodal language models into interpretable features.

multimodalsafety

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Aug 10, 2026

Diandian Zhang, Tingyu Song, Lin Fu et al.

Visual realism in video generation doesn't guarantee scientific correctness: models that look good often fail at capturing accurate scientific and causal dynamics, a gap that needs targeted evaluation and improvement.

Sci-VBench is a benchmark with 1,253 expert-annotated examples for evaluating how well AI models generate videos that require scientific knowledge and reasoning across 60 subjects in science, healthcare, humanities, and engineering.

evaluationmultimodalreasoning
evaluationmultimodal

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Aug 6, 2026

Sarvesh Baskar, Zikui Cai, Shayan Shabihi et al.

Video models fail not because they can't see events, but because they can't reliably track and count them over time—adding more frames helps slightly but doesn't fix the core temporal reasoning problem.

Video language models struggle with counting events in videos, especially when events happen frequently or repeatedly.

evaluationreasoningmultimodal

The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images

Aug 6, 2026

Zhiheng Wang, Bo Peng, Lai Wei et al.

Visual tool-use in multimodal models often creates an illusion of improvement: while aggregate accuracy may increase, the returned visual evidence frequently has no causal effect on the final answer, making these expensive operations ineffective for most queries.

This paper investigates whether multimodal AI models actually benefit from visual tool-use operations like cropping and zooming. Using causal analysis, the authors discover that these operations often don't meaningfully improve answers despite higher computational cost—models either ignore the visual evidence they retrieve or use it incoherently.

evaluationmultimodalreasoning

Depth-Guided Video Object Counting in Crowded Scenes

Aug 6, 2026

Yuanjing Xu, Xinyan Liu, Weidong Chen et al.

Adding depth information significantly improves object counting in crowded scenes—the method reduces counting errors by 62% compared to RGB-only approaches, showing that 3D spatial cues are crucial for handling occlusions.

This paper tackles counting objects in crowded, occluded video scenes by combining RGB images with depth information. The method uses a depth-guided detector with cross-attention between color and depth data, plus occlusion prediction, to better identify individual objects even when they overlap. The authors also release a new RGB-D dataset for this task.

multimodalevaluationapplications

From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks

Aug 6, 2026

Christo Kurisummoottil Thomas, Omar Hashash, Walid Saad

Networks can become active intelligence coordinators for physical AI by deploying reasoning agents (digital twins) that share spatiotemporal context through causal reasoning and transmit only beliefs with cognitive value, rather than optimizing for throughput alone.

This paper proposes holonic digital twins (HDT-Nets)—intelligent network agents that actively reason about their environment rather than passively mirroring physical systems.

agentsreasoningmultimodal

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Aug 5, 2026

Ayoub Kirouane, Christos Petrocheilos

Domain-specific fine-tuning dramatically improves retrieval and generation for underrepresented languages: a modest 1B embedder outperforms multilingual models when trained on just 65K in-domain Greek examples, showing that language adaptation is more important than model size for specialized tasks.

This paper adapts NVIDIA's Nemotron retrieval system for Modern Greek across legal, energy, financial, and medical domains. The authors mine Greek corpora, train specialized retrieval models, and create HERA—the first large-scale Greek RAG benchmark.

trainingmultimodalapplications

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

Aug 5, 2026

Aniri, Jinhe Bi, Peng Liao et al.

Modality imbalance (text drowning out vision) is a real bottleneck in multimodal reasoning. By detecting when visual input is being ignored and training only on those critical tokens, you can make self-distillation much more effective.

This paper identifies and addresses modality imbalance in multimodal language models—where text dominates over visual information during reasoning. OPD-V uses positive and negative teacher models with modified images to detect when the model isn't properly using visual input, then selectively applies self-distillation only on tokens where visual information matters most.

multimodaltrainingefficiency

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

Aug 5, 2026

Yuezhang Peng, Yuxin Liu, Changfeng Gao et al.

Framing spoken language understanding as structured function calling—rather than slot-filling—lets audio language models generalize to new tasks without retraining, similar to how code models handle function calls.

This paper introduces Spoken Function Calling (SFC), a new way to understand spoken language that treats semantic extraction like function calls with structured definitions.

multimodalapplicationstraining

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Aug 4, 2026

Yang Yang, Qinyu Zhao, Mouxiang Chen et al.

By sharing backbone parameters across parallel branches with task-specific computation allocation, you can improve multimodal model performance without increasing model size or inference latency.

ParVL introduces a framework for scaling multimodal AI models by running multiple parallel vision and language processing branches that share the same core parameters. Instead of making models bigger or slower, it reuses existing components more efficiently and lets different tasks use different amounts of vision vs.

architectureefficiencymultimodal

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Aug 4, 2026

Zhen Fang, Yu Zeng, Wenxuan Huang et al.

Multimodal agents need explicit architectural constraints to use visual tools and external knowledge rather than defaulting to text search and internal memory; decoupled perception-exploration pipelines with staged tool unlocking significantly improve performance.

Video-DeepResearch extends multimodal AI agents to handle continuous video streams with web search integration. The system addresses two key problems: agents ignoring visual information in favor of text search, and relying on memorized knowledge instead of actually using tools.

multimodalagentsreasoning
evaluationmultimodalreasoning