OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. Authors: Xianyun Sun, Chaoyou Fu, Zhengye Zhang.
Why it matters
Read this for the paper's specific claim in Artificial Intelligence / Machine Learning: Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals.