Benchmark Radar
AI BENCHMARK PROFILE

TABVERSE

General AIMultimodal Perception

TABVERSE is a controlled multimodal benchmark that aligns the same table content across HTML, Markdown, LaTeX, and rendered images, with question category and difficulty tags. It evaluates LLMs and VLMs on question answering, structural understanding, and structure reconstruction.

Released
2026-06-08
Readiness
Paper only
Primary field
General AI

Why it matters

Existing table benchmarks conflate content, format, and modality, obscuring the impact of representation choice. TABVERSE enables isolation of representation effects, providing practical guidance for selecting robust table formats for downstream applications and evaluation.

Motivation

Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.