TABVERSE
TABVERSE is a controlled multimodal benchmark that aligns the same table content across HTML, Markdown, LaTeX, and rendered images, with question category and difficulty tags. It evaluates LLMs and VLMs on question answering, structural understanding, and structure reconstruction.
- Released
- 2026-06-08
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing table benchmarks conflate content, format, and modality, obscuring the impact of representation choice. TABVERSE enables isolation of representation effects, providing practical guidance for selecting robust table formats for downstream applications and evaluation.
Motivation
Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation remains under-explored.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.