Researchers introduced the Office Comprehension Benchmark (OCB) to assess how well large language models understand Word, Excel, and PowerPoint files in their native formats. The benchmark features two tracks: one testing structural fidelity for elements like tables and charts, and another evaluating expert-level reasoning across 12 professional domains. Responses are graded using atomic claims and an ensemble of LLM judges to ensure precise evaluation.
- First public benchmark targeting native .docx, .xlsx, and .pptx file comprehension.
- Tests visual and structural perception of complex artifacts like embedded charts and formulas.
- Evaluates multi-step reasoning and synthesis across 12 distinct industry domains.
- Uses atomic claim decomposition and LLM judge ensembles for granular scoring.