Abstract

Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find that the two views disagree on 32% of required items, mostly code that no frame shows, and that fixed views miss small close-up contents. We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds. From one prompt describing a floating voxel island with ten biomes, physics, and day/night and seasonal cycles, the judge explores the running world, controlling its clock, orbiting it, and sending a navigator agent to frame each biome, and reads the code for what it sees. Neither channel is trusted on its own: a code quote counts only if it is text the source contains, and visual claims are checked against measured pixels where the property is measurable. A mutation test, in which we remove features by construction, shows that code-only judging gives full credit to four of five removed features, because their code remains in the file. Our judge cuts the points kept on removed features by a third (5.44 to 3.55 of 7.11), and what it still credits is mostly code that exists but never runs. We evaluate five frontier models: Claude Fable 5.1, GPT-6 Astra, Kimi K3, Grok 4.7 and Gemini 3.1 Pro. Code, prompt, tests and judge configuration are available at https://github.com/KrishBakshi/worldbench

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Bakshi, K. (2026). WorldBench: Evaluating LLMs on Three.js Voxel World Generation. https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation

MLA 9

Bakshi, Krish. "WorldBench: Evaluating LLMs on Three.js Voxel World Generation." https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation.

Chicago (author–date)

Bakshi, Krish. 2026. "WorldBench: Evaluating LLMs on Three.js Voxel World Generation." https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation.

Harvard

Bakshi, K. (2026) 'WorldBench: Evaluating LLMs on Three.js Voxel World Generation', Available at: https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation.

Vancouver

Bakshi K. WorldBench: Evaluating LLMs on Three.js Voxel World Generation. https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation

IEEE

K. Bakshi, "WorldBench: Evaluating LLMs on Three.js Voxel World Generation," https://omanscience.com/en/articles/worldbench-evaluating-llms-on-three-js-voxel-world-generation.