A generated multi-story villa with connected rooms and physically arranged furniture

Text-to-3D indoor scene generation

Text2Villa

Text2Villa connects macro-level multi-story layout generation with micro-level, physics-aware asset arrangement to create villa-scale 3D indoor environments from natural-language descriptions.

Xiang Tang1,2 Ruotong Li2 Xiaopeng Fan3,2,4

1Harbin Institute of Technology, Shenzhen 2Peng Cheng Laboratory 3Harbin Institute of Technology 4Harbin Institute of Technology, Suzhou Research Institute

arXiv

Why Text2Villa

A building is more than a stack of rooms.

Existing text-to-3D methods predominantly operate on single rooms with rectangular boundaries. They struggle with vertical connectivity, arbitrary building footprints, and the physical constraints that keep objects supported, contained, and collision-free.

Text2Villa introduces a hierarchical framework that first generates connected, polygonal multi-story foundations and then instantiates each room through an affordance-driven physical-semantic scene graph (A-PSSG) and closed-loop optimization.

Multi-story Connected foundations with usable vertical circulation.
Polygonal Flexible footprints beyond predefined rectangular rooms.
Physics-aware Support, containment, collision, and semantic constraints.

Method

Generate globally. Refine locally.

The pipeline separates architectural planning from asset instantiation, then closes the loop with geometric verification and multimodal semantic feedback.

01

Architectural layout

A fine-tuned autoregressive generator converts text into parameterized multi-story foundations.

02

A-PSSG construction

Support surfaces, containment cavities, and semantic relations become explicit graph constraints.

03

Closed-loop refinement

An MLLM and a geometry engine iteratively observe, evaluate, and modify each scene.

Stage 1 creates the building foundation, Stage 2 initializes room assets and constraints, and Stage 3 resolves physical conflicts and semantic errors.

Results

From one story to three.

A single prompt becomes a connected hierarchy of floor layouts and room-scale scenes. The media slots are ready for final rotating floor videos.

Complete Villa

Complete three-story villa with each floor exposed

Floor 1 Preview

Floor 1 Camera Trajectory

Virtual Tour

Comparison

Prompt-aligned at villa, floor, and room scale.

Text2Villa follows detailed spatial instructions while enforcing physical and semantic constraints, from global room topology to fine-grained furniture arrangement.

Villa-scale generation connects multi-story layouts with coherent room-level synthesis.

Applications

Explicit 3D scenes for downstream systems.

Generated buildings and furniture remain explicit meshes that can move into content engines for rapid authoring and into interactive environments for embodied-agent training.

Text2Villa also supports stylized text prompts, translating artistic descriptions into explicit, coherent 3D scenes whose architecture, furniture, and materials preserve the requested visual style without sacrificing spatial consistency.

In-game community authoring and an interactive environment for embodied agents
Stylized indoor scenes generated from stylized text prompts

Reference

Citation

@article{tang2026text2villa,
  title={Text2Villa: Hierarchical Generation of 3D Indoor Environments
         with Physics-Aware Analysis-by-Synthesis},
  author={Tang, Xiang and Li, Ruotong and Fan, Xiaopeng},
  journal={arXiv preprint arXiv:2607.17145},
  year={2026}
}