Text to 3D Scene Generation with Rich Lexical Grounding
Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning
- Year
- 2015
- Citations
- 62
- Access
- Open access
Abstract
The ability to map descriptions of scenes to 3D geometric representations has many applications in areas such as art, educa-tion, and robotics. However, prior work on the text to 3D scene generation task has used manually specified object cate-gories and language that identifies them. We introduce a dataset of 3D scenes an-notated with natural language descriptions and learn from this data how to ground tex-tual descriptions to physical objects. Our method successfully grounds a variety of lexical terms to concrete referents, and we show quantitatively that our method im-proves 3D scene generation over previ-ous work using purely rule-based meth-ods. We evaluate the fidelity and plau-sibility of 3D scenes generated with our grounding approach through human judg-ments. To ease evaluation on this task, we also introduce an automated metric that strongly correlates with human judgments. 1
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991