Papers

1

Total Citations

4

H-Index

1

About

Roy Lee is a rising researcher at the intersection of computer vision and natural language processing, with a primary focus on multimodal learning and visual question answering (VQA). His most cited work, "Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts" (2023, 4 citations), introduces a novel approach to enhancing VQA systems by integrating external knowledge through enriched prompts. This contribution addresses a critical limitation in standard multimodal models—their reliance solely on visual and textual inputs—by enabling them to leverage structured knowledge bases for more accurate and context-aware answers. Lee’s research pushes the boundaries of how machines interpret complex visual scenes and natural language queries, with potential applications in assistive technologies, automated customer support, and educational tools. Though early in his career, his work demonstrates a clear commitment to advancing human-AI interaction. As the field of multimodal AI rapidly evolves, Lee’s focus on knowledge-enriched prompting positions him as a promising voice in making AI systems not just more capable, but more intelligently responsive to the nuanced questions people ask.

Research Focus

Key Achievements

1
H-Index
1
Papers
4
Total Citations
4
Avg Citations/Paper
🏆 Most Cited Paper
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts
4 citations · 2023
📈 Most Prolific Year: 2023 (1 Papers)
🤝 Key Collaborators: 4
🏛 Institutions: Singapore University of Technology and Design

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 11 days ago